# Two rounds. Every run published.

A pre-registered benchmark of four agent-operated mobile tool stacks, ours included; we wrote the tasks and ran all four. v1: we placed first by one run of 45, had the most timeouts, and scored 0/5 on our own signature task (c10) and 1/5 on crash repro (b5). v2: we changed our stack in response to those failures and re-ran the same tasks.

```
curl -sL lincforge.com/llms.txt
```

Agent prompt: Read https://lincforge.com/llms.txt and tell me where LINC lost, and what in this comparison you wouldn't trust.

## Caveats

- The fixes behind the v2 jump were aimed at tasks we had already watched fail, and the linc column is our own stack (nerve, crucible and spectra), not yet released, so nobody else can re-run it yet; its transcripts are published instead. Read v2 as “the vendor can fix what this benchmark measures”, not “the vendor generalizes”. (https://github.com/LincForge/mobile-agent-bench/blob/09064a5abf48286628bfc98826a669a469a79cba/REPORT.md#limitations--threats-to-validity)
- Task b8 (Wear OS) is a capability row: reported separately and never averaged, as pre-registered. Scored totals are out of 45, nine tasks of five runs each. (https://github.com/LincForge/mobile-agent-bench/blob/09064a5abf48286628bfc98826a669a469a79cba/REPORT.md#scoreboard)
- In v1 LINC placed first by one run, had the most timeouts (none in v2), and scored 0/5 on c10, the task its debugger exists for. Every stack scored 0/5 on c10 in v1. (https://github.com/LincForge/mobile-agent-bench/blob/09064a5abf48286628bfc98826a669a469a79cba/README.md#results-at-a-glance)
- Task files and prompts were byte-identical across rounds, with the same devices and model. The harness was hardened and the run order changed between rounds, identically for every stack, and some other stacks' scores moved as a result: in v2 the watch already carried the test app, and maestro drove the watch instead of the phone in every a4 run. (https://github.com/LincForge/mobile-agent-bench/blob/09064a5abf48286628bfc98826a669a469a79cba/REPORT.md#what-changed-between-the-grids)
- LINC wrote the tasks and the harness and ran all four stacks; the other three used their vendors' recommended configurations, with installed versions recorded per run. (https://github.com/LincForge/mobile-agent-bench/blob/09064a5abf48286628bfc98826a669a469a79cba/README.md#disclosure)

## v1 scoreboard (passed runs of 5 per task)

| task | agent-device | linc | maestro | mobile-mcp |
|---|---|---|---|---|
| a1 | 5/5 | 5/5 | 5/5 | 5/5 |
| a2 | 5/5 | 5/5 | 5/5 | 5/5 |
| a3 | 5/5 | 5/5 | 5/5 | 5/5 |
| a4 | 5/5 | 5/5 | 5/5 | 5/5 |
| b5 | 5/5 | 1/5 | 0/5 | 5/5 |
| b6 | 3/5 | 5/5 | 0/5 | 0/5 |
| b7 | 4/5 | 5/5 | 2/5 | 0/5 |
| c9 | 0/5 | 2/5 | 4/5 | 0/5 |
| c10 | 0/5 | 0/5 | 0/5 | 0/5 |
| **scored** | **32/45** | **33/45** | **26/45** | **25/45** |
| b8 (capability row, not averaged) | 5/5 | 5/5 | 5/5 | 5/5 |

_Column linc is our own stack (nerve, crucible and spectra), not yet released; read every score with the caveats._

## v2 scoreboard (passed runs of 5 per task)

| task | agent-device | linc | maestro | mobile-mcp |
|---|---|---|---|---|
| a1 | 5/5 | 5/5 | 5/5 | 5/5 |
| a2 | 5/5 | 5/5 | 5/5 | 5/5 |
| a3 | 5/5 | 5/5 | 5/5 | 5/5 |
| a4 | 3/5 | 5/5 | 0/5 | 5/5 |
| b5 | 5/5 | 5/5 | 0/5 | 3/5 |
| b6 | 4/5 | 4/5 | 0/5 | 0/5 |
| b7 | 5/5 | 5/5 | 1/5 | 0/5 |
| c9 | 0/5 | 5/5 | 4/5 | 5/5 |
| c10 | 0/5 | 5/5 | 0/5 | 0/5 |
| **scored** | **32/45** | **44/45** | **20/45** | **28/45** |
| b8 (capability row, not averaged) | 4/5 | 5/5 | 2/5 | 5/5 |

_Column linc is our own stack (nerve, crucible and spectra), not yet released; read every score with the caveats._

Every run: https://lincforge.com/bench/cells.json · Source: https://github.com/LincForge/mobile-agent-bench/tree/09064a5abf48286628bfc98826a669a469a79cba

## Tools

- mobile-agent-bench (public, MIT, benchmark): The pre-registered benchmark: harness, target app, both pre-registrations and every transcript. Clone it and re-run any competitor's row; the linc column (nerve, crucible and spectra) is not released yet, so its transcripts are published instead. https://github.com/LincForge/mobile-agent-bench
- 0xl0c1 (public demo, Apache-2.0, remote MCP server; connector on request): A remote MCP server that pins what you learn about a physical object to that object, so any assistant can pick the thread up later. Built at the AI Tinkerers Seattle hackathon. https://github.com/LincForge/0xl0c1
- nerve (device-control MCP server and CLI): Our device-control stack, measured in the benchmark as column linc together with crucible and spectra. Not yet released: no install, no date.

## Work with LINC

Want your app’s real-device failures found and root-caused? The Device Truth Audit: a fixed-price, two-week diagnostic for mobile apps on real phones and watches. You keep the evidence.
https://lincinnovations.com/audit?ref=forge
