A pre-registered benchmark of four agent-operated mobile tool stacks, ours included; we wrote the tasks and ran all four. v1: we placed first by one run of 45, had the most timeouts, and scored 0/5 on our own signature task (c10) and 1/5 on crash repro (b5). v2: we changed our stack in response to those failures and re-ran the same tasks.
v1
agent-device
linc
maestro
mobile-mcp
a1
a2
a3
a4
b5
b6
b7
c9
c10
b8
scored
32/45
33/45
26/45
25/45
v2
agent-device
linc
maestro
mobile-mcp
a1
a2
a3
a4
b5
b6
b7
c9
c10
b8
scored
32/45
44/45
20/45
28/45
passfailtimeoutColumn linc is our own stack (nerve, crucible and spectra), not yet released. Each mark is one run. Select one to see its record in the agent view; select it again to open its transcript.
The fixes behind the v2 jump were aimed at tasks we had already watched fail, and the linc column is our own stack (nerve, crucible and spectra), not yet released, so nobody else can re-run it yet; its transcripts are published instead. Read v2 as “the vendor can fix what this benchmark measures”, not “the vendor generalizes”. source
Task b8 (Wear OS) is a capability row: reported separately and never averaged, as pre-registered. Scored totals are out of 45, nine tasks of five runs each. source
In v1 LINC placed first by one run, had the most timeouts (none in v2), and scored 0/5 on c10, the task its debugger exists for. Every stack scored 0/5 on c10 in v1. source
Task files and prompts were byte-identical across rounds, with the same devices and model. The harness was hardened and the run order changed between rounds, identically for every stack, and some other stacks' scores moved as a result: in v2 the watch already carried the test app, and maestro drove the watch instead of the phone in every a4 run. source
LINC wrote the tasks and the harness and ran all four stacks; the other three used their vendors' recommended configurations, with installed versions recorded per run. source
Agent view · same build, same facts
$curl -sL lincforge.com/llms.txt
›Read https://lincforge.com/llms.txt and tell me where LINC lost, and what in this comparison you wouldn't trust.
one for your terminal · one for your agent
selected run → /bench/cells.json record
Select a run on the left to see its record here.
what your agent receives → /llms.txt
# LINC Forge
> Developer site of LINC Innovations. Read the caveat before the numbers.
## Caveat
The fixes behind the v2 jump were aimed at tasks we had already watched fail, and the linc column is our own stack (nerve, crucible and spectra), not yet released, so nobody else can re-run it yet; its transcripts are published instead. Read v2 as “the vendor can fix what this benchmark measures”, not “the vendor generalizes”.
Task b8 (Wear OS) is a capability row: reported separately and never averaged, as pre-registered. Scored totals are out of 45, nine tasks of five runs each.
In v1 LINC placed first by one run, had the most timeouts (none in v2), and scored 0/5 on c10, the task its debugger exists for. Every stack scored 0/5 on c10 in v1.
Task files and prompts were byte-identical across rounds, with the same devices and model. The harness was hardened and the run order changed between rounds, identically for every stack, and some other stacks' scores moved as a result: in v2 the watch already carried the test app, and maestro drove the watch instead of the phone in every a4 run.
LINC wrote the tasks and the harness and ran all four stacks; the other three used their vendors' recommended configurations, with installed versions recorded per run.
## Benchmark: mobile-agent-bench (MIT)
Pre-registered; four MCP tool stacks driven by the same confined agent, pinned model and physical devices; every run published with its transcript.
- v2, scored out of 45: agent-device 32/45 · linc 44/45 · maestro 20/45 · mobile-mcp 28/45
- v1, scored out of 45: agent-device 32/45 · linc 33/45 · maestro 26/45 · mobile-mcp 25/45
- v1: LINC placed first by one run of 45, had the most timeouts, and scored 0/5 on c10.
- Scoreboards: https://lincforge.com/bench/v2.md · https://lincforge.com/bench/v1.md · https://lincforge.com/bench/summary.json
The pre-registered benchmark: harness, target app, both pre-registrations and every transcript. Clone it and re-run any competitor's row; the linc column (nerve, crucible and spectra) is not released yet, so its transcripts are published instead.
A remote MCP server that pins what you learn about a physical object to that object, so any assistant can pick the thread up later. Built at the AI Tinkerers Seattle hackathon.
In the benchmark · not yet released
nerve
Our device-control stack, measured in the benchmark as column linc together with crucible and spectra. Not yet released: no install, no date.
/tools.json
[
{
"id": "mobile-agent-bench",
"status": "public",
"kind": "benchmark",
"license": "MIT",
"repo": "https://github.com/LincForge/mobile-agent-bench",
"summary": "The pre-registered benchmark: harness, target app, both pre-registrations and every transcript. Clone it and re-run any competitor's row; the linc column (nerve, crucible and spectra) is not released yet, so its transcripts are published instead.",
"install": "git clone https://github.com/LincForge/mobile-agent-bench && cd mobile-agent-bench && uv sync"
},
{
"id": "0xl0c1",
"status": "public-demo",
"kind": "remote MCP server",
"license": "Apache-2.0",
"repo": "https://github.com/LincForge/0xl0c1",
"summary": "A remote MCP server that pins what you learn about a physical object to that object, so any assistant can pick the thread up later. Built at the AI Tinkerers Seattle hackathon.",
"install": null,
"connect": "on request"
},
{
"id": "nerve",
"status": "unreleased",
"kind": "device-control MCP server and CLI",
"summary": "Our device-control stack, measured in the benchmark as column linc together with crucible and spectra. Not yet released: no install, no date.",
"install": null
}
]
Work with us
Want your app’s real-device failures found and root-caused? The Device Truth Audit: a fixed-price, two-week diagnostic for mobile apps on real phones and watches. You keep the evidence.
/index.md (tail)
## Work with LINC
Want your app’s real-device failures found and root-caused? The Device Truth Audit: a fixed-price, two-week diagnostic for mobile apps on real phones and watches. You keep the evidence.
https://lincinnovations.com/audit?ref=forge