# LINC Forge > Developer site of LINC Innovations. Read the caveat before the numbers. ## Caveat The fixes behind the v2 jump were aimed at tasks we had already watched fail, and the linc column is our own stack (nerve, crucible and spectra), not yet released, so nobody else can re-run it yet; its transcripts are published instead. Read v2 as “the vendor can fix what this benchmark measures”, not “the vendor generalizes”. Task b8 (Wear OS) is a capability row: reported separately and never averaged, as pre-registered. Scored totals are out of 45, nine tasks of five runs each. In v1 LINC placed first by one run, had the most timeouts (none in v2), and scored 0/5 on c10, the task its debugger exists for. Every stack scored 0/5 on c10 in v1. Task files and prompts were byte-identical across rounds, with the same devices and model. The harness was hardened and the run order changed between rounds, identically for every stack, and some other stacks' scores moved as a result: in v2 the watch already carried the test app, and maestro drove the watch instead of the phone in every a4 run. LINC wrote the tasks and the harness and ran all four stacks; the other three used their vendors' recommended configurations, with installed versions recorded per run. ## Benchmark: mobile-agent-bench (MIT) Pre-registered; four MCP tool stacks driven by the same confined agent, pinned model and physical devices; every run published with its transcript. - v2, scored out of 45: agent-device 32/45 · linc 44/45 · maestro 20/45 · mobile-mcp 28/45 - v1, scored out of 45: agent-device 32/45 · linc 33/45 · maestro 26/45 · mobile-mcp 25/45 - v1: LINC placed first by one run of 45, had the most timeouts, and scored 0/5 on c10. - Scoreboards: https://lincforge.com/bench/v2.md · https://lincforge.com/bench/v1.md · https://lincforge.com/bench/summary.json - Every run (400 records with transcript links): https://lincforge.com/bench/cells.json - Source at the pinned commit: https://github.com/LincForge/mobile-agent-bench/tree/09064a5abf48286628bfc98826a669a469a79cba ## Tools - mobile-agent-bench (public, MIT, benchmark): The pre-registered benchmark: harness, target app, both pre-registrations and every transcript. Clone it and re-run any competitor's row; the linc column (nerve, crucible and spectra) is not released yet, so its transcripts are published instead. https://github.com/LincForge/mobile-agent-bench - 0xl0c1 (public demo, Apache-2.0, remote MCP server; connector on request): A remote MCP server that pins what you learn about a physical object to that object, so any assistant can pick the thread up later. Built at the AI Tinkerers Seattle hackathon. https://github.com/LincForge/0xl0c1 - nerve (device-control MCP server and CLI): Our device-control stack, measured in the benchmark as column linc together with crucible and spectra. Not yet released: no install, no date. ## Work with LINC - The Device Truth Audit: a fixed-price, two-week diagnostic for mobile apps on real phones and watches. You keep the evidence. https://lincinnovations.com/audit?ref=forge Human view: https://lincforge.com/ · Markdown view: https://lincforge.com/index.md