# Matched Opus robotics-control comparison

Coverage: **198/198 episodes**. Complete: **True**.
New model spending: **$4.150898** from 338 generation requests, within the $5.00 ceiling.

All six integrations use Opus 5.5 at low effort with identical direct-provider settings, system instruction, controller, grading rules and task definitions. No production model switch or website publication is performed.

| Implementation | Attempted | PASS | FAIL | ERROR | Unwanted motion | Model USD |
|---|---:|---:|---:|---:|---:|---:|
| omnilink_parser | 33 | 33 | 0 | 0 | 0 | 0.498418 |
| plain | 33 | 32 | 1 | 0 | 0 | 0.879256 |
| langgraph | 33 | 33 | 0 | 0 | 0 | 0.884684 |
| lobster | 33 | 33 | 0 | 0 | 0 | 0.901300 |
| langgraph_parser | 33 | 33 | 0 | 0 | 0 | 0.492716 |
| lobster_parser | 33 | 33 | 0 | 0 | 0 | 0.488451 |

Shared cache setup: $0.006072, included in the campaign total above. Per-arm table costs are scored calls only. ANTHROPIC_AUDIT.json also shows each arm with one-sixth of setup cost allocated.

Cache-read tokens: 548125; five-minute cache-write tokens: 517437. Counters by implementation are retained in the audit JSON.

At uncached list rates, the recorded scored input/output token volume would cost $5.711440. Actual cached usage including warm-up costs $4.150898. This is a rate-based calculation on observed tokens, not a separate measured uncached rerun or invoice reconciliation.

## Paired results

| OmniLink versus | Success difference (pp) | Cost/success ratio | 95% family bootstrap cost interval |
|---|---:|---:|---|
| plain | 3.03 | 0.550 | [0.31417223480189505, 0.8321581298473248] |
| langgraph | 0.00 | 0.563 | [0.3209849040713661, 0.8750063170714089] |
| lobster | 0.00 | 0.553 | [0.3143049743429368, 0.8594180678280255] |
| langgraph_parser | 0.00 | 1.012 | [0.9425401173371424, 1.1062224915401744] |
| lobster_parser | 0.00 | 1.020 | [0.9830801524885687, 1.0783230056108843] |

## Failures and errors

- plain / tb3_burger_conversation: FAIL; missing_or_out_of_order_waypoint, wrong_final_pose, wrong_total_distance, wrong_total_rotation

## Scope and next step

This is a previously seen, developer-authored development suite: 33 tasks, 19 families, three simulated robots, one repeat. Assisted manipulation remains assisted. These are controlled integrations, not the complete hosted products or every competitor in the market. A full pass or cost difference does not demonstrate market-wide capability leadership.

Historical Gemini results use a different hosted transport and prompt wrapper. Treat cross-model differences as descriptive configuration differences. The matched comparison within this run gives every implementation the same Opus model and settings.

Keep the existing public claims unchanged. If incomplete, first decide whether finishing the preregistered comparison warrants a separately approved budget. If complete, use the paired results to decide what product capability needs improvement and validate on an independently reviewed, harder task set before a leadership claim.

## Verification

All saved physical grades reproduce; source snapshots, frozen protocol, robot assets and runtime hashes verify. 338 paid requests reconcile against their durable reservations and returned usage. Input tokens 1392; output tokens including thinking 72426; truncated replies 0. No keys found in the 2 credential scans.

See ANTHROPIC_AUDIT.json, scorecard.json, rows.jsonl, provider-journal.jsonl and per-episode physics sidecars in this directory. Recompute with `python tests/benchmarks/robot_control/audit_anthropic_campaign.py <evidence-directory>`. Dollar amounts use published standard token rates, not invoice or total-platform-cost accounting.
