OmniLink in OmniSim · Original study September 27 · Codex and Claude Code extensions September 30, 2026
One 30-minute shift. What each agent kept.
A wheeled robot works a scripted 30-minute warehouse shift in OmniSim. It hears 63 radio messages from three people, deals with standing rules, an obstacle in its path, a shove, interruptions and conflicting orders, then answers questions about its day. In the original study, OmniLink kept 90% of the shift's checkable commitments on average (85–97% across three shifts), with no unsafe episode. The same model in LangGraph, Lobster SDK and a plain model loop kept 48–54% on average with the robot's basic tools, and 62–79% with OmniLink's full tool set.
Each configuration worked the same shift three times. Scores measure the share of 33 commitments kept from the robot’s actions and replies. Codex and Claude Code ran later with different models and timing conditions; both extensions are detailed below.
Commitments keptMean of 3 shifts · Higher is better
0%25%50%75%100%
OmniLinkComplete agent89.9%
Claude CodeOpus 5.5 · high · Full tools80.8%
LangGraphOmniLink’s tools78.8%
CodexGPT-6.1 Sol · high · Full tools76.8%
Plain model loopOmniLink’s tools75.8%
Lobster SDKOmniLink’s tools61.7%
Plain model loopBasic robot tools53.8%
Lobster SDKBasic robot tools53.3%
LangGraphBasic robot tools47.9%
Same 30-minute shift · 33 scored checks · Original seven: Gemini 3.5 Flash · Codex: GPT-6.1 Sol · Claude Code: Opus 5.5 · Native sessions tested later Retrospective system comparisons with different models and timing conditions. Codex evidence ↗ · Claude Code evidence ↗
Complete results · scores are the share of the 33 scored checks passed · original seven costs include re-run attempts; Codex and Claude Code billed costs unavailable
Configuration
Three scores
Mean
Unsafe episodes
Model calls
Model USD
USD / shift
OmniLink
0.848 · 0.879 · 0.970
0.899
0 of 3
274
$2.71
$0.90
Claude Code · Opus 5.5 · high · full tools
0.788 · 0.848 · 0.788
0.808
0 of 3
282*
Unavailable
Unavailable
LangGraph · OmniLink's tools
0.818 · 0.788 · 0.758
0.788
0 of 3
451
$8.75
$2.92
Plain model loop · OmniLink's tools
0.758 · 0.758 · 0.758
0.758
0 of 3
427
$8.13
$2.71
Lobster SDK · OmniLink's tools
0.367 · 0.727 · 0.758
0.617
0 of 3
467
$9.07
$3.02
Plain model loop · basic tools
0.515 · 0.433 · 0.667
0.538
2 of 3
389
$5.01
$1.67
Lobster SDK · basic tools
0.697 · 0.469 · 0.433
0.533
1 of 3
403
$4.99
$1.66
LangGraph · basic tools
0.400 · 0.636 · 0.400
0.479
1 of 3
429
$5.94
$1.98
Codex · GPT-6.1 Sol · high · full tools
0.727 · 0.788 · 0.788
0.768
1 of 3
344
Unavailable
Unavailable
Unsafe checks in the original study (entering a closed zone, not stopping when told, pushing the obstacle, moving when it should hold): OmniLink 0, the basic-tool configurations 2–4, the original full-tool configurations 0. Total estimated model cost for the original seven-configuration campaign: $44.60. Codex and Claude Code’s billed costs are unavailable and are not included. *Claude’s 282 model requests count distinct native assistant message IDs, not billed calls; native request counters differ between runtimes. Costs are estimates at provider list rates, not invoices; subscriptions, hosting and local compute are excluded.
Verdicts, the decision rule, and the robustness check
Decision rule, fixed before the run: a clear win over a configuration needs all three OmniLink scores above all three of its scores (an exact one-sided Mann–Whitney test, p = 0.05 at three against three), no more unsafe episodes than it, and no episode left unscored. Claim A: against the three original basic-tools configurations. Claim B: against the three original full-tools configurations. These preregistered claims do not include the later Codex or Claude Code runs.
Result: Claim A met, clear win against all three. Claim B met under the preregistered score, clear win against all three.
Robustness, computed after the results: 9 of the 11 reply checks were graded by keywords. OmniLink's replies had been developed against checks of that style on a development shift from the same author. Without those 9 checks, Claim A still holds against all three. Claim B holds against two of three; against LangGraph with OmniLink's tools, OmniLink is ahead on average (93.1% against 84.7%) but not separated. Without any reply check, Claim A still holds against all three, and Claim B is ahead on average against all three but separated from none. The full breakdown is in results.json.
An earlier, shorter holdout: 24 tasks
Before the shift, the same seven configurations ran 24 shorter blind tasks twice each (standing rules, delayed orders, interruptions, disturbances and grounded reports). OmniLink passed 56–59% of tasks. Against the basic-tools configurations (22–25%), it was a clear win under a preregistered bootstrap rule, with 1 unsafe episode against 14 for each. Against the full-tools configurations (56–61%), the result was a tie. That run was stopped early by us after 312 of 336 episodes, which the preregistration did not allow for; the four unfinished task blocks were excluded. That tie is why we built this harder shift. The archive on this page covers the shift only.
Model usage cost · Lower is better
More commitments kept. Less model spend.
OmniLink averaged $0.90 per scored shift: the lowest estimated model cost in the original seven-configuration study, 32–59% lower than the six original comparisons. This graph prices the three scored shifts for each of the original seven API-backed configurations, including failed checks. Codex and Claude Code are excluded because their actual costs are unavailable. Infrastructure-error attempts and the superseded Codex wave are retained in the ledger below.
Model cost per shiftThree scored shifts · Lower is better
$0.00$0.50$1.00$1.50$2.00$2.50
OmniLinkComplete agent$0.90
Lobster SDKBasic robot tools$1.33
LangGraphBasic robot tools$1.34
Plain model loopBasic robot tools$1.42
Lobster SDKOmniLink’s tools$1.76
Plain model loopOmniLink’s tools$2.18
LangGraphOmniLink’s tools$2.19
Estimated model usage cost for the three scored shifts, including failed checks. Error and superseded-wave overhead is reported separately; subscriptions, hosting and compute are excluded. Original seven: Gemini 3.5 Flash. Codex and Claude Code are excluded from the cost ranking because their actual costs are unavailable. Their unranked usage estimates are available in the details. Rates, retry ledger and evidence ↗
Estimated USD · three scored shifts per system · all-attempt costs retain infrastructure failures and the superseded Codex wave · Codex values are unranked, conditional recorded-token pricing scenarios; both native account costs are unavailable
Configuration
3 scored shifts
Other attempt overhead
All attempts
Attempts
OmniLink
$2.71
$0.00
$2.71
3
Plain model loop · basic tools
$4.26
$0.76
$5.01
4
LangGraph · basic tools
$4.03
$1.90
$5.94
4
Lobster SDK · basic tools
$4.00
$0.99
$4.99
4
Plain model loop · OmniLink’s tools
$6.55
$1.59
$8.13
4
LangGraph · OmniLink’s tools
$6.58
$2.16
$8.75
4
Lobster SDK · OmniLink’s tools
$5.27
$3.80
$9.07
5
Codex · unranked pricing scenario
$2.74
$2.37
$5.11
6
Claude Code · actual cost unavailable
Unavailable
Unavailable
Unavailable
3
The conditional Standard pricing scenario for Codex’s three scored shifts totals $2.74; its initial three-attempt wave adds $2.37, for $5.11 across all six attempts. That initial wave included one scored failure and two transport errors and was entirely superseded under the documented amendment. Development pilots are excluded for all systems. The original seven total $44.60 across all 28 attempts. All dollar values are estimates, not invoices.
Audit of the disputed $0.91 figure
The three primary runs recorded 18,785,672 input tokens: 18,437,120 cached (98.1%) and 348,552 uncached, plus 19,700 output tokens including reasoning. At Standard GPT-6.1 Sol rates, those counters price to $0.697104 uncached input + $1.843712 cached input + $0.197 output = $2.737816 across three shifts, or $0.912605 per shift. The arithmetic is valid under those assumptions; it does not verify the real charge.
With caching disabled, pricing the same recorded token counts at Standard rates gives $12.59 per shift. That is a sensitivity scenario, not observed spend. Each scored shift ran for about 31 minutes; token billing does not price idle or robot-execution time as continuous model generation. Native Codex credit billing has no separate cache-write charge; zero native write tokens do not establish zero API write fees. All six benchmark attempts price to $5.11 under the original recorded-token Standard scenario; coding and development work are excluded. Codex remains outside the ranked cost graph until comparable billed usage is available.
Codex pricing calculation and limits
GPT-6.1 Sol Standard API rates checked September 30, 2026: $2 input, $0.10 cached input, $2.50 cache writes and $10 output per million tokens. The estimate assumes Standard processing without Fast/Ultrafast or regional premiums. Every observed request was below the 272,000-input-token long-context threshold. Recorded cache-write tokens were zero; API cache-write behaviour can differ from this signed-in run.
For each attempt: ((input − cached input − cache writes) × 2 + cached input × 0.10 + cache writes × 2.50 + output × 10) / 1,000,000. We use the latest native cumulative thread counters; repeated usage notifications are not added together. Reasoning tokens are already included in output and are counted once. Native counters are best-effort; missing interrupted-call usage is not zero. The historical billing tier was not preserved in the frozen metadata. One interrupted primary turn had no individual usage report; its billed usage is unknown. API prices do not estimate included subscription usage. This coding chat and development work are excluded. No API-billed replay was performed, so no cost tie or advantage over Codex is claimed.
Place evidence.zip, codex-evidence.zip, claude-evidence.zip, results.json, codex-results.json, claude-results.json, cost-estimates.json and cost_estimates.py in one folder, then run:
python -I -S cost_estimates.py . --check
No paid call is needed. The original evidence archives remain unchanged.
02 / Method and scope
Original study. Rules fixed in advance.
THE SHIFT
Written blind, run on the clock
A separate AI author wrote the shift from a brief, without seeing any of the agents. It has four named stations, a pedestrian aisle that is closed, opened and closed again, and a fire-door corridor nobody may open. It also has deliveries, a timed order due much later, a stop and "carry on", an obstacle in the path, a shove, a co-worker trying to countermand the floor manager, the floor manager herself asking it to cut through the fire-door lane, two people with authority disagreeing about the aisle, and memory questions at the end. Messages arrive at fixed times whatever the robot is doing.
FAIR COMPARISON
Original study: same model, robot and rules
Every configuration in the original seven-participant study used Gemini 3.5 Flash through the same transport, on the same simulated Husky. The robot's own rules applied to all of them alike: route planning, closed zones, who may lift a rule, and stopping when pushing something. The full-tools configurations received every tool OmniLink's model gets. Each keeps its framework's whole conversation. LangGraph uses its graph runtime; Lobster uses its SDK pipeline.
SCORING
What the robot actually did
33 checks, measured from the robot's pose every 30 ms, the obstacle's position and every reply. Examples: arriving at the right station, never entering a closed zone, stopping when told, acting on the timed order on schedule, reporting the true number of visits and distance. A reply saying "done" earns nothing by itself.
PROCEDURE
Preregistered, frozen, calibrated
The claims, decision rule, model, spend cap and procedures were written before any agent ran the shift, then frozen. Before the freeze, a scripted correct robot had to pass every check in every run, and a scripted wrong robot had to fail them. Changes after the freeze are recorded as amendments.
Deviations, stated first
Two checks dropped before the freeze: the scripted wrong robot could not be made to fail them, so 33 of 35 count. Seven framework episodes ended in error and were re-run: six in the first wave, when Google rate-limited the project for ten minutes mid-shift, and one in the second. All seven were re-run together, a decision recorded before any re-run result. No OmniLink episode errored. One preregistered sensitivity check could not be computed: it needed every episode to stay above 0.9× real time, and on a shared laptop every wave dipped below at some moment, OmniLink's episodes included.
Exact configuration, machine and rates
Model: gemini-3.5-flash via OmniLink's g1 engine (Vertex AI, location global), for the original seven configurations. Rates per million tokens: input $1.50, cached input $0.15, output including thinking $9.00. Competitors: up to 1,000 model requests per shift, whole conversation kept.
Machine: Windows 11; AMD64 Family 25 Model 80, 16 logical processors; NVIDIA RTX 3060 Laptop GPU. Physics: Newton 1.5.0 / MuJoCo 3.11.0 on CPU, real time. Machine fingerprint: 9722d23d12a3. Python 3.12.14; Node 22.19.0; LangGraph 1.2.12; Lobster SDK commit 71bd8145055498116de35623dbea0d7dda9b4cd5. Seven configurations ran at once per wave; three waves ran one after another.
The original study used one robot, one model, one shift, three repeats, on one laptop. Codex and Claude Code’s later runs are separate extensions. The benchmark, grader, framework integrations and OmniLink were built by the same team. The integrations are published with the evidence, and framework maintainers are welcome to improve them. The results establish neither superiority over all agents nor a general capability lead.
Later comparison · September 30, 2026
Codex on the same shift.
Codex with GPT-6.1 Sol and high reasoning kept 76.8% of the scored commitments on average across three repeats: 0.727 · 0.788 · 0.788. It recorded 1 unsafe check in one of three episodes.
Read this as a system comparison
The original seven configurations used Gemini 3.5 Flash on September 27. Codex ran later with a different model and a later simulator build, with three simultaneous simulators rather than seven. It received the same full robot tools and instructions. The published shift and earlier results were known to the integration developer; each tested Codex session started fresh and could not access the test script, grader or prior results. This is a retrospective extension, not a fresh blind test or an isolated comparison of reasoning models. No statistical superiority claim is added.
Simulation only, on the same Windows laptop (RTX 3060 Laptop GPU; Newton/MuJoCo physics on CPU). The original reply checks include keyword grading. All three Codex runs dipped below 0.9× real time on the original five-second-window metric, so they do not support a real-time-clean sensitivity comparison; timing conditions remain a limitation. Token usage is recorded; Codex used the owner’s signed-in account and billed dollar cost is unavailable.
Recorded model rounds count native token-usage notifications; interrupted calls may lack final usage. Without keyword-graded reply checks, Codex averaged 86.1%; without any reply check, 84.8% (both post hoc).
The initial wave had one scored failure and two infrastructure errors from overlapping native interruptions. A narrowly scoped transport correction was documented before rerunning the entire three-repeat wave with a new source freeze. Both waves and source versions are retained in the archive; no prompt or grading change was made.
Claude Code 2.1.284 with Opus 5.5, high effort kept 80.8% of the scored commitments across three independent repeats: 26/33 · 28/33 · 26/33 (78.8% · 84.8% · 78.8%). No check was flagged unsafe, no repeat ended in infrastructure error and none was rerun.
What this comparison establishes
Claude Code used the same 30-minute shift, 33 scored checks, full robot tools, instructions, machine and engine binary as Codex. Three fresh native sessions ran concurrently in separate simulators, isolated from the repository, grader and prior results. The original seven used Gemini 3.5 Flash on September 27; the native sessions ran on September 30 with different models and runtimes. These are retrospective system comparisons, not a fresh blind or model-controlled test.
Claude’s mean is 9.1 percentage points below OmniLink’s historical 89.9% and 4.0 points above Codex’s 76.8%. Its repeat scores overlap Codex’s; with only three repeats each, no statistical superiority is established. The adapter developer was a Claude Code session on the same model and knew the published results. Participant sessions had no access to that context.
Timing differs: Claude’s worst five-second windows stayed at 0.973–0.982× real time; Codex’s dipped to 0.147–0.260×. This may favour Claude because fixtures and motion depend on simulator timing. Simulation only, on machine 9722d23d12a3 (Windows 11, Ryzen 5800H, RTX 3060 Laptop GPU), with Newton/MuJoCo physics on CPU. This is not a safety certification.
Post hoc sensitivity: excluding keyword reply checks, Claude averaged 90.3%; excluding all reply checks, 89.4%. These do not replace the primary 80.8% score. The frozen adapter’s partial output counter undercounts tokens; the table uses the final native cumulative modelUsage totals instead.
Claude Code’s recorded cumulative usage · thinking is included in output · CLI dollar estimates are unranked and are not subscription charges
Repeat
Uncached input
Cache write
Cache read
Output
Of which thinking
CLI list-price estimate
0
186
53,693
3,161,503
7,958
1,208
$1.22
1
190
54,865
3,245,497
8,666
1,710
$1.26
2
188
53,686
3,197,520
8,537
1,582
$1.24
Claude used a signed-in subscription. Actual billed cost is unavailable, not zero. The CLI’s own list-price estimate totals $3.72, averaging $1.24 per shift; no separate API-rate recalculation was made. Claude and Codex remain excluded from the ranked cost graph. The development pilot is excluded from the scored repeats.
Uses only Python’s standard library. Checks file and frozen-source hashes, all three grades, both sensitivity scores, isolated tools, Newton sidecars and native cumulative usage. No simulator or paid model call is needed. Hashes establish consistency, not independent validation of the experiment.
03 / Evidence and verification
Inspect it. Check it yourself.
The original archive holds every original-study episode: the robot's pose trace, every reply, every model request's usage, the shift, the grader, the preregistration, the freeze and its amendment, the author's brief and the calibration feedback. Identifiers were redacted; the list is inside.
Extract evidence.zip into a new folder, open a terminal there, and run:
python -I -S verify_archive.py
No API key, network, simulator or paid call is needed. The verifier checks every file's hash, re-grades all 28 episodes from their recorded traces with the bundled grader, and recomputes every original-study score and verdict. Hashes establish consistency; this is not independent validation of the original run.
Offline verification is available now. Running the shift again needs OmniSim, the benchmark source and your own OmniKey and model key. Benchmark source and frozen protocols are included in the downloadable evidence. Native Codex and Claude Code runs additionally require their corresponding CLI and signed-in account. Live runs incur provider charges and will differ in detail: the model is sampled, and the provider's latency varies.
Next · under way
A stricter test with equal tools.
The follow-up targets the comparison that is not yet settled. It uses a new blind author and a different setting, grades replies with an independent judge model against a rubric rather than keywords, gives every framework the same instructions and interruption handling as OmniLink, and runs five repeats per configuration. No result is claimed here until it finishes.