How 20 Vercel AI SDK tickets drove the agent cost gap
Claude Code on Opus 5.5 and Codex on GPT 6.1 Sol attempted the same 49 tickets. The recorded cost gap came entirely from high-scope work.

Codex's lower cost estimate came entirely from high-scope tickets. Across the same 49 pull requests from vercel/ai, Decant, our tool for analyzing agent session logs, estimated $44.00 for Claude Code on Opus 5.5 and $27.29 for Codex on GPT 6.1 Sol. On low- and mid-scope work combined, Codex's estimate was slightly higher.
A single cost ratio hides that pattern. The requested changes stayed the same, but each agent chose its own path through the repository and how much verification to do.
These are costs of attempts. We didn't score the patches for correctness. Codex's logs report zero cache writes, so we bounded that charge rather than assume it was zero. It adds at most $1.86, which would put Codex at $29.15. The totals reproduce recorded usage at pinned prices, so they don't establish an API bill or a 34% to 38% saving on equally good solutions.
This experiment follows Your coding agent budget pays for context, not code, which Michael Mangus and I wrote about a different repository and models. We used the same PR replay framework to ask where the money goes when two agents receive the same work.
The Claude Code vs Codex cost gap comes from high-scope tickets
Codex's recorded estimate was 0.62 times Claude Code's across all 49 pairs. We grouped tasks by how much of the repository someone needs to understand to make the change. The scope bins show where that difference arose.
| Scope | PRs | Median lines added | Claude Code total | Codex total | Codex / Claude Code | Claude Code median | Codex median |
|---|---|---|---|---|---|---|---|
| Low | 9 | 28 | $2.70 | $3.45 | 1.28× | $0.27 | $0.35 |
| Mid | 20 | 129 | $9.34 | $9.58 | 1.02× | $0.47 | $0.48 |
| High | 20 | 396 | $31.96 | $14.26 | 0.45× | $0.92 | $0.56 |
| All | 49 | — | $44.00 | $27.29 | 0.62× | — | — |
- Totals and medians describe recorded cost estimates.
- Lines added come from the reference PRs.
- Each ratio divides the two bin totals.
- Adding Codex's maximum cache-write premium of $1.86 entirely to the high bin would raise that bin's ratio to 0.50.
High-scope tasks account for a $17.70 difference, while Codex cost an estimated $0.98 more across the other two bins. Claude Code's median rose 3.4 times from low to high scope. Codex's rose 1.6 times. Expensive Claude Code runs widen the total gap, but the high-scope medians differ too, at $0.92 and $0.56.
On #19701, a nine-line model addition to the zai and gateway providers, the attempts cost an estimated $0.26 on Claude Code and $0.63 on Codex. That was the largest paired Codex-to-Claude ratio, at 2.48 times.
On high-scope work, Codex had a lower estimate for 17 of 20 PRs, including all four of Claude Code's most expensive attempts. A DNS-protection fix in provider-utils, #20809, came to $3.43 on Claude Code and $0.39 on Codex. Adding Opus 5.5 support to the Anthropic provider, #21286, came to $4.23 and $1.26.
Across all 49 pairs, Codex had the lower estimate in 33. The median of the 49 paired cost ratios was 0.83. That statistic describes a typical paired ratio, while the 0.62 ratio of totals weights expensive tickets more heavily.
Opus 5.5 assigned the scope bins, so we also sorted the PRs into thirds by lines added. The ratios of Codex to Claude Code totals were 1.18, 0.79, and 0.45, from the smallest third to the largest. Both splits locate the lower Codex estimates in larger tasks. Neither separates comprehension scope from change size.

- All panels share the same log scales, and the diagonal marks equal cost.
- Total ratio divides Codex's bin total by Claude Code's.
- Costs are API-equivalent estimates from Decant.
- Patches weren't scored for correctness.
- Three discarded Codex pilot attempts are excluded.
Cached input, not output, drives most of each agent's cost
A session's estimate combines each token type's volume with its pinned price. Prompt caching reduces the cost of reading a matching input prefix, but cache reads still have a per-token price.
| Token type | Claude Code on Opus 5.5 | Codex on GPT 6.1 Sol |
|---|---|---|
| Output | $11.73 (26.6%) | $5.22 (19.1%) |
| Other input at base rate | $0.01 (0.0%) | $7.45 (27.3%) |
| Cache reads | $13.18 (30.0%) | $14.61 (53.6%) |
| Cache writes | $19.09 (43.4%) | $0 recorded, at most $1.86 |
| Recorded total | $44.00 | $27.29 |
| Total with cache-write bound | — | At most $29.15 |
- Decant's pinned prices per million tokens for Opus 5.5 were $4 input, $0.20 cache reads, $8 one-hour cache writes, and $20 output.
- GPT 6.1 Sol used $2 input, $0.10 cache reads, $2.50 cache writes, and $10 output.
- Rounded components can differ from the rounded total.
Claude Code's logs identify one-hour cache writes. Codex's logs report zero cache writes in all 49 sessions. The Codex issue discussion shows that subscription accounts report the cache-write field as zero, and it hasn't settled whether that means no writes occurred. We don't treat the zeros as a measurement.
We can bound the charge instead. The API reports cache writes as part of input tokens, and a token written to the cache wasn't read from it. Any cache writes therefore fall within the 3.7 million non-read input tokens that Decant already priced at the $2 base rate. Billing all of them at the $2.50 cache-write rate adds $1.86 and raises Codex's estimate to $29.15, or 0.66 times Claude Code's. If the whole premium fell on high-scope tickets, Codex's total there would still be $15.84 below Claude Code's.
Codex recorded 146 million cache-read tokens, about three million per task, against Claude Code's 66 million. Yet the estimated cache-read components were close, at $14.61 and $13.18, because the pinned cache-read prices differed. Token volume and token price both shape the bill.

- Codex's logs report zero cache writes. Its dashed bar is an upper bound that bills every non-read input token at the $2.50 cache-write rate instead of the $2 base rate.
- Claude Code's cache writes are measured.
- Percentages use each agent's recorded total and are rounded.
- Estimates use Decant token counts at pinned API rates.
Claude Code spends more on planning, Codex on code and tests
Token types tell you what Decant priced. They don't tell you what the agent was doing when it spent the money. For that, Decant sorts every session into the four buckets you see in its Analytics view. Context covers file reads and code searches. Planning covers reasoning about what to do next. Code covers edits, builds, and tests, along with their output. Communicating is whatever the agent says back to you.
These shares are attributed, not billed. Decant splits each session's output cost by its share of generated tokens and its input cost by estimated context share, then adds those allocations up across sessions. That's how planning can reach 39.4% of Claude Code's attributed cost even though all of its output tokens come to 26.6% of the token-type estimate.
| Activity | Claude Code share of cost | Codex share of cost |
|---|---|---|
| Context | 37.4% | 34.7% |
| Planning | 39.4% | 10.0% |
| Code, including builds and tests | 12.1% | 52.7% |
| Communicating | 11.1% | 2.6% |
- Shares divide each activity's summed attributed cost by that agent's total across all 49 runs.
- Displayed values can differ from 100% because of rounding.
The biggest swing is in code. Codex put 52.7% of its attributed cost there, against 12.1% for Claude Code, and much of that is verification. Claude Code's logs show test commands in 46 of 49 runs. A conservative scan of Codex's literal commands found a test command in all 49 runs and pnpm install in 32. In this monorepo, testing one package can mean installing dependencies and building the packages it depends on first, and all of that counts as code. A test command in the log shows the agent tried to verify its work, not that the suite finished or the patch passed.
Getting those labels right for Codex took some work. Codex wraps many actions inside exec, including shell commands and file edits, so we taught Decant to look inside those calls, inside compound shell commands, and inside the short Python scripts GPT 6.1 Sol wrote to edit files. A read-only chain counts as context, and a build or file write counts as code. Commands assembled at runtime can still slip past that inspection.
Decant also estimates active time from the gaps between messages. Each gap is capped at five minutes, credited to the activity of the message that follows it, and skipped when it belongs to a user reply. That gives you a log-based estimate, not a stopwatch reading or model-compute time.

- Each agent's shares use its own totals across all 49 runs.
- Code includes edits, builds, and tests.
- Decant allocates session cost by generation and estimated context share, so shares aren't charges incurred during each activity.
- Time comes from message gaps capped at five minutes, not stopwatch or model-compute time.
- Activity labels depend on tool-call classification.
Put cost and time side by side, and the code bars stand out. A build can hold up the timeline while generating very few tokens, so a slow stretch in your own sessions isn't necessarily an expensive one, and a large cost share doesn't mean the agent was slow.
Both agents spend most of their cost after the first edit
In Michael's earlier experiment, both agents spent roughly half their budget orienting, the stretch before the first line of code. On these 49 tickets, the balance tipped the other way.
Orienting runs until the first detected file edit, and implementing covers everything after it. An agent can still gather context or run a build once it's implementing. Decant attributed 40.8% of Claude Code's estimate and 31.8% of Codex's to orienting, using the same allocation as the activity shares.
The scope bins show where the two agents differ. On low- and mid-scope tickets, Claude Code spent 57% to 59% of its cost before the first edit. On high-scope tickets that share fell to 34%, as the extra cost piled up after editing began. Codex held between 31% and 33% at every scope.

- Percentages are attributed shares before the first edit, using summed costs within each scope bin.
- Decant attributes phase cost using generation and estimated context proportions.
- The first-edit boundary doesn't measure how much understanding the agent acquired.
- Opus 5.5 assigned scope, and the low-scope bin contains only nine PRs.
Active time leans even further toward implementing. Orienting took 11.4% of Claude Code's attributed time and 5.4% of Codex's.

- Orienting ends at the first detected file edit.
- Time shares use the same capped message-gap method as Figure 3.
- Each agent's shares use its own totals across all 49 runs.
So when a small ticket costs more on Codex, start with the work after its first edit, where Decant put most of Codex's cost. Reading those requests tells you whether the agent kept investigating or was verifying a change that needed it.
How we replayed 49 Vercel AI SDK pull requests
Each ticket started from the base commit of a merged pull request. The agent got the PR's title and description as its ticket, and no diff. From there it had to find the relevant code, make a change, and decide how much to build and test before stopping. Our checks found no run that fetched the reference PR.
The tickets come from Vercel's AI SDK, a TypeScript monorepo with about 30 packages. We filtered its last 1,000 merges down to human-authored work on main, then had Opus 5.5 read 192 candidate diffs and rate each one's comprehension scope, meaning how much of the codebase you'd need to understand to make the change. Only nine candidates landed at the lowest scope, so the sample includes all nine, alongside 20 mid-scope and 20 high-scope PRs. With nine tickets, the low-scope bin is the least certain of the three.
Both agents ran unattended with their default settings, Claude Code 2.1.280 on Opus 5.5 and Codex 0.159.3 on GPT 6.1 Sol. Each could run up to three sessions at once and install dependencies over the network. Three Codex pilot attempts hit package-manager permission errors, so we left those out and reran the tickets.
We stripped user-level settings and plugins but kept each repository's instructions. Codex loaded the root AGENTS.md, and Claude Code's startup records listed the 12 repository skills in .claude/skills/. Codex also found nine user-level skills, about 1,200 tokens by a character-count estimate, and opened the documentation-lookup skill in eight runs.
That makes this a comparison of two tool-and-model bundles, instructions and verification habits included, rather than of the models alone. Both ran on subscription plans and never hit a plan limit or paid overage in the included runs. Decant priced their recorded tokens at pinned API rates, so the totals reflect API-equivalent costs.
How reruns and configuration change Claude Code and Codex costs
Configuration shows up before an agent reads a single file. In a three-PR pilot with user-level configuration turned on, each Claude Code session's first request was about 10,000 tokens larger, 32,000 instead of 22,000. Those three attempts came to an estimated $2.53, against $1.53 with the configuration stripped. The extra starting tokens are a direct measurement. The full 65% difference rests on three runs, though, and one of them ran tests only with the configuration loaded, so instructions don't account for all of it.
Reruns move the numbers too, and Codex moves more. We replayed the Claude Code median-cost ticket from each scope bin three more times on both agents. Across those four attempts per ticket, Claude Code's cost varied by 7% to 10%, measured as the coefficient of variation, and Codex's by 18% to 41%. Claude Code's most expensive attempt came to at most 1.3 times its cheapest. For Codex, that spread ran from 1.6 to 2.9 times.
Each repeat pass started at least 65 minutes after the previous one ended, long enough for Claude Code's one-hour cache to expire, and Codex got the same spacing. Each repeat's first request showed only the fixed cached prefix. On two of the three tickets, the original run was the cheapest of the four for both agents. Three tickets aren't enough to say whether that's a pattern, and the most expensive attempts weren't repeated.
Scope and patch size also travel together in this sample. Every low-scope PR added fewer than 80 lines, and Claude Code's cost had a rank correlation of about 0.8 with both lines added and scope. The data shows where the cost rises, but it can't yet separate how much an agent has to understand from how much it has to write.
Context is about a third of each agent's cost
Context, the reads and searches an agent runs to find its way around, accounted for 37.4% of Claude Code's attributed cost and 34.7% of Codex's. That's the work shared repository knowledge is meant to shorten, and these 49 tickets give us the baseline to measure it against.
Two agents given the same 49 tickets landed $16.72 apart, and the 20 high-scope tickets account for all of it. A single cost ratio would have hidden that. The session records show where the money went: how much each agent paid to re-read cached context, how much went to builds and tests, and which tickets drove the total.
To see how your own agents spend their time and money, Decant is our open source, local-only token accounting tool that powered this analysis. Try it on your own agent logs:
npx @dosu/decant # account your own Claude Code / Codex sessions
Start with the input components, then follow their cost into activities and the work after the first edit.
If you want your agents to spend less of each session rediscovering your codebase, Dosu provides knowledge infrastructure for agents. Start using Dosu today.
Appendix: per-PR costs for all 49 Vercel AI SDK tickets
| PR | Title | Scope | +lines | Claude Code | Codex |
|---|---|---|---|---|---|
| #21286 | feat(provider/anthropic): add Claude Opus 5.5 support | High | +2023 | $4.23 | $1.26 |
| #21162 | feat: add telemetry support to `experimental_evaluate` | High | +1531 | $3.91 | $0.96 |
| #20809 | fix(provider-utils): preserve download DNS protection under wrapped fetch | High | +437 | $3.43 | $0.39 |
| #19450 | fix(batch): align result parsing, request counts, and behavior across providers | High | +990 | $3.21 | $1.38 |
| #20875 | feat(ai): add experimental evaluation model registry support | High | +850 | $2.33 | $1.12 |
| #20027 | feat(harness): add `createBridgeToken()` and `withBridgeToken()` helpers for bridge backed harness adapters | High | +131 | $1.96 | $0.56 |
| #21231 | feat(harness): add `AbortSignal` to `SandboxChannel.connect`, plus a `sleep()` helper function | High | +734 | $1.85 | $1.05 |
| #21216 | feat(harness): allow consumers of bridge backed harnesses to configure `reconnect` timeout | High | +384 | $1.81 | $0.51 |
| #20652 | fix(harness-pi): support explicit custom provider models | High | +150 | $1.31 | $0.55 |
| #20885 | feat(gateway): add experimental evaluation model support | High | +741 | $0.94 | $0.43 |
1–10 of 49 pull requests
All token accounting and pricing in this report is done by Decant v0.9.0, plus a first-edit detection fix, from each tool's own local session transcript. Costs are modeled at standard API rates, so they reflect what an API-billed session would cost rather than subscription usage. Bucketing rules are documented in Decant's analytics methodology. Codex's cache-write charge is bounded at $1.86 across all 49 runs rather than measured. Sessions are not scored for correctness, so these are cost measurements, not a benchmark of which agent works better. The replays ran on our experiment framework using PRs merged between August 20 and September 22, 2026, and we can't rule out training-data contamination.
Found this article helpful?
Share it with your network to help others discover valuable insights.
Want more like this? Subscribe via RSS
Related Articles
Procedural Memory for Coding Agents: Turning Knowledge Into Action
Oct 2, 2026 / 16 min read
Learn how to turn a coding agent's discovery into a shared procedure, choose where the instructions belong, and check that the next agent uses them correctly.
Semantic memory for coding agents needs a history
Oct 1, 2026 / 10 min read
Your coding agent needs to know which facts still apply. Our Celery-to-DBOS migration shows why semantic memory needs scope, dates, and source evidence.
Episodic memory for coding agents: What survives your sessions
Sep 29, 2026 / 14 min read
Episodic memory is a coding agent's record of past sessions. Saved lessons can outlive the transcripts behind them, so keep each lesson tied to its source.
September Drop: Dosu goes back to school
Sep 23, 2026 / 6 min read
Dosu studies coding sessions, shows what your agent reads, and makes your credits go further.
