claude code vs codex for agents: we metered both

claude code vs codex for agents: we metered both


Same $20/month.

We run Claude Code and Codex on the same task board. Based on each vendor’s own weekly usage meter, we measured roughly:

  • Claude Pro: ~36 agent-hours/week
  • ChatGPT Plus + Codex: ~3 to 3.5 agent-hours/week

About a 10× difference.

why

Agent workloads are mostly cached context reads.

In our Claude run:

usageshare
cache reads97.7%
cache writes1.8%
output0.47%
fresh input0.001%

Over 17.6 hours and 15 tasks, fresh input was only 3,956 tokens.

The Codex task looked almost identical: 98.2% cache reads and 0.26% output.

That is the important part. Long-running coding agents keep a large context and re-read it every turn. For subscription limits, the useful mental model is:

turns × context size

Not output tokens.

what we measured

Claude Pro used 49% of its weekly limit in 17.6 hours, which extrapolates to about 36 hours. We caught 44 separate ticks of the weekly meter on the way, enough that the slope carries about ±1% rather than the ±33% you get from reading two.

Codex used 9% of the ChatGPT Plus weekly limit in 18 minutes, implying roughly 3 to 3.5 hours at the same rate.

The difference is not API cache pricing. Both vendors price cache reads at about one tenth of fresh input.

The difference is how much cached usage each subscription allows before the weekly limit is reached.

We compare hours rather than tokens or dollars on purpose. OpenAI’s allowance is denominated in credits and requests, so there is no token pool to extrapolate, and a dollar figure would force a choice about whether to count cache reads at list. Share of the week needs neither.

caveats

This is real production usage, not a controlled benchmark.

The workloads were comparable, not identical: 15 Claude tasks versus one Codex task. ChatGPT’s meter also reports whole percentages, so the Codex estimate is much noisier.

Codex was measured on September 7 and Claude on September 15. Claude’s temporary +50% weekly-limit boost had already ended, so the Claude result reflects the reduced September limits.

takeaway

If your agents are also around 97 to 98% cache reads, compare subscriptions by how much repeated context they actually let you consume, not by output limits or headline token prices.

We track measured and modelled plan value on the 5dive model value board.

One practical detail: keep context stable. Editing something early in the prompt can turn the next turn into a cache write for the whole context. In our Claude run, cache writes were only 1.8% of quota usage but 22.7% of the API-equivalent value.