OpenAI doesn't tell you how much Codex a ChatGPT subscription buys. You get a percentage bar and a reset date, and that's it. That makes it hard to plan, hard to budget, and impossible to compare against just paying for the API.
I run several AI coding agents in parallel, all day, on three machines and one ChatGPT account. Codex logs every turn's tokens, and it also logs the usage bar itself, so I lined the two up across two plans and several models. Here is what each plan is actually worth.
The short version:
- The usage bar is a dollar meter, not a token counter. Priced at API list rates, every 1 % of the bar costs about the same. Confirmed for GPT-6 Sol and GPT-6 Astra; the rate shifted when the model generation changed.
- Pro $200 (new 10× tier): about $650 of API-priced usage a week. Measured in two clean windows at $6.41 and $6.70 per 1 %.
- Pro $100: about $330–380 a week, $3.31 and $3.77 per 1 % in two clean windows. Roughly half, as OpenAI's 5× vs 10× multipliers predict.
- Every plan buys roughly 14× its price in API-priced usage, if you use it. Below about 7 % of your allowance, the API is cheaper.

Measure your own plan with the free Codex Usage Tracker: load your Codex session folders and it runs this whole analysis in your browser, without uploading anything.
If you need the basics of the 5-hour and weekly windows first, start with Codex CLI usage and rate limits.
What each plan buys
Measured rows are marked; the rest scale the measurements by OpenAI's plan multipliers.
| Plan | vs Plus | API-priced usage per 1 % | Per week | Per month | Value vs price | Status |
|---|---|---|---|---|---|---|
| Plus $20 | 1× | ~$0.66 | ~$66 | ~$285 | ~14× | Projected (one near-clean window: $0.62) |
| Pro $100 | 5× | $3.31–$3.77 | ~$330–380 | ~$1,440–1,640 | ~14–16× | Measured (2 windows) |
| Pro $200 (new) | 10× | $6.41–$6.70 | ~$640–670 | ~$2,800–2,900 | ~14× | Measured (2 windows) |
| Pro $200 (grandfathered, until Oct 29) | 20× | ~$13 | ~$1,300 | ~$5,700 | ~28× | Projected |
| Pro $500 | ~25× (reported) | ~$16.40 | ~$1,640 | ~$7,100 | ~14× | Projected |
Per month = per week × 4.35. "Value vs price" = API-priced usage per month ÷ subscription price.
In tokens, it depends on your model mix and cache rate, because cached input is cheap and output is expensive. On my workload (94 % GPT-6 Sol, 96 % of input cached), Pro $200 came to about 92M counted tokens a week (uncached input plus output, the "tokens used" figure Codex prints), or about 2.1 billion tokens including cache. Scale by the multipliers for other plans: about 9M a week on Plus and 46M on Pro $100.
Two limits to keep in mind: Plus also has a five-hour window, so a Plus user may not be able to spend the whole weekly allowance, and the grandfathered 20× rate ends Oct 29, after which every $200 plan is 10×.
Which plan should you buy?
The break-even rule is simple once the bar is measured in dollars: a subscription beats the API when your monthly usage, priced at API rates, exceeds the subscription price. Since every plan buys about 14× its price, that's about 7 % of the allowance on any tier.
In tokens, on a Sol workload with heavy caching (about $7.04 of API cost per million counted tokens):
| Plan | Break-even per month | ≈ counted Sol tokens / month | ≈ per day |
|---|---|---|---|
| Plus $20 | $20 of API-priced usage | 2.8M | 0.09M |
| Pro $100 | $100 | 14.2M | 0.47M |
| Pro $200 | $200 | 28.4M | 0.93M |
| Pro $500 | $500 | 71.0M | 2.3M |
- Pick the smallest plan you'd actually fill. Below break-even, the API is cheaper. Above the plan's allowance, you're waiting for the reset.
- Heavy Sol or Astra users: a subscription wins by a wide margin. One week at my pace passes the $200 plan's monthly break-even several times over.
- Light or Luna-heavy users: consider the API. Luna costs about $0.19 per million counted tokens on the API with this caching pattern, about 40× less than Sol, so a Luna-only workload needs about 105M counted tokens a month just to break even on Plus.
For the auth side of switching, see ChatGPT login vs API key in Codex CLI.
How I measured it
Every Codex turn writes a line to a session file with its token counts and the current state of the usage bar:
"rate_limits": {
"limit_id": "codex",
"primary": { "used_percent": 20.0, "window_minutes": 10080, "resets_at": 1791580394 },
"plan_type": "pro"
}
window_minutes: 10080 is the weekly window (7 × 24 × 60). plan_type records the plan: plus, prolite (Pro $100) or pro (Pro $200). My billing history covers all three: Plus until June and again Sep 13–22, Pro $100 from July to late September, and Pro $200 since Sep 29. I used windows from August onward; July's use models I don't have prices for.
For each weekly window I found the moment the bar first reached each whole percentage, priced every turn in between at API list rates (uncached input, cached input and output separately, by model), and divided by the percentage points crossed. If the bar counts some quantity, that quantity should be about the same for every 1 % step. I merged the session files from all three machines I run Codex on, because usage anywhere moves the same bar. My first pass used only two of them. Adding the third raised the October windows by 2–5 % and turned five noisy windows clean, which is exactly the undercount the warning further down is about.
The bar counts dollars, not tokens
In the Oct 2–4 window (Pro $200; two of the three machines: 387 runs, 9,594 bar readings), I measured how much of each candidate quantity filled each 1 % step, and how much that amount varied from step to step (coefficient of variation, CV; lower means steadier):
| What the bar might be counting | Amount per 1 % step | Variation (CV) |
|---|---|---|
| Counted tokens (uncached input + output), all models | 0.85M | 0.25 |
| Counted tokens, Sol only | 0.79M | 0.20 |
| All tokens, including cached | 20.7M | 0.19 |
| Uncached input only | 0.75M | 0.25 |
| The same usage priced at API list rates | $6.21 | 0.10 |
This is a first-pass analysis of one window on two machines; the step method below, on all three, puts the same window at $6.41.
API-priced dollars are about twice as steady as any token count. A one-weight fit on API dollars predicts 32.5 % of the bar for that window, against 33 % actual.
The step method then holds up across plans and models. These are the clean GPT-6 windows (steady steps, and no steps bought with suspiciously little spend), all three machines merged:
| Window | Plan | Usage, priced at API rates | $ per 1 % | Step CV |
|---|---|---|---|---|
| Reset Sep 30 | Pro $100 | Sol $343, Astra $29 | $3.77 | 0.11 |
| Reset Oct 3 | Pro $100 | Sol $261, Astra $70 | $3.31 | 0.07 |
| Reset Oct 6 | Pro $200 | Sol $425, Astra $217 | $6.70 | 0.04 |
| Reset Oct 9 (Oct 2–4 so far) | Pro $200 | Sol $210, Astra $1, Luna ~$0 | $6.41 | 0.08 |
- Pro $100 measures at about half of Pro $200 (Pro $200 comes out 1.7–2.0× Pro $100), in line with the 5× vs 10× multipliers.
- Astra counts at about its API price too. A third of the spend in the window resetting Oct 6 was GPT-6 Astra ($10 / $1 / $50 per million, five times Sol's price), and that window landed at $6.70 per 1 %, close to the almost Sol-only $6.41. Two Astra-only Pro $100 windows on Sep 29 came out at $3.49 and $3.56, in the same range as the Sol-heavy ones, though too uneven to count as clean. If the bar counted tokens, Astra-heavy windows would show far more dollars per 1 %.
- That explains the 14×. It isn't a quirk of my workload: each plan sells API-priced usage at about 1/14 of list price.
The rate changed with the model generation
Merging the third machine also cleaned up four August windows, all Pro $100 and almost all GPT-5.6 Sol (priced at $5 input / $30 output per million). They're steady too, but at a different rate:
| Window | Plan | Usage, priced at API rates | $ per 1 % | Step CV |
|---|---|---|---|---|
| Reset Aug 3 | Pro $100 | 5.6-Sol $113 | $5.40 | 0.07 |
| Reset Aug 7 | Pro $100 | 5.6-Sol $382, 5.6-Luna $5 | $5.95 | 0.09 |
| Reset Aug 15 | Pro $100 | 5.6-Sol $451, 5.6-Luna $16 | $6.31 | 0.08 |
| Reset Aug 17 | Pro $100 | 5.6-Sol $108, 5.6-Luna $17 | $5.68 | 0.12 |
Under GPT-5.6, each 1 % of the Pro $100 bar was worth about $5.80 of API-priced usage, about 1.6× the GPT-6 rate. So the bar tracks API prices within a model generation, but it isn't literally API dollars. When OpenAI cut API prices for GPT-6 (Sol input fell from $5 to $2, output from $30 to $10), the bar didn't get cheaper by the same amount. Two other explanations fit too: GPT-5.6's cached-input price may have been lower than the $0.50 I assumed, or Pro $100's allowance itself changed in September.
The practical upshot: re-measure after every model change. Under GPT-5.6, Pro $100 was worth about 25× its price at API rates; under GPT-6 it's about 15×. That's why the meter only uses clean windows from the last three weeks.
Why "tokens used" undercounts
Here's what the Oct 2–4 window (33 % of the bar, all three machines) looked like in tokens:
| Model | Runs | Input (total) | of which cached | Uncached input | Output (incl. reasoning) |
|---|---|---|---|---|---|
| GPT-6 Sol | 427 | 681.4M | 656.0M | 25.4M | 3.0M |
| GPT-6 Luna | 50 | 12.4M | 10.7M | 1.7M | 0.13M |
| GPT-6 Astra | 1 | 0.37M | 0.31M | 0.06M | <0.01M |
| Total | 478 | 694.2M | 667.1M (96 %) | 27.2M | 3.1M |
- "Tokens used" is not "tokens processed." Codex's end-of-run number is uncached input plus output: about 30M here. The model actually read about 694M tokens, but 96 % came from cache. Agentic coding re-sends the whole conversation every turn, so caching does almost all the heavy lifting.
- Output is tiny. About 3M output tokens against 694M of input. Coding agents mostly read.
Subscription vs API: the math
Current API list prices (see the full LLM API price comparison for other vendors):
| Model | Input | Cached input | Output |
|---|---|---|---|
| GPT-6 Sol | $2.00 / M | $0.20 / M | $10.00 / M |
| GPT-6 Luna | $0.10 / M | $0.01 / M | $0.50 / M |
| GPT-6 Astra | $10.00 / M | $1.00 / M | $50.00 / M |
The Oct 2–4 window priced at those rates:
| API cost | |
|---|---|
| Sol: 25.4M uncached × $2 + 656.0M cached × $0.20 + 3.0M output × $10 | $212.21 |
| Luna | $0.34 |
| Astra | $1.01 |
| Total for 33 % of the weekly bar | ≈ $214 |
That's about $650 for a full week, or about $2,800 a month, against $200 for the subscription:
| Measure | Subscription (Pro $200) | API, same workload | Ratio |
|---|---|---|---|
| Per counted token (uncached input + output) | $0.50 / M | $7.04 / M | 14.1× |
| Per processed token (including cache) | $0.022 / M | $0.31 / M | 14.0× |
The catch: unused allowance is worth nothing, while the API only bills for what you use.
How to measure your own allowance
The easy way: open the Codex Usage Tracker and choose your .codex/sessions folder. Using more than one computer? Add each one's folder, or a .zip / .tar.gz of it, one after another; they're merged and duplicates are skipped. It reads the files in your browser tab without uploading them, finds your clean windows, and shows your plan's rate, this week's burn, and which plan or the API is cheapest for your usage. You can also try it with my own data first. The steps below are what it does, if you'd rather do it yourself.
This only works if all of your Codex usage is in the files you measure. The usage bar counts everything on your ChatGPT account: every computer you run Codex on, every user account on those computers (including root on a server), Codex cloud tasks, and anyone sharing the account. Each computer only keeps session files for its own runs. Any usage you leave out makes each 1 % of the bar look cheaper than it really is, and your allowance look smaller.
- Collect the session files from every machine:
~/.codex/sessions/**/*.jsonlon Mac and Linux,%USERPROFILE%\.codex\sessionson Windows (or$CODEX_HOME/sessionsif you set it). - Keep only the main weekly limit: readings where
limit_idiscodexandwindow_minutesis10080. Group them byresets_at. - Price every turn at API list rates for its model: uncached input (input minus cached), cached input, and output, which already includes reasoning tokens. Count each turn once;
token_countevents repeat when only the bar updates. - Find the first moment the bar reaches each whole percentage, sum the dollars spent between levels, and divide by the points crossed.
- Check the steps are steady. A clean window has a CV around 0.1. Steps bought with much less than the typical spend (I flag anything under 40 % of the median) mean usage you don't have logs for. Don't trust that window.
Things the session files will trip you up on:
- There's more than one limit. Besides
codex, my readings include a separatecodex_bengalfoxlimit (labelled "GPT-5.3-Codex-Spark") and a rarepremiumlimit. Mixing them corrupts the steps. plan_typeisn't always right. A few late-September readings logged no plan, orpro, while billing still showed Pro $100. Check against your billing history.- Plan changes start a fresh window. My upgrade to Pro $200 on Sep 29 reset the bar to 0 %, with a new reset exactly 7 days later.
- Older windows need older prices. My July–September windows used GPT-5.6 Sol ($5 input / $30 output per million), not GPT-6.
How to stretch your allowance
What I changed after measuring this:
- Budget daily, in API dollars or counted tokens. I target about 10M counted tokens a day, measured over a moving 24-hour window plus a projection from the last 8 hours.
- Scale agents to the budget. As the projection climbs past set steps, the number of parallel agents drops, and it rises again when usage is under target.
- Move routine work off the expensive model. Since the bar charges API prices, model choice is the biggest lever. Log checks, verification and test triage moved to Luna or to plain deterministic scripts. About a fifth of the "AI" steps in my pipeline didn't need a model at all.
- Use Astra deliberately. At five times Sol's price, it burns the bar five times faster.
- Keep long sessions cache-friendly. With 96 % of input served from cache, anything that breaks the cache (reshuffling context, restarting sessions needlessly) is the expensive mistake.
What's still unknown
- Luna's weight. Luna has never been more than a small share of a clean window. A few hours of Luna-only work would settle whether it counts at its very low API price or carries a minimum weight. Moving the bar 5 % on Pro $200 would take about $33 of Luna usage, roughly 170M counted tokens.
- Reasoning effort. High and medium effort cost the same per token. Does effort only matter through the extra output tokens it produces?
- Plus and Pro $500. Neither has a clean window yet; my one near-clean Plus window (Sep 26, $0.62 per 1 %) is close to the projected $0.66. Readings from other users would pin those rows down.
- Why the rate moved with GPT-6. A smaller-than-API price cut, a wrong cached-price assumption for GPT-5.6, or a changed Pro $100 allowance all fit the August data. The next model change, measured as it happens, should tell them apart.
- The Oct 30 change. When Pro $200 drops from 20× to 10× for grandfathered subscribers, continuous logging should show it as a measured jump.
Caveats
- Four clean GPT-6 windows, one workload: long agentic coding runs with heavy caching. A chat-style workload with less caching will look different in tokens, though the dollar figures should hold.
- Only Pro $100 and Pro $200 are measured, two clean windows each under GPT-6. The other rows assume OpenAI applies its multipliers to the same dollar meter, which it doesn't document.
- The multipliers come from third-party write-ups quoting OpenAI's pricing page and Help Center. The $500 tier's ~25× is only reported, not published by OpenAI.
- The dollar rate is per model generation. August's GPT-5.6 windows ran at a different rate (see above), and GPT-5.6's cached-input price is my assumption. July is excluded.
- API prices are list prices after the recent GPT-6 Sol/Luna price cut. Batch or priority pricing would change the comparison.
If you've measured your own allowance, especially on Plus or Pro $500, I'd love to compare numbers.
Appendix: every window since August
Prices used ($ per million tokens: input / cached input / output):
| Model | Input | Cached input | Output |
|---|---|---|---|
| GPT-6 Sol | $2.00 | $0.20 | $10.00 |
| GPT-6 Luna | $0.10 | $0.01 | $0.50 |
| GPT-6 Astra | $10.00 | $1.00 | $50.00 |
| GPT-5.6 Sol | $5.00 | $0.50 (assumed, 10 % of input) | $30.00 |
| GPT-5.6 Luna | $0.20 | $0.02 | $1.20 |
Every window with at least 5 % of bar movement since August (plan from plan_type). All three machines merged, analysed with the Codex Usage Tracker. A "low step" is a step bought with less than 40 % of the window's median spend, the sign of usage on a machine whose files are missing. "Too short" means under 20 full steps.
| Reset | Plan | Bar | Steps | $ per 1 % | Step CV | Low steps | Usage, priced at API rates | Status |
|---|---|---|---|---|---|---|---|---|
| Aug 1 | Pro $100 | 0–31 % | 31 | $5.05 | 0.16 | 0 | 5.6-Sol $157 | Uneven |
| Aug 3 | Pro $100 | 0–21 % | 21 | $5.40 | 0.07 | 0 | 5.6-Sol $113 | Clean |
| Aug 4 | Pro $100 | 0–100 % | 100 | $5.78 | 0.18 | 1 | 5.6-Sol $571, 5.6-Luna $6 | Unlogged usage |
| Aug 7 | Pro $100 | 0–65 % | 65 | $5.95 | 0.09 | 0 | 5.6-Sol $382, 5.6-Luna $5 | Clean |
| Aug 14 | Pro $100 | 0–6 % | 6 | $5.75 | 0.04 | 0 | 5.6-Sol $32, 5.6-Luna $2 | Too short |
| Aug 15 | Pro $100 | 0–74 % | 74 | $6.31 | 0.08 | 0 | 5.6-Sol $451, 5.6-Luna $16 | Clean |
| Aug 17 | Pro $100 | 0–22 % | 22 | $5.68 | 0.12 | 0 | 5.6-Sol $108, 5.6-Luna $17 | Clean |
| Aug 19 | Pro $100 | 0–98 % | 88 | $6.62 | 0.33 | 2 | 5.6-Sol $644, 5.6-Luna $4 | Unlogged usage |
| Aug 26 | Pro $100 | 0–45 % | 40 | $6.72 | 0.34 | 3 | 5.6-Sol $302 | Unlogged usage |
| Aug 30 | Pro $100 | 0–12 % | 12 | $6.07 | 0.04 | 0 | 5.6-Sol $73 | Too short |
| Sep 1 | Pro $100 | 0–40 % | 40 | $7.61 | 0.21 | 0 | 5.6-Sol $305 | Uneven |
| Sep 3 | Pro $100 | 0–95 % | 92 | $7.27 | 0.21 | 1 | 5.6-Sol $691 | Unlogged usage |
| Sep 5 | Pro $100 | 0–8 % | 8 | $6.33 | 0.09 | 0 | 5.6-Sol $51 | Too short |
| Sep 6 | Pro $100 | 0–77 % | 71 | $4.49 | 0.46 | 11 | 5.6-Sol $346 | Unpriced models |
| Sep 13 | Pro $100 | 0–40 % | 40 | $8.66 | 0.24 | 0 | 5.6-Sol $275, Astra $72 | Uneven |
| Sep 14 | Pro $100 | 0–100 % | 99 | $4.31 | 0.36 | 1 | Astra $237, 5.6-Sol $194 | Unlogged usage |
| Sep 16 | Pro $100 | 0–89 % | 34 | $3.24 | 0.42 | 3 | Astra $288 | Too short |
| Sep 19 | Mixed | 0–100 % | 81 | $0.68 | 0.82 | 4 | Astra $68 | Unlogged usage |
| Sep 26 | Plus | 0–100 % | 99 | $0.62 | 0.18 | 3 | Astra $62 | Unlogged usage |
| Sep 29 | Pro $100 | 0–100 % | 100 | $3.56 | 0.22 | 0 | Astra $356 | Uneven |
| Sep 29 | Pro $100 | 0–100 % | 100 | $3.49 | 0.18 | 0 | Astra $349 | Uneven |
| Sep 30 | Pro $100 | 0–99 % | 99 | $3.77 | 0.11 | 0 | Sol $343, Astra $29 | Clean |
| Oct 3 | Pro $100 | 0–100 % | 100 | $3.31 | 0.07 | 0 | Sol $261, Astra $70 | Clean |
| Oct 6 | Pro $200 | 0–96 % | 96 | $6.70 | 0.04 | 0 | Sol $425, Astra $217 | Clean |
| Oct 9 | Pro $200 | 0–33 % | 33 | $6.41 | 0.08 | 0 | Sol $210, Astra $1 | Clean |
The two Sep 29 rows are separate reset times seen in the same period. The Sep 26 Plus window and the Sep 19 mixed-plan window come from my plan changes in late September. On two machines, several August windows looked like Luna-only usage at $0.08–$0.21 per 1 %; with the third machine merged they turn out to be GPT-5.6 Sol windows at $5–6 per 1 %, which is how much a missing machine can distort the picture.
Sources
- VentureBeat: OpenAI introduces ChatGPT Pro $100 tier with 5× Codex usage
- ThreatFrontier: ChatGPT Pro $100 vs $200 vs $500 usage limits
- CodePick: Codex plan value analysis
- SimpleMetrics: ChatGPT Codex limits 2026
- Yotta Labs: GPT-6 Sol and Luna pricing
- VentureBeat: GPT-6 Sol and Luna slash API costs
- Yotta Labs: GPT-6 Astra pricing
- Layer3 Labs: GPT-5.6 pricing
