evalstate's picture
|
download
raw
5.53 kB
# Codex CLI 0.159.2: API key vs ChatGPT OAuth reasoning effort
**Question.** Does the same client, sending the same request, get the same amount of
reasoning from OpenAI models when it authenticates with an **API key** as when it uses
a **ChatGPT (Codex) OAuth login**?
**Answer (so far).** No, not for `gpt-6-luna`. With the official Codex CLI as the
client, the OAuth route produced about **0.39×** the reasoning tokens of the API route
at medium effort (n=50 per arm), and was slower. `gpt-6.1-sol` showed no difference at
medium effort and a smaller gap (about 0.78×) at high effort (n=8).
![Luna medium reasoning tokens](results/luna-medium-n50/plots/reasoning_tokens.png)
## Method
- Client: Codex CLI `0.159.2`, `codex exec --json --ephemeral --ignore-user-config
--ignore-rules --skip-git-repo-check -s read-only -m <model> -c
'model_reasoning_effort="<effort>"'`. The prompt is read from stdin, and the CLI runs
in an empty working directory.
- The only difference between arms is authentication. Each arm runs in its own
throwaway `CODEX_HOME`:
- **api:** `codex login --with-api-key`. Standard tier, with no service-tier override.
- **oauth:** a copy of a ChatGPT-mode `~/.codex/auth.json`.
- Provider environment variables are removed from `codex exec`, so neither arm can
fall back to the other's credentials.
- Prompt: a single number-theory proof task with a known answer: all (x, y) with
1 ≤ x ≤ y ≤ 100 where xy divides x² + y² + 1. The answer is
`(1,1),(1,2),(2,5),(5,13),(13,34),(34,89)`. The prompt is in each results folder as
`prompt.txt` (sha256 `4adea627…`).
- Runs alternate between the arms (`api,oauth`, then `oauth,api`, and so on). The n=50
run used 4 parallel workers, shared by both arms. All runs took place on 2 October
2026 between 08:03 and 08:47 UTC.
- Metrics come from the CLI's `turn.completed` usage events: input, cached input,
output and `reasoning_output_tokens`. Output includes reasoning. Wall-clock time
includes CLI start-up.
## Results
| run | model / effort | n per arm | API reasoning (median) | OAuth reasoning (median) | OAuth/API (95% CI) | Mann-Whitney p | correct API / OAuth | median wall s API / OAuth |
|---|---|---|---|---|---|---|---|---|
| [luna-medium-n50](results/luna-medium-n50/report.md) | gpt-6-luna medium | 50 | 2,674 | 1,034 | **0.39** (0.35–0.41) | 7e-18 | 50 / 49 | 29.3 / 35.9 |
| [luna-medium-n4](results/luna-medium-n4/report.md) | gpt-6-luna medium | 4 | 2,287 | 1,325 | 0.58 (0.39–0.85) | 0.06 | 4 / 4 | 30.4 / 46.3 |
| [luna-high-n8](results/luna-high-n8/report.md) | gpt-6-luna high | 8 | 3,445 | 2,053 | 0.60 (0.42–0.71) | 0.0009 | 8 / 8 | 39.7 / 52.7 |
| [sol61-medium-n4](results/sol61-medium-n4/report.md) | gpt-6.1-sol medium | 4 | 423 | 427 | 1.01 (0.82–1.24) | 1.0 | 4 / 4 | 23.5 / 50.7 |
| [sol61-high-n8](results/sol61-high-n8/report.md) | gpt-6.1-sol high | 8 | 957 | 745 | 0.78 (0.66–0.87) | 0.004 | 8 / 8 | 37.1 / 63.7 |
- Confidence intervals are bootstrap intervals for the ratio of medians.
- The OAuth arm consistently sends about 15% more input tokens (11.8k against 13.6k
for Luna). The CLI adds different instructions when logged in with ChatGPT.
- The OAuth route is slower even though it produces fewer tokens. For Luna at n=50,
the median throughput was 111 output tokens per wall-clock second on the API arm
and 42 on the OAuth arm.
- The one wrong answer (luna-medium-n50, OAuth, round 39) used **0** reasoning tokens.
## Caveats
- Only one prompt was used, and the small runs (n=4 and n=8) are indicative only.
The n=50 Luna run is the robust result.
- The OAuth arm reflects one ChatGPT account and plan, observed at one point in time.
Server-side policy can change.
- The comparison is API standard tier against the ChatGPT/Codex backend. The
difference is attributed to the server side because the client, prompt, model and
requested effort are identical. This matches an earlier independent observation with
a different client (fast-agent): about 0.40× reasoning on the Codex route for the
same prompt.
## Layout
```
scripts/
codex_route_ab.py runner (stdlib only; needs `codex` on PATH)
codex_route_ab_report.py report.md: medians, p10/p90, ratios, bootstrap CIs, p-values
codex_route_ab_plot.py plots/reasoning_tokens.png, plots/wall_seconds.png (uv run; matplotlib)
results/<model>-<effort>-n<runs per arm>/
report.md summary tables and plots
summary.json run metadata (CLI version, model, effort, prompt sha256, auth method) and medians
results.jsonl one row per generation: route, round, exit, seconds, usage, answer, correct
prompt.txt exact prompt
plots/ reasoning-token and wall-time distributions (kernel density, median, one tick per run)
raw/ per-generation `codex exec --json` event logs and final answers
```
## Reproduce
```bash
export OPENAI_API_KEY=... # API arm
python3 scripts/codex_route_ab.py --model gpt-6-luna --effort medium --rounds 50 --parallel 4 \
--oauth-codex-auth ~/.codex/auth.json # OAuth arm (copied; refuses logins >7 days old)
python3 scripts/codex_route_ab_report.py codex-route-ab-<timestamp>
uv run scripts/codex_route_ab_plot.py codex-route-ab-<timestamp>
```
Add `--login-only` to check both logins without calling a model. Use `--prompt-file` and
`--expected` to run a different task. Credentials reach the CLI only through stdin and
are deleted with the temporary homes when the run ends.

Xet Storage Details

Size:
5.53 kB
·
Xet hash:
bb35fe66d94d2e677d094d78cec163d2079c637318190f3abc02eedba32fcf90

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.