Buckets:
| # Codex CLI 0.159.2: API key vs ChatGPT OAuth reasoning effort | |
| **Question.** Does the same client, sending the same request, get the same amount of | |
| reasoning from OpenAI models when it authenticates with an **API key** as when it uses | |
| a **ChatGPT (Codex) OAuth login**? | |
| **Answer (so far).** No, not for `gpt-6-luna`. With the official Codex CLI as the | |
| client, the OAuth route produced about **0.39×** the reasoning tokens of the API route | |
| at medium effort (n=50 per arm), and was slower. `gpt-6.1-sol` showed no difference at | |
| medium effort and a smaller gap (about 0.78×) at high effort (n=8). | |
|  | |
| ## Method | |
| - Client: Codex CLI `0.159.2`, `codex exec --json --ephemeral --ignore-user-config | |
| --ignore-rules --skip-git-repo-check -s read-only -m <model> -c | |
| 'model_reasoning_effort="<effort>"'`. The prompt is read from stdin, and the CLI runs | |
| in an empty working directory. | |
| - The only difference between arms is authentication. Each arm runs in its own | |
| throwaway `CODEX_HOME`: | |
| - **api:** `codex login --with-api-key`. Standard tier, with no service-tier override. | |
| - **oauth:** a copy of a ChatGPT-mode `~/.codex/auth.json`. | |
| - Provider environment variables are removed from `codex exec`, so neither arm can | |
| fall back to the other's credentials. | |
| - Prompt: a single number-theory proof task with a known answer: all (x, y) with | |
| 1 ≤ x ≤ y ≤ 100 where xy divides x² + y² + 1. The answer is | |
| `(1,1),(1,2),(2,5),(5,13),(13,34),(34,89)`. The prompt is in each results folder as | |
| `prompt.txt` (sha256 `4adea627…`). | |
| - Runs alternate between the arms (`api,oauth`, then `oauth,api`, and so on). The n=50 | |
| run used 4 parallel workers, shared by both arms. All runs took place on 2 October | |
| 2026 between 08:03 and 08:47 UTC. | |
| - Metrics come from the CLI's `turn.completed` usage events: input, cached input, | |
| output and `reasoning_output_tokens`. Output includes reasoning. Wall-clock time | |
| includes CLI start-up. | |
| ## Results | |
| | run | model / effort | n per arm | API reasoning (median) | OAuth reasoning (median) | OAuth/API (95% CI) | Mann-Whitney p | correct API / OAuth | median wall s API / OAuth | | |
| |---|---|---|---|---|---|---|---|---| | |
| | [luna-medium-n50](results/luna-medium-n50/report.md) | gpt-6-luna medium | 50 | 2,674 | 1,034 | **0.39** (0.35–0.41) | 7e-18 | 50 / 49 | 29.3 / 35.9 | | |
| | [luna-medium-n4](results/luna-medium-n4/report.md) | gpt-6-luna medium | 4 | 2,287 | 1,325 | 0.58 (0.39–0.85) | 0.06 | 4 / 4 | 30.4 / 46.3 | | |
| | [luna-high-n8](results/luna-high-n8/report.md) | gpt-6-luna high | 8 | 3,445 | 2,053 | 0.60 (0.42–0.71) | 0.0009 | 8 / 8 | 39.7 / 52.7 | | |
| | [sol61-medium-n4](results/sol61-medium-n4/report.md) | gpt-6.1-sol medium | 4 | 423 | 427 | 1.01 (0.82–1.24) | 1.0 | 4 / 4 | 23.5 / 50.7 | | |
| | [sol61-high-n8](results/sol61-high-n8/report.md) | gpt-6.1-sol high | 8 | 957 | 745 | 0.78 (0.66–0.87) | 0.004 | 8 / 8 | 37.1 / 63.7 | | |
| - Confidence intervals are bootstrap intervals for the ratio of medians. | |
| - The OAuth arm consistently sends about 15% more input tokens (11.8k against 13.6k | |
| for Luna). The CLI adds different instructions when logged in with ChatGPT. | |
| - The OAuth route is slower even though it produces fewer tokens. For Luna at n=50, | |
| the median throughput was 111 output tokens per wall-clock second on the API arm | |
| and 42 on the OAuth arm. | |
| - The one wrong answer (luna-medium-n50, OAuth, round 39) used **0** reasoning tokens. | |
| ## Caveats | |
| - Only one prompt was used, and the small runs (n=4 and n=8) are indicative only. | |
| The n=50 Luna run is the robust result. | |
| - The OAuth arm reflects one ChatGPT account and plan, observed at one point in time. | |
| Server-side policy can change. | |
| - The comparison is API standard tier against the ChatGPT/Codex backend. The | |
| difference is attributed to the server side because the client, prompt, model and | |
| requested effort are identical. This matches an earlier independent observation with | |
| a different client (fast-agent): about 0.40× reasoning on the Codex route for the | |
| same prompt. | |
| ## Layout | |
| ``` | |
| scripts/ | |
| codex_route_ab.py runner (stdlib only; needs `codex` on PATH) | |
| codex_route_ab_report.py report.md: medians, p10/p90, ratios, bootstrap CIs, p-values | |
| codex_route_ab_plot.py plots/reasoning_tokens.png, plots/wall_seconds.png (uv run; matplotlib) | |
| results/<model>-<effort>-n<runs per arm>/ | |
| report.md summary tables and plots | |
| summary.json run metadata (CLI version, model, effort, prompt sha256, auth method) and medians | |
| results.jsonl one row per generation: route, round, exit, seconds, usage, answer, correct | |
| prompt.txt exact prompt | |
| plots/ reasoning-token and wall-time distributions (kernel density, median, one tick per run) | |
| raw/ per-generation `codex exec --json` event logs and final answers | |
| ``` | |
| ## Reproduce | |
| ```bash | |
| export OPENAI_API_KEY=... # API arm | |
| python3 scripts/codex_route_ab.py --model gpt-6-luna --effort medium --rounds 50 --parallel 4 \ | |
| --oauth-codex-auth ~/.codex/auth.json # OAuth arm (copied; refuses logins >7 days old) | |
| python3 scripts/codex_route_ab_report.py codex-route-ab-<timestamp> | |
| uv run scripts/codex_route_ab_plot.py codex-route-ab-<timestamp> | |
| ``` | |
| Add `--login-only` to check both logins without calling a model. Use `--prompt-file` and | |
| `--expected` to run a different task. Credentials reach the CLI only through stdin and | |
| are deleted with the temporary homes when the run ends. | |
Xet Storage Details
- Size:
- 5.53 kB
- Xet hash:
- bb35fe66d94d2e677d094d78cec163d2079c637318190f3abc02eedba32fcf90
·
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.