Buckets:
1.35 MB
330 files
Updated about 4 hours ago
Ctrl+K
Codex CLI 0.159.2: API key vs ChatGPT OAuth reasoning effort
Question. Does the same client, sending the same request, get the same amount of reasoning from OpenAI models when it authenticates with an API key as when it uses a ChatGPT (Codex) OAuth login?
Answer (so far). No, not for gpt-6-luna. With the official Codex CLI as the
client, the OAuth route produced about 0.39× the reasoning tokens of the API route
at medium effort (n=50 per arm), and was slower. gpt-6.1-sol showed no difference at
medium effort and a smaller gap (about 0.78×) at high effort (n=8).
Method
- Client: Codex CLI
0.159.2,codex exec --json --ephemeral --ignore-user-config --ignore-rules --skip-git-repo-check -s read-only -m <model> -c 'model_reasoning_effort="<effort>"'. The prompt is read from stdin, and the CLI runs in an empty working directory. - The only difference between arms is authentication. Each arm runs in its own
throwaway
CODEX_HOME:- api:
codex login --with-api-key. Standard tier, with no service-tier override. - oauth: a copy of a ChatGPT-mode
~/.codex/auth.json. - Provider environment variables are removed from
codex exec, so neither arm can fall back to the other's credentials.
- api:
- Prompt: a single number-theory proof task with a known answer: all (x, y) with
1 ≤ x ≤ y ≤ 100 where xy divides x² + y² + 1. The answer is
(1,1),(1,2),(2,5),(5,13),(13,34),(34,89). The prompt is in each results folder asprompt.txt(sha2564adea627…). - Runs alternate between the arms (
api,oauth, thenoauth,api, and so on). The n=50 run used 4 parallel workers, shared by both arms. All runs took place on 2 October 2026 between 08:03 and 08:47 UTC. - Metrics come from the CLI's
turn.completedusage events: input, cached input, output andreasoning_output_tokens. Output includes reasoning. Wall-clock time includes CLI start-up.
Results
| run | model / effort | n per arm | API reasoning (median) | OAuth reasoning (median) | OAuth/API (95% CI) | Mann-Whitney p | correct API / OAuth | median wall s API / OAuth |
|---|---|---|---|---|---|---|---|---|
| luna-medium-n50 | gpt-6-luna medium | 50 | 2,674 | 1,034 | 0.39 (0.35–0.41) | 7e-18 | 50 / 49 | 29.3 / 35.9 |
| luna-medium-n4 | gpt-6-luna medium | 4 | 2,287 | 1,325 | 0.58 (0.39–0.85) | 0.06 | 4 / 4 | 30.4 / 46.3 |
| luna-high-n8 | gpt-6-luna high | 8 | 3,445 | 2,053 | 0.60 (0.42–0.71) | 0.0009 | 8 / 8 | 39.7 / 52.7 |
| sol61-medium-n4 | gpt-6.1-sol medium | 4 | 423 | 427 | 1.01 (0.82–1.24) | 1.0 | 4 / 4 | 23.5 / 50.7 |
| sol61-high-n8 | gpt-6.1-sol high | 8 | 957 | 745 | 0.78 (0.66–0.87) | 0.004 | 8 / 8 | 37.1 / 63.7 |
- Confidence intervals are bootstrap intervals for the ratio of medians.
- The OAuth arm consistently sends about 15% more input tokens (11.8k against 13.6k for Luna). The CLI adds different instructions when logged in with ChatGPT.
- The OAuth route is slower even though it produces fewer tokens. For Luna at n=50, the median throughput was 111 output tokens per wall-clock second on the API arm and 42 on the OAuth arm.
- The one wrong answer (luna-medium-n50, OAuth, round 39) used 0 reasoning tokens.
Caveats
- Only one prompt was used, and the small runs (n=4 and n=8) are indicative only. The n=50 Luna run is the robust result.
- The OAuth arm reflects one ChatGPT account and plan, observed at one point in time. Server-side policy can change.
- The comparison is API standard tier against the ChatGPT/Codex backend. The difference is attributed to the server side because the client, prompt, model and requested effort are identical. This matches an earlier independent observation with a different client (fast-agent): about 0.40× reasoning on the Codex route for the same prompt.
Layout
scripts/
codex_route_ab.py runner (stdlib only; needs `codex` on PATH)
codex_route_ab_report.py report.md: medians, p10/p90, ratios, bootstrap CIs, p-values
codex_route_ab_plot.py plots/reasoning_tokens.png, plots/wall_seconds.png (uv run; matplotlib)
results/<model>-<effort>-n<runs per arm>/
report.md summary tables and plots
summary.json run metadata (CLI version, model, effort, prompt sha256, auth method) and medians
results.jsonl one row per generation: route, round, exit, seconds, usage, answer, correct
prompt.txt exact prompt
plots/ reasoning-token and wall-time distributions (kernel density, median, one tick per run)
raw/ per-generation `codex exec --json` event logs and final answers
Reproduce
export OPENAI_API_KEY=... # API arm
python3 scripts/codex_route_ab.py --model gpt-6-luna --effort medium --rounds 50 --parallel 4 \
--oauth-codex-auth ~/.codex/auth.json # OAuth arm (copied; refuses logins >7 days old)
python3 scripts/codex_route_ab_report.py codex-route-ab-<timestamp>
uv run scripts/codex_route_ab_plot.py codex-route-ab-<timestamp>
Add --login-only to check both logins without calling a model. Use --prompt-file and
--expected to run a different task. Credentials reach the CLI only through stdin and
are deleted with the temporary homes when the run ends.
- Total size
- 1.35 MB
- Files
- 330
- Last updated
- Oct 2
- Pre-warmed CDN
- US EU US EU
