evalstate's picture
|
download
raw
5.53 kB

Codex CLI 0.159.2: API key vs ChatGPT OAuth reasoning effort

Question. Does the same client, sending the same request, get the same amount of reasoning from OpenAI models when it authenticates with an API key as when it uses a ChatGPT (Codex) OAuth login?

Answer (so far). No, not for gpt-6-luna. With the official Codex CLI as the client, the OAuth route produced about 0.39× the reasoning tokens of the API route at medium effort (n=50 per arm), and was slower. gpt-6.1-sol showed no difference at medium effort and a smaller gap (about 0.78×) at high effort (n=8).

Luna medium reasoning tokens

Method

  • Client: Codex CLI 0.159.2, codex exec --json --ephemeral --ignore-user-config --ignore-rules --skip-git-repo-check -s read-only -m <model> -c 'model_reasoning_effort="<effort>"'. The prompt is read from stdin, and the CLI runs in an empty working directory.
  • The only difference between arms is authentication. Each arm runs in its own throwaway CODEX_HOME:
    • api: codex login --with-api-key. Standard tier, with no service-tier override.
    • oauth: a copy of a ChatGPT-mode ~/.codex/auth.json.
    • Provider environment variables are removed from codex exec, so neither arm can fall back to the other's credentials.
  • Prompt: a single number-theory proof task with a known answer: all (x, y) with 1 ≤ x ≤ y ≤ 100 where xy divides x² + y² + 1. The answer is (1,1),(1,2),(2,5),(5,13),(13,34),(34,89). The prompt is in each results folder as prompt.txt (sha256 4adea627…).
  • Runs alternate between the arms (api,oauth, then oauth,api, and so on). The n=50 run used 4 parallel workers, shared by both arms. All runs took place on 2 October 2026 between 08:03 and 08:47 UTC.
  • Metrics come from the CLI's turn.completed usage events: input, cached input, output and reasoning_output_tokens. Output includes reasoning. Wall-clock time includes CLI start-up.

Results

run model / effort n per arm API reasoning (median) OAuth reasoning (median) OAuth/API (95% CI) Mann-Whitney p correct API / OAuth median wall s API / OAuth
luna-medium-n50 gpt-6-luna medium 50 2,674 1,034 0.39 (0.35–0.41) 7e-18 50 / 49 29.3 / 35.9
luna-medium-n4 gpt-6-luna medium 4 2,287 1,325 0.58 (0.39–0.85) 0.06 4 / 4 30.4 / 46.3
luna-high-n8 gpt-6-luna high 8 3,445 2,053 0.60 (0.42–0.71) 0.0009 8 / 8 39.7 / 52.7
sol61-medium-n4 gpt-6.1-sol medium 4 423 427 1.01 (0.82–1.24) 1.0 4 / 4 23.5 / 50.7
sol61-high-n8 gpt-6.1-sol high 8 957 745 0.78 (0.66–0.87) 0.004 8 / 8 37.1 / 63.7
  • Confidence intervals are bootstrap intervals for the ratio of medians.
  • The OAuth arm consistently sends about 15% more input tokens (11.8k against 13.6k for Luna). The CLI adds different instructions when logged in with ChatGPT.
  • The OAuth route is slower even though it produces fewer tokens. For Luna at n=50, the median throughput was 111 output tokens per wall-clock second on the API arm and 42 on the OAuth arm.
  • The one wrong answer (luna-medium-n50, OAuth, round 39) used 0 reasoning tokens.

Caveats

  • Only one prompt was used, and the small runs (n=4 and n=8) are indicative only. The n=50 Luna run is the robust result.
  • The OAuth arm reflects one ChatGPT account and plan, observed at one point in time. Server-side policy can change.
  • The comparison is API standard tier against the ChatGPT/Codex backend. The difference is attributed to the server side because the client, prompt, model and requested effort are identical. This matches an earlier independent observation with a different client (fast-agent): about 0.40× reasoning on the Codex route for the same prompt.

Layout

scripts/
  codex_route_ab.py          runner (stdlib only; needs `codex` on PATH)
  codex_route_ab_report.py   report.md: medians, p10/p90, ratios, bootstrap CIs, p-values
  codex_route_ab_plot.py     plots/reasoning_tokens.png, plots/wall_seconds.png (uv run; matplotlib)
results/<model>-<effort>-n<runs per arm>/
  report.md       summary tables and plots
  summary.json    run metadata (CLI version, model, effort, prompt sha256, auth method) and medians
  results.jsonl   one row per generation: route, round, exit, seconds, usage, answer, correct
  prompt.txt      exact prompt
  plots/          reasoning-token and wall-time distributions (kernel density, median, one tick per run)
  raw/            per-generation `codex exec --json` event logs and final answers

Reproduce

export OPENAI_API_KEY=...                         # API arm
python3 scripts/codex_route_ab.py --model gpt-6-luna --effort medium --rounds 50 --parallel 4 \
  --oauth-codex-auth ~/.codex/auth.json           # OAuth arm (copied; refuses logins >7 days old)
python3 scripts/codex_route_ab_report.py codex-route-ab-<timestamp>
uv run scripts/codex_route_ab_plot.py codex-route-ab-<timestamp>

Add --login-only to check both logins without calling a model. Use --prompt-file and --expected to run a different task. Credentials reach the CLI only through stdin and are deleted with the temporary homes when the run ends.

Xet Storage Details

Size:
5.53 kB
·
Xet hash:
bb35fe66d94d2e677d094d78cec163d2079c637318190f3abc02eedba32fcf90

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.