|
Download perf-tests/README.md from SaylorTwift/gemini-cli: direct link, hf CLI and curl.
- Browser
- Download file 3.74 kB
-
https://huggingface.co/SaylorTwift/gemini-cli/resolve/main/perf-tests/README.md
- Command line
-
hf download hf://SaylorTwift/gemini-cli/perf-tests/README.md
-
curl -L -o README.md https://huggingface.co/SaylorTwift/gemini-cli/resolve/main/perf-tests/README.md
3.74 kB
CPU Performance Integration Test Harness
Overview
This directory contains performance/CPU integration tests for the Gemini CLI. These tests measure wall-clock time, CPU usage, and event loop responsiveness to detect regressions across key scenarios.
CPU performance is inherently noisy, especially in CI. The harness addresses this with:
- IQR outlier filtering β discards anomalous samples
- Median sampling β takes N runs, reports the median after filtering
- Warmup runs β discards the first run to mitigate JIT compilation noise
- 15% default tolerance β won't panic at slight regressions
Running
# Run tests (compare against committed baselines)
npm run test:perf
# Update baselines (after intentional changes)
npm run test:perf:update-baselines
# Verbose output
VERBOSE=true npm run test:perf
# Keep test artifacts for debugging
KEEP_OUTPUT=true npm run test:perf
How It Works
Measurement Primitives
The PerfTestHarness class (in packages/test-utils) provides:
performance.now()β high-resolution wall-clock timingprocess.cpuUsage()β user + system CPU microseconds (delta between start/stop)perf_hooks.monitorEventLoopDelay()β event loop delay histogram (p50/p95/p99/max)
Noise Reduction
- Warmup: First run is discarded to mitigate JIT compilation artifacts
- Multiple samples: Each scenario runs N times (default 5)
- IQR filtering: Samples outside Q1β1.5ΓIQR and Q3+1.5ΓIQR are discarded
- Median: The median of remaining samples is used for comparison
Baseline Management
Baselines are stored in baselines.json in this directory. Each scenario has:
{
"cold-startup-time": {
"wallClockMs": 1234.5,
"cpuTotalUs": 567890,
"eventLoopDelayP99Ms": 12.3,
"timestamp": "2026-04-08T..."
}
}
Tests fail if the measured value exceeds baseline Γ 1.15 (15% tolerance).
To recalibrate after intentional changes:
npm run test:perf:update-baselines
# then commit baselines.json
Report Output
After all tests, the harness prints an ASCII summary:
βββββββββββββββββββββββββββββββββββββββββββββββββββ
PERFORMANCE TEST REPORT
βββββββββββββββββββββββββββββββββββββββββββββββββββ
cold-startup-time: 1234.5 ms (Baseline: 1200.0 ms, Delta: +2.9%) β
idle-cpu-usage: 2.1 % (Baseline: 2.0 %, Delta: +5.0%) β
skill-loading-time: 1567.8 ms (Baseline: 1500.0 ms, Delta: +4.5%) β
Architecture
perf-tests/
βββ README.md β you are here
βββ baselines.json β committed baseline values
βββ globalSetup.ts β test environment setup
βββ perf-usage.test.ts β test scenarios
βββ perf.*.responses β fake API responses per scenario
βββ tsconfig.json β TypeScript config
βββ vitest.config.ts β vitest config (serial, isolated)
packages/test-utils/src/
βββ perf-test-harness.ts β PerfTestHarness class
βββ index.ts β re-exports
CI Integration
These tests are excluded from preflight and designed for nightly CI:
- name: Performance regression tests
run: npm run test:perf
Adding a New Scenario
- Add a fake response file:
perf.<scenario-name>.responses - Add a test case in
perf-usage.test.tsusingharness.runScenario() - Run
npm run test:perf:update-baselinesto establish initial baseline - Commit the updated
baselines.json