tracinginsights/2026 / experiments
9.3 GB
51,008 files
Updated 27 days ago
Name
Size
README.md9.09 kB
xet
benchmark.py8.55 kB
xet
builder.py9.31 kB
xet
compare_outputs.py2.6 kB
xet
exp2_merge_once.py51.4 kB
xet
exp3_workers.py51.4 kB
xet
exp4_fastslice.py52.1 kB
xet
exp5_saturate.py49.6 kB
xet
exp6_numpy_drvhead.py57.9 kB
xet
exp7_numpy_merge.py69.7 kB
xet
make_exp6.py11.2 kB
xet
make_exp7.py19.5 kB
xet
make_results.py6.72 kB
xet
prof_lap.py3.15 kB
xet
progress_chart.png121 kB
xet
results.json3.42 kB
xet
results.md755 Bytes
xet
run_par.py49.6 kB
xet
run_seq.py42.4 kB
xet
README.md

Rp.py Optimization Campaign (GEPA "optimize anything" style)

This folder contains an iterative, GEPA-style optimization campaign applied to Rp.py (the parallel telemetry-extraction variant) benchmarked against R.py (the sequential original).

Hard requirement: byte-identical outputs

The user constraint is that every variant must produce byte-identical output to R.py. This is verified (see compare_outputs.py) by running each variant into a fresh output directory and diffing every *.json file against the R.py reference byte-for-byte.

Key finding: Rp.py's per-lap path (_lap_telemetry_or_none) is a line-for-line mirror of FastF1's Lap.get_telemetry(), so the parallel baseline is already byte-identical to R.py (all 1368 JSON files match).

Methodology (mirroring optimize_anything)

This repository does not commit an API key or provider configuration, so the saved campaign is a GEPA-style artifact/evaluator search, not a claim that an LLM proposer was run in this checkout. Each mutation is retained as a separate text artifact, evaluated by the same executable benchmark, and accepted only when its diagnostic gate (byte identity) passes. The optional official gepa package can be installed when running an LLM-guided proposer; the deterministic benchmark and verification harness do not require it.

  1. Seed: start from the existing implementations (R.py sequential, Rp.py parallel) as baseline candidates.
  2. Evaluate: run each candidate end-to-end on the 2026 Austrian Grand Prix - Race session, writing outputs into a fresh directory (EXP_ROOT) so the "file already exists => skip regeneration" guard never short-circuits the work. Reuse the warm FastF1 cache/ so no network is needed. Metric = wall-clock seconds (lower is better), gated by byte-identity vs R.py.
  3. Propose / mutate: apply a single, measurable optimization per experiment (each saved as its own file), re-benchmark, and verify byte-identity.
  4. Accept: adopt the best byte-identical candidate as the base.

Each experiment is a separate, self-contained file (no overwrite of Rp.py).

Results (Austrian GP Race, 1,368 JSON files / 1,341 telemetry files each)

Experiment File Time (s) vs R vs Rp Byte-identical?
exp7_numpy_merge.py [BYTE-IDENTICAL BEST] exp7_numpy_merge.py 6.53 60.7x 10.6x YES
exp6_numpy_drvhead.py exp6_numpy_drvhead.py 18.50 21.4x 3.7x YES
exp5_saturate.py exp5_saturate.py 51.57 7.7x 1.3x YES
Rp.py (parallel baseline) run_par.py 68.97 5.7x 1.0x YES
R.py (sequential baseline) run_seq.py 396.13 1.0x 0.2x YES
exp4_fastslice.py (NOT byte-identical) exp4_fastslice.py 11.11 35.7x 6.2x NO
exp3_workers.py (NOT byte-identical) exp3_workers.py 27.94 14.2x 2.5x NO
exp2_merge_once.py (NOT byte-identical) exp2_merge_once.py 38.94 10.2x 1.8x NO

Chart: progress_chart.png - raw data: results.json / results.md.

What was optimized (and what is / is not byte-identical)

The dominant per-lap cost in this pipeline is the FastF1 per-lap get_telemetry chain. Profiling showed its cost is concentrated in three places, each of which was replaced by an exact numpy/scipy replica in a separate byte-identical experiment:

  1. Parallelize (Rp vs R) - 5.7x: process drivers in a ProcessPoolExecutor sharing the loaded session via fork (copy-on-write) instead of one driver at a time. Byte-identical. OK
  2. exp5 saturate workers - 1.3x vs Rp: keep the exact per-lap get_telemetry but raise the worker cap from 16 to 32 so all 22 drivers run in a single wave. Byte-identical. OK
  3. exp6 numpy driver-ahead - 2.8x vs exp5: FastF1's add_driver_ahead (~180-280 ms/lap, ~75% of per-lap time; it integrates distance for all ~19 other drivers and runs an outer join + argmin) is re-implemented in numpy. The distance integration reproduces pandas' exact float sequence (/1e9 seconds on int64 ns, Speed / 3.6, float64 cumsum); the pandas "outer join on SessionTime" + per-row argmin matrix math is reproduced element-for-element. Byte-identical (all 1368 JSON files match R.py). OK
  4. exp7 numpy merge+fill+slice - 2.4x vs exp6, the byte-identical champion: the remaining per-lap pandas machinery - add_distance / add_relative_distance, the two merge_channels(frequency='original') calls (car+driver-ahead and pos+car) and the slice_by_lap(interpolate_edges=True) edge merge - is replaced with numpy/scipy replicas:
    • add_distance: the pandas dt sequence is reproduced exactly, including the scalar Timedelta.total_seconds() semantics for the first row (days*86400 + seconds + microseconds/1e6 with floor-based components, nanoseconds dropped) which differs from the vectorized .dt.total_seconds used for the diff.
    • merge_channels: outer join on the int64-ns Date union with per-channel interpolation dispatch matching pandas exactly - np.interp for 'index' channels (Speed, RPM, Throttle, DistanceToDriverAhead), scipy.interpolate.interp1d(kind='quadratic', fill_value='extrapolate') with pandas' leading-NaN preservation for 'quadratic' channels (Distance, RelativeDistance, X, Y, Z), and ffill().ffill().bfill() for 'discrete' channels (nGear, Brake, DRS, DriverAhead, Status).
    • slice_by_lap: the 2-row edges merge plus the [start, end] mask and the Time = SessionTime - start_time shift. Byte-identical (all 1368 JSON files match R.py). OK

The earlier exp2/exp3/exp4 experiments tried to remove the repeated per-lap merge and are dramatically faster, but they are NOT byte-identical and are therefore rejected by the campaign's correctness gate:

  • exp2 / exp3 (merge-once) build the telemetry ONCE per driver and slice laps from it. The driver-wide merge timebase differs from the per-lap merge timebase, so the interpolated distance / X / Y / Z / RelativeDistance samples differ from R.py. NO
  • exp4 (fast boolean-mask slice) additionally drops the 2 interpolated boundary rows per lap, so the time arrays have 2 fewer samples per lap. NO

They are kept in this folder as exploration data (they show where the speed ceiling is if you relax the byte-identity requirement), but the final accepted best is exp7_numpy_merge.py, which is verified byte-identical.

Reproducing

# 1) (re)generate the experiment scripts from R.py/Rp.py if needed
python experiments/builder.py          # exp2..exp5 + run_seq/run_par
python experiments/make_exp6.py        # exp6_numpy_drvhead.py
python experiments/make_exp7.py        # exp7_numpy_merge.py

# 2) run a single variant into a fresh output dir (warm ./cache)
rm -rf /tmp/exp_out && mkdir -p /tmp/exp_out
EXP_ROOT=/tmp/exp_out no_proxy='*' python experiments/exp7_numpy_merge.py

# 3) benchmark accepted candidates in fresh, isolated output directories and
#    record timings + byte-identity results in one JSON report. The output root
#    is resolved to an absolute path before child processes start.
.venv/bin/python experiments/benchmark.py --output-root /tmp/rp-benchmark

# Include the intentionally rejected, non-byte-identical explorations too.
.venv/bin/python experiments/benchmark.py --include-exploratory \
    --output-root /tmp/rp-benchmark-exploratory

# 4) or verify one candidate against an existing R.py-derived reference
#    (run_seq.py is the benchmark adapter generated from R.py)
.venv/bin/python experiments/compare_outputs.py /tmp/ref /tmp/exp_out

# 5) validate the fast replicas in-process (asserts every lap matches pandas)
FF1_COMPARE_FAST=1 EXP_ROOT=/tmp/exp_out no_proxy='*' \
    python experiments/exp6_numpy_drvhead.py
FF1_COMPARE_FAST=1 EXP_ROOT=/tmp/exp_out2 no_proxy='*' \
    python experiments/exp7_numpy_merge.py

# 6) regenerate the checked-in summary + chart
.venv/bin/python experiments/make_results.py

The benchmark checks that all generated candidates, including Rf.py, are freshly regenerated from the current R.py/Rp.py source chain before starting, then compares every candidate against the new run_seq.py output. By default it runs only accepted byte-identical candidates; --include-exploratory adds the deliberately rejected exp2/exp3/exp4 variants and returns nonzero if any candidate fails the byte-identity gate.

Notes: environments without os.fork fall back to spawn workers (each worker reloads the session once from cache). Use the repo .venv (Python 3.14) which has fastf1, orjson, numpy, pandas, scipy, psutil, requests and matplotlib; all are declared in requirements.txt. Run one heavy variant at a time - running several 22-process experiments concurrently exhausts RAM and triggers OOM kills. Timings above are best-of/representative single-run figures from one shared session; absolute numbers vary with machine load, relative ordering does not.

Total size
9.3 GB
Files
51,008
Last updated
Sep 7
Pre-warmed CDN
US EU US EU

Contributors