|
Download docs/performance.md from HWresearch/GNN4Colliders: direct link, hf CLI and curl.
- Browser
- Download file 2.23 kB
-
https://huggingface.co/HWresearch/GNN4Colliders/resolve/main/docs/performance.md
- Command line
-
hf download hf://HWresearch/GNN4Colliders/docs/performance.md
-
curl -L -o performance.md https://huggingface.co/HWresearch/GNN4Colliders/resolve/main/docs/performance.md
2.23 kB
| # Performance notes | |
| The scripts in [`../benchmarks/`](../benchmarks/) separate setup, warmup, and | |
| steady-state timings and emit JSON lines with execution metadata. Float32 eager | |
| execution remains the correctness reference path. | |
| `fully_connected_edges` has a bounded 32-entry cache keyed by node count, | |
| self-loop policy, and device. It preserves source-major ordering and only | |
| reuses topology indices; event-dependent edge features are always recomputed. | |
| The cache is graph-specific and does not alter scientific behavior. | |
| Training already uses `zero_grad(set_to_none=True)` and inference already uses | |
| `torch.inference_mode()` with detached CPU accumulation. | |
| Mixed precision, `torch.compile`, custom kernels, aggressive worker defaults, | |
| and cache-format replacement were not retained without target-machine | |
| measurements. The main known bottleneck is the quadratic graph workload | |
| `N * (N - 1)` and associated DGL message passing; size-aware batching and | |
| streaming prediction remain follow-up work because they affect ordering or | |
| output semantics. | |
| ## Graph preparation benchmark and parallelism | |
| Benchmark the complete ROOT-to-cache path on a deterministic temporary sample: | |
| ```bash | |
| uv run python benchmarks/benchmark_prepare.py --events 256 --nodes 32 --workers 0 | |
| uv run python benchmarks/benchmark_prepare.py --events 256 --nodes 32 --workers 4 | |
| ``` | |
| The existing training-step benchmark reports throughput, device metadata, and | |
| optional profiler traces: | |
| ```bash | |
| uv run python benchmarks/benchmark_training.py --iterations 50 --warmup 10 | |
| uv run python benchmarks/benchmark_training.py --device cuda --profile | |
| ``` | |
| Preparation uses contiguous event ranges and writes one temporary cache shard | |
| per worker before merging shards in source order. Set `data.num_workers` to a | |
| positive value to enable local process parallelism: | |
| ```bash | |
| uv run gnn4colliders prepare \ | |
| --config-name config_hf_smoke \ | |
| data.num_workers=4 | |
| ``` | |
| Workers never write the final cache concurrently. The merge preserves event | |
| ordering, labels, named metadata, and graph tensors. Small samples can be | |
| slower with workers because process startup and cache merging dominate; use | |
| the benchmark on the target dataset and machine before selecting a value. | |