Spaces:
Running
Running
|
Download README.md from ArchitSharma/InferScale-Sim: direct link, hf CLI and curl.
- Browser
- Download file 10.6 kB
-
https://huggingface.co/spaces/ArchitSharma/InferScale-Sim/resolve/main/README.md
- Command line
-
hf download hf://spaces/ArchitSharma/InferScale-Sim/README.md
-
curl -L -o README.md https://huggingface.co/spaces/ArchitSharma/InferScale-Sim/resolve/main/README.md
10.6 kB
| title: InferScale-Sim | |
| colorFrom: indigo | |
| colorTo: blue | |
| sdk: static | |
| app_file: index.html | |
| pinned: false | |
| license: mit | |
| short_description: Interactive LLM serving simulator and SLO planner | |
| # InferScale-Sim | |
| A browser-based research workbench for studying LLM serving systems without provisioning a GPU. | |
| The simulator is written in Python. Hugging Face serves a static site; Pyodide runs the same `src/inferscale` package inside a Web Worker on the visitor's CPU. There is no backend service, model download, API key, or server-side accelerator. | |
| > **Timing scope.** The bundled L4/A10G/A100 profiles are analytical references, not measured hardware benchmarks. Queueing, scheduling, KV-cache, transfer, SLO, and policy behavior is simulated live. Hardware claims require imported measurements and held-out validation. | |
| ## Research question | |
| InferScale asks how serving-policy choices interact with workload shape and resource constraints: | |
| - when continuous batching or chunked prefill changes tail latency; | |
| - when prefill/decode disaggregation helps enough to justify extra accelerators and KV transfer; | |
| - when reusable prefixes improve goodput versus consume scarce cache capacity; | |
| - how stateful agent sessions change KV-retention and routing decisions; | |
| - whether online prediction improves serving outcomes, not merely prediction accuracy; | |
| - which conclusions remain stable across matched workload seeds and timing uncertainty. | |
| The motivation is practical. Microsoft's [Vidur](https://arxiv.org/abs/2405.05465) reported an example LLaMA2-70B configuration search that took roughly one CPU-hour through simulation versus an estimated 42,000 GPU-hours (~$218K) using deployment-based exploration. | |
| ## What is implemented | |
| ### Stateless serving | |
| - open-loop constant, Poisson, bursty, and exact trace-replay workloads; | |
| - static and continuous batching; | |
| - FCFS, shortest-job-first, least-slack/SLO-aware scheduling; | |
| - chunked prefill; | |
| - paged KV-cache accounting and memory admission; | |
| - controlled shared-prefix reuse; | |
| - colocated and prefill/decode-disaggregated topologies; | |
| - independent P/D worker counts, accelerator profiles, and serialized KV-transfer cost; | |
| - TTFT, TPOT, E2E, queueing, throughput, goodput, SLO attainment, and KV telemetry. | |
| ### Stateful sessions | |
| - multi-turn session dependencies separated by sampled tool gaps; | |
| - HBM retention, TTL expiry, host-memory offload/restore, and pressure eviction; | |
| - least-load, strict-affinity, and bounded-affinity routing; | |
| - cross-turn cache hits, history recomputation, HBM/host GB-seconds, and session SLOs; | |
| - online global/per-tool EWMA tool-gap prediction under a controlled distribution shift. | |
| ### Execution learning | |
| - online first-order role-transition learning with no future-trace lookahead; | |
| - cumulative and decayed transition models; | |
| - confidence-gated top-1 prefetch; | |
| - multi-step transition rollout and top-k prefetch; | |
| - utility-aware prefetch using transfer and forecast-weighted eviction costs; | |
| - Brier score, log loss, ECE, reliability diagrams, future-role recall@K, prefetch utilization, and speculative-transfer waste; | |
| - confidence, forgetting-rate, forecast-horizon, and cache-budget studies. | |
| ### Research consolidation | |
| - common-random-number A/B experiments; | |
| - paired bootstrap intervals; | |
| - timing-model sensitivity analysis; | |
| - repeated-seed policy ranking; | |
| - Pareto stability across latency, HBM residency, and unused speculative transfer; | |
| - a bounded full-trace serving oracle over a declared policy/action-plan family; | |
| - per-policy regret to that bounded reference; | |
| - vLLM/SGLang-style measurement import, calibration, and held-out validation hooks; | |
| - Markdown/JSON/CSV/PNG research-artifact export. | |
| ## Reference result | |
| The repository includes a deterministic 12-seed analytical-reference consolidation study in [`reports/reference_consolidation.md`](reports/reference_consolidation.md). | |
| | Policy | Median p95 TTFT | TTFT wins | Pareto stability | Median oracle regret | | |
| |---|---:|---:|---:|---:| | |
| | Top-1 decayed | 6,474.9 ms | 41.7% | 75.0% | 59.3 ms | | |
| | No prefetch | 6,490.7 ms | 41.7% | **91.7%** | **34.7 ms** | | |
| | Utility-aware multi-step | 6,507.5 ms | 8.3% | 58.3% | 133.1 ms | | |
| | Multi-step top-k | 6,640.3 ms | 8.3% | 83.3% | 215.0 ms | | |
| `Top-1 decayed` is the robust TTFT winner under the configured ranking protocol, while `No prefetch` is Pareto-stable on more seeds and has lower median regret to the bounded oracle. The point is not to collapse the study to one score: prediction quality, latency, speculative transfer, and HBM pressure can prefer different policies. | |
| The bounded oracle searches 19 candidate plans per seed and has median p95 TTFT **6,408.6 ms** in this reference study. It is an information upper bound over that declared family, not a proof of globally optimal cache scheduling. | |
| ## Interface map | |
| The public Space is organized as a research tool rather than a product dashboard: | |
| | Section | Purpose | | |
| |---|---| | |
| | **Serving** | Inspect one workload and request-level behavior. | | |
| | **Schedulers** | Compare colocated schedulers on an identical trace. | | |
| | **Capacity** | Find the highest repeated SLO-compliant offered load. | | |
| | **P/D + Cache** | Compare colocated/P-D serving with and without prefix reuse. | | |
| | **Design Space** | Sweep bounded configurations and inspect performance/efficiency Pareto fronts. | | |
| | **A/B Studies** | Run paired bootstrap and timing-sensitivity experiments. | | |
| | **Stateful Sessions** | Study KV retention, host offload, affinity, and tool-gap prediction. | | |
| | **Execution Model** | Study transition learning, multi-step prefetch, calibration, and cache pressure. | | |
| | **Evidence** | Run repeated-seed consolidation and import external measurements. | | |
| | **Methods** | Read the simulator assumptions and research lineage. | | |
| ## Architecture | |
| ```text | |
| Hugging Face Static Space | |
| | | |
| | serves HTML / CSS / JS / Python sources | |
| v | |
| Browser | |
| | | |
| +-- UI + Chart.js | |
| | | |
| +-- Web Worker | |
| | | |
| +-- Pyodide | |
| | | |
| +-- src/inferscale | |
| | | |
| +-- discrete-event simulation | |
| +-- research protocols | |
| +-- calibration / reporting | |
| ``` | |
| The browser mirror under `py/inferscale/` is generated from `src/inferscale/`; `scripts/release_check.py` fails if the copies diverge. | |
| ## Run locally | |
| ```bash | |
| python3 -m venv .venv | |
| source .venv/bin/activate | |
| python -m pip install -U pip | |
| pip install -e '.[dev]' | |
| pytest -q | |
| python scripts/release_check.py | |
| ``` | |
| For the static UI, serve the repository root with any local HTTP server: | |
| ```bash | |
| python -m http.server 8000 | |
| ``` | |
| The deployed Space needs no secrets. | |
| ## Command-line examples | |
| Run a single simulation: | |
| ```bash | |
| python scripts/run_simulation.py examples/balanced.json | |
| ``` | |
| Import a serving-benchmark artifact: | |
| ```bash | |
| python scripts/import_measurements.py path/to/benchmark.json --format auto | |
| ``` | |
| Calibrate analytical timing scales and evaluate held-out cases: | |
| ```bash | |
| python scripts/calibrate_profiles.py measurements.json --holdout 0.25 | |
| ``` | |
| Generate a research report: | |
| ```bash | |
| python scripts/generate_report.py robust-policy-study.json \ | |
| --calibration calibration_result.json | |
| ``` | |
| ## Repository layout | |
| ```text | |
| src/inferscale/ canonical Python simulator and research code | |
| py/inferscale/ browser mirror loaded by Pyodide | |
| index.html static workbench structure | |
| styles.css project-specific design system | |
| app.js UI, charts, exports, worker orchestration | |
| worker.mjs Pyodide bootstrap and Python action bridge | |
| scripts/ release, import, calibration, report utilities | |
| tests/ deterministic unit/integration coverage | |
| docs/ architecture, methodology, validation, research notes | |
| examples/ workload and validation schemas | |
| reports/ reference analytical study | |
| ``` | |
| ## Validation | |
| The release suite checks both simulation behavior and deployment invariants, including: | |
| - stateless and P/D smoke tests; | |
| - trace replay; | |
| - prefix reuse and design-space Pareto logic; | |
| - stateful-session, host-tier, HBM-pressure, and affinity experiments; | |
| - predictive tiering and execution-learning studies; | |
| - paired bootstrap, timing sensitivity, repeated-seed consolidation, oracle regret; | |
| - measurement import/calibration/report generation; | |
| - source/browser Python parity; | |
| - HF metadata and public provenance guardrails; | |
| - DOM-reference integrity and chart-export controls. | |
| Run: | |
| ```bash | |
| pytest -q | |
| python scripts/release_check.py | |
| python -m compileall -q src scripts | |
| node --check app.js | |
| node --check worker.mjs | |
| ``` | |
| ## Design and implementation choices | |
| The interface deliberately avoids a product/SaaS visual language. It uses one restrained accent, flat surfaces, square geometry, system typography, and editorial hierarchy instead of gradients, floating cards, badge-heavy status UI, or decorative marketing sections. The rationale and reusable tokens are documented in [`docs/design.md`](docs/design.md). | |
| The simulator core is dependency-light Python rather than a simulation framework. Events, requests, schedulers, caches, predictors, and studies remain inspectable from source and testable outside the browser. | |
| ## Limitations | |
| 1. Bundled device timings are analytical reference profiles. | |
| 2. The agent-session simulator intentionally isolates state/routing effects from the full dynamic-batching model used in stateless serving. | |
| 3. Transfer models are simplified serialized bandwidth + base-latency abstractions, not packet/NCCL/NIXL simulators. | |
| 4. Oracle policies are bounded information references over declared action families. | |
| 5. Calibration can correct global timing bias; it does not establish fidelity on unseen models, hardware, schedulers, or workload regimes. | |
| These limitations are surfaced in the UI and exported reports rather than hidden. | |
| ## Research lineage | |
| InferScale is informed by work on serving simulation, disaggregation, stateful agent workloads, and predictive cache management. See [`docs/research.md`](docs/research.md) for the annotated list and the precise distinction between implemented abstractions and cited systems. | |
| Key starting points include: | |
| - [Vidur](https://arxiv.org/abs/2405.05465) — predictive profiling and workload-aware serving simulation. | |
| - [SGLang / RadixAttention](https://arxiv.org/abs/2312.07104) — structured prefix reuse and scheduling motivation. | |
| - Recent 2026 work discussed in `docs/research.md` on P/D disaggregation, GPU-free emulation, agent-session serving, KV tiering, routing locality, and predictive prefetch. | |
| ## License | |
| MIT. | |