InferScale-Sim / README.md
ArchitSharma's picture
Finalize InferScale-Sim
4649014
|
Raw History Blame Contribute Delete
10.6 kB
metadata
title: InferScale-Sim
colorFrom: indigo
colorTo: blue
sdk: static
app_file: index.html
pinned: false
license: mit
short_description: Interactive LLM serving simulator and SLO planner

InferScale-Sim

A browser-based research workbench for studying LLM serving systems without provisioning a GPU.

The simulator is written in Python. Hugging Face serves a static site; Pyodide runs the same src/inferscale package inside a Web Worker on the visitor's CPU. There is no backend service, model download, API key, or server-side accelerator.

Timing scope. The bundled L4/A10G/A100 profiles are analytical references, not measured hardware benchmarks. Queueing, scheduling, KV-cache, transfer, SLO, and policy behavior is simulated live. Hardware claims require imported measurements and held-out validation.

Research question

InferScale asks how serving-policy choices interact with workload shape and resource constraints:

  • when continuous batching or chunked prefill changes tail latency;
  • when prefill/decode disaggregation helps enough to justify extra accelerators and KV transfer;
  • when reusable prefixes improve goodput versus consume scarce cache capacity;
  • how stateful agent sessions change KV-retention and routing decisions;
  • whether online prediction improves serving outcomes, not merely prediction accuracy;
  • which conclusions remain stable across matched workload seeds and timing uncertainty.

The motivation is practical. Microsoft's Vidur reported an example LLaMA2-70B configuration search that took roughly one CPU-hour through simulation versus an estimated 42,000 GPU-hours (~$218K) using deployment-based exploration.

What is implemented

Stateless serving

  • open-loop constant, Poisson, bursty, and exact trace-replay workloads;
  • static and continuous batching;
  • FCFS, shortest-job-first, least-slack/SLO-aware scheduling;
  • chunked prefill;
  • paged KV-cache accounting and memory admission;
  • controlled shared-prefix reuse;
  • colocated and prefill/decode-disaggregated topologies;
  • independent P/D worker counts, accelerator profiles, and serialized KV-transfer cost;
  • TTFT, TPOT, E2E, queueing, throughput, goodput, SLO attainment, and KV telemetry.

Stateful sessions

  • multi-turn session dependencies separated by sampled tool gaps;
  • HBM retention, TTL expiry, host-memory offload/restore, and pressure eviction;
  • least-load, strict-affinity, and bounded-affinity routing;
  • cross-turn cache hits, history recomputation, HBM/host GB-seconds, and session SLOs;
  • online global/per-tool EWMA tool-gap prediction under a controlled distribution shift.

Execution learning

  • online first-order role-transition learning with no future-trace lookahead;
  • cumulative and decayed transition models;
  • confidence-gated top-1 prefetch;
  • multi-step transition rollout and top-k prefetch;
  • utility-aware prefetch using transfer and forecast-weighted eviction costs;
  • Brier score, log loss, ECE, reliability diagrams, future-role recall@K, prefetch utilization, and speculative-transfer waste;
  • confidence, forgetting-rate, forecast-horizon, and cache-budget studies.

Research consolidation

  • common-random-number A/B experiments;
  • paired bootstrap intervals;
  • timing-model sensitivity analysis;
  • repeated-seed policy ranking;
  • Pareto stability across latency, HBM residency, and unused speculative transfer;
  • a bounded full-trace serving oracle over a declared policy/action-plan family;
  • per-policy regret to that bounded reference;
  • vLLM/SGLang-style measurement import, calibration, and held-out validation hooks;
  • Markdown/JSON/CSV/PNG research-artifact export.

Reference result

The repository includes a deterministic 12-seed analytical-reference consolidation study in reports/reference_consolidation.md.

Policy Median p95 TTFT TTFT wins Pareto stability Median oracle regret
Top-1 decayed 6,474.9 ms 41.7% 75.0% 59.3 ms
No prefetch 6,490.7 ms 41.7% 91.7% 34.7 ms
Utility-aware multi-step 6,507.5 ms 8.3% 58.3% 133.1 ms
Multi-step top-k 6,640.3 ms 8.3% 83.3% 215.0 ms

Top-1 decayed is the robust TTFT winner under the configured ranking protocol, while No prefetch is Pareto-stable on more seeds and has lower median regret to the bounded oracle. The point is not to collapse the study to one score: prediction quality, latency, speculative transfer, and HBM pressure can prefer different policies.

The bounded oracle searches 19 candidate plans per seed and has median p95 TTFT 6,408.6 ms in this reference study. It is an information upper bound over that declared family, not a proof of globally optimal cache scheduling.

Interface map

The public Space is organized as a research tool rather than a product dashboard:

Section Purpose
Serving Inspect one workload and request-level behavior.
Schedulers Compare colocated schedulers on an identical trace.
Capacity Find the highest repeated SLO-compliant offered load.
P/D + Cache Compare colocated/P-D serving with and without prefix reuse.
Design Space Sweep bounded configurations and inspect performance/efficiency Pareto fronts.
A/B Studies Run paired bootstrap and timing-sensitivity experiments.
Stateful Sessions Study KV retention, host offload, affinity, and tool-gap prediction.
Execution Model Study transition learning, multi-step prefetch, calibration, and cache pressure.
Evidence Run repeated-seed consolidation and import external measurements.
Methods Read the simulator assumptions and research lineage.

Architecture

Hugging Face Static Space
        |
        | serves HTML / CSS / JS / Python sources
        v
Browser
  |
  +-- UI + Chart.js
  |
  +-- Web Worker
        |
        +-- Pyodide
              |
              +-- src/inferscale
                    |
                    +-- discrete-event simulation
                    +-- research protocols
                    +-- calibration / reporting

The browser mirror under py/inferscale/ is generated from src/inferscale/; scripts/release_check.py fails if the copies diverge.

Run locally

python3 -m venv .venv
source .venv/bin/activate
python -m pip install -U pip
pip install -e '.[dev]'

pytest -q
python scripts/release_check.py

For the static UI, serve the repository root with any local HTTP server:

python -m http.server 8000

The deployed Space needs no secrets.

Command-line examples

Run a single simulation:

python scripts/run_simulation.py examples/balanced.json

Import a serving-benchmark artifact:

python scripts/import_measurements.py path/to/benchmark.json --format auto

Calibrate analytical timing scales and evaluate held-out cases:

python scripts/calibrate_profiles.py measurements.json --holdout 0.25

Generate a research report:

python scripts/generate_report.py robust-policy-study.json \
  --calibration calibration_result.json

Repository layout

src/inferscale/         canonical Python simulator and research code
py/inferscale/          browser mirror loaded by Pyodide
index.html              static workbench structure
styles.css              project-specific design system
app.js                  UI, charts, exports, worker orchestration
worker.mjs              Pyodide bootstrap and Python action bridge
scripts/                release, import, calibration, report utilities
tests/                  deterministic unit/integration coverage
docs/                   architecture, methodology, validation, research notes
examples/                workload and validation schemas
reports/                 reference analytical study

Validation

The release suite checks both simulation behavior and deployment invariants, including:

  • stateless and P/D smoke tests;
  • trace replay;
  • prefix reuse and design-space Pareto logic;
  • stateful-session, host-tier, HBM-pressure, and affinity experiments;
  • predictive tiering and execution-learning studies;
  • paired bootstrap, timing sensitivity, repeated-seed consolidation, oracle regret;
  • measurement import/calibration/report generation;
  • source/browser Python parity;
  • HF metadata and public provenance guardrails;
  • DOM-reference integrity and chart-export controls.

Run:

pytest -q
python scripts/release_check.py
python -m compileall -q src scripts
node --check app.js
node --check worker.mjs

Design and implementation choices

The interface deliberately avoids a product/SaaS visual language. It uses one restrained accent, flat surfaces, square geometry, system typography, and editorial hierarchy instead of gradients, floating cards, badge-heavy status UI, or decorative marketing sections. The rationale and reusable tokens are documented in docs/design.md.

The simulator core is dependency-light Python rather than a simulation framework. Events, requests, schedulers, caches, predictors, and studies remain inspectable from source and testable outside the browser.

Limitations

  1. Bundled device timings are analytical reference profiles.
  2. The agent-session simulator intentionally isolates state/routing effects from the full dynamic-batching model used in stateless serving.
  3. Transfer models are simplified serialized bandwidth + base-latency abstractions, not packet/NCCL/NIXL simulators.
  4. Oracle policies are bounded information references over declared action families.
  5. Calibration can correct global timing bias; it does not establish fidelity on unseen models, hardware, schedulers, or workload regimes.

These limitations are surfaced in the UI and exported reports rather than hidden.

Research lineage

InferScale is informed by work on serving simulation, disaggregation, stateful agent workloads, and predictive cache management. See docs/research.md for the annotated list and the precise distinction between implemented abstractions and cited systems.

Key starting points include:

  • Vidur — predictive profiling and workload-aware serving simulation.
  • SGLang / RadixAttention — structured prefix reuse and scheduling motivation.
  • Recent 2026 work discussed in docs/research.md on P/D disaggregation, GPU-free emulation, agent-session serving, KV tiering, routing locality, and predictive prefetch.

License

MIT.