InferScale-Sim / docs /architecture.md
ArchitSharma's picture
Finalize InferScale-Sim
4649014
|
Raw History Blame Contribute Delete
5.98 kB

Architecture

InferScale-Sim separates serving-system logic, analytical latency estimation, and research methodology.

Main Python modules

  1. workloads.py creates deterministic synthetic workloads or exact trace-replay requests.
  2. simulator.py implements the colocated serving loop.
  3. disaggregated.py implements separate prefill/decode worker pools plus KV transfer.
  4. kv_cache.py handles memory/admission and shared-prefix allocation.
  5. latency.py predicts reference prefill/decode operation durations and exposes sensitivity scales.
  6. metrics.py derives TTFT, TPOT, E2E, queueing, throughput, goodput, and SLO attainment.
  7. diagnostics.py converts simulated telemetry into explicit heuristic bottleneck labels.
  8. optimizer.py implements capacity search, scheduler comparison, topology/cache comparison, and Pareto sweeps.
  9. research.py implements paired common-seed A/B studies, bootstrap intervals, and analytical-model sensitivity analysis.
  10. agentic.py models multi-turn programs, tool gaps, session routing, KV retention/TTL eviction, host tiering, and online tool-gap prediction.
  11. execution.py models online agent-role transition learning, calibration, multi-step forecast planning, and bounded static-prefix prefetch under a shifting workflow distribution.
  12. consolidation.py repeats execution-policy studies across matched seeds, computes bootstrap uncertainty and Pareto stability, and evaluates regret to a bounded full-trace oracle family.
  13. measurements.py normalizes external vLLM/SGLang-style benchmark artifacts and fits simple train-only timing-scale calibration.
  14. validation.py compares baseline or calibrated simulator predictions against externally supplied measured cases.
  15. reports.py exports the consolidated study and optional empirical calibration as Markdown.
  16. api.py exposes JSON-like actions to local Python and Pyodide.

Colocated path

arrival -> waiting -> prefill -> active decode batch -> complete

P/D path

arrival
  -> prefill queue
  -> prefill worker batch
  -> KV-transfer link
  -> decode-ready queue
  -> continuous decode worker
  -> complete

The P/D path is a genuine discrete-event loop: prefill workers, decode workers, and the transfer link can overlap in virtual time.

Stateful agent-session path

session arrival
  -> turn ready
  -> route to replica
  -> [KV hit: append prefill | KV miss: full-history prefill]
  -> decode
  -> retain / TTL / evict KV
  -> tool gap
  -> next turn ready
  -> ...
  -> session complete

Each replica is intentionally a serial service station in this mode. This isolates state residency, routing locality, and tool-gap effects from the dynamic-batching questions already covered by the request-level simulators.

Research path

base configuration
      |
      +--> paired A/B study --> shared seeds --> paired deltas --> bootstrap CI
      |
      +--> sensitivity study --> shared latency perturbations --> ranking/SLO stability
      |
      +--> repeated-seed policy study
      |       |
      |       +--> TTFT win rate / bootstrap CI / worst seed
      |       +--> Pareto stability
      |       `--> bounded full-trace oracle --> policy regret
      |
      +--> external serving artifacts
      |       |
      |       +--> normalize cases
      |       +--> train-only timing-scale fit
      |       `--> held-out residuals / MAPE
      |
      `--> Markdown research report

The simulator and statistical layer are separate on purpose: research conclusions are derived from repeated simulations rather than from one displayed run. The bounded oracle is exhaustive only over a declared family of deployable and clairvoyant future-set plans; it is not a proof of globally optimal cache scheduling.

Browser execution

Canonical source lives in src/inferscale. scripts/sync_web_python.py mirrors it into py/inferscale. A module Web Worker loads Pyodide, writes those files into the virtual filesystem, imports inferscale.api, and exchanges JSON messages with the UI.

The simulator therefore runs away from the browser main thread and uses the same Python implementation as the local tests.

Extension boundary

AnalyticalLatencyModel can later be replaced by an empirical profile interpolator exposing:

  • prefill_seconds(token_counts)
  • decode_step_seconds(context_lengths)
  • model_weight_gb
  • kv_bytes_per_token()

without rewriting workload generation, scheduling, cache logic, P/D orchestration, SLO metrics, or research protocols.

Agent memory tiering

Stateful Stateful Sessions has an additional memory path that is independent from the stateless/P-D simulator:

turn completes
   |
   +-- retain HBM --------------------------+
   |                                        |
   +-- TTL -> expire / pressure evict       | next turn
   |                                        |
   +-- host offload -> host KV -> restore --+
   |                                        |
   +-- evict -> history recomputation ------+

A global host tier models capacity, residency, offload/restore volume, and transfer latency. A bounded-affinity router can trade cached-replica locality against estimated queue imbalance.

Execution-learning path

current agent role
  -> predict next role from observed transition history
  -> confidence gate
  -> [optional host -> HBM static-prefix prefetch]
  -> tool gap
  -> next role becomes observable
  -> update transition model
  -> run next step with prefix hit/miss

This path is intentionally separate from session-KV retention. It studies cross-workflow reuse of static agent prefixes, transition-model adaptation, prefetch precision/coverage, cache pollution, and transfer waste. The default workflow generator changes its transition matrix partway through the trace so cumulative and forgetting-based learners can be compared under non-stationarity.