Spaces:
Running
Running
File size: 5,975 Bytes
0c6c82c ce2d64b 0c6c82c 44745f2 0c6c82c ce2d64b 44745f2 ce2d64b 44745f2 ce2d64b 8f91935 9916edb 0c6c82c 44745f2 ce2d64b e5c4ee4 ce2d64b 9916edb ce2d64b 9916edb 0c6c82c ce2d64b 0c6c82c 44745f2 0c6c82c 44745f2 0c6c82c ce2d64b 94910ac 4649014 94910ac 8f91935 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 | # Architecture
InferScale-Sim separates **serving-system logic**, **analytical latency estimation**, and **research methodology**.
## Main Python modules
1. `workloads.py` creates deterministic synthetic workloads or exact trace-replay requests.
2. `simulator.py` implements the colocated serving loop.
3. `disaggregated.py` implements separate prefill/decode worker pools plus KV transfer.
4. `kv_cache.py` handles memory/admission and shared-prefix allocation.
5. `latency.py` predicts reference prefill/decode operation durations and exposes sensitivity scales.
6. `metrics.py` derives TTFT, TPOT, E2E, queueing, throughput, goodput, and SLO attainment.
7. `diagnostics.py` converts simulated telemetry into explicit heuristic bottleneck labels.
8. `optimizer.py` implements capacity search, scheduler comparison, topology/cache comparison, and Pareto sweeps.
9. `research.py` implements paired common-seed A/B studies, bootstrap intervals, and analytical-model sensitivity analysis.
10. `agentic.py` models multi-turn programs, tool gaps, session routing, KV retention/TTL eviction, host tiering, and online tool-gap prediction.
11. `execution.py` models online agent-role transition learning, calibration, multi-step forecast planning, and bounded static-prefix prefetch under a shifting workflow distribution.
12. `consolidation.py` repeats execution-policy studies across matched seeds, computes bootstrap uncertainty and Pareto stability, and evaluates regret to a bounded full-trace oracle family.
13. `measurements.py` normalizes external vLLM/SGLang-style benchmark artifacts and fits simple train-only timing-scale calibration.
14. `validation.py` compares baseline or calibrated simulator predictions against externally supplied measured cases.
15. `reports.py` exports the consolidated study and optional empirical calibration as Markdown.
16. `api.py` exposes JSON-like actions to local Python and Pyodide.
## Colocated path
```text
arrival -> waiting -> prefill -> active decode batch -> complete
```
## P/D path
```text
arrival
-> prefill queue
-> prefill worker batch
-> KV-transfer link
-> decode-ready queue
-> continuous decode worker
-> complete
```
The P/D path is a genuine discrete-event loop: prefill workers, decode workers, and the transfer link can overlap in virtual time.
## Stateful agent-session path
```text
session arrival
-> turn ready
-> route to replica
-> [KV hit: append prefill | KV miss: full-history prefill]
-> decode
-> retain / TTL / evict KV
-> tool gap
-> next turn ready
-> ...
-> session complete
```
Each replica is intentionally a serial service station in this mode. This isolates state residency, routing locality, and tool-gap effects from the dynamic-batching questions already covered by the request-level simulators.
## Research path
```text
base configuration
|
+--> paired A/B study --> shared seeds --> paired deltas --> bootstrap CI
|
+--> sensitivity study --> shared latency perturbations --> ranking/SLO stability
|
+--> repeated-seed policy study
| |
| +--> TTFT win rate / bootstrap CI / worst seed
| +--> Pareto stability
| `--> bounded full-trace oracle --> policy regret
|
+--> external serving artifacts
| |
| +--> normalize cases
| +--> train-only timing-scale fit
| `--> held-out residuals / MAPE
|
`--> Markdown research report
```
The simulator and statistical layer are separate on purpose: research conclusions are derived from repeated simulations rather than from one displayed run. The bounded oracle is exhaustive only over a declared family of deployable and clairvoyant future-set plans; it is not a proof of globally optimal cache scheduling.
## Browser execution
Canonical source lives in `src/inferscale`. `scripts/sync_web_python.py` mirrors it into `py/inferscale`. A module Web Worker loads Pyodide, writes those files into the virtual filesystem, imports `inferscale.api`, and exchanges JSON messages with the UI.
The simulator therefore runs away from the browser main thread and uses the same Python implementation as the local tests.
## Extension boundary
`AnalyticalLatencyModel` can later be replaced by an empirical profile interpolator exposing:
- `prefill_seconds(token_counts)`
- `decode_step_seconds(context_lengths)`
- `model_weight_gb`
- `kv_bytes_per_token()`
without rewriting workload generation, scheduling, cache logic, P/D orchestration, SLO metrics, or research protocols.
## Agent memory tiering
Stateful Stateful Sessions has an additional memory path that is independent from the stateless/P-D simulator:
```text
turn completes
|
+-- retain HBM --------------------------+
| |
+-- TTL -> expire / pressure evict | next turn
| |
+-- host offload -> host KV -> restore --+
| |
+-- evict -> history recomputation ------+
```
A global host tier models capacity, residency, offload/restore volume, and transfer latency. A bounded-affinity router can trade cached-replica locality against estimated queue imbalance.
## Execution-learning path
```text
current agent role
-> predict next role from observed transition history
-> confidence gate
-> [optional host -> HBM static-prefix prefetch]
-> tool gap
-> next role becomes observable
-> update transition model
-> run next step with prefix hit/miss
```
This path is intentionally separate from session-KV retention. It studies cross-workflow reuse of static agent prefixes, transition-model adaptation, prefetch precision/coverage, cache pollution, and transfer waste. The default workflow generator changes its transition matrix partway through the trace so cumulative and forgetting-based learners can be compared under non-stationarity.
|