File size: 5,975 Bytes
0c6c82c
 
ce2d64b
0c6c82c
44745f2
0c6c82c
ce2d64b
44745f2
 
 
ce2d64b
 
44745f2
ce2d64b
 
8f91935
9916edb
 
 
 
 
 
0c6c82c
44745f2
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
ce2d64b
 
e5c4ee4
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
ce2d64b
 
 
 
 
 
 
 
 
9916edb
 
 
 
 
 
 
 
 
 
 
 
 
ce2d64b
 
9916edb
0c6c82c
 
 
ce2d64b
0c6c82c
44745f2
0c6c82c
 
 
44745f2
0c6c82c
 
 
 
 
 
ce2d64b
94910ac
 
 
4649014
94910ac
 
 
 
 
 
 
 
 
 
 
 
 
 
8f91935
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
# Architecture

InferScale-Sim separates **serving-system logic**, **analytical latency estimation**, and **research methodology**.

## Main Python modules

1. `workloads.py` creates deterministic synthetic workloads or exact trace-replay requests.
2. `simulator.py` implements the colocated serving loop.
3. `disaggregated.py` implements separate prefill/decode worker pools plus KV transfer.
4. `kv_cache.py` handles memory/admission and shared-prefix allocation.
5. `latency.py` predicts reference prefill/decode operation durations and exposes sensitivity scales.
6. `metrics.py` derives TTFT, TPOT, E2E, queueing, throughput, goodput, and SLO attainment.
7. `diagnostics.py` converts simulated telemetry into explicit heuristic bottleneck labels.
8. `optimizer.py` implements capacity search, scheduler comparison, topology/cache comparison, and Pareto sweeps.
9. `research.py` implements paired common-seed A/B studies, bootstrap intervals, and analytical-model sensitivity analysis.
10. `agentic.py` models multi-turn programs, tool gaps, session routing, KV retention/TTL eviction, host tiering, and online tool-gap prediction.
11. `execution.py` models online agent-role transition learning, calibration, multi-step forecast planning, and bounded static-prefix prefetch under a shifting workflow distribution.
12. `consolidation.py` repeats execution-policy studies across matched seeds, computes bootstrap uncertainty and Pareto stability, and evaluates regret to a bounded full-trace oracle family.
13. `measurements.py` normalizes external vLLM/SGLang-style benchmark artifacts and fits simple train-only timing-scale calibration.
14. `validation.py` compares baseline or calibrated simulator predictions against externally supplied measured cases.
15. `reports.py` exports the consolidated study and optional empirical calibration as Markdown.
16. `api.py` exposes JSON-like actions to local Python and Pyodide.

## Colocated path

```text
arrival -> waiting -> prefill -> active decode batch -> complete
```

## P/D path

```text
arrival
  -> prefill queue
  -> prefill worker batch
  -> KV-transfer link
  -> decode-ready queue
  -> continuous decode worker
  -> complete
```

The P/D path is a genuine discrete-event loop: prefill workers, decode workers, and the transfer link can overlap in virtual time.


## Stateful agent-session path

```text
session arrival
  -> turn ready
  -> route to replica
  -> [KV hit: append prefill | KV miss: full-history prefill]
  -> decode
  -> retain / TTL / evict KV
  -> tool gap
  -> next turn ready
  -> ...
  -> session complete
```

Each replica is intentionally a serial service station in this mode. This isolates state residency, routing locality, and tool-gap effects from the dynamic-batching questions already covered by the request-level simulators.

## Research path

```text
base configuration
      |
      +--> paired A/B study --> shared seeds --> paired deltas --> bootstrap CI
      |
      +--> sensitivity study --> shared latency perturbations --> ranking/SLO stability
      |
      +--> repeated-seed policy study
      |       |
      |       +--> TTFT win rate / bootstrap CI / worst seed
      |       +--> Pareto stability
      |       `--> bounded full-trace oracle --> policy regret
      |
      +--> external serving artifacts
      |       |
      |       +--> normalize cases
      |       +--> train-only timing-scale fit
      |       `--> held-out residuals / MAPE
      |
      `--> Markdown research report
```

The simulator and statistical layer are separate on purpose: research conclusions are derived from repeated simulations rather than from one displayed run. The bounded oracle is exhaustive only over a declared family of deployable and clairvoyant future-set plans; it is not a proof of globally optimal cache scheduling.

## Browser execution

Canonical source lives in `src/inferscale`. `scripts/sync_web_python.py` mirrors it into `py/inferscale`. A module Web Worker loads Pyodide, writes those files into the virtual filesystem, imports `inferscale.api`, and exchanges JSON messages with the UI.

The simulator therefore runs away from the browser main thread and uses the same Python implementation as the local tests.

## Extension boundary

`AnalyticalLatencyModel` can later be replaced by an empirical profile interpolator exposing:

- `prefill_seconds(token_counts)`
- `decode_step_seconds(context_lengths)`
- `model_weight_gb`
- `kv_bytes_per_token()`

without rewriting workload generation, scheduling, cache logic, P/D orchestration, SLO metrics, or research protocols.

## Agent memory tiering

Stateful Stateful Sessions has an additional memory path that is independent from the stateless/P-D simulator:

```text
turn completes
   |
   +-- retain HBM --------------------------+
   |                                        |
   +-- TTL -> expire / pressure evict       | next turn
   |                                        |
   +-- host offload -> host KV -> restore --+
   |                                        |
   +-- evict -> history recomputation ------+
```

A global host tier models capacity, residency, offload/restore volume, and transfer latency. A bounded-affinity router can trade cached-replica locality against estimated queue imbalance.


## Execution-learning path

```text
current agent role
  -> predict next role from observed transition history
  -> confidence gate
  -> [optional host -> HBM static-prefix prefetch]
  -> tool gap
  -> next role becomes observable
  -> update transition model
  -> run next step with prefix hit/miss
```

This path is intentionally separate from session-KV retention. It studies cross-workflow reuse of static agent prefixes, transition-model adaptation, prefetch precision/coverage, cache pollution, and transfer waste. The default workflow generator changes its transition matrix partway through the trace so cumulative and forgetting-based learners can be compared under non-stationarity.