InferScale-Sim
Loading Python runtime...

LLM serving simulator

Discrete-event Python model for queueing, batching, KV state, prefill/decode disaggregation, and stateful agent workloads. Device timings use analytical reference profiles unless measurements are imported.

Python/WASM / no backend / deterministic replay / JSON, CSV, PNG, and Markdown export

Timing note. L4, A10G, and A100 latencies are analytical predictions, not measured GPU benchmarks. Queueing, cache, transfer, and SLO behavior is simulated live.

Run summary

Waiting

No run yet

The Python simulator executes in a background Web Worker and returns request-level virtual timestamps.