InferScale-Sim / index.html
ArchitSharma's picture
Finalize InferScale-Sim
ca14639
Raw History Blame Contribute Delete
84.2 kB
<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="utf-8"/>
<meta content="width=device-width, initial-scale=1" name="viewport"/>
<title>InferScale-Sim - LLM serving systems simulator</title>
<meta content="Browser-based Python simulator for LLM serving, KV-cache policies, P/D disaggregation, and SLO-constrained experiments." name="description"/>
<link crossorigin="" href="https://cdn.jsdelivr.net" rel="preconnect"/>
<link href="styles.css" rel="stylesheet"/>
<script src="https://cdn.jsdelivr.net/npm/chart.js@4.5.1/dist/chart.umd.min.js"></script>
</head>
<body data-interface="research-workbench"><a class="skip-link" href="#lab">Skip to simulator</a>
<header class="topbar" id="topbar">
<div class="brand-wrap"><div class="brand">InferScale-Sim</div></div>
<div class="runtime-pill" id="runtimePill"><span class="dot"></span><span aria-live="polite" id="runtimeText">Loading Python runtime...</span></div>
</header>
<main class="shell">
<section class="intro"><div class="intro-copy"><h1>LLM serving simulator</h1><p>Discrete-event Python model for queueing, batching, KV state, prefill/decode disaggregation, and stateful agent workloads. Device timings use analytical reference profiles unless measurements are imported.</p><p class="intro-meta">Python/WASM / no backend / deterministic replay / JSON, CSV, PNG, and Markdown export</p></div></section>
<div class="reference-note"><strong>Timing note.</strong> L4, A10G, and A100 latencies are analytical predictions, not measured GPU benchmarks. Queueing, cache, transfer, and SLO behavior is simulated live.</div>
<nav aria-label="InferScale sections" class="tabs" role="tablist">
<button aria-controls="lab" aria-selected="true" class="tab active" data-tab="lab" id="tab-lab" role="tab" tabindex="0" type="button">Serving</button>
<button aria-controls="arena" aria-selected="false" class="tab" data-tab="arena" id="tab-arena" role="tab" tabindex="-1" type="button">Schedulers</button>
<button aria-controls="planner" aria-selected="false" class="tab" data-tab="planner" id="tab-planner" role="tab" tabindex="-1" type="button">Capacity</button>
<button aria-controls="modern" aria-selected="false" class="tab" data-tab="modern" id="tab-modern" role="tab" tabindex="-1" type="button">P/D + Cache</button>
<button aria-controls="design" aria-selected="false" class="tab" data-tab="design" id="tab-design" role="tab" tabindex="-1" type="button">Design Space</button>
<button aria-controls="research" aria-selected="false" class="tab" data-tab="research" id="tab-research" role="tab" tabindex="-1" type="button">A/B Studies</button>
<button aria-controls="agent" aria-selected="false" class="tab" data-tab="agent" id="tab-agent" role="tab" tabindex="-1" type="button">Stateful Sessions</button>
<button aria-controls="execution" aria-selected="false" class="tab" data-tab="execution" id="tab-execution" role="tab" tabindex="-1" type="button">Execution Model</button>
<button aria-controls="consolidation" aria-selected="false" class="tab" data-tab="consolidation" id="tab-consolidation" role="tab" tabindex="-1" type="button">Evidence</button>
<button aria-controls="method" aria-selected="false" class="tab" data-tab="method" id="tab-method" role="tab" tabindex="-1" type="button">Methods</button>
</nav>
<section aria-hidden="false" aria-labelledby="tab-lab" class="tab-panel active" id="lab" role="tabpanel">
<div class="workspace">
<aside class="panel controls-panel">
<div class="panel-title-row"><h2>Experiment</h2><span class="tag">Deterministic seed</span></div>
<div class="field-grid two">
<label>Model<select id="model"><option>Qwen2.5-3B</option><option>Llama-3.1-8B</option><option>Mistral-7B-v0.3</option></select></label>
<label>Weight precision<select id="quantization"><option value="fp16">FP16</option><option selected="" value="int8">INT8 scenario</option><option value="int4">INT4 scenario</option></select></label>
</div>
<div class="field-grid two">
<label>Topology<select id="topology"><option value="colocated">Colocated</option><option value="disaggregated_pd">P/D disaggregated</option></select></label>
<label id="colocatedAcceleratorLabel">Accelerator<select id="accelerator"><option value="L4">NVIDIA L4</option><option value="A10G">NVIDIA A10G</option><option value="A100-40GB">NVIDIA A100 40GB</option></select></label>
</div>
<div class="field-grid two">
<label>Scheduler<select id="scheduler"><option value="continuous_fcfs">Continuous - FCFS</option><option value="continuous_sjf">Continuous - SJF</option><option value="continuous_slo">Continuous - SLO-aware</option><option value="chunked_slo">Chunked prefill - SLO-aware</option><option value="static_fcfs">Static batching - FCFS</option></select></label>
<label>Prefix reuse<select id="prefixCache"><option value="off">Off</option><option value="on">On</option></select></label>
</div>
<div class="subcontrols hidden" id="prefixControls">
<div class="section-kicker">Shared-prefix scenario</div>
<div class="field-grid two">
<label>Shared prefix<input id="sharedPrefix" min="0" step="32" type="number" value="256"/><span class="unit">tokens</span></label>
<label>Reuse fraction<input id="prefixReuse" max="1" min="0" step="0.05" type="number" value="0.60"/><span class="unit">fraction</span></label>
</div>
</div>
<div class="subcontrols hidden" id="pdControls">
<div class="section-kicker">P/D topology</div>
<div class="field-grid two">
<label>Prefill accelerator<select id="prefillAccelerator"><option value="L4">NVIDIA L4</option><option value="A10G">NVIDIA A10G</option><option value="A100-40GB">NVIDIA A100 40GB</option></select></label>
<label>Decode accelerator<select id="decodeAccelerator"><option value="L4">NVIDIA L4</option><option value="A10G">NVIDIA A10G</option><option value="A100-40GB">NVIDIA A100 40GB</option></select></label>
<label>Prefill workers<input id="prefillWorkers" max="8" min="1" step="1" type="number" value="1"/></label>
<label>Decode workers<input id="decodeWorkers" max="8" min="1" step="1" type="number" value="1"/></label>
<label>Interconnect<input id="interconnect" min="1" step="1" type="number" value="50"/><span class="unit">GB/s</span></label>
<label>Transfer base<input id="transferBase" min="0" step="0.05" type="number" value="0.20"/><span class="unit">ms</span></label>
</div>
</div>
<hr/>
<div class="section-kicker">Workload</div>
<div class="field-grid two">
<label>Arrival process<select id="arrival"><option value="poisson">Poisson</option><option value="constant">Constant</option><option value="bursty">Bursty</option><option value="trace">Trace replay</option></select></label>
<label>Request rate<input id="rate" min="0.1" step="0.1" type="number" value="4"/><span class="unit">req/s</span></label>
<label>Duration<input id="duration" min="2" step="1" type="number" value="30"/><span class="unit">sim s</span></label>
<label>Seed<input id="seed" step="1" type="number" value="7"/></label>
</div>
<div class="field-grid two hidden" id="burstControls">
<label>Burst multiplier<input id="burstMultiplier" min="1" step="0.25" type="number" value="3"/><span class="unit">x</span></label>
<label>Burst period<input id="burstPeriod" min="0.5" step="0.5" type="number" value="10"/><span class="unit">s</span></label>
</div>
<div class="subcontrols hidden" id="traceControls">
<div class="section-kicker">Exact trace replay</div>
<label class="file-label">Workload file<input accept=".csv,.json,text/csv,application/json" id="traceFile" type="file"/></label>
<div class="trace-row"><span id="traceStatus">No trace loaded</span><button class="mini-button" id="traceClearBtn" type="button">Clear</button></div>
<p class="control-help">CSV or JSON rows: arrival_time, prompt_tokens, output_tokens. Trace replay preserves exact arrivals and token lengths; the Capacity Planner is disabled because offered rate is fixed by the trace.</p>
</div>
<div class="field-grid two">
<label>Prompt mean<input id="promptMean" min="16" type="number" value="512"/><span class="unit">tokens</span></label>
<label>Prompt CV<input id="promptCv" max="2" min="0" step="0.05" type="number" value="0.50"/></label>
<label>Output mean<input id="outputMean" min="1" type="number" value="64"/><span class="unit">tokens</span></label>
<label>Output CV<input id="outputCv" max="2" min="0" step="0.05" type="number" value="0.60"/></label>
</div>
<hr/>
<div class="section-kicker">Serving controls</div>
<div class="field-grid two">
<label>Max batch size<input id="maxBatch" max="128" min="1" type="number" value="16"/></label>
<label>Max batch tokens<input id="maxBatchTokens" min="128" step="128" type="number" value="8192"/></label>
<label>Prefill chunk<input id="chunkSize" min="64" step="64" type="number" value="512"/><span class="unit">tokens</span></label>
<label>KV block<input id="kvBlock" min="1" type="number" value="16"/><span class="unit">tokens</span></label>
</div>
<hr/>
<div class="section-kicker">SLO</div>
<div class="field-grid two">
<label>TTFT limit<input id="sloTtft" min="1" type="number" value="500"/><span class="unit">ms</span></label>
<label>E2E limit<input id="sloE2e" min="100" type="number" value="5000"/><span class="unit">ms</span></label>
</div>
<button class="primary" disabled="" id="runBtn" type="button">Run simulation</button>
<div class="button-row"><button class="secondary" disabled="" id="copyResultBtn" type="button">Copy JSON</button><button class="secondary" disabled="" id="exportBtn" type="button">Download JSON</button></div>
</aside>
<div class="results-column">
<section class="panel result-panel">
<div class="panel-title-row"><h2>Run summary</h2><span aria-live="polite" class="tag neutral" id="runState">Waiting</span></div>
<div class="empty-state" id="emptyState"><h3>No run yet</h3><p>The Python simulator executes in a background Web Worker and returns request-level virtual timestamps.</p></div>
<div class="hidden" id="resultContent">
<div class="metric-grid">
<div class="metric"><span>p95 TTFT</span><strong id="mTtft">N/A</strong></div>
<div class="metric"><span>p95 E2E</span><strong id="mE2e">N/A</strong></div>
<div class="metric"><span>Goodput</span><strong id="mGoodput">N/A</strong></div>
<div class="metric"><span>SLO attainment</span><strong id="mSlo">N/A</strong></div>
<div class="metric"><span>Throughput</span><strong id="mReq">N/A</strong></div>
<div class="metric"><span>Peak KV</span><strong id="mKv">N/A</strong></div>
</div>
<section class="diagnostic-card" id="diagnosticCard">
<div class="diagnostic-head"><span>Simulator diagnosis</span><strong id="mBottleneck">N/A</strong></div>
<p id="mDiagnosis">Run a simulation to generate a bottleneck explanation.</p>
<p class="diagnostic-action" id="mRecommendation"></p>
<div class="evidence-row" id="mEvidence"></div>
</section>
<div class="chart-grid">
<div class="chart-card" data-chart-card="" data-chart-name="latency-percentiles">
<div class="chart-head"><div class="chart-title">Latency percentiles</div><div class="chart-actions"><button class="chart-download" type="button">Download PNG</button><button class="chart-expand" type="button">Expand</button></div></div>
<div class="chart-body"><canvas aria-label="Latency percentiles" id="latencyChart" role="img"></canvas></div>
</div>
<div class="chart-card" data-chart-card="" data-chart-name="queue-decode-kv-timeline">
<div class="chart-head"><div class="chart-title">Queue, decode, and KV timeline</div><div class="chart-actions"><button class="chart-download" type="button">Download PNG</button><button class="chart-expand" type="button">Expand</button></div></div>
<div class="chart-body"><canvas aria-label="Queue, decode, and KV timeline" id="timelineChart" role="img"></canvas></div>
</div>
</div>
<div class="chart-card full" data-chart-card="" data-chart-name="ttft-vs-prompt-length">
<div class="chart-head"><div class="chart-title">Request TTFT vs prompt length</div><div class="chart-actions"><button class="chart-download" type="button">Download PNG</button><button class="chart-expand" type="button">Expand</button></div></div>
<div class="chart-body"><canvas aria-label="Request TTFT vs prompt length" id="scatterChart" role="img"></canvas></div>
</div>
<div class="warnings hidden" id="warnings"></div>
</div>
</section>
</div>
</div>
</section>
<section aria-hidden="true" aria-labelledby="tab-arena" class="tab-panel" id="arena" role="tabpanel">
<div class="panel wide-panel">
<div class="panel-title-row arena-title-row">
<div><h2>Scheduler comparison</h2><p class="muted">Run every colocated scheduler against the current Serving Lab workload and rank by SLO attainment, then goodput.</p></div>
<button class="primary compact" disabled="" id="arenaBtn" type="button">Compare schedulers</button>
</div>
<div class="empty-state small" id="arenaEmpty"><h3>No comparison yet</h3><p>Your Serving Lab workload and model controls are reused automatically.</p></div>
<div class="hidden" id="arenaContent">
<div class="chart-card full" data-chart-card="" data-chart-name="scheduler-throughput-comparison">
<div class="chart-head"><div class="chart-title">Useful vs raw request throughput</div><div class="chart-actions"><button class="chart-download" type="button">Download PNG</button><button class="chart-expand" type="button">Expand</button></div></div>
<div class="chart-body large"><canvas aria-label="Useful vs raw request throughput" id="arenaChart" role="img"></canvas></div>
</div>
<div class="table-toolbar"><span>Scheduler comparison</span><div><button class="mini-button" disabled="" id="arenaCopyBtn" type="button">Copy table</button><button class="mini-button" disabled="" id="arenaCsvBtn" type="button">Download CSV</button></div></div>
<div class="table-wrap"><table><thead><tr><th>Scheduler</th><th>Goodput</th><th>SLO attainment</th><th>p95 TTFT</th><th>p95 E2E</th><th>KV peak</th><th>Unfinished</th><th>Diagnosis</th></tr></thead><tbody id="arenaRows"></tbody></table></div>
</div>
</div>
</section>
<section aria-hidden="true" aria-labelledby="tab-planner" class="tab-panel" id="planner" role="tabpanel">
<div class="workspace planner-grid">
<aside class="panel controls-panel">
<div class="panel-title-row"><h2>Capacity search</h2><span class="tag">Robust binary search</span></div>
<p class="muted">Find the highest offered load where every repetition meets the SLO target and fully drains.</p>
<label>Required SLO attainment<input id="targetSlo" max="1" min="0.5" step="0.001" type="number" value="0.99"/></label>
<div class="field-grid two">
<label>Minimum rate<input id="minRate" min="0.05" step="0.1" type="number" value="0.25"/><span class="unit">req/s</span></label>
<label>Maximum rate<input id="maxRate" min="0.1" step="1" type="number" value="20"/><span class="unit">req/s</span></label>
<label>Search iterations<input id="searchIter" max="12" min="2" type="number" value="7"/></label>
<label>Repetitions / rate<input id="repetitions" max="5" min="1" type="number" value="2"/></label>
</div>
<label>Safety headroom<input id="headroom" max="0.8" min="0" step="0.05" type="number" value="0.20"/><span class="unit">fraction</span></label>
<button class="primary" disabled="" id="capacityBtn" type="button">Find sustainable capacity</button>
</aside>
<section class="panel result-panel">
<div class="panel-title-row"><h2>Planner result</h2><span aria-live="polite" class="tag neutral" id="plannerState">Waiting</span></div>
<div class="empty-state" id="plannerEmpty"><h3>No search yet</h3><p>The planner repeatedly runs the simulator at different offered loads.</p></div>
<div class="hidden" id="plannerContent">
<div class="metric-grid four">
<div class="metric"><span>Estimated capacity</span><strong id="pCapacity">N/A</strong></div>
<div class="metric"><span>Recommended load</span><strong id="pRecommended">N/A</strong></div>
<div class="metric"><span>Safety headroom</span><strong id="pHeadroom">N/A</strong></div>
<div class="metric"><span>Status</span><strong id="pStatus">N/A</strong></div>
</div>
<div class="planner-note" id="plannerCriterion">A rate passes only if every repetition meets the target and drains all generated requests.</div>
<div class="chart-card full" data-chart-card="" data-chart-name="capacity-slo-curve">
<div class="chart-head"><div class="chart-title">SLO attainment across searched rates</div><div class="chart-actions"><button class="chart-download" type="button">Download PNG</button><button class="chart-expand" type="button">Expand</button></div></div>
<div class="chart-body large"><canvas aria-label="SLO attainment across searched rates" id="capacityChart" role="img"></canvas></div>
</div>
<div class="table-toolbar"><span>Search trace</span><div><button class="mini-button" disabled="" id="capacityCopyBtn" type="button">Copy table</button><button class="mini-button" disabled="" id="capacityCsvBtn" type="button">Download CSV</button></div></div>
<div class="table-wrap"><table><thead><tr><th>Rate</th><th>Pass</th><th>Mean SLO</th><th>Worst repetition</th><th>Target</th><th>SLO range</th><th>Goodput</th><th>p95 TTFT</th><th>p95 E2E</th></tr></thead><tbody id="capacityRows"></tbody></table></div>
</div>
</section>
</div>
</section>
<section aria-hidden="true" aria-labelledby="tab-modern" class="tab-panel" id="modern" role="tabpanel">
<div class="panel wide-panel">
<div class="panel-title-row arena-title-row">
<div><h2>Topology and cache comparison</h2><p class="muted">Compare colocated and prefill/decode-disaggregated serving, each with and without the current shared-prefix reuse scenario. P/D worker counts and interconnect settings come from Serving Lab.</p></div>
<button class="primary compact" disabled="" id="topologyBtn" type="button">Compare 4 scenarios</button>
</div>
<div class="empty-state small" id="topologyEmpty"><h3>No comparison yet</h3><p>Set shared-prefix and P/D parameters in Serving Lab, then run this controlled comparison.</p></div>
<div class="hidden" id="topologyContent">
<div class="chart-card full" data-chart-card="" data-chart-name="topology-goodput-ttft">
<div class="chart-head"><div class="chart-title">Goodput vs p95 TTFT</div><div class="chart-actions"><button class="chart-download" type="button">Download PNG</button><button class="chart-expand" type="button">Expand</button></div></div>
<div class="chart-body large"><canvas aria-label="Goodput vs p95 TTFT" id="topologyChart" role="img"></canvas></div>
</div>
<div class="table-toolbar"><span>Topology and cache comparison</span><div><button class="mini-button" disabled="" id="topologyCopyBtn" type="button">Copy table</button><button class="mini-button" disabled="" id="topologyCsvBtn" type="button">Download CSV</button></div></div>
<div class="table-wrap"><table><thead><tr><th>Scenario</th><th>GPU instances</th><th>Goodput</th><th>Goodput / GPU</th><th>SLO attainment</th><th>p95 TTFT</th><th>p95 E2E</th><th>p95 KV transfer</th><th>Cache hit</th><th>Prefill saved</th><th>Diagnosis</th></tr></thead><tbody id="topologyRows"></tbody></table></div>
</div>
</div>
</section>
<section aria-hidden="true" aria-labelledby="tab-design" class="tab-panel" id="design" role="tabpanel">
<div class="panel wide-panel">
<div class="panel-title-row arena-title-row">
<div><h2>Design-space sweep</h2><p class="muted">Sweep 10 colocated candidates plus six P/D/cache variants, then identify the non-dominated goodput/TTFT frontier. This is an interactive design study, not an exhaustive optimizer.</p></div>
<div class="action-stack"><label class="checkline"><input checked="" id="includePd" type="checkbox"/> Include P/D candidates</label><button class="primary compact" disabled="" id="designBtn" type="button">Explore design space</button></div>
</div>
<div class="empty-state small" id="designEmpty"><h3>No sweep yet</h3><p>The current workload, model, SLO, and P/D parameters are reused from Serving Lab.</p></div>
<div class="hidden" id="designContent">
<div class="metric-grid four"><div class="metric"><span>Candidates</span><strong id="dCandidates">N/A</strong></div><div class="metric"><span>Performance Pareto</span><strong id="dPareto">N/A</strong></div><div class="metric"><span>Efficiency Pareto</span><strong id="dEfficiencyPareto">N/A</strong></div><div class="metric"><span>Profile</span><strong class="metric-small">Analytical</strong></div></div>
<div class="chart-grid">
<div class="chart-card" data-chart-card="" data-chart-name="design-performance-pareto-frontier">
<div class="chart-head"><div class="chart-title">Raw goodput vs p95 TTFT</div><div class="chart-actions"><button class="chart-download" type="button">Download PNG</button><button class="chart-expand" type="button">Expand</button></div></div>
<div class="chart-body large"><canvas aria-label="Raw goodput vs p95 TTFT" id="designChart" role="img"></canvas></div>
</div>
<div class="chart-card" data-chart-card="" data-chart-name="design-efficiency-pareto-frontier">
<div class="chart-head"><div class="chart-title">Goodput per accelerator vs p95 TTFT</div><div class="chart-actions"><button class="chart-download" type="button">Download PNG</button><button class="chart-expand" type="button">Expand</button></div></div>
<div class="chart-body large"><canvas aria-label="Goodput per accelerator vs p95 TTFT" id="designEfficiencyChart" role="img"></canvas></div>
</div>
</div>
<div class="table-toolbar"><span>Candidate configurations</span><div><button class="mini-button" disabled="" id="designCopyBtn" type="button">Copy table</button><button class="mini-button" disabled="" id="designCsvBtn" type="button">Download CSV</button></div></div>
<div class="table-wrap"><table><thead><tr><th>Candidate</th><th>Perf Pareto</th><th>Efficiency Pareto</th><th>SLO pass</th><th>GPU instances</th><th>Goodput</th><th>Goodput / GPU</th><th>p95 TTFT</th><th>p95 E2E</th><th>KV peak</th><th>Diagnosis</th></tr></thead><tbody id="designRows"></tbody></table></div>
</div>
</div>
</section>
<section aria-hidden="true" aria-labelledby="tab-research" class="tab-panel" id="research" role="tabpanel">
<div class="research-grid">
<aside class="panel research-controls">
<div class="panel-title-row"><h2>Experiment settings</h2><span class="tag">Paired seeds</span></div>
<p class="muted">Test one systems hypothesis repeatedly on matched workloads, then stress the conclusion against analytical-profile uncertainty.</p>
<label>Hypothesis<select id="studyPreset">
<option value="prefix_cache">Prefix reuse: off vs on</option>
<option value="pd_vs_colocated">Topology: colocated vs P/D</option>
<option value="chunked_vs_fcfs">Scheduling: FCFS vs chunked prefill</option>
<option value="slo_vs_fcfs">Scheduling: FCFS vs least-slack</option>
</select></label>
<p class="control-help">The selected hypothesis controls both the paired study and the model-uncertainty stress test below.</p>
<hr/>
<div class="section-kicker">Paired Monte Carlo</div>
<div class="field-grid two">
<label>Repetitions<input id="studyReps" max="64" min="2" type="number" value="12"/></label>
<label>Bootstrap samples<input id="studyBootstrap" max="5000" min="50" step="50" type="number" value="500"/></label>
</div>
<button class="primary" disabled="" id="pairedStudyBtn" type="button">Run paired study</button>
<hr/>
<div class="section-kicker">Model uncertainty stress test</div>
<p class="control-help">Uses the Hypothesis selected above. For P/D robustness, choose Topology: colocated vs P/D, set the perturbation controls here, then click the button below.</p>
<div class="field-grid two">
<label>Samples<input id="robustSamples" max="96" min="4" type="number" value="32"/></label>
<label>Latency uncertainty<input id="robustUncertainty" max="0.75" min="0" step="0.05" type="number" value="0.20"/><span class="unit">fraction</span></label>
</div>
<button class="primary" disabled="" id="robustStudyBtn" type="button">Stress-test selected hypothesis</button>
</aside>
<div class="research-results">
<section class="panel research-panel">
<div class="panel-title-row"><div><h2>Paired A/B result</h2></div><span aria-live="polite" class="tag neutral" id="pairedState">Waiting</span></div>
<div class="empty-state small" id="pairedEmpty"><h3>No paired result yet</h3><p>Baseline and treatment use the same random seed on every repetition. The reported confidence interval is over paired deltas, not unrelated runs.</p></div>
<div class="hidden" id="pairedContent">
<div class="study-summary" id="pairedSummary"></div>
<div class="chart-card full" data-chart-card="" data-chart-name="paired-study-relative-effects">
<div class="chart-head"><div class="chart-title">Treatment effect by metric</div><div class="chart-actions"><button class="chart-download" type="button">Download PNG</button><button class="chart-expand" type="button">Expand</button></div></div>
<div class="chart-body research-chart"><canvas aria-label="Treatment effect by metric" id="pairedChart" role="img"></canvas></div>
</div>
<div class="table-toolbar"><span>Paired effect estimates</span><div><button class="mini-button" disabled="" id="pairedCopyBtn" type="button">Copy table</button><button class="mini-button" disabled="" id="pairedCsvBtn" type="button">Download CSV</button><button class="mini-button" disabled="" id="pairedJsonBtn" type="button">Download JSON</button></div></div>
<div class="table-wrap"><table><thead><tr><th>Metric</th><th>Baseline mean</th><th>Treatment mean</th><th>Mean delta</th><th>95% bootstrap CI</th><th>Relative change</th><th>Treatment win rate</th><th>CI excludes zero</th></tr></thead><tbody id="pairedRows"></tbody></table></div>
</div>
</section>
<section class="panel research-panel">
<div class="panel-title-row"><div><h2>Sensitivity result</h2></div><span aria-live="polite" class="tag neutral" id="robustState">Waiting</span></div>
<div class="empty-state small" id="robustEmpty"><h3>No sensitivity result yet</h3><p>Shared multiplicative perturbations are applied to prefill, decode, and transfer latency proxies. This tests sensitivity to profile error; it is not a probability distribution over real hardware.</p></div>
<div class="hidden" id="robustContent">
<div class="metric-grid four">
<div class="metric"><span>TTFT wins</span><strong id="rTtftWins">N/A</strong></div>
<div class="metric"><span>Goodput wins</span><strong id="rGoodputWins">N/A</strong></div>
<div class="metric"><span>Treatment SLO pass</span><strong id="rTreatmentPass">N/A</strong></div>
<div class="metric"><span>Baseline SLO pass</span><strong id="rBaselinePass">N/A</strong></div>
</div>
<div class="study-summary" id="robustSummary"></div>
<div class="table-toolbar"><span>Perturbation summary</span><div><button class="mini-button" disabled="" id="robustCopyBtn" type="button">Copy table</button><button class="mini-button" disabled="" id="robustCsvBtn" type="button">Download CSV</button><button class="mini-button" disabled="" id="robustJsonBtn" type="button">Download JSON</button></div></div>
<div class="table-wrap"><table><thead><tr><th>Sample</th><th>Prefill scale</th><th>Decode scale</th><th>Transfer scale</th><th>Baseline SLO</th><th>Treatment SLO</th><th>Goodput delta</th><th>TTFT delta</th><th>E2E delta</th></tr></thead><tbody id="robustRows"></tbody></table></div>
</div>
</section>
</div>
</div>
</section>
<section aria-hidden="true" aria-labelledby="tab-agent" class="tab-panel" id="agent" role="tabpanel">
<div class="agent-layout">
<aside class="panel agent-controls">
<div class="panel-title-row"><div><h2>Session configuration</h2></div><span class="tag">Common program trace</span></div>
<p class="muted">Model multi-turn programs separated by tool calls. Cross-turn KV can be evicted, retained, or held under a TTL, while routing can prioritize load balance or session locality.</p>
<div class="field-grid two">
<label>Model<select id="agentModel"><option>Qwen2.5-3B</option><option>Llama-3.1-8B</option><option>Mistral-7B-v0.3</option></select></label>
<label>Accelerator<select id="agentAccelerator"><option value="L4">NVIDIA L4</option><option value="A10G">NVIDIA A10G</option><option value="A100-40GB">NVIDIA A100 40GB</option></select></label>
<label>Weight precision<select id="agentQuantization"><option value="fp16">FP16</option><option selected="" value="int8">INT8 scenario</option><option value="int4">INT4 scenario</option></select></label>
<label>Replicas<input id="agentReplicas" max="8" min="1" type="number" value="2"/></label>
</div>
<hr/>
<div class="section-kicker">Program arrivals</div>
<div class="field-grid two">
<label>Session rate<input id="agentRate" min="0.01" step="0.05" type="number" value="0.20"/><span class="unit">sessions/s</span></label>
<label>Arrival horizon<input id="agentDuration" min="5" step="5" type="number" value="30"/><span class="unit">s</span></label>
<label>Mean turns<input id="agentTurns" min="2" step="0.5" type="number" value="4"/></label>
<label>Turns CV<input id="agentTurnsCv" max="1.5" min="0" step="0.05" type="number" value="0.25"/></label>
<label>Initial prompt<input id="agentInitialPrompt" min="32" step="32" type="number" value="640"/><span class="unit">tokens</span></label>
<label>Per-turn append<input id="agentAppend" min="1" step="16" type="number" value="180"/><span class="unit">tokens</span></label>
<label>Mean output<input id="agentOutput" min="1" step="8" type="number" value="72"/><span class="unit">tokens</span></label>
<label>Token CV<input id="agentTokenCv" max="1.5" min="0" step="0.05" type="number" value="0.35"/></label>
<label>Tool gap mean<input id="agentToolGap" min="0" step="0.25" type="number" value="1.5"/><span class="unit">s</span></label>
<label>Tool gap CV<input id="agentToolGapCv" max="2" min="0" step="0.05" type="number" value="0.75"/></label>
<label>Seed<input id="agentSeed" step="1" type="number" value="7"/></label>
</div>
<hr/>
<div class="section-kicker">State policy</div>
<div class="field-grid two">
<label>KV retention<select id="agentRetention"><option value="evict">Evict after every turn</option><option selected="" value="ttl">TTL retention</option><option value="retain">Retain until session ends</option><option value="offload">Offload idle KV to host</option><option value="gap_aware">Oracle gap-aware HBM / host tiering</option><option value="adaptive">Adaptive tool-gap predictor</option></select></label>
<label>Routing<select id="agentRouting"><option value="least_load">Least-load</option><option selected="" value="session_affinity">Strict session affinity</option><option value="bounded_affinity">Bounded affinity</option></select></label>
<label id="agentTtlLabel">KV / adaptive safety TTL<input id="agentTtl" min="0" step="0.25" type="number" value="3"/><span class="unit">s</span></label>
<label>KV memory fraction<input id="agentKvFraction" max="0.98" min="0.01" step="0.05" type="number" value="0.85"/><span class="unit">fraction</span></label>
<label id="agentAffinitySlackLabel">Affinity slack<input id="agentAffinitySlack" min="0" step="25" type="number" value="150"/><span class="unit">ms</span></label>
<label id="agentGapThresholdLabel">Gap-aware HBM threshold<input id="agentGapThreshold" min="0" step="0.25" type="number" value="1.5"/><span class="unit">s</span></label>
<label id="agentHostMemoryLabel">Host KV capacity<input id="agentHostMemory" min="0" step="4" type="number" value="32"/><span class="unit">GB</span></label>
<label id="agentHostBandwidthLabel">Host transfer bandwidth<input id="agentHostBandwidth" min="0.1" step="1" type="number" value="32"/><span class="unit">GB/s</span></label>
<label id="agentHostBaseLabel">Transfer base latency<input id="agentHostBase" min="0" step="0.05" type="number" value="0.15"/><span class="unit">ms</span></label>
<label id="agentPredictorScopeLabel">Adaptive predictor<select id="agentPredictorScope"><option selected="" value="per_tool_ema">Per-tool EWMA</option><option value="global_ema">Global EWMA</option></select></label>
<label id="agentPredictorAlphaLabel">EWMA alpha<input id="agentPredictorAlpha" max="1" min="0.01" step="0.05" type="number" value="0.30"/><span class="unit">fraction</span></label>
<label id="agentPredictorMinLabel">Warm-up observations<input id="agentPredictorMin" max="20" min="1" step="1" type="number" value="2"/></label>
</div>
<div class="field-grid two">
<label>Turn TTFT SLO<input id="agentTtftSlo" min="1" type="number" value="500"/><span class="unit">ms</span></label>
<label>Session E2E SLO<input id="agentSessionSlo" min="1000" step="1000" type="number" value="30000"/><span class="unit">ms</span></label>
</div>
<button class="primary" disabled="" id="agentRunBtn" type="button">Run stateful session simulation</button>
</aside>
<div class="agent-results">
<section class="panel research-panel">
<div class="panel-title-row"><div><h2>Session run</h2></div><span aria-live="polite" class="tag neutral" id="agentRunState">Waiting</span></div>
<div class="empty-state small" id="agentRunEmpty"><h3>No session run yet</h3><p>Each session has ordered LLM turns separated by sampled tool gaps. A replica processes one turn at a time in this mode so cache/routing effects remain isolated from the batching experiments in Serving Lab.</p></div>
<div class="hidden" id="agentRunContent">
<div class="metric-grid eight-agent">
<div class="metric"><span>p95 turn TTFT</span><strong id="agentTtft">N/A</strong></div>
<div class="metric"><span>p95 session E2E</span><strong id="agentSessionE2e">N/A</strong></div>
<div class="metric"><span>Session SLO</span><strong id="agentSessionSloValue">N/A</strong></div>
<div class="metric"><span>Cross-turn reuse</span><strong id="agentCacheHit">N/A</strong></div>
<div class="metric"><span>Recomputed history</span><strong id="agentRecompute">N/A</strong></div>
<div class="metric"><span>Peak HBM KV</span><strong id="agentPeakKv">N/A</strong></div>
<div class="metric"><span>Mean host KV</span><strong id="agentHostMeanKv">N/A</strong></div>
<div class="metric"><span>p95 host transfer</span><strong id="agentHostTransfer">N/A</strong></div>
</div>
<div class="study-summary" id="agentRunSummary"></div>
<div class="table-toolbar"><span>Run artifact</span><div><button class="mini-button" disabled="" id="agentRunCopyJson" type="button">Copy JSON</button><button class="mini-button" disabled="" id="agentRunJson" type="button">Download JSON</button></div></div>
<div class="chart-grid">
<div class="chart-card" data-chart-card="" data-chart-name="agent-session-kv-residency-timeline">
<div class="chart-head"><div class="chart-title">KV residency and replica activity</div><div class="chart-actions"><button class="chart-download" type="button">Download PNG</button><button class="chart-expand" type="button">Expand</button></div></div>
<div class="chart-body large"><canvas aria-label="KV residency and replica activity" id="agentTimelineChart" role="img"></canvas></div>
</div>
<div class="chart-card" data-chart-card="" data-chart-name="agent-turn-ttft-by-turn-index">
<div class="chart-head"><div class="chart-title">Turn TTFT across session depth</div><div class="chart-actions"><button class="chart-download" type="button">Download PNG</button><button class="chart-expand" type="button">Expand</button></div></div>
<div class="chart-body large"><canvas aria-label="Turn TTFT across session depth" id="agentTurnChart" role="img"></canvas></div>
</div>
</div>
</div>
</section>
<section class="panel research-panel">
<div class="panel-title-row"><div><h2>Session policy comparison</h2><p class="muted">All candidates replay the identical generated program trace.</p></div><button class="primary compact" disabled="" id="agentCompareBtn" type="button">Compare 4 policies</button></div>
<div class="empty-state small" id="agentCompareEmpty"><h3>No policy comparison yet</h3><p>Compare stateless load balancing, retention without affinity, TTL with affinity, and full retention with affinity.</p></div>
<div class="hidden" id="agentCompareContent">
<div class="chart-card full" data-chart-card="" data-chart-name="agent-policy-latency-memory-tradeoff">
<div class="chart-head"><div class="chart-title">p95 turn TTFT vs mean resident KV</div><div class="chart-actions"><button class="chart-download" type="button">Download PNG</button><button class="chart-expand" type="button">Expand</button></div></div>
<div class="chart-body research-chart"><canvas aria-label="p95 turn TTFT vs mean resident KV" id="agentCompareChart" role="img"></canvas></div>
</div>
<div class="table-toolbar"><span>Common-trace policy results</span><div><button class="mini-button" disabled="" id="agentCompareCopy" type="button">Copy table</button><button class="mini-button" disabled="" id="agentCompareCsv" type="button">Download CSV</button></div></div>
<div class="table-wrap"><table><thead><tr><th>Policy</th><th>p95 turn TTFT</th><th>p95 session E2E</th><th>Cache hit</th><th>Route locality</th><th>Recomputed history</th><th>Mean KV</th><th>Peak KV</th><th>HBM GB-s</th><th>Evictions</th></tr></thead><tbody id="agentCompareRows"></tbody></table></div>
</div>
</section>
<section class="panel research-panel">
<div class="panel-title-row"><div><h2>TTL frontier</h2><p class="muted">Sweep the KV retention horizon on the same program trace and expose the latency vs memory-residency frontier.</p></div><button class="primary compact" disabled="" id="agentTtlBtn" type="button">Run TTL sweep</button></div>
<div class="empty-state small" id="agentTtlEmpty"><h3>No TTL sweep yet</h3><p>The sweep varies retention from immediate eviction through long-lived state while keeping session arrivals, turns, tool gaps, and token lengths fixed.</p></div>
<div class="hidden" id="agentTtlContent">
<div class="chart-card full" data-chart-card="" data-chart-name="agent-ttl-retention-frontier">
<div class="chart-head"><div class="chart-title">Mean resident KV vs p95 turn TTFT</div><div class="chart-actions"><button class="chart-download" type="button">Download PNG</button><button class="chart-expand" type="button">Expand</button></div></div>
<div class="chart-body research-chart"><canvas aria-label="Mean resident KV vs p95 turn TTFT" id="agentTtlChart" role="img"></canvas></div>
</div>
<div class="table-toolbar"><span>TTL sweep</span><div><button class="mini-button" disabled="" id="agentTtlCopy" type="button">Copy table</button><button class="mini-button" disabled="" id="agentTtlCsv" type="button">Download CSV</button></div></div>
<div class="table-wrap"><table><thead><tr><th>TTL</th><th>Pareto</th><th>p95 turn TTFT</th><th>p95 session E2E</th><th>Cache hit</th><th>Recomputed history</th><th>Mean KV</th><th>HBM GB-s</th><th>TTL evictions</th><th>Pressure evictions</th></tr></thead><tbody id="agentTtlRows"></tbody></table></div>
</div>
</section>
<section class="panel research-panel">
<div class="panel-title-row"><div><h2>Memory policy comparison</h2><p class="muted">Replay one program trace across stateless, TTL, bounded-affinity, host-offload, and gap-aware policies, then deliberately shrink the per-replica HBM KV budget to expose eviction pressure.</p></div><div class="stacked-actions"><button class="primary compact" disabled="" id="agentMemoryCompareBtn" type="button">Compare memory policies</button><button class="primary compact" disabled="" id="agentBudgetBtn" type="button">Stress HBM budget</button></div></div>
<div class="empty-state small" id="agentMemoryEmpty"><h3>No memory comparison yet</h3><p>Host offload preserves reusable KV outside HBM and pays an explicit transfer cost on restore. Gap-aware tiering is reported as an oracle upper bound because it uses the realized simulated tool gap.</p></div>
<div class="hidden" id="agentMemoryContent">
<div class="hidden" id="agentMemoryCompareBlock">
<div class="chart-card full" data-chart-card="" data-chart-name="agent-memory-policy-tradeoff">
<div class="chart-head"><div class="chart-title">Mean HBM residency vs p95 turn TTFT</div><div class="chart-actions"><button class="chart-download" type="button">Download PNG</button><button class="chart-expand" type="button">Expand</button></div></div>
<div class="chart-body research-chart"><canvas aria-label="Mean HBM residency vs p95 turn TTFT" id="agentMemoryChart" role="img"></canvas></div>
</div>
<div class="table-toolbar"><span>Memory-policy comparison</span><div><button class="mini-button" disabled="" id="agentMemoryCopy" type="button">Copy table</button><button class="mini-button" disabled="" id="agentMemoryCsv" type="button">Download CSV</button></div></div>
<div class="table-wrap"><table><thead><tr><th>Policy</th><th>p95 turn TTFT</th><th>p95 session E2E</th><th>Session SLO</th><th>Reuse</th><th>HBM hit</th><th>Host hit</th><th>Mean HBM</th><th>Mean host</th><th>Transfer p95</th><th>Recomputed</th></tr></thead><tbody id="agentMemoryRows"></tbody></table></div>
</div>
<div class="hidden experiment-subsection" id="agentBudgetBlock">
<div class="chart-card full" data-chart-card="" data-chart-name="agent-hbm-budget-stress">
<div class="chart-head"><div class="chart-title">Finite HBM budget stress</div><div class="chart-actions"><button class="chart-download" type="button">Download PNG</button><button class="chart-expand" type="button">Expand</button></div></div>
<div class="chart-body research-chart"><canvas aria-label="Finite HBM budget stress" id="agentBudgetChart" role="img"></canvas></div>
</div>
<div class="table-toolbar"><span id="agentBudgetCaption">HBM budget sweep</span><div><button class="mini-button" disabled="" id="agentBudgetCopy" type="button">Copy table</button><button class="mini-button" disabled="" id="agentBudgetCsv" type="button">Download CSV</button></div></div>
<div class="table-wrap"><table><thead><tr><th>Policy</th><th>Budget / replica</th><th>Reference x</th><th>p95 TTFT</th><th>Session SLO</th><th>Reuse</th><th>Mean HBM</th><th>Mean host</th><th>Pressure evictions</th><th>Failed turns</th></tr></thead><tbody id="agentBudgetRows"></tbody></table></div>
</div>
</div>
</section>
<section class="panel research-panel">
<div class="panel-title-row"><div><h2>Affinity trade-off</h2><p class="muted">Bounded affinity keeps a session on its cached replica only while the estimated queue penalty stays within a configurable slack. Sweep that slack on one common program trace.</p></div><button class="primary compact" disabled="" id="agentAffinityBtn" type="button">Run affinity sweep</button></div>
<div class="empty-state small" id="agentAffinityEmpty"><h3>No routing sweep yet</h3><p>Low slack behaves closer to least-load routing; high slack increasingly prioritizes KV locality. This exposes when cache reuse begins to overload a hot replica.</p></div>
<div class="hidden" id="agentAffinityContent">
<div class="chart-card full" data-chart-card="" data-chart-name="agent-affinity-routing-frontier">
<div class="chart-head"><div class="chart-title">Routing slack vs latency and locality</div><div class="chart-actions"><button class="chart-download" type="button">Download PNG</button><button class="chart-expand" type="button">Expand</button></div></div>
<div class="chart-body research-chart"><canvas aria-label="Routing slack vs latency and locality" id="agentAffinityChart" role="img"></canvas></div>
</div>
<div class="table-toolbar"><span>Bounded-affinity sweep</span><div><button class="mini-button" disabled="" id="agentAffinityCopy" type="button">Copy table</button><button class="mini-button" disabled="" id="agentAffinityCsv" type="button">Download CSV</button></div></div>
<div class="table-wrap"><table><thead><tr><th>Affinity slack</th><th>p95 turn TTFT</th><th>p95 session E2E</th><th>Route locality</th><th>Cache hit</th><th>Recomputed</th><th>Session SLO</th><th>Pressure evictions</th></tr></thead><tbody id="agentAffinityRows"></tbody></table></div>
</div>
</section>
<section class="panel research-panel">
<div class="panel-title-row"><div><h2>Predictive tiering</h2><p class="muted">Replace the oracle tool-gap policy with online estimates learned only from completed tool calls. The study changes slow-tool latency partway through one common program trace, then measures whether global and per-tool predictors adapt without future-duration leakage.</p></div></div>
<div class="predictive-controls field-grid four">
<label>Study horizon<input id="predictiveHorizon" max="600" min="30" step="30" type="number" value="120"/><span class="unit">sim s</span></label>
<label>Shift point<input id="predictiveShift" max="0.9" min="0.1" step="0.05" type="number" value="0.55"/><span class="unit">fraction</span></label>
<label>Slow-tool multiplier<input id="predictiveMultiplier" max="6" min="1" step="0.25" type="number" value="2.5"/><span class="unit">x</span></label>
<label>Adaptive alpha<input id="predictiveAlpha" max="1" min="0.01" step="0.05" type="number" value="0.30"/><span class="unit">fraction</span></label>
</div>
<div class="stacked-actions horizontal-actions"><button class="primary compact" disabled="" id="predictiveCompareBtn" type="button">Compare predictive policies</button><button class="primary compact" disabled="" id="predictiveAlphaBtn" type="button">Sweep adaptation rate</button></div>
<div class="empty-state small" id="predictiveEmpty"><h3>No predictive comparison yet</h3><p>The adaptive policies use exponentially weighted tool-duration history. The oracle candidate is retained only as an upper bound; it can see the realized future gap while adaptive candidates cannot.</p></div>
<div class="hidden" id="predictiveContent">
<div class="hidden" id="predictiveCompareBlock">
<div class="metric-grid four">
<div class="metric"><span>Best non-oracle p95 TTFT</span><strong id="predictiveBestTtft">N/A</strong></div>
<div class="metric"><span>Per-tool oracle agreement</span><strong id="predictiveAgreement">N/A</strong></div>
<div class="metric"><span>Post-shift MAE</span><strong id="predictivePostMae">N/A</strong></div>
<div class="metric"><span>Shift point</span><strong id="predictiveShiftValue">N/A</strong></div>
</div>
<div class="chart-grid">
<div class="chart-card" data-chart-card="" data-chart-name="agent-predictive-policy-tradeoff">
<div class="chart-head"><div class="chart-title">Mean HBM residency vs p95 turn TTFT</div><div class="chart-actions"><button class="chart-download" type="button">Download PNG</button><button class="chart-expand" type="button">Expand</button></div></div>
<div class="chart-body research-chart"><canvas aria-label="Mean HBM residency vs p95 turn TTFT" id="predictivePolicyChart" role="img"></canvas></div>
</div>
<div class="chart-card" data-chart-card="" data-chart-name="agent-tool-gap-learning-curve">
<div class="chart-head"><div class="chart-title">Online tool-gap prediction error</div><div class="chart-actions"><button class="chart-download" type="button">Download PNG</button><button class="chart-expand" type="button">Expand</button></div></div>
<div class="chart-body research-chart"><canvas aria-label="Online tool-gap prediction error" id="predictiveLearningChart" role="img"></canvas></div>
</div>
</div>
<div class="table-toolbar"><span>Common shifted-trace policy results</span><div><button class="mini-button" disabled="" id="predictiveCopy" type="button">Copy table</button><button class="mini-button" disabled="" id="predictiveCsv" type="button">Download CSV</button><button class="mini-button" disabled="" id="predictiveJson" type="button">Download JSON</button></div></div>
<div class="table-wrap"><table><thead><tr><th>Policy</th><th>p95 turn TTFT</th><th>p95 session E2E</th><th>Session SLO</th><th>Reuse</th><th>Mean HBM</th><th>Mean host</th><th>Prediction MAE</th><th>Oracle agreement</th><th>Post-shift MAE</th><th>Recomputed</th></tr></thead><tbody id="predictiveRows"></tbody></table></div>
</div>
<div class="hidden experiment-subsection" id="predictiveAlphaBlock">
<div class="chart-card full" data-chart-card="" data-chart-name="agent-adaptation-rate-sweep">
<div class="chart-head"><div class="chart-title">EWMA alpha vs prediction error before and after shift</div><div class="chart-actions"><button class="chart-download" type="button">Download PNG</button><button class="chart-expand" type="button">Expand</button></div></div>
<div class="chart-body research-chart"><canvas aria-label="EWMA alpha vs prediction error before and after shift" id="predictiveAlphaChart" role="img"></canvas></div>
</div>
<div class="table-toolbar"><span id="predictiveAlphaCaption">Adaptation-rate sweep</span><div><button class="mini-button" disabled="" id="predictiveAlphaCopy" type="button">Copy table</button><button class="mini-button" disabled="" id="predictiveAlphaCsv" type="button">Download CSV</button><button class="mini-button" disabled="" id="predictiveAlphaJson" type="button">Download JSON</button></div></div>
<div class="table-wrap"><table><thead><tr><th>Alpha</th><th>Pre-shift MAE</th><th>Post-shift MAE</th><th>Oracle agreement</th><th>p95 TTFT</th><th>p95 session E2E</th><th>Session SLO</th><th>Mean HBM</th><th>Mean host</th></tr></thead><tbody id="predictiveAlphaRows"></tbody></table></div>
</div>
</div>
</section>
</div>
</div>
</section>
<section aria-hidden="true" aria-labelledby="tab-execution" class="tab-panel" id="execution" role="tabpanel">
<div class="agent-layout">
<aside class="panel agent-controls">
<div class="panel-title-row"><div><h2>Execution configuration</h2></div><span class="tag">No lookahead</span></div>
<p class="muted">Learn first-order agent execution transitions online and use them to prefetch reusable static agent prefixes during tool gaps. A deliberate mid-trace workflow shift tests whether forgetting helps adaptation.</p>
<div class="field-grid two">
<label>Model<select id="execModel"><option>Qwen2.5-3B</option><option>Llama-3.1-8B</option><option>Mistral-7B-v0.3</option></select></label>
<label>Accelerator<select id="execAccelerator"><option value="L4">NVIDIA L4</option><option value="A10G">NVIDIA A10G</option><option value="A100-40GB">NVIDIA A100 40GB</option></select></label>
<label>Weight precision<select id="execQuantization"><option value="fp16">FP16</option><option selected="" value="int8">INT8 scenario</option><option value="int4">INT4 scenario</option></select></label>
<label>Seed<input id="execSeed" step="1" type="number" value="7"/></label>
</div>
<hr/>
<div class="section-kicker">Workflow stream</div>
<div class="field-grid two">
<label>Workflow rate<input id="execRate" min="0.01" step="0.02" type="number" value="0.18"/><span class="unit">flows/s</span></label>
<label>Study horizon<input id="execDuration" min="30" step="10" type="number" value="120"/><span class="unit">s</span></label>
<label>Max agent steps<input id="execMaxSteps" max="16" min="2" step="1" type="number" value="8"/></label>
<label>Regime shift<input id="execShift" max="0.95" min="0.05" step="0.05" type="number" value="0.55"/><span class="unit">fraction</span></label>
<label>Dynamic prompt<input id="execPrompt" min="32" step="32" type="number" value="320"/><span class="unit">tokens</span></label>
<label>Mean output<input id="execOutput" min="1" step="8" type="number" value="64"/><span class="unit">tokens</span></label>
<label>Tool gap mean<input id="execGap" min="0" step="0.1" type="number" value="1.2"/><span class="unit">s</span></label>
<label>Prefix cache budget<input id="execCacheBudget" max="1.5" min="0.1" step="0.05" type="number" value="0.62"/><span class="unit">working-set x</span></label>
</div>
<hr/>
<div class="section-kicker">Prediction and prefetch</div>
<div class="field-grid two">
<label>Prefetch policy<select id="execPolicy"><option value="none">No prefetch</option><option value="cumulative">Cumulative transitions</option><option selected="" value="decayed">Decayed top-1</option><option value="multistep">Multi-step top-k</option><option value="utility">Utility-aware multi-step</option><option value="oracle">Oracle next-role upper bound</option><option value="oracle_horizon">Oracle future-set upper bound</option></select></label>
<label>Transition decay<input id="execDecay" max="1" min="0.2" step="0.05" type="number" value="0.85"/><span class="unit">fraction</span></label>
<label>Confidence threshold<input id="execThreshold" max="1" min="0" step="0.05" type="number" value="0.55"/><span class="unit">top-1</span></label>
<label>Host bandwidth<input id="execBandwidth" min="0.1" step="1" type="number" value="32"/><span class="unit">GB/s</span></label>
<label id="execHorizonLabel">Forecast horizon<input id="execHorizon" max="6" min="1" step="1" type="number" value="3"/><span class="unit">steps</span></label>
<label id="execTopKLabel">Prefetch top-k<input id="execTopK" max="5" min="1" step="1" type="number" value="2"/><span class="unit">roles</span></label>
<label id="execDiscountLabel">Forecast discount<input id="execDiscount" max="1" min="0.05" step="0.05" type="number" value="0.75"/><span class="unit">fraction</span></label>
<label id="execMinScoreLabel">Min forecast score<input id="execMinScore" max="1" min="0" step="0.05" type="number" value="0.10"/><span class="unit">normalized</span></label>
<label id="execUtilityLabel">Utility threshold<input id="execUtility" step="1" type="number" value="0"/><span class="unit">ms</span></label>
</div>
<p class="muted tiny-note">Multi-step policies roll the learned transition matrix forward without reading the future trace. Utility-aware planning estimates saved prefill time minus host-transfer and forecast-weighted eviction cost.</p>
<button class="primary" disabled="" id="execRunBtn" type="button">Run execution-learning simulation</button>
<div class="button-row"><button class="secondary" disabled="" id="execRunCopyJson" type="button">Copy JSON</button><button class="secondary" disabled="" id="execRunJson" type="button">Download JSON</button></div>
</aside>
<div class="agent-results">
<section class="panel research-panel">
<div class="panel-title-row"><div><h2>Execution-learning run</h2></div><span aria-live="polite" class="tag neutral" id="execRunState">Waiting</span></div>
<div class="empty-state small" id="execRunEmpty"><h3>No execution run yet</h3><p>Agent-role transitions are revealed only when the next workflow step becomes ready. Learned policies may prefetch likely static prefixes from the modeled host tier before that happens.</p></div>
<div class="hidden" id="execRunContent">
<div class="metric-grid eight-agent">
<div class="metric"><span>p95 step TTFT</span><strong id="execTtft">N/A</strong></div>
<div class="metric"><span>p95 workflow E2E</span><strong id="execE2e">N/A</strong></div>
<div class="metric"><span>Transition top-1</span><strong id="execAccuracy">N/A</strong></div>
<div class="metric"><span>Post-shift top-1</span><strong id="execPostAccuracy">N/A</strong></div>
<div class="metric"><span>Prefix hit</span><strong id="execPrefixHit">N/A</strong></div>
<div class="metric"><span>Prefetch precision</span><strong id="execPrecision">N/A</strong></div>
<div class="metric"><span>Prefetch coverage</span><strong id="execCoverage">N/A</strong></div>
<div class="metric"><span>Wrong-step prefetch</span><strong id="execWaste">N/A</strong></div>
<div class="metric"><span>Calibration ECE</span><strong id="execEce">N/A</strong></div>
<div class="metric"><span>Future-role recall@K</span><strong id="execForecastRecall">N/A</strong></div>
<div class="metric"><span>Prefetch utilization</span><strong id="execUtilization">N/A</strong></div>
<div class="metric"><span>Unused prefetch</span><strong id="execUnused">N/A</strong></div>
</div>
<div class="study-summary" id="execRunSummary"></div>
<div class="chart-grid">
<div class="chart-card" data-chart-card="" data-chart-name="execution-prefix-cache-timeline">
<div class="chart-head"><div class="chart-title">Prefix-cache occupancy and queued steps</div><div class="chart-actions"><button class="chart-download" type="button">Download PNG</button><button class="chart-expand" type="button">Expand</button></div></div>
<div class="chart-body research-chart"><canvas aria-label="Prefix-cache occupancy and queued steps" id="execTimelineChart" role="img"></canvas></div>
</div>
<div class="chart-card" data-chart-card="" data-chart-name="execution-transition-learning-curve">
<div class="chart-head"><div class="chart-title">Rolling next-role prediction accuracy</div><div class="chart-actions"><button class="chart-download" type="button">Download PNG</button><button class="chart-expand" type="button">Expand</button></div></div>
<div class="chart-body research-chart"><canvas aria-label="Rolling next-role prediction accuracy" id="execLearningChart" role="img"></canvas></div>
</div>
</div>
<div class="chart-card full" data-chart-card="" data-chart-name="execution-transition-calibration">
<div class="chart-head"><div class="chart-title">Transition confidence reliability</div><div class="chart-actions"><button class="chart-download" type="button">Download PNG</button><button class="chart-expand" type="button">Expand</button></div></div>
<div class="chart-body research-chart"><canvas aria-label="Transition confidence reliability" id="execCalibrationChart" role="img"></canvas></div>
</div>
</div>
</section>
<section class="panel research-panel">
<div class="panel-title-row"><div><h2>Prefetch policy comparison</h2><p class="muted">Compare no prefetch, cumulative transition counts, exponentially decayed counts, and a clairvoyant next-role upper bound on exactly the same workflows.</p></div><button class="primary compact" disabled="" id="execCompareBtn" type="button">Compare 4 policies</button></div>
<div class="empty-state small" id="execCompareEmpty"><h3>No policy comparison yet</h3><p>The study separates transition-model quality from serving outcomes such as TTFT, cache hits, HBM residency, and wrong-step prefetch traffic.</p></div>
<div class="hidden" id="execCompareContent">
<div class="chart-card full" data-chart-card="" data-chart-name="execution-accuracy-vs-ttft">
<div class="chart-head"><div class="chart-title">Post-shift transition accuracy vs p95 TTFT</div><div class="chart-actions"><button class="chart-download" type="button">Download PNG</button><button class="chart-expand" type="button">Expand</button></div></div>
<div class="chart-body research-chart"><canvas aria-label="Post-shift transition accuracy vs p95 TTFT" id="execCompareChart" role="img"></canvas></div>
</div>
<div class="table-toolbar"><span>Common-trace policy results</span><div><button class="mini-button" disabled="" id="execCompareCopy" type="button">Copy table</button><button class="mini-button" disabled="" id="execCompareCsv" type="button">Download CSV</button><button class="mini-button" disabled="" id="execCompareJson" type="button">Download JSON</button></div></div>
<div class="table-wrap"><table><thead><tr><th>Policy</th><th>p95 TTFT</th><th>p95 workflow E2E</th><th>Post-shift top-1</th><th>Prefix hit</th><th>Prefetch precision</th><th>Coverage</th><th>Mean HBM</th><th>Wrong prefetch</th><th>Saved prefill</th></tr></thead><tbody id="execCompareRows"></tbody></table></div>
</div>
</section>
<section class="panel research-panel">
<div class="panel-title-row"><div><h2>Confidence-threshold sweep</h2><p class="muted">Sweep when a learned next-role prediction is confident enough to justify prefetch. Lower thresholds raise coverage but can pollute a small prefix cache with wrong-step transfers.</p></div><button class="primary compact" disabled="" id="execThresholdBtn" type="button">Sweep confidence threshold</button></div>
<div class="empty-state small" id="execThresholdEmpty"><h3>No threshold sweep yet</h3><p>The same learned transition model is replayed at multiple decision thresholds on one common shifted workflow trace.</p></div>
<div class="hidden" id="execThresholdContent">
<div class="metric-grid four">
<div class="metric"><span>Coverage vs TTFT r</span><strong id="execCoverageCorr">N/A</strong></div>
<div class="metric"><span>Precision vs wrong-prefetch r</span><strong id="execPrecisionCorr">N/A</strong></div>
<div class="metric"><span>Lowest p95 TTFT</span><strong id="execThresholdBestTtft">N/A</strong></div>
<div class="metric"><span>Lowest wrong-prefetch</span><strong id="execThresholdBestWaste">N/A</strong></div>
</div>
<div class="chart-card full" data-chart-card="" data-chart-name="execution-prefetch-confidence-frontier">
<div class="chart-head"><div class="chart-title">Confidence threshold vs latency and wrong-step traffic</div><div class="chart-actions"><button class="chart-download" type="button">Download PNG</button><button class="chart-expand" type="button">Expand</button></div></div>
<div class="chart-body research-chart"><canvas aria-label="Confidence threshold vs latency and wrong-step traffic" id="execThresholdChart" role="img"></canvas></div>
</div>
<div class="table-toolbar"><span>Confidence-threshold sweep</span><div><button class="mini-button" disabled="" id="execThresholdCopy" type="button">Copy table</button><button class="mini-button" disabled="" id="execThresholdCsv" type="button">Download CSV</button><button class="mini-button" disabled="" id="execThresholdJson" type="button">Download JSON</button></div></div>
<div class="table-wrap"><table><thead><tr><th>Threshold</th><th>p95 TTFT</th><th>Post-shift top-1</th><th>Precision</th><th>Coverage</th><th>Prefix hit</th><th>Wrong prefetch</th><th>Mean HBM</th></tr></thead><tbody id="execThresholdRows"></tbody></table></div>
</div>
</section>
<section class="panel research-panel">
<div class="panel-title-row"><div><h2>Transition-decay sweep</h2><p class="muted">Sweep exponential forgetting in the transition matrix after the workflow changes. The goal is not merely the most accurate predictor, but the best downstream serving behavior under cache and transfer constraints.</p></div><button class="primary compact" disabled="" id="execDecayBtn" type="button">Sweep transition decay</button></div>
<div class="empty-state small" id="execDecayEmpty"><h3>No forgetting-rate sweep yet</h3><p>Decay = 1.0 is cumulative history. Smaller values forget old transitions faster and can react more quickly to the regime shift.</p></div>
<div class="hidden" id="execDecayContent">
<div class="metric-grid four">
<div class="metric"><span>Best post-shift accuracy decay</span><strong id="execBestAccuracyDecay">N/A</strong></div>
<div class="metric"><span>Best TTFT decay</span><strong id="execBestTtftDecay">N/A</strong></div>
<div class="metric"><span>Accuracy vs TTFT r</span><strong id="execAccuracyTtftCorr">N/A</strong></div>
<div class="metric"><span>Accuracy vs prefix-hit r</span><strong id="execAccuracyHitCorr">N/A</strong></div>
</div>
<div class="chart-card full" data-chart-card="" data-chart-name="execution-transition-forgetting-sweep">
<div class="chart-head"><div class="chart-title">Transition forgetting vs post-shift accuracy and p95 TTFT</div><div class="chart-actions"><button class="chart-download" type="button">Download PNG</button><button class="chart-expand" type="button">Expand</button></div></div>
<div class="chart-body research-chart"><canvas aria-label="Transition forgetting vs post-shift accuracy and p95 TTFT" id="execDecayChart" role="img"></canvas></div>
</div>
<div class="table-toolbar"><span>Forgetting-rate sweep</span><div><button class="mini-button" disabled="" id="execDecayCopy" type="button">Copy table</button><button class="mini-button" disabled="" id="execDecayCsv" type="button">Download CSV</button><button class="mini-button" disabled="" id="execDecayJson" type="button">Download JSON</button></div></div>
<div class="table-wrap"><table><thead><tr><th>Decay</th><th>Pre-shift top-1</th><th>Post-shift top-1</th><th>p95 TTFT</th><th>Prefix hit</th><th>Precision</th><th>Coverage</th><th>Wrong prefetch</th></tr></thead><tbody id="execDecayRows"></tbody></table></div>
</div>
</section>
<section class="panel research-panel">
<div class="panel-title-row"><div><h2>Prefetch planning comparison</h2><p class="muted">Compare one-step prediction with multi-step forecasting, a utility-aware planner, and a clairvoyant future-set information bound on one common shifted workflow trace.</p></div><button class="primary compact" disabled="" id="execPlanningBtn" type="button">Compare planning policies</button></div>
<div class="empty-state small" id="execPlanningEmpty"><h3>No planning study yet</h3><p>This experiment separates next-role accuracy from forecast-set recall, cache utilization, unused prefetch traffic, and actual serving latency.</p></div>
<div class="hidden" id="execPlanningContent">
<div class="metric-grid four">
<div class="metric"><span>Lowest p95 TTFT</span><strong id="execPlanningBestTtft">N/A</strong></div>
<div class="metric"><span>Best prefill / HBM</span><strong id="execPlanningBestEfficiency">N/A</strong></div>
<div class="metric"><span>Forecast horizon</span><strong id="execPlanningHorizon">N/A</strong></div>
<div class="metric"><span>Top-k</span><strong id="execPlanningTopK">N/A</strong></div>
</div>
<div class="chart-card full" data-chart-card="" data-chart-name="execution-planning-tradeoff">
<div class="chart-head"><div class="chart-title">Mean HBM residency vs p95 TTFT</div><div class="chart-actions"><button class="chart-download" type="button">Download PNG</button><button class="chart-expand" type="button">Expand</button></div></div>
<div class="chart-body research-chart"><canvas aria-label="Mean HBM residency vs p95 TTFT" id="execPlanningChart" role="img"></canvas></div>
</div>
<div class="table-toolbar"><span>Common-trace planning results</span><div><button class="mini-button" disabled="" id="execPlanningCopy" type="button">Copy table</button><button class="mini-button" disabled="" id="execPlanningCsv" type="button">Download CSV</button><button class="mini-button" disabled="" id="execPlanningJson" type="button">Download JSON</button></div></div>
<div class="table-wrap"><table><thead><tr><th>Policy</th><th>p95 TTFT</th><th>Future-role recall@K</th><th>Prefix hit</th><th>Utilization</th><th>Next-step precision</th><th>Mean HBM</th><th>Unused prefetch</th><th>Saved prefill</th><th>ECE</th></tr></thead><tbody id="execPlanningRows"></tbody></table></div>
</div>
</section>
<section class="panel research-panel">
<div class="panel-title-row"><div><h2>Forecast-horizon sweep</h2><p class="muted">Sweep a utility-aware planner from one-step prediction to longer rollout horizons. Longer forecasts can reveal reuse but can also waste bandwidth or evict nearer-term prefixes.</p></div><button class="primary compact" disabled="" id="execHorizonBtn" type="button">Sweep forecast horizon</button></div>
<div class="empty-state small" id="execHorizonEmpty"><h3>No horizon sweep yet</h3><p>Each horizon replays the exact same shifted workflow trace and uses the same cache budget and learned transition model.</p></div>
<div class="hidden" id="execHorizonContent">
<div class="metric-grid four">
<div class="metric"><span>Best TTFT horizon</span><strong id="execHorizonBestTtft">N/A</strong></div>
<div class="metric"><span>Best utilization horizon</span><strong id="execHorizonBestUtil">N/A</strong></div>
<div class="metric"><span>Configured top-k</span><strong id="execHorizonTopK">N/A</strong></div>
<div class="metric"><span>Cache budget</span><strong id="execHorizonBudget">N/A</strong></div>
</div>
<div class="chart-card full" data-chart-card="" data-chart-name="execution-forecast-horizon">
<div class="chart-head"><div class="chart-title">Forecast horizon vs latency and prefetch utilization</div><div class="chart-actions"><button class="chart-download" type="button">Download PNG</button><button class="chart-expand" type="button">Expand</button></div></div>
<div class="chart-body research-chart"><canvas aria-label="Forecast horizon vs latency and prefetch utilization" id="execHorizonChart" role="img"></canvas></div>
</div>
<div class="table-toolbar"><span>Horizon sweep</span><div><button class="mini-button" disabled="" id="execHorizonCopy" type="button">Copy table</button><button class="mini-button" disabled="" id="execHorizonCsv" type="button">Download CSV</button><button class="mini-button" disabled="" id="execHorizonJson" type="button">Download JSON</button></div></div>
<div class="table-wrap"><table><thead><tr><th>Horizon</th><th>p95 TTFT</th><th>Future-role recall@K</th><th>Prefix hit</th><th>Utilization</th><th>Unused prefetch</th><th>Mean HBM</th><th>Saved prefill</th></tr></thead><tbody id="execHorizonRows"></tbody></table></div>
</div>
</section>
<section class="panel research-panel">
<div class="panel-title-row"><div><h2>Prefix-cache budget sweep</h2><p class="muted">Compare top-1, multi-step, and utility-aware planning while shrinking or expanding the shared prefix-cache budget. The useful policy can change when speculative prefixes compete for HBM.</p></div><button class="primary compact" disabled="" id="execBudgetBtn" type="button">Sweep cache budget</button></div>
<div class="empty-state small" id="execBudgetEmpty"><h3>No cache-budget sweep yet</h3><p>The study executes 12 controlled simulations: three planning policies across four working-set-relative HBM budgets.</p></div>
<div class="hidden" id="execBudgetContent">
<div class="chart-card full" data-chart-card="" data-chart-name="execution-cache-budget-policy">
<div class="chart-head"><div class="chart-title">Cache budget vs p95 TTFT by planning policy</div><div class="chart-actions"><button class="chart-download" type="button">Download PNG</button><button class="chart-expand" type="button">Expand</button></div></div>
<div class="chart-body research-chart"><canvas aria-label="Cache budget vs p95 TTFT by planning policy" id="execBudgetChart" role="img"></canvas></div>
</div>
<div class="table-toolbar"><span>Cache-budget policy sweep</span><div><button class="mini-button" disabled="" id="execBudgetCopy" type="button">Copy table</button><button class="mini-button" disabled="" id="execBudgetCsv" type="button">Download CSV</button><button class="mini-button" disabled="" id="execBudgetJson" type="button">Download JSON</button></div></div>
<div class="table-wrap"><table><thead><tr><th>Budget</th><th>Policy</th><th>p95 TTFT</th><th>Future-role recall@K</th><th>Prefix hit</th><th>Utilization</th><th>Mean HBM</th><th>Unused prefetch</th><th>Pressure evictions</th></tr></thead><tbody id="execBudgetRows"></tbody></table></div>
</div>
</section>
</div>
</div>
</section>
<section aria-hidden="true" aria-labelledby="tab-consolidation" class="tab-panel" id="consolidation" role="tabpanel">
<div class="research-grid">
<aside class="panel research-controls">
<h2>Evidence settings</h2>
<p class="muted">Repeat the execution-policy comparison across matched workload seeds, quantify uncertainty, track Pareto stability, and measure regret to a bounded full-trace serving oracle.</p>
<div class="field-grid two">
<label>Matched seeds<input id="consSeeds" max="24" min="4" step="1" type="number" value="12"/></label>
<label>Bootstrap samples<input id="consBootstrap" max="4000" min="100" step="100" type="number" value="600"/></label>
</div>
<button class="primary" disabled="" id="consRunBtn" type="button">Run robust policy study</button>
<p class="control-help">The offline reference searches the same completed trace over deployable policies plus clairvoyant future-set plans H=1..5, K=1..3 under the same cache and transfer constraints. It is a bounded upper bound, not a global-optimality claim.</p>
<hr/>
<div class="section-kicker">External measurements</div>
<h3>Import + calibrate</h3>
<p class="muted">Import current vLLM <code>bench serve</code> JSON, SGLang <code>bench_serving</code> JSONL, or an InferScale case bundle. Artifact token lengths/request rate override the current Serving Lab workload when present; model, accelerator, and precision otherwise come from Serving Lab.</p>
<label>Benchmark artifact<input accept=".json,.jsonl,application/json,text/plain" id="measurementFile" type="file"/></label>
<div class="field-grid two">
<label>Format<select id="measurementSource"><option value="auto">Auto detect</option><option value="vllm">vLLM</option><option value="sglang">SGLang</option><option value="case_bundle">InferScale cases</option></select></label>
<label>Holdout fraction<input id="measurementHoldout" max="0.8" min="0" step="0.05" type="number" value="0.33"/><span class="unit">fraction</span></label>
</div>
<div class="trace-row" id="measurementStatus"><span>No measurement file loaded</span></div>
<button class="primary" disabled="" id="measurementCalibrateBtn" type="button">Import and calibrate</button>
<p class="control-help">With fewer than three cases, validation falls back to resubstitution and says so explicitly. Held-out results are preferred.</p>
<hr/>
<div class="section-kicker">Report export</div>
<h3>Report export</h3>
<p class="muted">Generate a Markdown report from the robust study and, when available, external-measurement calibration.</p>
<button class="secondary" disabled="" id="consReportBtn" type="button">Download Markdown report</button>
</aside>
<div class="research-results">
<section class="panel research-panel">
<div class="panel-title-row"><div><h2>Repeated-seed policy ranking</h2><p class="muted">Nominal winners can move with one workload trace. This study reports bootstrap uncertainty, per-seed TTFT wins, Pareto stability, worst-seed latency, and regret to the bounded offline oracle.</p></div><span aria-live="polite" class="tag" id="consState">Waiting</span></div>
<div class="empty-state small" id="consEmpty"><h3>No repeated-seed result yet</h3><p>Run the matched-seed experiment from the left panel. The default executes four deployable policies plus an offline oracle search on each seed.</p></div>
<div class="hidden" id="consContent">
<div class="metric-grid four">
<div class="metric"><span>Robust winner</span><strong class="metric-small" id="consWinner">-</strong></div>
<div class="metric"><span>First-seed winner</span><strong class="metric-small" id="consNominal">-</strong></div>
<div class="metric"><span>Matched seeds</span><strong id="consSeedCount">-</strong></div>
<div class="metric"><span>Median oracle TTFT</span><strong id="consOracle">-</strong></div>
</div>
<div class="chart-card full" data-chart-card="" data-chart-name="robust-policy-ranking">
<div class="chart-head"><div class="chart-title">Pareto stability vs median oracle regret</div><div class="chart-actions"><button class="chart-download" type="button">Download PNG</button><button class="chart-expand" type="button">Expand</button></div></div>
<div class="chart-body research-chart"><canvas aria-label="Pareto stability vs median oracle regret" id="consChart" role="img"></canvas></div>
</div>
<div class="table-toolbar"><span>Repeated-seed policy summary</span><div><button class="mini-button" disabled="" id="consCopy" type="button">Copy table</button><button class="mini-button" disabled="" id="consCsv" type="button">Download CSV</button><button class="mini-button" disabled="" id="consJson" type="button">Download JSON</button></div></div>
<div class="table-wrap"><table><thead><tr><th>Rank</th><th>Policy</th><th>Median TTFT</th><th>95% CI of mean</th><th>TTFT wins</th><th>Pareto stable</th><th>Oracle regret</th><th>Worst seed</th><th>Unused prefetch</th><th>Mean HBM</th></tr></thead><tbody id="consRows"></tbody></table></div>
<div class="planner-note" id="consNote"></div>
</div>
</section>
<section class="panel research-panel">
<div class="panel-title-row"><div><h2>Held-out calibration</h2><p class="muted">Fit robust global prefill/decode timing multipliers on training measurements, then evaluate the analytical simulator on held-out benchmark cases. This can correct global scale bias; it cannot by itself validate unseen schedulers or hardware.</p></div><span aria-live="polite" class="tag" id="measurementState">No data</span></div>
<div class="empty-state small" id="measurementEmpty"><h3>No external measurements loaded</h3><p>This public demo ships no fake benchmark truth. Import your own vLLM/SGLang measurements or an InferScale validation-case bundle.</p></div>
<div class="hidden" id="measurementContent">
<div class="metric-grid four">
<div class="metric"><span>Cases</span><strong id="measurementCases">-</strong></div>
<div class="metric"><span>Prefill scale</span><strong id="measurementPrefill">-</strong></div>
<div class="metric"><span>Decode scale</span><strong id="measurementDecode">-</strong></div>
<div class="metric"><span>Held-out MAPE</span><strong id="measurementMape">-</strong></div>
</div>
<div class="table-toolbar"><span id="measurementCaption">Validation residuals</span><div><button class="mini-button" disabled="" id="measurementCopy" type="button">Copy table</button><button class="mini-button" disabled="" id="measurementCsv" type="button">Download CSV</button><button class="mini-button" disabled="" id="measurementJson" type="button">Download JSON</button></div></div>
<div class="table-wrap"><table><thead><tr><th>Case</th><th>Metric</th><th>Measured</th><th>Baseline pred.</th><th>Calibrated pred.</th><th>Baseline APE</th><th>Calibrated APE</th></tr></thead><tbody id="measurementRows"></tbody></table></div>
<div class="planner-note" id="measurementNote"></div>
</div>
</section>
</div>
</div>
</section>
<section aria-hidden="true" aria-labelledby="tab-method" class="tab-panel" id="method" role="tabpanel">
<div class="method-grid">
<article class="panel prose"><h2>What is actually simulated?</h2><p>Requests are generated from deterministic workload distributions and advanced through virtual time. Colocated runs model admission, prefill, paged KV allocation, dynamic decode batches, and completion. P/D runs use separate prefill and decode worker pools plus an explicit serialized KV-transfer link.</p><div class="formula">request -&gt; queue -&gt; prefill -&gt; [KV transfer] -&gt; decode -&gt; completion</div></article>
<article class="panel prose"><h2>Cache without pretending to implement a radix tree</h2><p>The simulator models a single shared prompt prefix with configurable length and reuse fraction. Cache hits avoid redundant prefill work and share one persistent KV allocation. It is deliberately a controlled what-if abstraction, not a claim to reproduce SGLang's full RadixAttention policy.</p></article>
<article class="panel prose"><h2>Role-specific resources and transfer cost</h2><p>Prefill and decode have separate accelerator profiles and worker counts. Prompt KV state crosses a modeled interconnect before decode admission. The simulator reports role utilization, transfer latency, and transfer pressure so the benefit of isolation can be weighed against data-movement overhead.</p></article>
<article class="panel prose"><h2>Tool gaps turn KV into a residency and routing decision</h2><p>Agent Sessions preserves program identity and turn order, materializes tool-induced gaps, and tracks HBM retention, TTL expiry, host-memory offload/restore, recomputation, and pressure eviction. Strict affinity, least-load, and bounded-affinity routing expose the tension between cache locality and hot-replica queueing. The per-replica service model remains intentionally serial so state-management effects are not confounded with dynamic batching.</p></article>
<article class="panel prose"><h2>Prediction without future-gap leakage</h2><p>Adaptive tiering learns an exponentially weighted estimate from tool calls only after they complete. Global and per-tool predictors can be evaluated on the same non-stationary program trace, while an oracle policy that sees the realized future gap is kept separate as an upper bound. A shift experiment changes slower external-tool durations mid-trace to expose the stability-versus-adaptation trade-off.</p></article>
<article class="panel prose"><h2>Learn transitions, forecast multiple steps, then spend cache carefully</h2><p>Execution Learning updates a first-order role-transition model only after the next role is observed. The learned matrix can be rolled forward for multi-step forecasts without future-trace access. A utility-aware planner weighs expected prefill savings against host-transfer cost and forecast-weighted eviction cost, while calibration, horizon, confidence, forgetting, and cache-budget studies expose where better predictions do or do not improve serving outcomes.</p></article>
<article class="panel prose"><h2>Pareto, not one magic configuration</h2><p>The Design Explorer reports both a raw-performance frontier and a resource-normalized frontier using goodput per accelerator. That prevents a multi-GPU P/D layout from looking unconditionally better merely because it uses more simulated hardware.</p></article>
<article class="panel prose"><h2>Generated load or exact replay</h2><p>Poisson, constant, and bursty workloads are open-loop: arrivals are scheduled independently of response completion, so queueing delay is not hidden by client backpressure. Exact CSV/JSON traces can be replayed with their original arrival times and token lengths.</p></article>
<article class="panel prose"><h2>Paired conclusions, not one lucky seed</h2><p>Research Studies use common random numbers: baseline and treatment receive identical seeds, reducing workload variance in the paired difference. Research Summary extends that idea across matched seeds, bootstrap uncertainty, Pareto stability, and regret to a bounded full-trace serving oracle.</p></article>
<article class="panel prose"><h2>Analytical profiles are falsifiable</h2><p>The measurement bridge imports serving-benchmark artifacts, fits simple prefill/decode scale factors on training cases only, and evaluates held-out residuals. Current vLLM serving benchmarks can save JSON results, while SGLang bench_serving supports JSONL output; InferScale normalizes both into explicit validation cases rather than treating imported measurements as hidden truth.</p></article>
<article class="panel prose wide-method"><h2>Research lineage and scope</h2><p>Vidur established the value of simulation for avoiding expensive deployment sweeps. Recent systems have pushed toward heterogeneous and disaggregated serving, communication-aware modeling, stateful workloads, trace replay, and SLA-dependent design-space exploration. InferScale-Sim remains intentionally smaller and inspectable, with paired experiments and sensitivity analysis built into the workflow.</p><div class="paper-grid"><div><strong>Vidur / 2024</strong><span>Predictive profiling, workload-aware serving simulation, configuration search.</span></div><div><strong>TokenSim / 2025</strong><span>Extensible scheduling and memory-management simulation.</span></div><div><strong>Revati / 2026</strong><span>GPU-free time-warp emulation of serving control logic.</span></div><div><strong>LLMServingSim 2.0 / 2026</strong><span>Heterogeneous and disaggregated infrastructure, memory and communication.</span></div><div><strong>Frontier / May 2026</strong><span>P/D disaggregation, runtime optimizations, stateful workloads, Pareto exploration.</span></div><div><strong>HeteroPanacea / Aug 2026</strong><span>Heterogeneous stage specialization motivates resource-aware P/D comparison.</span></div><div><strong>Vanguard / Jun 2026</strong><span>Open-loop replay avoids coordinated omission when studying latency under load.</span></div><div><strong>AgentServeSim / Jun 2026</strong><span>Stateful multi-turn serving motivates session-aware workload modeling.</span></div><div><strong>IdleKV / Jun 2026</strong><span>Tool-call idle windows motivate explicit HBM-to-host KV offload experiments.</span></div><div><strong>SMetric / Jul 2026</strong><span>Cache-local routing can overload hot replicas, motivating bounded affinity.</span></div><div><strong>Continuum / May 2026</strong><span>Per-tool duration history and bounded TTL motivate adaptive retention without clairvoyance.</span></div><div><strong>CacheScout / Jul 2026</strong><span>Online learning of agent execution transitions and between-step prefetch motivates the Execution Learning experiments.</span></div><div><strong>PBKV / May 2026</strong><span>Multi-step prediction for dynamic workflows motivates forecast-horizon and utility-aware prefetch experiments.</span></div><div><strong>Predictive KV Memory / Aug 2026 rev.</strong><span>Bayesian reuse prediction and multi-tier placement motivate explicit prediction-quality experiments.</span></div><div><strong>SGLang / RadixAttention</strong><span>Automatic shared-prefix KV reuse motivates the controlled cache scenario.</span></div></div></article>
</div>
</section>
</main>
<div class="toast" id="toast" role="status"></div>
<script src="app.js"></script>
</body>
</html>