openjev / code /serving /V100.md
AlexWortega's picture
Add image Decisions serving and explicit option probabilities
26de23c verified
|
Raw History Blame Contribute Delete
9.61 kB
# OpenJEV on SGLang / V100
This uses `sglang_openjev/` from
https://huggingface.co/AlexWortega/openjev/tree/main/code/sglang_openjev.
`HF_SOURCE.json` records the verified source revision and hashes. The adapter
registers a Qwen3.5 classification head inside SGLang; `/classify` returns three
raw logits in contradiction, entailment, neutral order.
The separate `code/v100` offline scorer is not this HTTP serving engine.
## Isolated runtime
Work directory on eva01: `~/storage/sglang-openjev-v100` (the large storage mount).
The existing evaluation environment and system CUDA installation are unchanged.
Runtime source: https://github.com/haohervchb/sglang-V100,
revision `dca488908ee4e3f1bc676c3bf5dcd26ff049cfc3`.
FlashInfer source: https://github.com/haohervchb/flashinfer,
revision `c3c40a7b90b792fc59f90f8f55c9e2de9c1b6833`, with the runtime's
`patches/flashinfer-sm70.patch` applied.
Use PyTorch 2.9.1 **cu126**, torchvision 0.24.1, and a private CUDA 12.6.3
compiler. Official cu128 PyTorch wheels omit SM70. The installer explicitly
checks `torch.cuda.get_arch_list()` and performs an FP16 GPU matmul before
installing the runtime. `prepare_v100_cuda.py ROOT --version 12.6.3` downloads
NVIDIA toolkit components and verifies their manifest hashes.
After cloning the pinned runtime, preparing the private toolkit at
`$OPENJEV_V100_ROOT/cuda126`, and installing the cu126 wheels in its `venv`, run
`install_v100_runtime.sh`. It builds the minimal SM70 `sgl-kernel` and installs
the patched FlashInfer. This is a local source build, not the fork's system
installation script.
The pinned SGLang version also builds a Rust gRPC extension during installation.
Provision Rust in `$OPENJEV_V100_ROOT/rustup` with cargo caches under
`$OPENJEV_V100_ROOT/cargo`; the installer uses these paths. Rust 1.98.1 and
`protoc` 33.6 (under `$OPENJEV_V100_ROOT/protoc`) were provisioned on eva01.
The host's older protoc does not accept `--experimental_allow_proto3_optional`.
Neither a system-wide Rust default nor a CUDA driver update is required.
## Launch and validation
The training checkpoint lacks image/video processor configs, which SGLang's
Qwen3.5 architecture requires even for text classification. Run
`prepare_v100_model.py CHECKPOINT $OPENJEV_V100_ROOT/model` to create a symlink
overlay. It adds only the two processor configs from the pinned base
`Qwen/Qwen3.5-4B` revision, checks their hashes, and preserves the checkpoint's
weights, tokenizer and model config.
`serve_sglang_v100.sh CHECKPOINT 31000` launches on GPU 0 and binds localhost.
Set `CUDA_VISIBLE_DEVICES` to choose another free GPU. Initial settings use FP16,
TileLang V100 attention and GDN prefill, an 8192-token server context, and disabled CUDA
graphs. Pass `--disable-radix-cache` for the initial uncached baseline.
The benchmark still truncates at 4096 tokens, matching the reference; this
runtime rejects inputs exactly equal to the configured server context length.
The launcher selects GCC 10 for C++20 JIT headers and PyTorch sampling (the
classifier does not sample output tokens). Since the September 28 deployment,
it caps running requests at 16 and uses LPM scheduling by default; override with
`MAX_RUNNING_REQUESTS` and `SCHEDULE_POLICY`. The original September 26 results
below used 8 running requests and FCFS. Prefix caching uses
`extra_buffer` and a Mamba-to-KV memory ratio of 3. This provides 503 recurrent
state slots and 135,008 KV tokens at the same 80% GPU memory budget. The original
ratio of 0.9 provided 317 slots and suffered cache eviction on the larger set.
`--disable-radix-cache` selects the uncached `no_buffer` baseline automatically.
All JIT caches use the large storage mount.
Apply `patches/dynamic-paged.patch` to the pinned runtime before launching.
It makes the batch, physical-page count and block-table width runtime dimensions
of the existing TileLang paged attention kernel. Static specialization on these
shapes caused repeated compilation during real request batches. The math and
tiling remain unchanged. `SGLANG_V100_DYNAMIC_PAGED=0` restores static selection.
`check_dynamic_paged.py` checks ragged causal attention with and without cached
prefixes against a dense FP32 reference and asserts that changing all three
dimensions reuses one kernel. The GPU check passed with maximum absolute error
0.001073; this kernel check alone does not establish full-model agreement.
`bench_classify_v100.py` submits exact token IDs to `/classify` and compares the
responses to saved reference logits. It records input and token hashes,
entailment probability differences, NLI argmax changes, HTTP wall time and
per-batch latencies. Warmup is excluded. `--cold-cache` flushes SGLang's cache
before each iteration. Cold and warm-cache results must be reported separately.
## Verified result, 2026-09-26
On eva01, one V100-SXM2-32GB, FP16, HTTP `/classify`, batches of 32 inputs with
four client workers and eight running sequences. Each cell is the median of
three measured repetitions after one initial pass:
| Pairs | Prefix cache disabled | Prefix cache cleared before each pass | Repeated inputs in warm prefix cache |
|---|---:|---:|---:|
| G: 303 | 24.928 s | 13.417 s | 2.697 s |
| T: 564 | 47.618 s | 29.635 s | 4.732 s |
The cache therefore improves these fresh-set passes by 1.86x / 1.61x and exact
repeated-set passes by 9.24x / 10.06x. Warm-prefix numbers are not fresh-request
throughput. Changing the Mamba memory ratio from 0.9 to 3 reduced the larger
set's warm-cache median from 28.801 s to 4.732 s.
All 867 NLI argmax labels agree with an independently generated, unmodified
Transformers FP16 reference. Input/token hashes were checked. The largest
entailment probability difference across the selected cached runs was 0.012851.
In the saved final pass of each profile, among 291 same-premise/instruction
option groups (793 rows), no strict reference winner changed. Two cold-cache
choices changed where the reference probabilities were exactly tied. This is agreement testing, not a held-out Decision Index
accuracy result; singleton and cross-document aggregation are not covered by
that grouped check.
First-use JIT remains expensive, especially after changing the GDN state-pool
size: the selected profile's first G pass took 202.709 s. Initial passes are
excluded from medians; additional JIT outliers in measured repetitions are
retained, including 45.085 s for uncached G and 15.064 s for warm-cache T.
These are controlled throughput measurements, not production p99 guarantees.
Receipts and raw arrays: `results_v100/sglang/summary.json`,
`runtime_receipt.json`, `ratio3_server_config.json`, individual benchmark JSONs,
and reference/candidate `.npy` files. The remote directory is
`~/storage/sglang-openjev-v100/results`.
The endpoint is `http://127.0.0.1:31000/classify` on eva01. Since September 28
it is a systemd socket-activated gateway over two replicas; see
[`DEPLOYMENT-eva01.md`](DEPLOYMENT-eva01.md). Do not launch another standalone
server on that port. For a separate development instance on a free GPU/port:
```bash
CUDA_VISIBLE_DEVICES=3 bash ~/storage/sglang-openjev-v100/app/serve_sglang_v100.sh \
~/storage/sglang-openjev-v100/model 31003
```
The model overlay defaults to `~/storage/sglang-openjev-v100/model`. The old
standalone PID record is archived as `server.legacy-20260926.pid`; current PIDs
and logs are managed by systemd. `server_cached_ratio3.log` is historical.
## Further scheduling comparison, 2026-09-27
`tune_sglang_v100.py` compares FCFS with 8, 16 and 32 running requests, then
LPM with 16. It uses a separate GPU and port, waits for 60 seconds without GPU
processes before each launch, and preserves the existing endpoint. Each profile
runs both input sets with cleared and warm prefix cache and the same FP16
references. It saves server configuration, warmup, all repetitions, logits and
correctness checks. It does not automatically promote a profile.
The benchmark now audits same-premise/instruction option groups on every pass,
including warmup, and retains all logits in `.passes.npy`. Selecting outside the
set of reference winners fails the benchmark; selecting another exactly tied
winner is recorded separately. The guard reproduces the previous saved-pass
audit: 291 groups, no choices outside the reference winners, two exact ties in
cold T. Singleton and cross-document ranking remain outside this check.
```bash
cd ~/storage/sglang-openjev-v100
venv/bin/python app/tune_sglang_v100.py --gpu 1 --port 31001 \
--out "$PWD/results/tuning_20260927"
```
The output directory must be new. `run.json` records live phase, server PID and
completed measurements. A daemon invocation is currently recorded in
`tuning_20260927.pid`; check that process and its log before starting another.
`bench_classify_v100.py --exclusive-gpu INDEX --server-pid PID` samples physical
GPU process ownership every 0.5 seconds. Only the named server and its descendants
are allowed; PID start times protect against reuse. Competing processes or failed
monitoring invalidate the report and return a nonzero exit status. Sub-sampling
interval competition is still possible; this is sampled evidence, not a GPU lease.
The initial September 27 refresh was contaminated by simultaneous v6 evaluation
jobs: warm G took 6.28–7.30 seconds instead of the prior 2.70 seconds. The candidate
also received only 192 Mamba state slots because another job allocated GPU memory
during startup. That candidate was stopped. These timings must not be used to
rank configurations; the next valid comparison must run on an exclusive GPU.