openjev / code /serving /V100.md
AlexWortega's picture
Add image Decisions serving and explicit option probabilities
26de23c verified
|
Raw History Blame Contribute Delete
9.61 kB

OpenJEV on SGLang / V100

This uses sglang_openjev/ from https://huggingface.co/AlexWortega/openjev/tree/main/code/sglang_openjev. HF_SOURCE.json records the verified source revision and hashes. The adapter registers a Qwen3.5 classification head inside SGLang; /classify returns three raw logits in contradiction, entailment, neutral order.

The separate code/v100 offline scorer is not this HTTP serving engine.

Isolated runtime

Work directory on eva01: ~/storage/sglang-openjev-v100 (the large storage mount). The existing evaluation environment and system CUDA installation are unchanged.

Runtime source: https://github.com/haohervchb/sglang-V100, revision dca488908ee4e3f1bc676c3bf5dcd26ff049cfc3. FlashInfer source: https://github.com/haohervchb/flashinfer, revision c3c40a7b90b792fc59f90f8f55c9e2de9c1b6833, with the runtime's patches/flashinfer-sm70.patch applied.

Use PyTorch 2.9.1 cu126, torchvision 0.24.1, and a private CUDA 12.6.3 compiler. Official cu128 PyTorch wheels omit SM70. The installer explicitly checks torch.cuda.get_arch_list() and performs an FP16 GPU matmul before installing the runtime. prepare_v100_cuda.py ROOT --version 12.6.3 downloads NVIDIA toolkit components and verifies their manifest hashes.

After cloning the pinned runtime, preparing the private toolkit at $OPENJEV_V100_ROOT/cuda126, and installing the cu126 wheels in its venv, run install_v100_runtime.sh. It builds the minimal SM70 sgl-kernel and installs the patched FlashInfer. This is a local source build, not the fork's system installation script.

The pinned SGLang version also builds a Rust gRPC extension during installation. Provision Rust in $OPENJEV_V100_ROOT/rustup with cargo caches under $OPENJEV_V100_ROOT/cargo; the installer uses these paths. Rust 1.98.1 and protoc 33.6 (under $OPENJEV_V100_ROOT/protoc) were provisioned on eva01. The host's older protoc does not accept --experimental_allow_proto3_optional. Neither a system-wide Rust default nor a CUDA driver update is required.

Launch and validation

The training checkpoint lacks image/video processor configs, which SGLang's Qwen3.5 architecture requires even for text classification. Run prepare_v100_model.py CHECKPOINT $OPENJEV_V100_ROOT/model to create a symlink overlay. It adds only the two processor configs from the pinned base Qwen/Qwen3.5-4B revision, checks their hashes, and preserves the checkpoint's weights, tokenizer and model config.

serve_sglang_v100.sh CHECKPOINT 31000 launches on GPU 0 and binds localhost. Set CUDA_VISIBLE_DEVICES to choose another free GPU. Initial settings use FP16, TileLang V100 attention and GDN prefill, an 8192-token server context, and disabled CUDA graphs. Pass --disable-radix-cache for the initial uncached baseline. The benchmark still truncates at 4096 tokens, matching the reference; this runtime rejects inputs exactly equal to the configured server context length.

The launcher selects GCC 10 for C++20 JIT headers and PyTorch sampling (the classifier does not sample output tokens). Since the September 28 deployment, it caps running requests at 16 and uses LPM scheduling by default; override with MAX_RUNNING_REQUESTS and SCHEDULE_POLICY. The original September 26 results below used 8 running requests and FCFS. Prefix caching uses extra_buffer and a Mamba-to-KV memory ratio of 3. This provides 503 recurrent state slots and 135,008 KV tokens at the same 80% GPU memory budget. The original ratio of 0.9 provided 317 slots and suffered cache eviction on the larger set. --disable-radix-cache selects the uncached no_buffer baseline automatically. All JIT caches use the large storage mount.

Apply patches/dynamic-paged.patch to the pinned runtime before launching. It makes the batch, physical-page count and block-table width runtime dimensions of the existing TileLang paged attention kernel. Static specialization on these shapes caused repeated compilation during real request batches. The math and tiling remain unchanged. SGLANG_V100_DYNAMIC_PAGED=0 restores static selection. check_dynamic_paged.py checks ragged causal attention with and without cached prefixes against a dense FP32 reference and asserts that changing all three dimensions reuses one kernel. The GPU check passed with maximum absolute error 0.001073; this kernel check alone does not establish full-model agreement.

bench_classify_v100.py submits exact token IDs to /classify and compares the responses to saved reference logits. It records input and token hashes, entailment probability differences, NLI argmax changes, HTTP wall time and per-batch latencies. Warmup is excluded. --cold-cache flushes SGLang's cache before each iteration. Cold and warm-cache results must be reported separately.

Verified result, 2026-09-26

On eva01, one V100-SXM2-32GB, FP16, HTTP /classify, batches of 32 inputs with four client workers and eight running sequences. Each cell is the median of three measured repetitions after one initial pass:

Pairs Prefix cache disabled Prefix cache cleared before each pass Repeated inputs in warm prefix cache
G: 303 24.928 s 13.417 s 2.697 s
T: 564 47.618 s 29.635 s 4.732 s

The cache therefore improves these fresh-set passes by 1.86x / 1.61x and exact repeated-set passes by 9.24x / 10.06x. Warm-prefix numbers are not fresh-request throughput. Changing the Mamba memory ratio from 0.9 to 3 reduced the larger set's warm-cache median from 28.801 s to 4.732 s.

All 867 NLI argmax labels agree with an independently generated, unmodified Transformers FP16 reference. Input/token hashes were checked. The largest entailment probability difference across the selected cached runs was 0.012851. In the saved final pass of each profile, among 291 same-premise/instruction option groups (793 rows), no strict reference winner changed. Two cold-cache choices changed where the reference probabilities were exactly tied. This is agreement testing, not a held-out Decision Index accuracy result; singleton and cross-document aggregation are not covered by that grouped check.

First-use JIT remains expensive, especially after changing the GDN state-pool size: the selected profile's first G pass took 202.709 s. Initial passes are excluded from medians; additional JIT outliers in measured repetitions are retained, including 45.085 s for uncached G and 15.064 s for warm-cache T. These are controlled throughput measurements, not production p99 guarantees.

Receipts and raw arrays: results_v100/sglang/summary.json, runtime_receipt.json, ratio3_server_config.json, individual benchmark JSONs, and reference/candidate .npy files. The remote directory is ~/storage/sglang-openjev-v100/results.

The endpoint is http://127.0.0.1:31000/classify on eva01. Since September 28 it is a systemd socket-activated gateway over two replicas; see DEPLOYMENT-eva01.md. Do not launch another standalone server on that port. For a separate development instance on a free GPU/port:

CUDA_VISIBLE_DEVICES=3 bash ~/storage/sglang-openjev-v100/app/serve_sglang_v100.sh \
  ~/storage/sglang-openjev-v100/model 31003

The model overlay defaults to ~/storage/sglang-openjev-v100/model. The old standalone PID record is archived as server.legacy-20260926.pid; current PIDs and logs are managed by systemd. server_cached_ratio3.log is historical.

Further scheduling comparison, 2026-09-27

tune_sglang_v100.py compares FCFS with 8, 16 and 32 running requests, then LPM with 16. It uses a separate GPU and port, waits for 60 seconds without GPU processes before each launch, and preserves the existing endpoint. Each profile runs both input sets with cleared and warm prefix cache and the same FP16 references. It saves server configuration, warmup, all repetitions, logits and correctness checks. It does not automatically promote a profile.

The benchmark now audits same-premise/instruction option groups on every pass, including warmup, and retains all logits in .passes.npy. Selecting outside the set of reference winners fails the benchmark; selecting another exactly tied winner is recorded separately. The guard reproduces the previous saved-pass audit: 291 groups, no choices outside the reference winners, two exact ties in cold T. Singleton and cross-document ranking remain outside this check.

cd ~/storage/sglang-openjev-v100
venv/bin/python app/tune_sglang_v100.py --gpu 1 --port 31001 \
  --out "$PWD/results/tuning_20260927"

The output directory must be new. run.json records live phase, server PID and completed measurements. A daemon invocation is currently recorded in tuning_20260927.pid; check that process and its log before starting another.

bench_classify_v100.py --exclusive-gpu INDEX --server-pid PID samples physical GPU process ownership every 0.5 seconds. Only the named server and its descendants are allowed; PID start times protect against reuse. Competing processes or failed monitoring invalidate the report and return a nonzero exit status. Sub-sampling interval competition is still possible; this is sampled evidence, not a GPU lease.

The initial September 27 refresh was contaminated by simultaneous v6 evaluation jobs: warm G took 6.28–7.30 seconds instead of the prior 2.70 seconds. The candidate also received only 192 Mamba state slots because another job allocated GPU memory during startup. That candidate was stopped. These timings must not be used to rank configurations; the next valid comparison must run on an exclusive GPU.