# OpenJEV on SGLang / V100 This uses `sglang_openjev/` from https://huggingface.co/AlexWortega/openjev/tree/main/code/sglang_openjev. `HF_SOURCE.json` records the verified source revision and hashes. The adapter registers a Qwen3.5 classification head inside SGLang; `/classify` returns three raw logits in contradiction, entailment, neutral order. The separate `code/v100` offline scorer is not this HTTP serving engine. ## Isolated runtime Work directory on eva01: `~/storage/sglang-openjev-v100` (the large storage mount). The existing evaluation environment and system CUDA installation are unchanged. Runtime source: https://github.com/haohervchb/sglang-V100, revision `dca488908ee4e3f1bc676c3bf5dcd26ff049cfc3`. FlashInfer source: https://github.com/haohervchb/flashinfer, revision `c3c40a7b90b792fc59f90f8f55c9e2de9c1b6833`, with the runtime's `patches/flashinfer-sm70.patch` applied. Use PyTorch 2.9.1 **cu126**, torchvision 0.24.1, and a private CUDA 12.6.3 compiler. Official cu128 PyTorch wheels omit SM70. The installer explicitly checks `torch.cuda.get_arch_list()` and performs an FP16 GPU matmul before installing the runtime. `prepare_v100_cuda.py ROOT --version 12.6.3` downloads NVIDIA toolkit components and verifies their manifest hashes. After cloning the pinned runtime, preparing the private toolkit at `$OPENJEV_V100_ROOT/cuda126`, and installing the cu126 wheels in its `venv`, run `install_v100_runtime.sh`. It builds the minimal SM70 `sgl-kernel` and installs the patched FlashInfer. This is a local source build, not the fork's system installation script. The pinned SGLang version also builds a Rust gRPC extension during installation. Provision Rust in `$OPENJEV_V100_ROOT/rustup` with cargo caches under `$OPENJEV_V100_ROOT/cargo`; the installer uses these paths. Rust 1.98.1 and `protoc` 33.6 (under `$OPENJEV_V100_ROOT/protoc`) were provisioned on eva01. The host's older protoc does not accept `--experimental_allow_proto3_optional`. Neither a system-wide Rust default nor a CUDA driver update is required. ## Launch and validation The training checkpoint lacks image/video processor configs, which SGLang's Qwen3.5 architecture requires even for text classification. Run `prepare_v100_model.py CHECKPOINT $OPENJEV_V100_ROOT/model` to create a symlink overlay. It adds only the two processor configs from the pinned base `Qwen/Qwen3.5-4B` revision, checks their hashes, and preserves the checkpoint's weights, tokenizer and model config. `serve_sglang_v100.sh CHECKPOINT 31000` launches on GPU 0 and binds localhost. Set `CUDA_VISIBLE_DEVICES` to choose another free GPU. Initial settings use FP16, TileLang V100 attention and GDN prefill, an 8192-token server context, and disabled CUDA graphs. Pass `--disable-radix-cache` for the initial uncached baseline. The benchmark still truncates at 4096 tokens, matching the reference; this runtime rejects inputs exactly equal to the configured server context length. The launcher selects GCC 10 for C++20 JIT headers and PyTorch sampling (the classifier does not sample output tokens). Since the September 28 deployment, it caps running requests at 16 and uses LPM scheduling by default; override with `MAX_RUNNING_REQUESTS` and `SCHEDULE_POLICY`. The original September 26 results below used 8 running requests and FCFS. Prefix caching uses `extra_buffer` and a Mamba-to-KV memory ratio of 3. This provides 503 recurrent state slots and 135,008 KV tokens at the same 80% GPU memory budget. The original ratio of 0.9 provided 317 slots and suffered cache eviction on the larger set. `--disable-radix-cache` selects the uncached `no_buffer` baseline automatically. All JIT caches use the large storage mount. Apply `patches/dynamic-paged.patch` to the pinned runtime before launching. It makes the batch, physical-page count and block-table width runtime dimensions of the existing TileLang paged attention kernel. Static specialization on these shapes caused repeated compilation during real request batches. The math and tiling remain unchanged. `SGLANG_V100_DYNAMIC_PAGED=0` restores static selection. `check_dynamic_paged.py` checks ragged causal attention with and without cached prefixes against a dense FP32 reference and asserts that changing all three dimensions reuses one kernel. The GPU check passed with maximum absolute error 0.001073; this kernel check alone does not establish full-model agreement. `bench_classify_v100.py` submits exact token IDs to `/classify` and compares the responses to saved reference logits. It records input and token hashes, entailment probability differences, NLI argmax changes, HTTP wall time and per-batch latencies. Warmup is excluded. `--cold-cache` flushes SGLang's cache before each iteration. Cold and warm-cache results must be reported separately. ## Verified result, 2026-09-26 On eva01, one V100-SXM2-32GB, FP16, HTTP `/classify`, batches of 32 inputs with four client workers and eight running sequences. Each cell is the median of three measured repetitions after one initial pass: | Pairs | Prefix cache disabled | Prefix cache cleared before each pass | Repeated inputs in warm prefix cache | |---|---:|---:|---:| | G: 303 | 24.928 s | 13.417 s | 2.697 s | | T: 564 | 47.618 s | 29.635 s | 4.732 s | The cache therefore improves these fresh-set passes by 1.86x / 1.61x and exact repeated-set passes by 9.24x / 10.06x. Warm-prefix numbers are not fresh-request throughput. Changing the Mamba memory ratio from 0.9 to 3 reduced the larger set's warm-cache median from 28.801 s to 4.732 s. All 867 NLI argmax labels agree with an independently generated, unmodified Transformers FP16 reference. Input/token hashes were checked. The largest entailment probability difference across the selected cached runs was 0.012851. In the saved final pass of each profile, among 291 same-premise/instruction option groups (793 rows), no strict reference winner changed. Two cold-cache choices changed where the reference probabilities were exactly tied. This is agreement testing, not a held-out Decision Index accuracy result; singleton and cross-document aggregation are not covered by that grouped check. First-use JIT remains expensive, especially after changing the GDN state-pool size: the selected profile's first G pass took 202.709 s. Initial passes are excluded from medians; additional JIT outliers in measured repetitions are retained, including 45.085 s for uncached G and 15.064 s for warm-cache T. These are controlled throughput measurements, not production p99 guarantees. Receipts and raw arrays: `results_v100/sglang/summary.json`, `runtime_receipt.json`, `ratio3_server_config.json`, individual benchmark JSONs, and reference/candidate `.npy` files. The remote directory is `~/storage/sglang-openjev-v100/results`. The endpoint is `http://127.0.0.1:31000/classify` on eva01. Since September 28 it is a systemd socket-activated gateway over two replicas; see [`DEPLOYMENT-eva01.md`](DEPLOYMENT-eva01.md). Do not launch another standalone server on that port. For a separate development instance on a free GPU/port: ```bash CUDA_VISIBLE_DEVICES=3 bash ~/storage/sglang-openjev-v100/app/serve_sglang_v100.sh \ ~/storage/sglang-openjev-v100/model 31003 ``` The model overlay defaults to `~/storage/sglang-openjev-v100/model`. The old standalone PID record is archived as `server.legacy-20260926.pid`; current PIDs and logs are managed by systemd. `server_cached_ratio3.log` is historical. ## Further scheduling comparison, 2026-09-27 `tune_sglang_v100.py` compares FCFS with 8, 16 and 32 running requests, then LPM with 16. It uses a separate GPU and port, waits for 60 seconds without GPU processes before each launch, and preserves the existing endpoint. Each profile runs both input sets with cleared and warm prefix cache and the same FP16 references. It saves server configuration, warmup, all repetitions, logits and correctness checks. It does not automatically promote a profile. The benchmark now audits same-premise/instruction option groups on every pass, including warmup, and retains all logits in `.passes.npy`. Selecting outside the set of reference winners fails the benchmark; selecting another exactly tied winner is recorded separately. The guard reproduces the previous saved-pass audit: 291 groups, no choices outside the reference winners, two exact ties in cold T. Singleton and cross-document ranking remain outside this check. ```bash cd ~/storage/sglang-openjev-v100 venv/bin/python app/tune_sglang_v100.py --gpu 1 --port 31001 \ --out "$PWD/results/tuning_20260927" ``` The output directory must be new. `run.json` records live phase, server PID and completed measurements. A daemon invocation is currently recorded in `tuning_20260927.pid`; check that process and its log before starting another. `bench_classify_v100.py --exclusive-gpu INDEX --server-pid PID` samples physical GPU process ownership every 0.5 seconds. Only the named server and its descendants are allowed; PID start times protect against reuse. Competing processes or failed monitoring invalidate the report and return a nonzero exit status. Sub-sampling interval competition is still possible; this is sampled evidence, not a GPU lease. The initial September 27 refresh was contaminated by simultaneous v6 evaluation jobs: warm G took 6.28–7.30 seconds instead of the prior 2.70 seconds. The candidate also received only 192 Mamba state slots because another job allocated GPU memory during startup. That candidate was stopped. These timings must not be used to rank configurations; the next valid comparison must run on an exclusive GPU.