Instructions to use AlexWortega/openjev with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use AlexWortega/openjev with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="AlexWortega/openjev")# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("AlexWortega/openjev", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Download code/serving/V100.md from AlexWortega/openjev: direct link, hf CLI and curl.
- Browser
- Download file 9.61 kB
-
https://huggingface.co/AlexWortega/openjev/resolve/main/code/serving/V100.md
- Command line
-
hf download hf://AlexWortega/openjev/code/serving/V100.md
-
curl -L -o V100.md https://huggingface.co/AlexWortega/openjev/resolve/main/code/serving/V100.md
OpenJEV on SGLang / V100
This uses sglang_openjev/ from
https://huggingface.co/AlexWortega/openjev/tree/main/code/sglang_openjev.
HF_SOURCE.json records the verified source revision and hashes. The adapter
registers a Qwen3.5 classification head inside SGLang; /classify returns three
raw logits in contradiction, entailment, neutral order.
The separate code/v100 offline scorer is not this HTTP serving engine.
Isolated runtime
Work directory on eva01: ~/storage/sglang-openjev-v100 (the large storage mount).
The existing evaluation environment and system CUDA installation are unchanged.
Runtime source: https://github.com/haohervchb/sglang-V100,
revision dca488908ee4e3f1bc676c3bf5dcd26ff049cfc3.
FlashInfer source: https://github.com/haohervchb/flashinfer,
revision c3c40a7b90b792fc59f90f8f55c9e2de9c1b6833, with the runtime's
patches/flashinfer-sm70.patch applied.
Use PyTorch 2.9.1 cu126, torchvision 0.24.1, and a private CUDA 12.6.3
compiler. Official cu128 PyTorch wheels omit SM70. The installer explicitly
checks torch.cuda.get_arch_list() and performs an FP16 GPU matmul before
installing the runtime. prepare_v100_cuda.py ROOT --version 12.6.3 downloads
NVIDIA toolkit components and verifies their manifest hashes.
After cloning the pinned runtime, preparing the private toolkit at
$OPENJEV_V100_ROOT/cuda126, and installing the cu126 wheels in its venv, run
install_v100_runtime.sh. It builds the minimal SM70 sgl-kernel and installs
the patched FlashInfer. This is a local source build, not the fork's system
installation script.
The pinned SGLang version also builds a Rust gRPC extension during installation.
Provision Rust in $OPENJEV_V100_ROOT/rustup with cargo caches under
$OPENJEV_V100_ROOT/cargo; the installer uses these paths. Rust 1.98.1 and
protoc 33.6 (under $OPENJEV_V100_ROOT/protoc) were provisioned on eva01.
The host's older protoc does not accept --experimental_allow_proto3_optional.
Neither a system-wide Rust default nor a CUDA driver update is required.
Launch and validation
The training checkpoint lacks image/video processor configs, which SGLang's
Qwen3.5 architecture requires even for text classification. Run
prepare_v100_model.py CHECKPOINT $OPENJEV_V100_ROOT/model to create a symlink
overlay. It adds only the two processor configs from the pinned base
Qwen/Qwen3.5-4B revision, checks their hashes, and preserves the checkpoint's
weights, tokenizer and model config.
serve_sglang_v100.sh CHECKPOINT 31000 launches on GPU 0 and binds localhost.
Set CUDA_VISIBLE_DEVICES to choose another free GPU. Initial settings use FP16,
TileLang V100 attention and GDN prefill, an 8192-token server context, and disabled CUDA
graphs. Pass --disable-radix-cache for the initial uncached baseline.
The benchmark still truncates at 4096 tokens, matching the reference; this
runtime rejects inputs exactly equal to the configured server context length.
The launcher selects GCC 10 for C++20 JIT headers and PyTorch sampling (the
classifier does not sample output tokens). Since the September 28 deployment,
it caps running requests at 16 and uses LPM scheduling by default; override with
MAX_RUNNING_REQUESTS and SCHEDULE_POLICY. The original September 26 results
below used 8 running requests and FCFS. Prefix caching uses
extra_buffer and a Mamba-to-KV memory ratio of 3. This provides 503 recurrent
state slots and 135,008 KV tokens at the same 80% GPU memory budget. The original
ratio of 0.9 provided 317 slots and suffered cache eviction on the larger set.
--disable-radix-cache selects the uncached no_buffer baseline automatically.
All JIT caches use the large storage mount.
Apply patches/dynamic-paged.patch to the pinned runtime before launching.
It makes the batch, physical-page count and block-table width runtime dimensions
of the existing TileLang paged attention kernel. Static specialization on these
shapes caused repeated compilation during real request batches. The math and
tiling remain unchanged. SGLANG_V100_DYNAMIC_PAGED=0 restores static selection.
check_dynamic_paged.py checks ragged causal attention with and without cached
prefixes against a dense FP32 reference and asserts that changing all three
dimensions reuses one kernel. The GPU check passed with maximum absolute error
0.001073; this kernel check alone does not establish full-model agreement.
bench_classify_v100.py submits exact token IDs to /classify and compares the
responses to saved reference logits. It records input and token hashes,
entailment probability differences, NLI argmax changes, HTTP wall time and
per-batch latencies. Warmup is excluded. --cold-cache flushes SGLang's cache
before each iteration. Cold and warm-cache results must be reported separately.
Verified result, 2026-09-26
On eva01, one V100-SXM2-32GB, FP16, HTTP /classify, batches of 32 inputs with
four client workers and eight running sequences. Each cell is the median of
three measured repetitions after one initial pass:
| Pairs | Prefix cache disabled | Prefix cache cleared before each pass | Repeated inputs in warm prefix cache |
|---|---|---|---|
| G: 303 | 24.928 s | 13.417 s | 2.697 s |
| T: 564 | 47.618 s | 29.635 s | 4.732 s |
The cache therefore improves these fresh-set passes by 1.86x / 1.61x and exact repeated-set passes by 9.24x / 10.06x. Warm-prefix numbers are not fresh-request throughput. Changing the Mamba memory ratio from 0.9 to 3 reduced the larger set's warm-cache median from 28.801 s to 4.732 s.
All 867 NLI argmax labels agree with an independently generated, unmodified Transformers FP16 reference. Input/token hashes were checked. The largest entailment probability difference across the selected cached runs was 0.012851. In the saved final pass of each profile, among 291 same-premise/instruction option groups (793 rows), no strict reference winner changed. Two cold-cache choices changed where the reference probabilities were exactly tied. This is agreement testing, not a held-out Decision Index accuracy result; singleton and cross-document aggregation are not covered by that grouped check.
First-use JIT remains expensive, especially after changing the GDN state-pool size: the selected profile's first G pass took 202.709 s. Initial passes are excluded from medians; additional JIT outliers in measured repetitions are retained, including 45.085 s for uncached G and 15.064 s for warm-cache T. These are controlled throughput measurements, not production p99 guarantees.
Receipts and raw arrays: results_v100/sglang/summary.json,
runtime_receipt.json, ratio3_server_config.json, individual benchmark JSONs,
and reference/candidate .npy files. The remote directory is
~/storage/sglang-openjev-v100/results.
The endpoint is http://127.0.0.1:31000/classify on eva01. Since September 28
it is a systemd socket-activated gateway over two replicas; see
DEPLOYMENT-eva01.md. Do not launch another standalone
server on that port. For a separate development instance on a free GPU/port:
CUDA_VISIBLE_DEVICES=3 bash ~/storage/sglang-openjev-v100/app/serve_sglang_v100.sh \
~/storage/sglang-openjev-v100/model 31003
The model overlay defaults to ~/storage/sglang-openjev-v100/model. The old
standalone PID record is archived as server.legacy-20260926.pid; current PIDs
and logs are managed by systemd. server_cached_ratio3.log is historical.
Further scheduling comparison, 2026-09-27
tune_sglang_v100.py compares FCFS with 8, 16 and 32 running requests, then
LPM with 16. It uses a separate GPU and port, waits for 60 seconds without GPU
processes before each launch, and preserves the existing endpoint. Each profile
runs both input sets with cleared and warm prefix cache and the same FP16
references. It saves server configuration, warmup, all repetitions, logits and
correctness checks. It does not automatically promote a profile.
The benchmark now audits same-premise/instruction option groups on every pass,
including warmup, and retains all logits in .passes.npy. Selecting outside the
set of reference winners fails the benchmark; selecting another exactly tied
winner is recorded separately. The guard reproduces the previous saved-pass
audit: 291 groups, no choices outside the reference winners, two exact ties in
cold T. Singleton and cross-document ranking remain outside this check.
cd ~/storage/sglang-openjev-v100
venv/bin/python app/tune_sglang_v100.py --gpu 1 --port 31001 \
--out "$PWD/results/tuning_20260927"
The output directory must be new. run.json records live phase, server PID and
completed measurements. A daemon invocation is currently recorded in
tuning_20260927.pid; check that process and its log before starting another.
bench_classify_v100.py --exclusive-gpu INDEX --server-pid PID samples physical
GPU process ownership every 0.5 seconds. Only the named server and its descendants
are allowed; PID start times protect against reuse. Competing processes or failed
monitoring invalidate the report and return a nonzero exit status. Sub-sampling
interval competition is still possible; this is sampled evidence, not a GPU lease.
The initial September 27 refresh was contaminated by simultaneous v6 evaluation jobs: warm G took 6.28–7.30 seconds instead of the prior 2.70 seconds. The candidate also received only 192 Mamba state slots because another job allocated GPU memory during startup. That candidate was stopped. These timings must not be used to rank configurations; the next valid comparison must run on an exclusive GPU.