Text Classification
Transformers
Safetensors
English
nli
cross-encoder
qwen3.5
reranker
image-text-to-text
Instructions to use AlexWortega/openjev with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use AlexWortega/openjev with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="AlexWortega/openjev")# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("AlexWortega/openjev", device_map="auto") - Notebooks
- Google Colab
- Kaggle
|
Download code/serving/V100.md from AlexWortega/openjev: direct link, hf CLI and curl.
- Browser
- Download file 9.61 kB
-
https://huggingface.co/AlexWortega/openjev/resolve/main/code/serving/V100.md
- Command line
-
hf download hf://AlexWortega/openjev/code/serving/V100.md
-
curl -L -o V100.md https://huggingface.co/AlexWortega/openjev/resolve/main/code/serving/V100.md
9.61 kB
| # OpenJEV on SGLang / V100 | |
| This uses `sglang_openjev/` from | |
| https://huggingface.co/AlexWortega/openjev/tree/main/code/sglang_openjev. | |
| `HF_SOURCE.json` records the verified source revision and hashes. The adapter | |
| registers a Qwen3.5 classification head inside SGLang; `/classify` returns three | |
| raw logits in contradiction, entailment, neutral order. | |
| The separate `code/v100` offline scorer is not this HTTP serving engine. | |
| ## Isolated runtime | |
| Work directory on eva01: `~/storage/sglang-openjev-v100` (the large storage mount). | |
| The existing evaluation environment and system CUDA installation are unchanged. | |
| Runtime source: https://github.com/haohervchb/sglang-V100, | |
| revision `dca488908ee4e3f1bc676c3bf5dcd26ff049cfc3`. | |
| FlashInfer source: https://github.com/haohervchb/flashinfer, | |
| revision `c3c40a7b90b792fc59f90f8f55c9e2de9c1b6833`, with the runtime's | |
| `patches/flashinfer-sm70.patch` applied. | |
| Use PyTorch 2.9.1 **cu126**, torchvision 0.24.1, and a private CUDA 12.6.3 | |
| compiler. Official cu128 PyTorch wheels omit SM70. The installer explicitly | |
| checks `torch.cuda.get_arch_list()` and performs an FP16 GPU matmul before | |
| installing the runtime. `prepare_v100_cuda.py ROOT --version 12.6.3` downloads | |
| NVIDIA toolkit components and verifies their manifest hashes. | |
| After cloning the pinned runtime, preparing the private toolkit at | |
| `$OPENJEV_V100_ROOT/cuda126`, and installing the cu126 wheels in its `venv`, run | |
| `install_v100_runtime.sh`. It builds the minimal SM70 `sgl-kernel` and installs | |
| the patched FlashInfer. This is a local source build, not the fork's system | |
| installation script. | |
| The pinned SGLang version also builds a Rust gRPC extension during installation. | |
| Provision Rust in `$OPENJEV_V100_ROOT/rustup` with cargo caches under | |
| `$OPENJEV_V100_ROOT/cargo`; the installer uses these paths. Rust 1.98.1 and | |
| `protoc` 33.6 (under `$OPENJEV_V100_ROOT/protoc`) were provisioned on eva01. | |
| The host's older protoc does not accept `--experimental_allow_proto3_optional`. | |
| Neither a system-wide Rust default nor a CUDA driver update is required. | |
| ## Launch and validation | |
| The training checkpoint lacks image/video processor configs, which SGLang's | |
| Qwen3.5 architecture requires even for text classification. Run | |
| `prepare_v100_model.py CHECKPOINT $OPENJEV_V100_ROOT/model` to create a symlink | |
| overlay. It adds only the two processor configs from the pinned base | |
| `Qwen/Qwen3.5-4B` revision, checks their hashes, and preserves the checkpoint's | |
| weights, tokenizer and model config. | |
| `serve_sglang_v100.sh CHECKPOINT 31000` launches on GPU 0 and binds localhost. | |
| Set `CUDA_VISIBLE_DEVICES` to choose another free GPU. Initial settings use FP16, | |
| TileLang V100 attention and GDN prefill, an 8192-token server context, and disabled CUDA | |
| graphs. Pass `--disable-radix-cache` for the initial uncached baseline. | |
| The benchmark still truncates at 4096 tokens, matching the reference; this | |
| runtime rejects inputs exactly equal to the configured server context length. | |
| The launcher selects GCC 10 for C++20 JIT headers and PyTorch sampling (the | |
| classifier does not sample output tokens). Since the September 28 deployment, | |
| it caps running requests at 16 and uses LPM scheduling by default; override with | |
| `MAX_RUNNING_REQUESTS` and `SCHEDULE_POLICY`. The original September 26 results | |
| below used 8 running requests and FCFS. Prefix caching uses | |
| `extra_buffer` and a Mamba-to-KV memory ratio of 3. This provides 503 recurrent | |
| state slots and 135,008 KV tokens at the same 80% GPU memory budget. The original | |
| ratio of 0.9 provided 317 slots and suffered cache eviction on the larger set. | |
| `--disable-radix-cache` selects the uncached `no_buffer` baseline automatically. | |
| All JIT caches use the large storage mount. | |
| Apply `patches/dynamic-paged.patch` to the pinned runtime before launching. | |
| It makes the batch, physical-page count and block-table width runtime dimensions | |
| of the existing TileLang paged attention kernel. Static specialization on these | |
| shapes caused repeated compilation during real request batches. The math and | |
| tiling remain unchanged. `SGLANG_V100_DYNAMIC_PAGED=0` restores static selection. | |
| `check_dynamic_paged.py` checks ragged causal attention with and without cached | |
| prefixes against a dense FP32 reference and asserts that changing all three | |
| dimensions reuses one kernel. The GPU check passed with maximum absolute error | |
| 0.001073; this kernel check alone does not establish full-model agreement. | |
| `bench_classify_v100.py` submits exact token IDs to `/classify` and compares the | |
| responses to saved reference logits. It records input and token hashes, | |
| entailment probability differences, NLI argmax changes, HTTP wall time and | |
| per-batch latencies. Warmup is excluded. `--cold-cache` flushes SGLang's cache | |
| before each iteration. Cold and warm-cache results must be reported separately. | |
| ## Verified result, 2026-09-26 | |
| On eva01, one V100-SXM2-32GB, FP16, HTTP `/classify`, batches of 32 inputs with | |
| four client workers and eight running sequences. Each cell is the median of | |
| three measured repetitions after one initial pass: | |
| | Pairs | Prefix cache disabled | Prefix cache cleared before each pass | Repeated inputs in warm prefix cache | | |
| |---|---:|---:|---:| | |
| | G: 303 | 24.928 s | 13.417 s | 2.697 s | | |
| | T: 564 | 47.618 s | 29.635 s | 4.732 s | | |
| The cache therefore improves these fresh-set passes by 1.86x / 1.61x and exact | |
| repeated-set passes by 9.24x / 10.06x. Warm-prefix numbers are not fresh-request | |
| throughput. Changing the Mamba memory ratio from 0.9 to 3 reduced the larger | |
| set's warm-cache median from 28.801 s to 4.732 s. | |
| All 867 NLI argmax labels agree with an independently generated, unmodified | |
| Transformers FP16 reference. Input/token hashes were checked. The largest | |
| entailment probability difference across the selected cached runs was 0.012851. | |
| In the saved final pass of each profile, among 291 same-premise/instruction | |
| option groups (793 rows), no strict reference winner changed. Two cold-cache | |
| choices changed where the reference probabilities were exactly tied. This is agreement testing, not a held-out Decision Index | |
| accuracy result; singleton and cross-document aggregation are not covered by | |
| that grouped check. | |
| First-use JIT remains expensive, especially after changing the GDN state-pool | |
| size: the selected profile's first G pass took 202.709 s. Initial passes are | |
| excluded from medians; additional JIT outliers in measured repetitions are | |
| retained, including 45.085 s for uncached G and 15.064 s for warm-cache T. | |
| These are controlled throughput measurements, not production p99 guarantees. | |
| Receipts and raw arrays: `results_v100/sglang/summary.json`, | |
| `runtime_receipt.json`, `ratio3_server_config.json`, individual benchmark JSONs, | |
| and reference/candidate `.npy` files. The remote directory is | |
| `~/storage/sglang-openjev-v100/results`. | |
| The endpoint is `http://127.0.0.1:31000/classify` on eva01. Since September 28 | |
| it is a systemd socket-activated gateway over two replicas; see | |
| [`DEPLOYMENT-eva01.md`](DEPLOYMENT-eva01.md). Do not launch another standalone | |
| server on that port. For a separate development instance on a free GPU/port: | |
| ```bash | |
| CUDA_VISIBLE_DEVICES=3 bash ~/storage/sglang-openjev-v100/app/serve_sglang_v100.sh \ | |
| ~/storage/sglang-openjev-v100/model 31003 | |
| ``` | |
| The model overlay defaults to `~/storage/sglang-openjev-v100/model`. The old | |
| standalone PID record is archived as `server.legacy-20260926.pid`; current PIDs | |
| and logs are managed by systemd. `server_cached_ratio3.log` is historical. | |
| ## Further scheduling comparison, 2026-09-27 | |
| `tune_sglang_v100.py` compares FCFS with 8, 16 and 32 running requests, then | |
| LPM with 16. It uses a separate GPU and port, waits for 60 seconds without GPU | |
| processes before each launch, and preserves the existing endpoint. Each profile | |
| runs both input sets with cleared and warm prefix cache and the same FP16 | |
| references. It saves server configuration, warmup, all repetitions, logits and | |
| correctness checks. It does not automatically promote a profile. | |
| The benchmark now audits same-premise/instruction option groups on every pass, | |
| including warmup, and retains all logits in `.passes.npy`. Selecting outside the | |
| set of reference winners fails the benchmark; selecting another exactly tied | |
| winner is recorded separately. The guard reproduces the previous saved-pass | |
| audit: 291 groups, no choices outside the reference winners, two exact ties in | |
| cold T. Singleton and cross-document ranking remain outside this check. | |
| ```bash | |
| cd ~/storage/sglang-openjev-v100 | |
| venv/bin/python app/tune_sglang_v100.py --gpu 1 --port 31001 \ | |
| --out "$PWD/results/tuning_20260927" | |
| ``` | |
| The output directory must be new. `run.json` records live phase, server PID and | |
| completed measurements. A daemon invocation is currently recorded in | |
| `tuning_20260927.pid`; check that process and its log before starting another. | |
| `bench_classify_v100.py --exclusive-gpu INDEX --server-pid PID` samples physical | |
| GPU process ownership every 0.5 seconds. Only the named server and its descendants | |
| are allowed; PID start times protect against reuse. Competing processes or failed | |
| monitoring invalidate the report and return a nonzero exit status. Sub-sampling | |
| interval competition is still possible; this is sampled evidence, not a GPU lease. | |
| The initial September 27 refresh was contaminated by simultaneous v6 evaluation | |
| jobs: warm G took 6.28–7.30 seconds instead of the prior 2.70 seconds. The candidate | |
| also received only 192 Mamba state slots because another job allocated GPU memory | |
| during startup. That candidate was stopped. These timings must not be used to | |
| rank configurations; the next valid comparison must run on an exclusive GPU. | |