Instructions to use Code4me2/apex-flash-1-abliterated-NVFP4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Code4me2/apex-flash-1-abliterated-NVFP4 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="Code4me2/apex-flash-1-abliterated-NVFP4") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("Code4me2/apex-flash-1-abliterated-NVFP4") model = AutoModelForMultimodalLM.from_pretrained("Code4me2/apex-flash-1-abliterated-NVFP4", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Code4me2/apex-flash-1-abliterated-NVFP4 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Code4me2/apex-flash-1-abliterated-NVFP4" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Code4me2/apex-flash-1-abliterated-NVFP4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/Code4me2/apex-flash-1-abliterated-NVFP4
- SGLang
How to use Code4me2/apex-flash-1-abliterated-NVFP4 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Code4me2/apex-flash-1-abliterated-NVFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Code4me2/apex-flash-1-abliterated-NVFP4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Code4me2/apex-flash-1-abliterated-NVFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Code4me2/apex-flash-1-abliterated-NVFP4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use Code4me2/apex-flash-1-abliterated-NVFP4 with Docker Model Runner:
docker model run hf.co/Code4me2/apex-flash-1-abliterated-NVFP4
apex-flash-1-abliterated-NVFP4
Preliminary model card. Quality evaluation is pending (see Evaluation).
This is an NVFP4 (W4A4) quantization of cantina-security/apex-flash-1-abliterated, in the compressed-tensors format. The checkpoint is about 205 GB, compared with about 643 GB for the BF16 source.
Base model
apex-flash-1-abliterated is an experimental derivative of apex-flash-1, Cantina Security's open-weights security model developed in partnership with Yeta (@yetalabs on X). apex-flash-1 is a reinforcement learning post-train of GLM-5.3-Flash (zai-org/GLM-5.3-Flash) for focused investigations: reading code, using tools, pursuing an exploit, and verifying its effect in a running target. The abliterated variant modifies refusal behavior broadly; the changes are not limited to security tasks.
Further reading from the base authors: Apex Flash release post, Explore Apex.
All credit for the model itself belongs to Cantina Security and Yeta. This repository only contains a quantized copy produced by a third party; it is not affiliated with or endorsed by them.
Intended use
As stated on the base card: authorized security research in environments the researcher owns or has permission to test.
Architecture
| Property | Value |
|---|---|
| Architecture | GLM-5.3-Flash (glm5_next, Glm5NextForConditionalGeneration) |
| Parameters | ~314B total |
| Layers | 45 decoder layers + 1 MTP layer (layer 45) |
| Experts | 288 routed (top-8) + 1 shared |
| Attention | Hybrid: 34 KDA linear-attention layers + 11 MLA/DSA sparse-attention layers (with lightning indexer) |
| Other | Manifold-constrained hyper-connections (mHC); 24-block ViT vision tower |
| Max positions | 1,048,576 |
| BF16 source size | ~643 GB |
Quantization
| Item | Value |
|---|---|
| Method | One-shot PTQ, QuantizationModifier(scheme="NVFP4") via llm-compressor oneshot |
| Format | nvfp4-pack-quantized (compressed-tensors) |
| Tooling | llm-compressor 0.14.0, compressed-tensors 0.19.0, transformers 5.17.0, torch 2.14.0 |
| Hardware | 8x NVIDIA A100-SXM4-80GB (Lambda), torchrun DDP |
| Calibration | moe_calibrate_all_experts=True (every expert sees every calibration token) |
| Checkpoint size | ~205 GB |
NVFP4 is W4A4. Weights are FP4 (E2M1) with group size 16, FP8 (E4M3) block scales and an FP32 per-tensor global scale. Activations use dynamic local FP8 block scales per 16 elements plus a static per-tensor input_global_scale calibrated from data.
What is quantized
Only the routed-expert gate/up/down projections in layers 3-44 are quantized (36,288 projections, the vast majority of parameters). Everything else stays BF16.
| Component | Precision |
|---|---|
| Routed experts, layers 3-44 (gate/up/down) | NVFP4 |
Vision tower, embeddings, lm_head |
BF16 |
| All attention (KDA, MLA/DSA, indexer) | BF16 |
MoE router mlp.gate and e_score_correction_bias |
BF16 |
| Shared experts; dense MLPs of layers 0-2 | BF16 |
| Hyper-connection parameters | BF16 |
| MTP layer 45 | BF16 (spliced from the source; llm-compressor 0.14.0 silently drops it for this architecture) |
Calibration data
512 samples, 501,470 tokens, every sample rendered through the model's own chat template.
| Domain | Source | Target share | Samples | Tokens |
|---|---|---|---|---|
| General chat | mlabonne/open-perfectblend | 30% | 154 | 98,634 |
| Cybersecurity instruction | Trendyol/Trendyol-Cybersecurity-Instruction-Tuning-Dataset | 20% | 102 | 69,059 |
| Vulnerable/fixed code review | CyberNative/Code_Vulnerability_Security_DPO | 15% | 77 | 27,616 |
Tool calling (re-rendered in GLM native <tool_call>/<arg_key> format) |
NousResearch/hermes-function-calling-v1 | 15% | 77 | 101,972 |
Long <think> code reasoning |
nvidia/OpenCodeReasoning | 20% | 102 | 204,189 |
| Total | 512 | 501,470 |
Context length used for calibration: at most 2048 tokens per sample (mean ~980). Calibration did not use long contexts. The impact is expected to be limited but is untested: weights are quantized independently of context length; activation block scales are computed dynamically at runtime; only the per-tensor activation global scale is static, and it was calibrated on sequences of 2048 tokens or fewer; attention and the KV path are BF16 and not quantized here. Long-context (>4K) quality has not been evaluated.
Fixes relative to the upstream llm-compressor GLM-5.3-Flash example
- MTP layer dropped: llm-compressor 0.14.0 only copies MTP tensors for configs with
num_mtp_layers/mtp_num_hidden_layersand anmtp.prefix, whereas this model usesnum_nextn_predict_layersandlayers.45.*. The layer is spliced back from the source (splice_mtp.py). - Target regex:
.*mlp\.experts\..*also matches layer 45, which would make loaders expect NVFP4 weights there. Targets are restricted to layers 3-44 with an explicit ignore for layer 45. - Missing dependencies in the upstream environment:
torchvisionandflash-linear-attention.
Structural verification
Result of verify_ckpt.py against the BF16 source: VERIFY OK. This confirms checkpoint structure only, not model quality.
| Check | Result |
|---|---|
| Tensors | 147,634 |
| Dtypes | bf16: 2,269; fp32: 72,789; uint8: 36,288; float8_e4m3fn: 36,288 |
| Quantized expert projections | 36,288 |
| Scales checked | 108,954 (all finite; global scales > 0) |
| MTP tensors | 889 |
| Size | 205.1 GB |
Evaluation
Pending. A BF16-vs-NVFP4 parity evaluation is in progress: a held-out set disjoint from the calibration data across the same five domains, plus 4096-token long samples and WikiText-2, measuring top-1 agreement, approximate KL divergence and perplexity. Results will be added to this card. No parity, accuracy-retention or benchmark claims are made at this time. The base model's reported results apply to the original apex-flash-1 BF16 checkpoint only.
Usage (untested on this checkpoint)
The compressed-tensors NVFP4 format is auto-detected by vLLM:
vllm serve Code4me2/apex-flash-1-abliterated-NVFP4 --tensor-parallel-size 2
MTP speculative decoding is available via:
--speculative-config '{"method":"mtp","num_speculative_tokens":1}'
Native NVFP4 kernels require Blackwell (SM100/SM12x). See the SM12x caveat below.
Known gaps and limitations
- Calibration length: limited to 2048 tokens per sample; long-context behavior is unevaluated.
- No multimodal calibration: no image or video data was used. The vision tower is BF16, but expert activations on image tokens were not calibrated. Image/video performance is unevaluated, as in the base.
- English-centric calibration: GLM is bilingual (zh/en); the calibration set is not.
- Plain round-to-nearest quantization with min-max observers: no GPTQ/AWQ-style error compensation and no MSE clipping search.
- All-expert calibration:
moe_calibrate_all_expertsexposes each expert to tokens it would not normally receive, which can make activation global scales conservative. - SM120/SM121 serving: stock vLLM 0.30 currently fails for all GLM-5.3-Flash checkpoints on SM120/SM121 (RTX PRO 6000, DGX Spark/GB10) with
pe_dim must be 64 for fp8_ds_mla(vllm-project/vllm#55773, #53963). Community patches exist (Libertai/vllm-sparse-mla-blackwell, tonyd2wild/DGX-Spark). - Inherited limitations: all limitations of the abliterated base apply. Refusal behavior is modified broadly, the variant was not separately evaluated, and image/video performance was not evaluated.
Reproducibility
The quantization recipe is in recipe.yaml. The logs quantize.log and verify.log and the calibration statistics calib_stats.json are included in this repository.
- Downloads last month
- 93
Model tree for Code4me2/apex-flash-1-abliterated-NVFP4
Base model
zai-org/GLM-5.3-Flash