apex-flash-1-abliterated-NVFP4

Preliminary model card. Quality evaluation is pending (see Evaluation).

This is an NVFP4 (W4A4) quantization of cantina-security/apex-flash-1-abliterated, in the compressed-tensors format. The checkpoint is about 205 GB, compared with about 643 GB for the BF16 source.

Base model

apex-flash-1-abliterated is an experimental derivative of apex-flash-1, Cantina Security's open-weights security model developed in partnership with Yeta (@yetalabs on X). apex-flash-1 is a reinforcement learning post-train of GLM-5.3-Flash (zai-org/GLM-5.3-Flash) for focused investigations: reading code, using tools, pursuing an exploit, and verifying its effect in a running target. The abliterated variant modifies refusal behavior broadly; the changes are not limited to security tasks.

Further reading from the base authors: Apex Flash release post, Explore Apex.

All credit for the model itself belongs to Cantina Security and Yeta. This repository only contains a quantized copy produced by a third party; it is not affiliated with or endorsed by them.

Intended use

As stated on the base card: authorized security research in environments the researcher owns or has permission to test.

Architecture

Property Value
Architecture GLM-5.3-Flash (glm5_next, Glm5NextForConditionalGeneration)
Parameters ~314B total
Layers 45 decoder layers + 1 MTP layer (layer 45)
Experts 288 routed (top-8) + 1 shared
Attention Hybrid: 34 KDA linear-attention layers + 11 MLA/DSA sparse-attention layers (with lightning indexer)
Other Manifold-constrained hyper-connections (mHC); 24-block ViT vision tower
Max positions 1,048,576
BF16 source size ~643 GB

Quantization

Item Value
Method One-shot PTQ, QuantizationModifier(scheme="NVFP4") via llm-compressor oneshot
Format nvfp4-pack-quantized (compressed-tensors)
Tooling llm-compressor 0.14.0, compressed-tensors 0.19.0, transformers 5.17.0, torch 2.14.0
Hardware 8x NVIDIA A100-SXM4-80GB (Lambda), torchrun DDP
Calibration moe_calibrate_all_experts=True (every expert sees every calibration token)
Checkpoint size ~205 GB

NVFP4 is W4A4. Weights are FP4 (E2M1) with group size 16, FP8 (E4M3) block scales and an FP32 per-tensor global scale. Activations use dynamic local FP8 block scales per 16 elements plus a static per-tensor input_global_scale calibrated from data.

What is quantized

Only the routed-expert gate/up/down projections in layers 3-44 are quantized (36,288 projections, the vast majority of parameters). Everything else stays BF16.

Component Precision
Routed experts, layers 3-44 (gate/up/down) NVFP4
Vision tower, embeddings, lm_head BF16
All attention (KDA, MLA/DSA, indexer) BF16
MoE router mlp.gate and e_score_correction_bias BF16
Shared experts; dense MLPs of layers 0-2 BF16
Hyper-connection parameters BF16
MTP layer 45 BF16 (spliced from the source; llm-compressor 0.14.0 silently drops it for this architecture)

Calibration data

512 samples, 501,470 tokens, every sample rendered through the model's own chat template.

Domain Source Target share Samples Tokens
General chat mlabonne/open-perfectblend 30% 154 98,634
Cybersecurity instruction Trendyol/Trendyol-Cybersecurity-Instruction-Tuning-Dataset 20% 102 69,059
Vulnerable/fixed code review CyberNative/Code_Vulnerability_Security_DPO 15% 77 27,616
Tool calling (re-rendered in GLM native <tool_call>/<arg_key> format) NousResearch/hermes-function-calling-v1 15% 77 101,972
Long <think> code reasoning nvidia/OpenCodeReasoning 20% 102 204,189
Total 512 501,470

Context length used for calibration: at most 2048 tokens per sample (mean ~980). Calibration did not use long contexts. The impact is expected to be limited but is untested: weights are quantized independently of context length; activation block scales are computed dynamically at runtime; only the per-tensor activation global scale is static, and it was calibrated on sequences of 2048 tokens or fewer; attention and the KV path are BF16 and not quantized here. Long-context (>4K) quality has not been evaluated.

Fixes relative to the upstream llm-compressor GLM-5.3-Flash example

  1. MTP layer dropped: llm-compressor 0.14.0 only copies MTP tensors for configs with num_mtp_layers/mtp_num_hidden_layers and an mtp. prefix, whereas this model uses num_nextn_predict_layers and layers.45.*. The layer is spliced back from the source (splice_mtp.py).
  2. Target regex: .*mlp\.experts\..* also matches layer 45, which would make loaders expect NVFP4 weights there. Targets are restricted to layers 3-44 with an explicit ignore for layer 45.
  3. Missing dependencies in the upstream environment: torchvision and flash-linear-attention.

Structural verification

Result of verify_ckpt.py against the BF16 source: VERIFY OK. This confirms checkpoint structure only, not model quality.

Check Result
Tensors 147,634
Dtypes bf16: 2,269; fp32: 72,789; uint8: 36,288; float8_e4m3fn: 36,288
Quantized expert projections 36,288
Scales checked 108,954 (all finite; global scales > 0)
MTP tensors 889
Size 205.1 GB

Evaluation

Pending. A BF16-vs-NVFP4 parity evaluation is in progress: a held-out set disjoint from the calibration data across the same five domains, plus 4096-token long samples and WikiText-2, measuring top-1 agreement, approximate KL divergence and perplexity. Results will be added to this card. No parity, accuracy-retention or benchmark claims are made at this time. The base model's reported results apply to the original apex-flash-1 BF16 checkpoint only.

Usage (untested on this checkpoint)

The compressed-tensors NVFP4 format is auto-detected by vLLM:

vllm serve Code4me2/apex-flash-1-abliterated-NVFP4 --tensor-parallel-size 2

MTP speculative decoding is available via:

--speculative-config '{"method":"mtp","num_speculative_tokens":1}'

Native NVFP4 kernels require Blackwell (SM100/SM12x). See the SM12x caveat below.

Known gaps and limitations

  • Calibration length: limited to 2048 tokens per sample; long-context behavior is unevaluated.
  • No multimodal calibration: no image or video data was used. The vision tower is BF16, but expert activations on image tokens were not calibrated. Image/video performance is unevaluated, as in the base.
  • English-centric calibration: GLM is bilingual (zh/en); the calibration set is not.
  • Plain round-to-nearest quantization with min-max observers: no GPTQ/AWQ-style error compensation and no MSE clipping search.
  • All-expert calibration: moe_calibrate_all_experts exposes each expert to tokens it would not normally receive, which can make activation global scales conservative.
  • SM120/SM121 serving: stock vLLM 0.30 currently fails for all GLM-5.3-Flash checkpoints on SM120/SM121 (RTX PRO 6000, DGX Spark/GB10) with pe_dim must be 64 for fp8_ds_mla (vllm-project/vllm#55773, #53963). Community patches exist (Libertai/vllm-sparse-mla-blackwell, tonyd2wild/DGX-Spark).
  • Inherited limitations: all limitations of the abliterated base apply. Refusal behavior is modified broadly, the variant was not separately evaluated, and image/video performance was not evaluated.

Reproducibility

The quantization recipe is in recipe.yaml. The logs quantize.log and verify.log and the calibration statistics calib_stats.json are included in this repository.

Downloads last month
93
Safetensors
Model size
181B params
Tensor type
F32
·
BF16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Code4me2/apex-flash-1-abliterated-NVFP4