Spaces:
Running on Zero
Running on Zero
File size: 1,357 Bytes
695f3b6 bc3f111 695f3b6 bc3f111 695f3b6 bc3f111 695f3b6 bc3f111 ed72f18 1a6347c 7973967 ed72f18 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 | ---
title: Embedding Quantization CUDA Control
emoji: 🔬
colorFrom: blue
colorTo: yellow
sdk: gradio
sdk_version: 6.5.1
app_file: app.py
pinned: false
license: apache-2.0
---
# Embedding Quantization CUDA Control
A bounded ZeroGPU control lane for the matched local MLX embedding
quantization experiment. It reproduces the frozen Qwen3-Embedding-0.6B BF16
vectors on CUDA and compares them with the saved MLX BF16 vectors. CUDA results
do not reproduce MLX/Metal performance.
The second milestone adds CUDA-native bitsandbytes INT8 and NF4 controls.
They are deliberately reported as separate quantizers and are not treated as
equivalents of MLX Q, oQ, or oQe formats.
Matched BF16, INT8, and NF4 controls are available for
Qwen3-Embedding-0.6B, GTE-Qwen2-1.5B, and Qwen3-Embedding-8B. Quantized CUDA
vectors are evaluated against their same-loading-path CUDA BF16 reference.
GTE-Qwen2-1.5B failed the loading-path and MLX/CUDA parity checks. Its CUDA
quantized results remain valid only as same-path comparisons against the direct
CUDA BF16 reference and must not be mixed into cross-runtime vector trends.
Frozen inputs, saved vectors, raw metrics, locked model revisions, and the
portable client are published separately in
[`TiGa-RCE/embedding-quant-repro-2026-07-28`](https://huggingface.co/datasets/TiGa-RCE/embedding-quant-repro-2026-07-28).
|