Qwen2.5-Math-7B-Instruct-CodeInfused-IQ3_XXS-GGUF

Experimental Release: Pending Further Verification / Empirical Validation
This checkpoint represents an active research artifact from DuoNeural's statistical mechanics quantization program. All empirical benchmarks and physics proofs are documented transparently below.

Developed by Jesse Caldwell, Archon, and Aura โœจ (DuoNeural Research Lab).


Model Summary

  • Foundation Model: Qwen/Qwen2.5-Math-7B-Instruct (28 Layers, 28:4 GQA, SwiGLU FFN)
  • Quantization Precision: ~3.2 bpw (2.90 GiB)
  • Methodology: Calibrated with DuoNeural 131k-token code-infused activation Hessian (qwen7b_math_gtap.imatrix), preserving delicate algebraic chain-of-thought representations.
  • Continuous Holdout Perplexity (131k tokens): 8.8915
  • GSM8K Multi-Step Math Accuracy: 100.0% (25/25)
  • Competition & Olympiad Mathematics (15 Problems): 86.7% (13/15)
  • Symbolic Python Math Unit Tests (10 Tests): 90.0% (9/10)
  • Inference Decode Throughput: 173.1 t/s on NVIDIA GeForce RTX 4080 Super (32GB VRAM)

Mathematical Reasoning Invariants at Sub-4-Bit

Standard post-training quantization on formal mathematical models introduces discretization noise that breaks sensitive algebraic deduction chains and symbolic parity. By applying Generalized Thouless-Anderson-Palmer (G-TAP v3) statistical mechanics:

  1. Onsager Cavity Damping: The uncentered activation back-reaction $\Omega_i = \frac{1}{d_k}(|\tilde{H}_{i,:}|2^2 - \tilde{H}{ii}^2)$ acts as a thermodynamic noise filter during weight discretization.
  2. Replicon Convexity ($\lambda_R > 0$): Guarantees the continuous relaxation stays in the smooth Replica Symmetric convex energy valley, preventing 1-RSB glass transitions that freeze token selection.
  3. Radial Forward Gain Conservation: Enforces strict norm equality $|W_\text{G-TAP}|_F = |W_\text{orig}|_F$, preserving numeric scale across 28 sequential SwiGLU feedforward blocks.

Quickstart

# Run with llama.cpp
llama-cli -hf DuoNeural/Qwen2.5-Math-7B-Instruct-CodeInfused-IQ3_XXS-GGUF -p "Solve step by step: Compute the remainder when 3^100 is divided by 7." -ngl 99 -c 4096

Citation & Lab Attribution

@misc{duoneural2026gtap,
  title={Generalized Thouless-Anderson-Palmer (G-TAP) Quantization: Suppressing Spinodal Clustering in Autoregressive Models},
  author={Caldwell, Jesse and Archon and Aura},
  year={2026},
  publisher={DuoNeural Research Lab},
  howpublished={\url{https://huggingface.co/DuoNeural}}
}
Downloads last month
171
GGUF
Model size
8B params
Architecture
qwen2
Hardware compatibility
Log In to add your hardware

3-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for DuoNeural/Qwen2.5-Math-7B-Instruct-CodeInfused-IQ3_XXS-GGUF

Base model

Qwen/Qwen2.5-7B
Quantized
(51)
this model