BasisServe-CALS Checkpoint Collection

This private repository contains factor-only research checkpoints produced by BasisServe-CALS. It publishes both the PaLU baseline used in our comparisons and checkpoints for our C1 communication-aware low-rank methods. It also includes one post-RoPE KQ-SVD experiment.

The collection does not redistribute full base-model weights. Download the matching base model and revision from its original Hugging Face repository, then install these factors with the BasisServe-CALS loaders. These directories are not directly loadable with AutoModelForCausalLM.from_pretrained().

Release contents

Family Base model Released variants Compressed object
PaLU Llama-3.1-8B V25 uniform, V25 Fisher Value projection/cache; Key remains dense
PaLU Llama-3.1-70B V25 uniform, V25 Fisher Value projection/cache; Key remains dense
PaLU Qwen3-8B-Base V25 uniform, V25 Fisher, V64 Fisher Value projection/cache; Key remains dense
PaLU Qwen3-32B V25 uniform, V25 Fisher, V64 Fisher Value projection/cache; Key remains dense
Attention C1 Qwen3-8B-Base uniform V64 ALS5, per-layer Global-KL avg64 ALS5 Value cache and attention-output TP communication
Attention C1 Qwen3-32B uniform V64 ALS5, per-layer Global-KL avg64 ALS5 Value cache and attention-output TP communication
Attention-output C1 DeepSeek-V2-Lite per-layer Global-KL avg128 ALS5 Dense attention-output AllGather; MLA latent KV is unchanged
KQ-SVD Qwen3-8B-Base post-RoPE K/Q rank64 Key cache; intended to pair with Qwen3-8B C1 V64
MLP C1 Qwen3-8B-Base per-layer Global-KL mean-DP avg2560 Row-parallel down_proj communication

Compression semantics

V25 means 25% reduction of the Value width, typically rank 96 from a physical head dimension of 128. Since Key remains dense, that is 12.5% total KV-cache reduction. V64 uses average Value rank 64, giving approximately 50% Value reduction and 25% total KV-cache reduction.

The Qwen3 Attention C1 V64 checkpoints keep Key dense, jointly fit all physical Value sources and the complete output decoder, and close the decoder after each ALS sweep. Their calibration uses 256 C4 fit documents and 64 held-out C4 documents, with all 2048 token positions contributing through streamed sufficient statistics.

The DeepSeek-V2-Lite C1 checkpoint does not compress MLA latent KV or alter the KV cache. It compresses only the post-attention tensor-parallel AllGather. Its average source rank is 128, corresponding to 50% attention-output communication reduction relative to dense AllGather. Its recorded WikiText-2 perplexity is 6.69518; the uniform rank-128 C1 control is 6.65635.

The MLP C1 mean-DP checkpoint has average rank 2560 and a theoretical 37.5% communication reduction for the row-parallel MLP output. The exported factors are the independently selected mean_dp schedule; the unused UCB factor copy is intentionally omitted.

The KQ-SVD release contains only deployment factors and its result manifest. The 144 MiB post-RoPE Gram/statistics artifact is intentionally omitted. When paired with Qwen3-8B C1 V64, KQ-SVD rank64 retains 50% of the full dense KV width; its measured WikiText-2 perplexity was 12.18063 versus 8.42927 for the same C1 V64 checkpoint with dense Key.

Directory layout

checkpoints/
  palu/
    llama31_8b_palu_m_v25/
    llama31_8b_palu_m_v25_fisher/
    llama31_70b_palu_m_v25/
    llama31_70b_palu_m_v25_fisher/
    qwen3_8b_palu_m_v25/
    qwen3_8b_palu_m_v25_fisher/
    qwen3_8b_palu_m_r64_fisher_c4_256x2048/
    qwen3_32b_palu_m_v25/
    qwen3_32b_palu_m_v25_fisher/
    qwen3_32b_palu_m_r64_fisher_c4_256x2048/
  attention_c1/
    qwen3_8b_uniform_v64_als5/
    qwen3_8b_layer_global_kl_avg64_als5/
    qwen3_32b_uniform_v64_als5/
    qwen3_32b_layer_global_kl_avg64_als5/
    deepseek_v2_lite_layer_global_kl_avg128_als5/
  kq_svd/
    qwen3_8b_post_rope_r64/
  mlp_c1/
    qwen3_8b_layer_global_kl_avg2560_mean_dp/

Artifact formats

  • PaLU directories contain one BF16 Safetensors artifact plus manifest.json.
  • Qwen3 Attention C1 directories contain one BF16 Safetensors file per decoder layer, layer diagnostics, and an authenticated result/summary manifest.
  • DeepSeek-V2-Lite contains selected per-layer factors, result.json, and a human-readable rank schedule.
  • KQ-SVD contains factors.safetensors and result.json.
  • MLP C1 contains one selected BF16 factor file per layer, its manifest, the Global-KL result, and the schedule summary.

Each manifest records the exact base-model revision and configuration hashes. Absolute path fields record the original generation environment; consumers should use the adjacent Hugging Face repository/revision fields instead.

Code and reproducibility

The corresponding implementation and evaluation workflows are in BasisServe-CALS, including:

  • PaLU whitening, Fisher allocation, checkpoint construction, and evaluation;
  • streamed C4 covariance collection and decoder-closed ALS;
  • terminal-logits Global-KL profiling and exact-budget dynamic programming;
  • Qwen3 and DeepSeek tensor-parallel runtimes;
  • post-RoPE KQ-SVD fitting and evaluation;
  • MLP C1 fitting and per-layer Global-KL allocation.

This release was prepared from Git commit e9252e7.

License and base-model terms

These files are derived factor checkpoints. Use of each checkpoint remains subject to the license and acceptable-use terms of its corresponding base model (Meta Llama, Qwen, or DeepSeek) as well as the BasisServe-CALS code license. No base-model weights are included here.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for alexz949/BasisServe-CALS

Base model

Qwen/Qwen3-32B
Finetuned
(602)
this model