Prism Q6 (blocks-q6-v1)

A first-attempt, zero-optimization 6-bit quantization of Tencent Hunyuan's Prism image-to-video-and-audio model. Produced and tested on a single RTX 5090. Shared for experimentation and hardware comparisons β€” quality is experimental (see Known limitations).

Format prism-custom-q6 β€” block Linear weights packed at 6 bits, group size 64, FP32 scales; all other weights retained
Size 27.1 GB across 52 Safetensors shards (source checkpoint: 65.3 GB BF16)
Source weights FrancisRing/Prism @ a4ea8f5f6a5d3b71df04a048ed05c4e916f8595e
Source code Tencent-Hunyuan/Prism @ eb943ff6079a8ea261247e965f9556703109c1c0
Manifest SHA-256 74eb0f622587458b4eb7c4f18d9de4ea95ee56af85141cd3824cc778602963e1

Announcement: https://x.com/atomtanstudio/status/2107415066299809977

Measured on RTX 5090 (single GPU, CPU offload)

  • 480p, 24 fps, 5 s clip (121 frames), 50 steps: **30 minutes**, runs in 16 GB VRAM (peak CUDA allocation 11.7 GB / reservation 13.3 GB), peak process RSS ~44 GB host RAM.
  • Native 720p, 61 frames, 50 steps: ~31-34 minutes.

Repository layout

  • manifest.json β€” versioned Q6 manifest: shard SHA-256s, tensor map, per-module quantization records, provenance.
  • model-000NN.safetensors β€” the 52 quantized shards.
  • receipts/ β€” per-shard conversion receipts produced during quantization.
  • loader/prism_quant/ β€” the loader/conversion package (quant_loader.py verifies every shard hash before loading; quant.py implements the Q6 format and Q6Linear; runtime.py is the research runtime integration it was validated with).
  • loader/convert.py β€” resumable script that regenerates these shards from the official BF16 checkpoint.

Using it

The format is custom Q6, not GGUF β€” stock Prism checkpoints loaders will not read it directly. Use prism_quant with the Prism inference code:

from prism_quant.quant_loader import read_manifest, verify_shards

path, manifest = read_manifest("Prism-Q6")   # validates format + shard hashes
verify_shards(path, manifest)                 # full SHA-256 verification

Requires PyTorch and safetensors. Shards are loaded block-wise with CPU offload; in testing the launcher required at least 45 GiB free host RAM and 24 GiB free GPU memory before starting a run.

Known limitations

This is the alpha blocks-q6-v1 artifact: a first attempt with zero optimizations. In completed test clips: restricted articulation, mouth and teeth artifacts, lip sync unverified, and audio that can sound distorted. Exact ASR transcription of prompts was verified, but that is not a quality approval. Treat this as an experimental preview, not a finished preset.

License

MIT, inherited from the base Prism release. This repository contains a quantized derivative of the FrancisRing/Prism weights and adds no additional restrictions.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for atomtanstudio/Prism-Q6

Finetuned
(3)
this model