Audio-Codec-LLM native audio projector (v3 smoke)

audio_projector.pt is the trainable Whisper-to-Qwen continuous prefix projector from the v3 RTX 4090 smoke run. Whisper-small and Qwen2.5-1.5B-Instruct were loaded as real frozen base models; two optimizer steps reduced the smoke loss from 3.31035 to 2.70892 (mean 3.00964), with 3,859.58 MiB peak VRAM.

This artifact proves the tensor path, backward pass, optimizer and serialization. Two steps are not a convergence experiment and do not establish enhancement quality. See run_manifest.json and the source repository.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for jatshi/Audio-Codec-LLM-Native-Audio-Projector-v3

Finetuned
(1750)
this model