Audio-Text-to-Text
Transformers
Safetensors
qwen2_5_omni
text-to-audio
audio
audio-question-answering
audio-classification
candidate-scoring
Instructions to use shlv/AudioJev with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use shlv/AudioJev with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("shlv/AudioJev") model = AutoModelForMultimodalLM.from_pretrained("shlv/AudioJev", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Add links to paper and code
#1
by nielsr HF Staff - opened
README.md
CHANGED
|
@@ -1,10 +1,9 @@
|
|
| 1 |
---
|
|
|
|
|
|
|
| 2 |
license: other
|
| 3 |
license_name: qwen-research
|
| 4 |
license_link: LICENSE
|
| 5 |
-
base_model: Qwen/Qwen2.5-Omni-3B
|
| 6 |
-
base_model_relation: finetune
|
| 7 |
-
library_name: transformers
|
| 8 |
pipeline_tag: audio-text-to-text
|
| 9 |
tags:
|
| 10 |
- audio
|
|
@@ -13,12 +12,17 @@ tags:
|
|
| 13 |
- candidate-scoring
|
| 14 |
- qwen2_5_omni
|
| 15 |
- safetensors
|
|
|
|
| 16 |
---
|
| 17 |
|
| 18 |
# AudioJev
|
| 19 |
|
| 20 |
**Direct audio decisions with order-calibrated candidate probabilities. Built with Qwen.**
|
| 21 |
|
|
|
|
|
|
|
|
|
|
|
|
|
| 22 |
AudioJev takes an audio waveform, a natural-language question, and a list of candidate descriptions. It returns a probability distribution over those candidates using the next-token logits of their position labels. This release contains the **random-derangement SKL (RD-SKL) model with lambda = 0.5 and training seed = 20261001**, fine-tuned from [Qwen2.5-Omni-3B](https://huggingface.co/Qwen/Qwen2.5-Omni-3B).
|
| 23 |
|
| 24 |
## Checkpoint
|
|
@@ -78,4 +82,4 @@ The released model is intended for audio-conditioned classification and question
|
|
| 78 |
|
| 79 |
This model is derived from Qwen2.5-Omni-3B and carries the **Qwen Research License Agreement**, supplied in [LICENSE](LICENSE). The upstream agreement grants use for research or evaluation and requires a separate license from Alibaba Cloud for commercial use. See [Notice](Notice) for attribution and a description of modified files.
|
| 80 |
|
| 81 |
-
**Built with Qwen.** The original pretrained weights were modified by AudioJev fine-tuning with RD-SKL, lambda 0.5, seed 20261001. This release does not grant rights beyond the applicable upstream terms.
|
|
|
|
| 1 |
---
|
| 2 |
+
base_model: Qwen/Qwen2.5-Omni-3B
|
| 3 |
+
library_name: transformers
|
| 4 |
license: other
|
| 5 |
license_name: qwen-research
|
| 6 |
license_link: LICENSE
|
|
|
|
|
|
|
|
|
|
| 7 |
pipeline_tag: audio-text-to-text
|
| 8 |
tags:
|
| 9 |
- audio
|
|
|
|
| 12 |
- candidate-scoring
|
| 13 |
- qwen2_5_omni
|
| 14 |
- safetensors
|
| 15 |
+
base_model_relation: finetune
|
| 16 |
---
|
| 17 |
|
| 18 |
# AudioJev
|
| 19 |
|
| 20 |
**Direct audio decisions with order-calibrated candidate probabilities. Built with Qwen.**
|
| 21 |
|
| 22 |
+
Paper: https://huggingface.co/papers/2610.01293
|
| 23 |
+
|
| 24 |
+
Code: https://github.com/SihanLv/AudioJev-Inference
|
| 25 |
+
|
| 26 |
AudioJev takes an audio waveform, a natural-language question, and a list of candidate descriptions. It returns a probability distribution over those candidates using the next-token logits of their position labels. This release contains the **random-derangement SKL (RD-SKL) model with lambda = 0.5 and training seed = 20261001**, fine-tuned from [Qwen2.5-Omni-3B](https://huggingface.co/Qwen/Qwen2.5-Omni-3B).
|
| 27 |
|
| 28 |
## Checkpoint
|
|
|
|
| 82 |
|
| 83 |
This model is derived from Qwen2.5-Omni-3B and carries the **Qwen Research License Agreement**, supplied in [LICENSE](LICENSE). The upstream agreement grants use for research or evaluation and requires a separate license from Alibaba Cloud for commercial use. See [Notice](Notice) for attribution and a description of modified files.
|
| 84 |
|
| 85 |
+
**Built with Qwen.** The original pretrained weights were modified by AudioJev fine-tuning with RD-SKL, lambda 0.5, seed 20261001. This release does not grant rights beyond the applicable upstream terms.
|