EdgeIn / README.md
chenjz24's picture
Upload folder using huggingface_hub
e8c2e4c verified
|
Raw History Blame Contribute Delete
2.84 kB
---
library_name: transformers
pipeline_tag: audio-text-to-text
tags:
- edgeinstant
- audio
- text-to-speech
---
# EdgeInstant AudioIn / Thinker / Talker
This model combines the Qwen3-ASR audio encoder, trained AudioIn projection and
Qwen3.5 Thinker, compact native-token Talker, speaker conditions, and Qwen3-TTS
waveform decoder. All weights, token mappings and input processors are included.
Inference uses one Transformers 5.12.1 environment.
Install the dependencies in `requirements.txt`. The package uses custom modeling
code, so load it with `trust_remote_code=True`.
```python
import soundfile as sf
from transformers import AutoModel, AutoProcessor
path = "/path/to/EdgeInstant-1.5b-hf"
processor = AutoProcessor.from_pretrained(path, trust_remote_code=True)
model = AutoModel.from_pretrained(
path, trust_remote_code=True, dtype="auto", device_map="cuda:0",
).eval()
waveform, sample_rate = sf.read("question.wav", dtype="float32")
inputs = processor(audio=waveform, sampling_rate=sample_rate).to("cuda:0")
reply = model.generate_speech(**inputs, max_new_tokens=256, do_sample=False)
print(processor.decode(reply["text_token_ids"], skip_special_tokens=True))
sf.write("reply.wav", reply["audio"].cpu().numpy(), reply["sampling_rate"])
```
Input audio must be mono, 16 kHz and at most 30 seconds. Output audio is 24 kHz.
The bundled speaker is fixed. Keep the loaded dtype: the AudioIn projector and
special-token deltas use FP32; the other weights use BF16.
Text-only input uses `processor(text="你好")`. For transcription, pass
`task="asr"` to the processor and call `model.generate(**inputs)`; generated
sequences include the prompt, following the Hugging Face causal LM convention.
`AutoModelForCausalLM` loads the same class.
`model.synthesize(token_ids, language="chinese")` speaks a sequence of native
Thinker tokens. `language` also accepts `"english"` and `"auto"`.
```python
inputs = processor(text="请用一句话介绍你自己。").to("cuda:0")
for event in model.stream_generate(**inputs):
if event["type"] == "text":
print(processor.decode([event["token_id"]]), end="", flush=True)
elif event["type"] == "audio":
# Consecutive float32 waveform chunks, ready for a 24 kHz audio sink.
audio_chunk = event["audio"].numpy()
```
Speech and streaming generation accept one conversation at a time. Text
generation and audio-conditioned forward support batches. `forward(...,
labels=...)` returns the Thinker language-model loss. Acoustic training remains
in the repository's Talker training scripts. Streaming uses greedy Thinker decoding and the
native Talker state machine, with configurable packet size and decoder context.
The model contains Qwen components and weights; their respective licenses apply.
The adapted Qwen3-TTS decoder code is covered by `LICENSE.codec`.