Papers
arxiv:2609.31892

NVAlign: Direct-Gradient Optimization for Non-Verbal Control in Continuous Autoregressive Flow Matching Text-to-Speech

Published on Sep 25
· Submitted by
Qiaolin Wang
on Sep 29
Authors:
,
,
,
,

Abstract

While modern text-to-speech (TTS) systems generate highly natural speech and support inline non-verbal vocalization (NVV) tags, accurate control over these events remains challenging. A key gap is the lack of established post-training methods for non-verbal control in continuous autoregressive flow-matching TTS. To this end, we present NVAlign, a direct-gradient post-training framework for NVV tag-following in this architecture. We first perform supervised fine-tuning (SFT) of TTS models and an NVV-aware automatic speech recognition (NV-ASR) model on NVV-annotated speech, then freeze the NV-ASR model to serve as the reward model for post-training. A two-step gradient surrogate enables efficient reward backpropagation through the flow-matching sampler to jointly update the autoregressive backbone and acoustic flow head. Fidelity penalties and reference-velocity regularization help preserve speaker similarity and speech quality. Results from NVV-SuperBench and human listening evaluations show that NVAlign improves tag-following accuracy over SFT and Flow-GRPO baselines. These findings demonstrate that direct reward-gradient optimization can improve non-verbal control in continuous autoregressive flow-matching TTS. Audio samples are available at https://nvalign.github.io/.

Community

Paper submitter

Ask a TTS model for a [laughs] or a [sigh] mid-sentence, and it often just… skips it.

Inline non-verbal tags are now common, but even after fine-tuning, models drop or confuse these sounds, especially rare ones with little training data. And continuous autoregressive flow-matching TTS (like VoxCPM2 and dots.tts) has no token probabilities, so GRPO-style RL doesn't directly apply.

NVAlign takes a direct route: we fine-tune an ASR model that can hear non-verbal sounds (NV-ASR), freeze it, and backpropagate its judgment straight through the flow-matching sampler with a two-step gradient surrogate, updating both the AR backbone and the flow head.

In blinded human listening tests, NVAlign follows non-verbal tags more reliably than SFT and Flow-GRPO on VoxCPM2, dots.tts, and an English production system, while keeping speaker similarity and naturalness.

The catch: the recognizer can be gamed. Without velocity regularization, its score keeps climbing while the sounds get worse. The constraints are what turn a higher score into sounds people actually hear.

NVAlign overview

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.31892
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.31892 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.31892 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.31892 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.