Audio Classification
Transformers
Safetensors
multilingual
wav2vec2-dual-hypersphere
audio-deepfake
deepfake-detection
deepfake
voice-cloning
anti-spoofing
asvspoof
wav2vec2
speech
audio
synthetic-voice
voice-conversion
tts-detection
trust-and-safety
security
SoTA
Modotte
custom_code
Instructions to use Modotte/AIRealNet-Audio with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Modotte/AIRealNet-Audio with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("audio-classification", model="Modotte/AIRealNet-Audio", trust_remote_code=True)# Load model directly from transformers import AutoModelForAudioClassification model = AutoModelForAudioClassification.from_pretrained("Modotte/AIRealNet-Audio", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
|
Download README.md from Modotte/AIRealNet-Audio: direct link, hf CLI and curl.
- Browser
- Download file 10.8 kB
-
https://huggingface.co/Modotte/AIRealNet-Audio/resolve/main/README.md
- Command line
-
hf download hf://Modotte/AIRealNet-Audio/README.md
-
curl -L -o README.md https://huggingface.co/Modotte/AIRealNet-Audio/resolve/main/README.md
10.8 kB
| license: mit | |
| pipeline_tag: audio-classification | |
| library_name: transformers | |
| language: | |
| - multilingual | |
| tags: | |
| - audio-deepfake | |
| - deepfake-detection | |
| - deepfake | |
| - voice-cloning | |
| - anti-spoofing | |
| - asvspoof | |
| - wav2vec2 | |
| - speech | |
| - audio | |
| - synthetic-voice | |
| - voice-conversion | |
| - tts-detection | |
| - trust-and-safety | |
| - security | |
| - SoTA | |
| - Modotte | |
| inference: | |
| parameters: | |
| function_to_apply: "softmax" | |
| ## Modotte | |
| <p align="center"> | |
| <img | |
| src="https://huggingface.co/Modotte/AIRealNet-Audio/resolve/main/assets/banner.jpeg" | |
| alt="AIRealNet-Audio Banner" | |
| width="95%" | |
| style="border-radius:15px;" | |
| /> | |
| </p> | |
| - [Live Demo](https://huggingface.co/spaces/sujalrajpoot/AIRealNetAudio) | |
| ## Overview | |
| > This is the future iteration of [AIRealNet](https://huggingface.co/Modotte/AIRealNet) | |
| In an era of rapidly advancing AI-generated speech, voice cloning, and audio deepfakes, the need for reliable detection tools has never been higher. **AIRealNet-Audio** is a binary audio classifier designed to distinguish AI-generated / spoofed audio from real human speech. | |
| It is built on a Wav2Vec-based audio encoder that processes audio in fixed **12-second chunks**. The model uses a default decision threshold of **50%**, which can be adjusted based on your use case. | |
| A key design choice addresses a common failure mode of public deepfake detectors: models quickly learn a shortcut based on embedding vector length (magnitude). To prevent this, all embeddings from the base model are **L1-normalized onto the unit hypersphere**, forcing the classifier to rely purely on angular (directional) information rather than magnitude. | |
| - **Class 0**: AI-generated / spoof audio | |
| - **Class 1**: Real human audio | |
| **Default decision threshold**: 0.5 (tune per use-case). | |
| ## Architecture | |
| This is a **Wav2Vec-based audio encoder** operating at **16 kHz**, paired with a projection layer that maps its base features of dimension **768** (pooled from shape `[T, 768]`) down to **256**. This 256-dimensional representation is the expected input shape for the classification head, which is a **single-layer MLP** producing 2 output logits (with softmax applied during inference). | |
| The core innovation of this architecture is the **dual L1 unit-norm** applied to the model features: | |
| 1. The base features are L1-normalized **once** when passed to the projector. | |
| 2. The projected features are L1-normalized **again** when passed to the classification head. | |
| This dual-hypersphere design completely eliminates size (magnitude) bias. | |
| ### Why Dual L1 Normalization? | |
| Most studies show that deepfake embeddings tend to cluster in a specific region of feature space **and** exhibit a characteristic vector magnitude. As a result, many detectors learn a shortcut: they simply detect the length of the embedding vector (or its presence in a particular cluster) instead of learning meaningful acoustic or spectral cues. | |
| To eliminate this shortcut, we force **every** feature vector onto the unit hypersphere via L1 normalization. Consequently: | |
| - The classification head has only one meaningful signal left to learn from — the **direction (θ)** of the vector. | |
| - Because all vectors lie on the hypersphere, the opportunity for models to exploit magnitude-based or region-specific clustering is drastically reduced. | |
| This approach is similar in spirit to **GenD**, which applied a single L1 normalization onto the unit hypersphere. AIRealNet-Audio extends the idea by applying the normalization **twice** (dual hypersphere), once before the projector and once before the head. | |
| ### Training Behavior | |
| When training **without** L1 normalization we observed a consistent pattern: the model performs well in the early stages, then begins to degrade after a fixed number of steps — even after extensive hyperparameter tuning and aggressive learning-rate reduction aimed at slower convergence. | |
| With the **dual-hypersphere** architecture the opposite occurs: | |
| - Steady upward trend in performance. | |
| - The model continues to improve with each successive step instead of collapsing after a few thousand steps. | |
| **Loss per step** | |
| <p align="center"> | |
| <img | |
| src="https://huggingface.co/Modotte/AIRealNet-Audio/resolve/main/assets/loss.png" | |
| alt="loss Banner" | |
| width="90%" | |
| style="border-radius:15px;" | |
| /> | |
| </p> | |
| **Accuracy per 1000 steps** | |
| <p align="center"> | |
| <img | |
| src="https://huggingface.co/Modotte/AIRealNet-Audio/resolve/main/assets/accuracy.png" | |
| alt="accuracy Banner" | |
| width="90%" | |
| style="border-radius:15px;" | |
| /> | |
| </p> | |
| --- | |
| ## Training Data | |
| - AI-generated speech produced by more than **100 different TTS / voice-cloning systems**. | |
| - Large collection of real human speech drawn from varied recording conditions and sources. | |
| - On-the-fly augmentations including multiple compression codecs, bitrate changes, and additive noise. | |
| - Separate balanced evaluation and development sets used for monitoring. | |
| The combination of high TTS diversity and aggressive augmentation is intended to force the model to learn genuine synthesis artifacts rather than dataset-specific fingerprints. | |
| --- | |
| ## Limitations | |
| - Very short utterances (significantly under 12 s) must be padded or repeated; performance on extremely short clips may be lower. | |
| - Highly adversarial or “nano-edit” modifications of real audio remain challenging. | |
| - Completely unseen generators or extreme domain shifts (e.g., heavy telephony distortion not seen during training) can still reduce accuracy. | |
| - The model is a probabilistic detector; it should not be used as the sole evidence in high-stakes forensic or legal settings without human review. | |
| ## Performance | |
| Across a range of standard and challenging evaluation sets the model demonstrates strong generalization. In most evaluation datasets AIRealNet-Audio maintains an **Equal error rate (EER) well below 6%**. | |
| > We eliminated audio files below 3sec as model is trained on 12-seconds. | |
| | Evaluation Dataset | EER (%) | | |
| |------------------------|---------| | |
| | ASVspoof5 | 1.21 | | |
| | Voxness | 1.48 | | |
| | Malaad | 2.11 | | |
| | ASVspoof2019 DF | 3.14 | | |
| | In-the-Wild | 3.64 | | |
| **Additional training-time observations** | |
| - After 2 epochs: ~**0.99 accuracy** on both the held-out evaluation set and the development set. | |
| - Subsequent checkpoints continue to show steady gains in accuracy (see accuracy curve above). | |
| - Training without the dual L1 normalization exhibits the classic early-peak-then-degradation pattern; the dual-hypersphere design removes this failure mode. | |
| *Note*: Extremely high numbers on controlled evaluation sets are expected. Real-world performance depends on the distribution of generators and acoustic conditions encountered at deployment time. Always calibrate the decision threshold on data that matches your target domain. | |
| --- | |
| ## Usage | |
| ### Sample Audio | |
| The following sample was generated by **Gemini-3.8 Flash TTS** (AI-generated speech): | |
| <audio controls src="https://huggingface.co/Modotte/AIRealNet-Audio/resolve/main/assets/assistant.wav"></audio> | |
| ### Quick Inference | |
| ```python | |
| from transformers import pipeline | |
| pipe = pipeline( | |
| "audio-classification", | |
| model="Modotte/AIRealNet-Audio", | |
| trust_remote_code=True | |
| ) | |
| result = pipe( | |
| "https://cdn-uploads.huggingface.co/production/uploads/677fcdf29b9a9863eba3f29f/C2llrmFhlWx-oryC9wF6f.wav" | |
| ) | |
| print(result) | |
| ``` | |
| **Expected output:** | |
| ```python | |
| [ | |
| {'score': 0.9979, 'label': 'AIVoice'}, | |
| {'score': 0.0021, 'label': 'HumanVoice'} | |
| ] | |
| ``` | |
| | Label | Meaning | | |
| |-------------|--------------------------| | |
| | `AIVoice` | AI-generated / spoof audio | | |
| | `HumanVoice`| Real human speech | | |
| Default decision threshold is **0.5**. You can adjust it according to your precision/recall needs. | |
| ### Notes for Production Use | |
| 1. Resample input to **16 kHz mono**. | |
| 2. Segment long recordings into **12-second chunks** (the training chunk size). | |
| 3. Aggregate chunk-level scores (mean, max, or calibrated fusion) as needed. | |
| --- | |
| ## Intended Use | |
| - Detection of AI-generated / spoofed speech on social media, messaging platforms, call centers, and research datasets. | |
| - Assistance for content moderators, journalists, fact-checkers, and platform trust-and-safety teams. | |
| - Research baseline for audio deepfake detection under a hyperspherical feature constraint. | |
| Not intended as the sole source of evidence in legal, forensic, or high-stakes verification scenarios without corroborating human analysis. | |
| --- | |
| ## Ethical Considerations | |
| - Training data construction followed the same privacy-first principles used for the original AIRealNet image model (no personal or sensitive recordings). | |
| - Users should treat model scores as one signal among many and always apply human review near the decision threshold. | |
| - The model card explicitly documents the dual-hypersphere design and the known length-bias failure mode so that downstream users understand both the strengths and the residual risks. | |
| --- | |
| ## How It Works | |
| 1. Input audio is resampled to 16 kHz and segmented into 12-second chunks. | |
| 2. A Wav2Vec encoder extracts frame-level features [T, 768]. | |
| 3. Features are mean-pooled, L1-normalized, and projected to 256 dimensions. | |
| 4. The 256-dimensional vector is L1-normalized a second time and passed through a linear classification head. | |
| 5. Softmax yields class probabilities; a threshold (default 0.5) produces the final binary decision. | |
| --- | |
| ## Future Work | |
| - Improve robustness to very short utterances and adversarial nano-edits. | |
| - Expand coverage to additional languages and emerging TTS / voice-conversion systems. | |
| - Investigate multi-modal (audio + video / metadata) detection. | |
| - Explore variable-length or streaming-friendly variants of the dual-hypersphere constraint. | |
| --- | |
| ## Citation | |
| ```bibtex | |
| @misc{Modotte_AIRealNet_Audio_2025, | |
| title = {AIRealNet-Audio: Dual-Hypersphere Constrained Wav2Vec for Detecting AI-Generated vs Real Speech}, | |
| author = {Parvesh Rawal}, | |
| year = {2025}, | |
| publisher = {Hugging Face}, | |
| url = {https://huggingface.co/Modotte/AIRealNet-Audio} | |
| } | |
| ``` | |
| --- | |
| ## Acknowledgments | |
| Special thanks to [Sujal](https://huggingface.co/sujalrajpoot) for performing all the model evaluations. | |
| --- | |
| ## References | |
| - Yermakov, A., Cech, J., Matas, J., & Fritz, M. (2026). *Deepfake Detection that Generalizes Across Benchmarks* (GenD). Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV). | |
| [arXiv:2508.06248](https://arxiv.org/abs/2508.06248) · [GitHub](https://github.com/yermandy/GenD) | |
| - Microsoft / Facebook Wav2Vec 2.0 and related self-supervised speech models. | |
| - [AIRealNet](https://huggingface.co/Modotte/AIRealNet) — the image counterpart that motivated this audio extension. | |