--- license: cc-by-4.0 library_name: pytorch pipeline_tag: audio-classification base_model: - openai/whisper-base - openai/whisper-small - laion/whisper-base-small-layered-audio-scores base_model_relation: finetune datasets: - laion/tts-scaling-ladder-de-en - laion/laions_got_talent_raw - mitermix/balanced_audio_snippets_40x3k - mitermix/voice-annotation-poc language: - multilingual tags: - whisper - encoder-only - emotion-recognition - vocal-burst-localization - speaker-embeddings - audio-quality - multitask-regression - gemini-distillation --- # Humaneness Ears: Base and Medium audio understanding **By Christoph Schuhmann · LAION · 7 October 2026 · CC BY 4.0** This repository releases the **best completed two-epoch fine-tunes** of LAION's encoder-only Whisper Base and Whisper Small multitask models. The release names are **Humaneness Ears Base** (Whisper Base backbone) and **Humaneness Ears Medium** (Whisper Small backbone). Given a recording, one forward pass predicts emotions, speaking style, sound quality, speaker-characteristic vectors and the location and type of non-speech vocal events such as laughter or crying. Both models are included: `base/model.safetensors` and `medium/model.safetensors`. The older training logs, evaluation JSON and the machine-readable listening manifest retain the `small` key as the historical Whisper backbone label; it refers to Humaneness Ears Medium. **These are audio encoders with task heads. There is no Whisper text decoder, ASR transcript, generated caption or language model.** Gemini transcripts and captions helped prepare annotations; captions were not model inputs or an autoregressive training objective. CPS is a scalar prediction rather than a transcript. The requested separate MOSS caption LoRAs are outside this release and were not trained as part of these two runs. **[Explore 1,000 real-audio examples in the Humaneness Ears atlas](https://laion-humaneness-ears-atlas.static.hf.space/)**: 50 timed vocal bursts, 20-query Orange timbre and identity neighbor pages, 40 EmoNet and 57 VoiceNet top-ten rankings, and five-tier quality pages. [Download the Gemini annotation dataset](https://huggingface.co/datasets/laion/whisper-gemini-voice-annotations). **[Open the full benchmark: Humaneness Ears Base/Medium versus CLAPv2 XS/M](https://laion-whisper-base-small-emotion-voice-burst.static.hf.space/benchmark.html)** — human-label results, all 192 targets, Orange speaker-vector cosine and vocal-burst timing. [Jump to the CLAP comparison](https://laion-whisper-base-small-emotion-voice-burst.static.hf.space/benchmark.html#clap-v2-full-ft). ## Results and related releases - **[Comprehensive English benchmark report](https://laion-whisper-base-small-emotion-voice-burst.static.hf.space/benchmark.html)**: 36 frozen embedding probes, original and Humaneness Ears Base/Medium after Gemini tuning, and the two full CLAPv2 fine-tunes; human-label emotion benchmarks, all 192 score targets, speaker vectors, validation, burst timing and limitations. - [Which model is best for each task?](https://laion-whisper-base-small-emotion-voice-burst.static.hf.space/benchmark.html#study-overview) includes genuineness and vocal-burst blend. - [Every target and validation result](https://laion-whisper-base-small-emotion-voice-burst.static.hf.space/benchmark.html#study-targets). - [Original layered Whisper weights](https://huggingface.co/laion/whisper-base-small-layered-audio-scores): the Base **S4** and original Whisper Small **S3** checkpoints are the actual starting states for this release. **Medium** has better Orange speaker-vector cosine and lower AudioBox/DNSMOS target error than Base, and a higher EmoNet intensity correlation. **Base** has slightly better CPS error, Flash frame F1 and VoiceNet-Emo mean rank correlation. A frozen probe can outperform both on other tasks: see the comprehensive report for all winners. This release does not claim one universal best model or statistical significance from small numerical differences. ## What the models output | Output | Shape / scale | Meaning | | --- | --- | --- | | EmoNet emotions | 40 scalars, reference 0–4 | Fine-grained emotion intensity; simultaneous regression, not a mutually exclusive class. | | VoiceNet dimensions | 57 scalars, mostly 0–6 | Delivery, register, timbre and style; Flash BKGN is 0–4 and EXPL 0–2. | | Genuineness | 1 scalar, 0–6 | The annotation rubric's perceived genuineness. | | Vocal-burst blend | 1 scalar, 0–10 | How a vocal event blends with surrounding speech; meaningful only when a burst is present. | | R_quality | 1 legacy scalar | Its distinct original teacher target and units; not overwritten with another quality construct. | | Empathic Insight Plus extras | 19 scalars | Voice traits, arousal/valence, recording/background quality and content enjoyment. | | AudioBox Aesthetics | 4 scalars | CE: content enjoyment; CU: content usefulness; PC: production complexity; PQ: production quality. | | DNSMOS | 7 scalars | SIG, BAK, OVRL, their raw versions and P808 MOS. | | Burst count | 1 scalar trained in log1p space | Number of annotated events; inference additionally reports expm1 of the nonnegative prediction. | | VoiceCLAP attributes | 61 scalars | Additional emotion, voice and quality targets; some emotion targets overlap the first 40. | | Characters per second (CPS) | 1 separate scalar | Unicode transcript characters, including spaces/punctuation, divided by total recording duration; bracketed nonlexical placeholders removed. | | Orange timbre | 128D unit vector | Regression to the Orange timbre teacher embedding. | | Orange identity | 250D unit vector | Regression to the Orange identity teacher embedding; not a person's name or an identification guarantee. | | Burst occupancy | 20 ms frame probabilities | Whether each encoder frame belongs to at least one vocal burst. | | Independent burst events | Start **and** end seconds | Onset plus duration proposals can overlap; maximum 32 at the default evaluation setting. | | Burst type | 53-way event-local class probabilities | The top three labels and uncalibrated softmax probabilities for each predicted interval. | There are **192 jointly predicted scalar outputs plus the separate CPS head**. A complete, ordered list of exact output keys and all their held-out results appears at the end of this card. The models return raw units, training-standardized values and display z-scores. Regression outputs are continuous and are not automatically clipped to the annotation's ordinal range. ## Quick start: local CPU or GPU inference Use Python 3.10 or newer. The measured production environment used Python 3.13.5, PyTorch 2.9.1, Transformers 5.14.1, SoundFile 0.14 and Safetensors. Install the appropriate PyTorch CPU/CUDA wheel for your machine; see `requirements.txt` for the other dependencies. Inference uses a custom `model.py`, not `pipeline('automatic-speech-recognition')` or `WhisperForConditionalGeneration.from_pretrained`. Download inference assets without the large optimizer checkpoints: ```python from huggingface_hub import snapshot_download snapshot_download( repo_id="laion/humaneness-ears-base-medium", local_dir="humaneness-ears", allow_patterns=["*.py", "requirements.txt", "*.json", "base/*.json", "medium/*.json", "base/model.safetensors", "medium/model.safetensors"], ) ``` ```bash cd humaneness-ears pip install -r requirements.txt python inference.py clip.wav --model medium --device cpu --threads 4 > medium_prediction.json python inference.py clip.wav --model base --device cpu --threads 1 > base_prediction.json python inference.py clip.wav --model medium --device cuda:0 --threads 4 \ --include-frame-probabilities > gpu_prediction.json ``` Input recordings must be **0.1–30 seconds**. SoundFile decodes the supplied audio, stereo is averaged to mono, and SciPy resamples to 16 kHz. The bundled Whisper frontend creates 80-bin log-mel features. Longer recordings need an explicit external chunking policy. No original OpenAI weights, teacher checkpoints or Gemini API access are required at inference. Reuse the model when processing multiple files: ```python from inference import Predictor predictor = Predictor(".", model_size="medium", device="cpu", threads=4) result = predictor.predict("clip.wav", include_frames=True) print(result["scores_raw"]["genuineness_0_6"]) print(result["scores_z"]["emo_Amusement"]) print(result["burst_event_proposals"]) # start_s, end_s, top3_classes ``` `--frame-threshold` and `--event-threshold` default to 0.5; `--max-events` defaults to 32. This matches the new report's proposal operating point. Lowering the event threshold can increase recall and false positives. Frame regions merge overlapping events, whereas the independent proposal head can represent them separately. Neither class softmax values nor sigmoid values have been calibrated as real-world confidence estimates. ## Architecture Base uses **six encoder layers, hidden width 512**, and **26,441,735 total parameters**. Medium uses **12 encoder layers, hidden width 768**, and **97,102,031 total parameters**. These counts include our heads and exclude the absent Whisper decoder. The encoder alone contains 20,590,592 / 88,154,112 parameters respectively. From the final encoder sequence we concatenate masked mean, minimum, maximum and standard deviation. Separately, each transformer layer contributes its masked mean vector. Shared per-family projection MLPs reduce these means to 32 or 64 dimensions, and a sample-specific softmax gate combines layer information. These features add residuals to the 192-score head. Speaker heads use 128D layer-mixture routes; CPS uses 32D. The frame head predicts occupancy; the proposal head predicts independent onset logits and log durations. A 64D layer-mixture route for event classification averages **only frames inside that event** from every layer, then combines them with the final-layer local mean and a 53-way classifier. `model.py` is the exact architecture used in training. Exported weights preserve every saved tensor without quantization or conversion. The inherited Whisper configuration still contains some unused decoder fields; `architectures` and `is_encoder_decoder` are set for this encoder-only release. No decoder tensors are present. The custom inference code must be used. ## Training data, annotation priority and splits The selected pool contains **66,199 exact-audio unique clips / 197.36 hours** after deduplication across the four selections. Valid Flash annotations exist for **65,163 clips / 194.17 hours**. The 1,036 invalid or provider-blocked responses are excluded from this tuning stage. They were not converted into zero-emotion or no-burst targets. | Split actually used | Clips | Purpose | | --- | ---: | --- | | Training | 58,380 | Gradient updates for both full fine-tunes. | | Validation | 3,239 | Full validation after each epoch; selects best checkpoint. | | Test | 3,544 | Matched before/after evaluation; not used for gradient updates or epoch selection. | Sources are the **balanced S1 100-hour, 400-bucket selection** from [TTS Scaling Ladder DE/EN](https://huggingface.co/datasets/laion/tts-scaling-ladder-de-en), [LAION's Got Talent raw](https://huggingface.co/datasets/laion/laions_got_talent_raw), [balanced audio snippets 40×3k](https://huggingface.co/datasets/mitermix/balanced_audio_snippets_40x3k) and [voice annotation POC](https://huggingface.co/datasets/mitermix/voice-annotation-poc). Final selections preserve all source/task memberships while annotating each exact compressed-audio SHA256 once. Source membership counts are nonexclusive. Deterministic 90/5/5 hashing groups explicit speaker IDs, otherwise known synthetic audio families, otherwise exact audio hashes; observed counts differ from exact percentages because groups are kept intact. `training/provenance/valid_split_summary.json` records valid counts, hours, language strings, memberships and the hash of the exact prepared target snapshot. The older `TRAINING_READY.json` also contains split counts including invalid responses; those are not the training counts above. Source revisions and original pool details are in `training/provenance/selected_pool.json`. The data, Gemini responses and teacher TARs are not redistributed by this model repository. The recorded annotation model identifier is **`gemini-3.8-flash`**. Schema-valid Flash outputs take priority for 40 emotions, 57 VoiceNet dimensions, genuineness, blend, event counts, event start/end times and transcript-derived CPS. Corresponding VoiceCLAP emotion attributes also take the Flash value. Distinct constructs such as DNSMOS BAK or the original authenticity axis keep their own teachers. A null blend on a no-burst clip is masked, not replaced by an old blend score. Missing values remain masked. Speech-only objectives use domain masks; Orange speaker losses require valid known single-speaker clips. Teacher vectors are preserved even when their training loss is masked. The prompt permits 78 vocal fine labels. Explicit aliases map compatible labels into the existing 53-class checkpoint vocabulary; unrepresented labels map to `Other / rare burst`. See `gemini_burst_mapping.json` for every mapping. Original fine labels and event captions remain in the annotation sidecars. End times may be clamped to the decoded duration within the accepted 0.3-second annotation tolerance, with original timing and the clamp flag retained in the data. ## Exact fine-tuning procedure Starting models were the lowest-validation-loss **Base S4** and **Whisper Small S3** snapshots from the original S1→S10 layered curriculum campaign. Each snapshot contains only the progress up to its own stage, rather than the final S10 weights: Base's original update was 266,238; Medium's was 237,434. The original campaign ran one epoch per nested stage with 10% P3 synthetic mixture exposures per stage and a continuous cosine schedule. Its full code and recorded run configs are included under `training/` and `training/original_curriculum/`. Gemini tuning **strictly restored all encoder and head weights**, started a fresh AdamW optimizer and cosine schedule, and unfroze the entire encoder, including its positional embedding table. This is full fine-tuning, not LoRA or a frozen-encoder probe. | Setting | Base | Medium | | --- | ---: | ---: | | Epochs | 2 | 2 | | Actual optimizer updates | 914 | 914 | | Best checkpoint | Epoch 2 | Epoch 2 | | Peak encoder learning rate | 1e-5 | 5e-6 | | Peak head learning rate | 1e-4 | 1e-4 | | Warmup | 5% (46 updates) | 5% (46 updates) | | Cosine minimum / peak | 10% | 10% | | Weight decay | 0.01 | 0.01 | | Global effective batch | 128 | 128 | | Microbatch / GPU | 32 | 16 | | Accumulation | 1 | 2 | | Data workers / GPU | 4 | 4 | | Hardware | 1 node, 4 GH200 GPUs | 1 node, 4 GH200 GPUs | | Precision | BF16 autocast; FP32 saved weights | BF16 autocast; FP32 saved weights | | Gradient clipping | Global norm 1.0 | Global norm 1.0 | | Seed | 20261005 | 20261005 | | Selected class-weighted validation loss | 1.7683769 | 1.6989245 | | Training Slurm job | 2199189 | 2199190 | | Recorded allocation elapsed time | 10m17s | 10m31s | There is **no P3 replay in this extra two-epoch stage**. A DDP sampler may pad a few examples to distribute batches equally; validation is evaluated completely on rank zero without duplicating its samples. The full config, class counts and epoch loss logs are under `training/`. ### Losses The implementation is `training/train_layered_curriculum.py::loss_terms`, reused unchanged by `training/gemini_finetune/train_whisper.py`: - Masked Huber regression, delta 1, on standardized scalar targets; per-family weights are the exact `GROUPS` constants in `training/multitask_p3_smoke.py`. - CPS: standardized Huber delta 1, weight 0.2. - Speaker vectors: 0.2 × (cosine distance + 0.1 × Huber delta 0.1), only valid single speakers. - Frames: weighted BCE (positive weight 5), weight 0.5, plus soft Dice weight 0.25. - Onsets: weighted BCE (positive weight 100), weight 0.2. - Event log duration: Huber delta 0.5, weight 0.1, evaluated at labeled starts. - Event class: categorical cross entropy, weight 0.2, label smoothing 0.02; training-only class weights sqrt(median event count / class count), clipped to 0.25–4. Every epoch saves latest and, when improved, best states. Published `base/checkpoint.pt` and `medium/checkpoint.pt` are the exact best files: model, AdamW, scheduler, four-rank RNG, completed epoch, update, normalization, taxonomy and original config. They support epoch-boundary continuation with the same DDP world size. Both runs already finished epoch 2; extending training requires an explicit new phase/schedule, not blindly resuming the completed two-epoch plan. Inference uses only Safetensors and never loads the `.pt` training pickle. ## Normalization ```text raw_score = standardized_network_output × training_std + training_mean display_z = (raw_score − original_training_teacher_median) / original_training_teacher_std CPS_raw = CPS_network_output × CPS_training_std + CPS_training_mean ``` `training_normalization.json` is the original frozen checkpoint normalization, preserved through Gemini tuning. It is required for correct raw units. `display_normalization.json` is the train-only median/std reference from the original release, retained for compatible visualization and rankings. `gemini_train_statistics.json` separately stores train-only Flash mean/std/counts; its inherited “provisional” description refers to the incremental parser, but the file here is the completed final export. It was not substituted into network decoding. Display z-scores are relative context, not class probabilities or confidence intervals. ## How the data were mounted and how to reproduce Training ran on **JUPITER Booster**. `/e/home`, `/e/project1`, `/e/scratch` and `/e/software` were existing shared HPC filesystems visible to every rank; no Docker bind mount or network download occurred inside the GPU training. Source audio and same-audio teacher outputs stayed in WebDataset TARs. `TarReader` caches a bounded set of open archives and reads named members; it does not extract millions of individual files. Teachers are cached annotations, not called inside the gradient loop. | Original absolute path | Contents / use | | --- | --- | | `/e/home/jusers/schuhmann1/jupiter/whisper_score_regression` | Production training and inference source. | | `/e/scratch/reformo/schuhmann1_moss/whisper_score_regression/hf_layered_release/model` | Original Base S4 / Whisper Small S3 states, frontend, normalization and classes. | | `/e/scratch/reformo/schuhmann1_moss/whisper_score_regression/gemini_multisource_100h_20261004` | Combined exact-audio selection and Gemini response sidecars. | | `/e/scratch/reformo/schuhmann1_moss/whisper_score_regression/gemini_finetune_preparation_20261005/prepared` | `targets.jsonl`, final training gate, normalization and burst mapping. | | Same preparation root, `teacher_sidecars/` | SHA-bound TAR companions with Orange vectors and all additional teacher outputs. | | Same preparation root, `training/whisper_{base,small}` | Exact best/latest states and per-epoch metrics. | | `/e/scratch/reformo/schuhmann1_moss/code/fastgen.sh` | Python/software paths for the original offline HPC environment. | `targets.jsonl` contains `audio_tar` / `audio_member` and embedding `tar` / `member` references. On another machine, obtain the same complete prepared data, copy TARs and rewrite the absolute TAR prefixes to your local mounts while preserving member names and audio hashes. Provide `checkpoint_normalization.json`, `gemini_burst_mapping.json` and `TRAINING_READY.json` beside it. The repository includes the parser, schema prompt and teacher backfill code; recreating its data requires the upstream datasets, teacher weights and your own annotation service access. There is no credential embedded here. Source availability and terms still apply. To reproduce the **two-epoch experiment**, download the original Base S4 or Whisper Small S3 `.pt` initializer from the previous repository, not this release's already-tuned state. Then: ```bash pip install -r requirements-training.txt python training/make_finetune_config.py --model base \ --data-root /your/prepared --initial-checkpoint /your/original/base/S4/checkpoint.pt \ --output-dir /your/new-base-run --config /your/base-run.json python -m torch.distributed.run --standalone --nproc_per_node=4 \ training/gemini_finetune/train_whisper.py --config /your/base-run.json ``` For the historical Whisper Small training run, use `--model small` and the original `small/S3/checkpoint.pt`. The config helper changes paths while preserving recorded hyperparameters. `training/gemini_finetune/train_whisper.sbatch` is the original submitted Slurm launcher, with original mount paths for reference. Evaluator scripts likewise preserve original study paths; adapt `study_paths.py` and benchmark data paths if reproducing the full comparison. Safe-weight export tools and verification receipts are included. ## Held-out Flash target agreement Both models are tested on the **same 3,544 valid Flash test clips**. These targets are machine annotations, not a fresh human panel. Speaker cosine uses 1,718 valid single-speaker clips; blend uses 2,391 clips with non-null blend. Family MAE is the unweighted mean over valid axes. | Flash test target / metric | Base | Medium | | --- | ---: | ---: | | Burst frame F1 ↑ | 0.6929 | 0.6892 | | Burst event F1, IoU ≥0.5 ↑ | 0.1945 | 0.1949 | | Burst start MAE on IoU ≥0.1 matches, seconds ↓ | 0.1201 | 0.1150 | | Burst end MAE on IoU ≥0.1 matches, seconds ↓ | 0.1973 | 0.2006 | | 53-class accuracy with reference spans ↑ | 0.4280 | 0.4483 | | 53-class accuracy on matched predicted spans ↑ | 0.4189 | 0.4370 | | Orange timbre cosine ↑ (1,718 valid single-speaker clips) | 0.9038 | 0.9223 | | Orange identity cosine ↑ (1,718 valid single-speaker clips) | 0.8541 | 0.8741 | | CPS raw MAE, characters/second ↓ | 0.9012 | 0.9309 | | Genuineness raw MAE on 0–6 scale ↓ (3,544 clips) | 0.4774 | 0.4704 | | Genuineness normalized MAE ↓ | 0.2971 | 0.2928 | | Blend raw MAE on 0–10 scale ↓ (2,391 valid clips) | 0.8773 | 0.8417 | | Blend normalized MAE ↓ | 0.3228 | 0.3097 | | 40 emotions, mean normalized MAE ↓ | 0.5873 | 0.5605 | | 57 VoiceNet dimensions, mean normalized MAE ↓ | 0.4967 | 0.4813 | | AudioBox four axes, mean normalized MAE ↓ | 0.1877 | 0.1739 | | DNSMOS seven outputs, mean normalized MAE ↓ | 0.2047 | 0.1975 | Normalized MAE is measured in frozen training standard deviations. Burst interval metrics use Hungarian matching of independent onset/duration proposals, not binary contiguous-frame regions. Boundary errors and predicted-span class accuracy are conditional on matches at IoU ≥0.1. Reference-span class accuracy gives the classifier the true interval; it is not end-to-end localization accuracy. Maximum proposals is 32, with a fixed 0.5 onset threshold. Event precision is low despite high recall; the report includes all IoU cutoffs and precision/recall. The validation tables in `evaluation/{base,small}/validation_metrics.json` (historical backbone keys) use the existing 3,239 checkpoint-selection clips, not an untouched test. Supplementary evaluator loss uses unit class weights, while the training selection loss above used the configured class weights. Different loss values under those two policies are expected. ## Public human-label benchmarks | Human-label benchmark / metric | Base | Medium | | --- | ---: | ---: | | EmoNet-Voice intensity, Pearson r ↑ | 0.4340 | 0.4512 | | EmoNet-Voice intensity, Spearman ρ ↑ | 0.4577 | 0.4679 | | VoiceNet-Emo, mean prompt Spearman ρ ↑ | 0.4288 | 0.4282 | | VoiceNet-Ext / emolia-dim, mean prompt Spearman ρ ↑ | 0.1566 | 0.1502 | | CREMA-D matched actor-CV score-MLP accuracy ↑ | 0.6310 | 0.6423 | | RAVDESS matched actor-CV score-MLP accuracy ↑ | 0.6597 | 0.6639 | EmoNet uses 12,000 mapped clips and fixed 0–4→0–10 endpoint scaling; two unsupported labels from the published 12,600-clip benchmark are omitted. VoiceNet-Emo uses 7,986 current-repository questions with at least two raters; VoiceNet-Ext uses 13,917. **EXT and DIM refer to the same benchmark**. Current snapshot coverage differs from paper coverage. The report includes unflagged, untruncated and alternate agreement cuts and explains optimistic oracle threshold numbers. CREMA-D and RAVDESS figures are **supervised matched nested actor-cross-validation score adapters**, not zero-shot paper numbers. Each adapter sees the model's 192 raw predictions, a train-fold StandardScaler and a 64-unit GELU MLP: 12,742 / 12,872 parameters respectively. Five outer actor folds and three inner folds choose learning rate 0.001/0.003 and 20/50 epochs. Outer test actors never fit that fold's scaler or adapter. These benchmark-only adapters are not the model's burst classifier and are not automatically applied by `inference.py`. Full fold records, metrics and confusions are included; the report compares the same adapter budget across all backbones. Upstream training/source overlap with the public benchmarks was not exhaustively audited. ## P3 retention and limitations On the identical fixed **2,000-clip P3 test audit**, frame F1 changes from **0.9622 to 0.8051** for Base and **0.9693 to 0.8392** for Medium after Gemini tuning. Flash test frame F1 rises from 0.3072/0.3171 to 0.6929/0.6892 respectively. This is a target/domain tradeoff: there was no P3 replay in this phase. All matched pre/post P3 validation/test JSONs are included. It is not valid to claim universal burst improvement from Flash agreement alone. The selected pool is not a representative random sample of all natural audio. Quality scores imitate AudioBox/DNSMOS/Empathic teachers rather than measuring new human MOS. Speaker cosine measures regression agreement with Orange, not verification accuracy or human identity. The burst classifier groups unsupported fine labels, event proposals have many false positives, and emotional or expressive clips can differ from the synthetic P3 construction labels. Recordings longer than 30 seconds and multi-speaker embeddings require additional policies. ## Annotation models, taxonomies and packages | Supervision / package | Primary reference | | --- | --- | | Original Whisper encoders | [OpenAI Whisper Base](https://huggingface.co/openai/whisper-base), [OpenAI Whisper Small](https://huggingface.co/openai/whisper-small), [Whisper code](https://github.com/openai/whisper). | | Gemini annotation schema | The exact 108 KB master prompt is `training/gemini_s10_smoke/gemini_voice_annotation_master_prompt_detailed.txt`; parser and validation code are included. | | 40 emotions | [EmoNet taxonomy](https://github.com/LAION-AI/voicenet/blob/main/taxonomy/emonet_taxonomy.md), [emotion annotations toolkit](https://github.com/LAION-AI/emotion-annotations). | | 57 voice dimensions | [VoiceNet taxonomy](https://github.com/LAION-AI/voicenet/blob/main/taxonomy/voicenet_taxonomy.md), [commercial dimension predictors](https://huggingface.co/laion/voicenet-dimension-predictors-commercial). | | Original genuineness teacher | [VoiceCLAP commercial genuineness](https://huggingface.co/laion/voiceclap-commercial-genuineness); matching targets superseded by valid Flash labels. | | Original blend teacher | [VoiceCLAP commercial vocalburst blend](https://huggingface.co/laion/voiceclap-commercial-vocalburst-blend); matching targets superseded by valid Flash labels. | | VoiceCLAP attributes / intermediate embedding | [Commercial VoiceCLAP](https://huggingface.co/laion/voiceclap-commercial), [attribute heads](https://huggingface.co/laion/voiceclap-commercial-attribute-heads); 768D intermediate embeddings are cached in data, not an output head. | | Empathic extras / original emotion teacher | [Empathic Insight Voice Plus](https://huggingface.co/laion/Empathic-Insight-Voice-Plus), [BUD-E Whisper](https://huggingface.co/laion/BUD-E-Whisper); 3072D teacher features are cached, not an output head. | | AudioBox | [Meta AudioBox Aesthetics](https://huggingface.co/facebook/audiobox-aesthetics). | | DNSMOS | [Microsoft DNS Challenge / DNSMOS](https://github.com/microsoft/DNS-Challenge/tree/master/DNSMOS). | | Orange timbre, 128D | [Orange Speaker-wavLM-tbr](https://huggingface.co/Orange/Speaker-wavLM-tbr), pinned annotation revision `b8d2608d56f18e2b6e27ba566cf69132e2c2ad6c`. | | Orange identity, 250D | [Orange Speaker-wavLM-id](https://huggingface.co/Orange/Speaker-wavLM-id), pinned revision `abbb3c7b8d220ceebc33d9bd3bb5aedb580342c7`. | | Timbre generation implementation | [LAION generation script](https://github.com/LAION-AI/emotion-annotations/blob/main/generate_timbre_embeddings.py). | | Vocal event taxonomy / locator | [LAION voice taxonomies](https://github.com/LAION-AI/voice-taxonomies), [vocalburst locator](https://huggingface.co/laion/vocalburst-locator), canonical classes and Flash mapping bundled here. | | Human emotion/style benchmark | [LAION emolia-bench](https://github.com/LAION-AI/emolia-bench), [EmoNet-Voice Bench](https://huggingface.co/datasets/t1a5anu-anon/emonet-voice-bench). | Teacher inference runs on the exact selected waveform. New scores and float16 target vectors are persisted in SHA-bound companion TARs; the gradient loop reuses those sidecars. Their model revisions and target-level provenance remain in the prepared annotations. ## Repository contents and export verification | Path | Contents | | --- | --- | | `base/model.safetensors`, `medium/model.safetensors` | Complete best Gemini-tuned encoder and head weights; no decoder. | | `base/checkpoint.pt`, `medium/checkpoint.pt` | Exact best resumable training states including optimizer, scheduler and per-rank RNG. | | `{base,medium}/config.json`, `preprocessor_config.json` | Corresponding encoder architecture and Whisper log-mel frontend. | | `model.py`, `inference.py`, `requirements.txt` | Complete standalone inference, cached Predictor API and CLI. | | `training/`, `vocal_burst_pool/p3/` | Actual training source and local dependency closure, target parser, prompt, backfill and evaluation code, configs, logs and reproduction helper. | | `training_normalization.json`, `display_normalization.json`, `gemini_train_statistics.json` | Frozen decoding, contextual display and separate final Flash train-only statistics. | | `classes.json`, `gemini_burst_mapping.json` | 53 canonical class names and all Flash aliases/fallbacks. | | `evaluation/{base,small}/` (historical keys) | Flash validation/test, pre/post P3 retention, public benchmarks and actor-adapter folds. | | `verification/{base,small}.json` (historical keys) | Bitwise safe-weight and model-output parity; real held-out audio API checks, dimensions and timestamp validity. | | `release_manifest.json` | File sizes and SHA256 hashes, original checkpoint hashes and provenance. | | `release_tools/` | Export, verification and publication source. | | `LICENSE`, `NOTICE.md` | CC BY 4.0 and upstream attribution. | No audio, source transcripts, raw Gemini responses or service credentials are published by this model repository. Model verification uses a real held-out waveform internally and saves only its hash and output-contract evidence. Evaluation JSONs contain aggregate benchmark results, prompts/classes and fold protocols; no listening audio is copied into this release. ## License and attribution The new fine-tuned weights, new code and documentation are released by **LAION** under **[Creative Commons Attribution 4.0 International](LICENSE)**. Credit **Christoph Schuhmann and LAION**, link this repository and indicate changes. This license does not relicense upstream source audio, teacher models or external packages. The original OpenAI Whisper components remain MIT; Orange's identity teacher retains its upstream CC BY-SA 3.0 terms. See `NOTICE.md` and upstream model/data cards for their notices. ```bibtex @misc{schuhmann_humaneness_ears_2026, title = {Humaneness Ears: Base and Medium Encoder-Only Audio Understanding}, author = {Schuhmann, Christoph and {LAION}}, year = {2026}, url = {https://huggingface.co/laion/humaneness-ears-base-medium} } ``` ## All 192 scalar targets: exact held-out results These are Flash test targets, with the remaining AudioBox/DNSMOS/Orange/Empathic families provided by their own teachers. Raw MAE uses each target's units. Norm. MAE uses its frozen training standard deviation. `r` is Pearson and `ρ` is Spearman. Missing metrics remain unavailable. The source JSON has separate counts for every model/target. | Index | Exact target key | Valid N | Base raw MAE | Base norm. MAE | Base r | Base ρ | Medium raw MAE | Medium norm. MAE | Medium r | Medium ρ | | ---: | --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | | 0 | `emo_Affection` | 3544 | 0.4034 | 0.7418 | 0.7163 | 0.5685 | 0.3911 | 0.7192 | 0.7345 | 0.5839 | | 1 | `emo_Amusement` | 3544 | 0.5187 | 0.5978 | 0.7413 | 0.6742 | 0.4879 | 0.5623 | 0.7606 | 0.6992 | | 2 | `emo_Anger` | 3544 | 0.3462 | 0.5968 | 0.6770 | 0.4799 | 0.3376 | 0.5819 | 0.6729 | 0.5020 | | 3 | `emo_Astonishment_Surprise` | 3544 | 0.4007 | 0.7257 | 0.6906 | 0.5951 | 0.3838 | 0.6950 | 0.7075 | 0.6081 | | 4 | `emo_Awe` | 3544 | 0.3259 | 0.7764 | 0.7652 | 0.5528 | 0.3002 | 0.7152 | 0.7946 | 0.5699 | | 5 | `emo_Bitterness` | 3544 | 0.3718 | 0.9136 | 0.6692 | 0.5244 | 0.3488 | 0.8571 | 0.6932 | 0.5548 | | 6 | `emo_Concentration` | 3544 | 0.4721 | 0.6463 | 0.6937 | 0.6930 | 0.4634 | 0.6343 | 0.7024 | 0.7005 | | 7 | `emo_Confusion` | 3544 | 0.3659 | 0.6860 | 0.6370 | 0.5124 | 0.3462 | 0.6490 | 0.6617 | 0.5392 | | 8 | `emo_Contemplation` | 3544 | 0.5538 | 0.9403 | 0.7380 | 0.7293 | 0.5367 | 0.9113 | 0.7552 | 0.7449 | | 9 | `emo_Contempt` | 3544 | 0.3103 | 0.6878 | 0.6823 | 0.4370 | 0.2999 | 0.6648 | 0.6994 | 0.4587 | | 10 | `emo_Contentment` | 3544 | 0.5967 | 1.1843 | 0.6973 | 0.6890 | 0.5675 | 1.1264 | 0.7270 | 0.7214 | | 11 | `emo_Disappointment` | 3544 | 0.4548 | 0.6485 | 0.7058 | 0.5985 | 0.4328 | 0.6171 | 0.7290 | 0.6164 | | 12 | `emo_Disgust` | 3544 | 0.2244 | 0.4576 | 0.6700 | 0.4164 | 0.2140 | 0.4365 | 0.7002 | 0.4420 | | 13 | `emo_Distress` | 3544 | 0.4677 | 0.5722 | 0.7701 | 0.6569 | 0.4445 | 0.5438 | 0.7917 | 0.6834 | | 14 | `emo_Doubt` | 3544 | 0.4587 | 0.9491 | 0.6336 | 0.6068 | 0.4360 | 0.9020 | 0.6652 | 0.6390 | | 15 | `emo_Elation` | 3544 | 0.3643 | 0.5071 | 0.7254 | 0.5958 | 0.3513 | 0.4890 | 0.7449 | 0.6097 | | 16 | `emo_Embarrassment` | 3544 | 0.1472 | 0.3367 | 0.5970 | 0.2711 | 0.1434 | 0.3282 | 0.6350 | 0.2784 | | 17 | `emo_Emotional_Numbness` | 3544 | 0.2516 | 0.2617 | 0.5849 | 0.4138 | 0.2434 | 0.2532 | 0.6081 | 0.4254 | | 18 | `emo_Fatigue_Exhaustion` | 3544 | 0.4472 | 0.8269 | 0.7700 | 0.6846 | 0.4385 | 0.8108 | 0.7746 | 0.6890 | | 19 | `emo_Fear` | 3544 | 0.2303 | 0.3764 | 0.7266 | 0.4491 | 0.2201 | 0.3597 | 0.7331 | 0.4586 | | 20 | `emo_Helplessness` | 3544 | 0.4093 | 0.5781 | 0.7839 | 0.6445 | 0.3889 | 0.5492 | 0.8065 | 0.6608 | | 21 | `emo_Hope_Enthusiasm_Optimism` | 3544 | 0.5715 | 0.7535 | 0.6933 | 0.6714 | 0.5412 | 0.7135 | 0.7238 | 0.7046 | | 22 | `emo_Impatience_and_Irritability` | 3544 | 0.5024 | 0.7343 | 0.6582 | 0.5313 | 0.4835 | 0.7068 | 0.6653 | 0.5517 | | 23 | `emo_Infatuation` | 3544 | 0.1289 | 0.2382 | 0.7034 | 0.3314 | 0.1205 | 0.2226 | 0.7270 | 0.3399 | | 24 | `emo_Interest` | 3544 | 0.5271 | 0.5766 | 0.7397 | 0.7085 | 0.5056 | 0.5531 | 0.7626 | 0.7317 | | 25 | `emo_Intoxication_Altered_States_of_Consciousness` | 3544 | 0.1044 | 0.1877 | 0.4826 | 0.2741 | 0.1046 | 0.1881 | 0.4474 | 0.2786 | | 26 | `emo_Jealousy_and_Envy` | 3544 | 0.0624 | 0.1400 | 0.6992 | 0.1091 | 0.0559 | 0.1254 | 0.7301 | 0.1062 | | 27 | `emo_Longing` | 3544 | 0.3539 | 0.5566 | 0.7124 | 0.5817 | 0.3355 | 0.5277 | 0.7392 | 0.5894 | | 28 | `emo_Malevolence_Malice` | 3544 | 0.1676 | 0.3200 | 0.6364 | 0.3145 | 0.1636 | 0.3124 | 0.6531 | 0.3299 | | 29 | `emo_Pain` | 3544 | 0.2795 | 0.4859 | 0.7490 | 0.5796 | 0.2580 | 0.4485 | 0.7785 | 0.5889 | | 30 | `emo_Pleasure_Ecstasy` | 3544 | 0.2679 | 0.4388 | 0.7381 | 0.5662 | 0.2471 | 0.4048 | 0.7743 | 0.5881 | | 31 | `emo_Pride` | 3544 | 0.5015 | 0.9645 | 0.6547 | 0.5704 | 0.4816 | 0.9263 | 0.6812 | 0.5928 | | 32 | `emo_Relief` | 3544 | 0.4468 | 0.6403 | 0.6572 | 0.5138 | 0.4189 | 0.6002 | 0.6932 | 0.5388 | | 33 | `emo_Sadness` | 3544 | 0.3507 | 0.4589 | 0.7617 | 0.5775 | 0.3287 | 0.4301 | 0.7960 | 0.5914 | | 34 | `emo_Sexual_Lust` | 3544 | 0.0958 | 0.1924 | 0.6895 | 0.2417 | 0.0930 | 0.1868 | 0.6770 | 0.2413 | | 35 | `emo_Shame` | 3544 | 0.0955 | 0.1822 | 0.6960 | 0.2302 | 0.0896 | 0.1710 | 0.7326 | 0.2385 | | 36 | `emo_Sourness` | 3544 | 0.3766 | 0.8710 | 0.6533 | 0.5006 | 0.3576 | 0.8271 | 0.6761 | 0.5236 | | 37 | `emo_Teasing` | 3544 | 0.4089 | 0.7599 | 0.6783 | 0.5625 | 0.4021 | 0.7473 | 0.6848 | 0.5687 | | 38 | `emo_Thankfulness_Gratitude` | 3544 | 0.2490 | 0.3602 | 0.7406 | 0.4788 | 0.2260 | 0.3269 | 0.7811 | 0.5085 | | 39 | `emo_Triumph` | 3544 | 0.3263 | 0.6180 | 0.6941 | 0.5031 | 0.3151 | 0.5967 | 0.7145 | 0.5123 | | 40 | `vn_AGEV_reg` | 3544 | 0.3185 | 0.2527 | 0.4460 | 0.4312 | 0.3093 | 0.2455 | 0.4414 | 0.4270 | | 41 | `vn_AROU_reg` | 3544 | 0.5437 | 0.4826 | 0.7243 | 0.6929 | 0.5284 | 0.4690 | 0.7376 | 0.7127 | | 42 | `vn_ARSH_reg` | 3544 | 0.4911 | 0.5604 | 0.3468 | 0.3323 | 0.4776 | 0.5450 | 0.3800 | 0.3731 | | 43 | `vn_ATCK_reg` | 3544 | 0.4661 | 0.3540 | 0.6870 | 0.6544 | 0.4635 | 0.3520 | 0.6932 | 0.6690 | | 44 | `vn_BKGN_reg` | 3544 | 0.3384 | 0.4472 | 0.6167 | 0.6001 | 0.3359 | 0.4439 | 0.6118 | 0.5989 | | 45 | `vn_BRGT_reg` | 3544 | 0.3844 | 0.4743 | 0.6861 | 0.6711 | 0.3833 | 0.4729 | 0.6862 | 0.6676 | | 46 | `vn_CHNK_reg` | 3544 | 0.4343 | 0.3484 | 0.7160 | 0.7076 | 0.4312 | 0.3459 | 0.7241 | 0.7166 | | 47 | `vn_CLRT_reg` | 3544 | 0.4416 | 0.4595 | 0.8017 | 0.6867 | 0.4334 | 0.4510 | 0.8035 | 0.6948 | | 48 | `vn_COGL_reg` | 3544 | 0.5413 | 0.5135 | 0.7345 | 0.7335 | 0.5232 | 0.4963 | 0.7542 | 0.7556 | | 49 | `vn_DARC_reg` | 3544 | 0.4769 | 0.3501 | 0.4287 | 0.3938 | 0.4663 | 0.3423 | 0.4556 | 0.4118 | | 50 | `vn_DFLU_reg` | 3544 | 0.5694 | 0.4820 | 0.8008 | 0.7930 | 0.5619 | 0.4757 | 0.8075 | 0.8025 | | 51 | `vn_EMPH_reg` | 3544 | 0.4545 | 0.2893 | 0.6617 | 0.6401 | 0.4437 | 0.2824 | 0.6713 | 0.6575 | | 52 | `vn_ESTH_reg` | 3544 | 0.4608 | 0.5917 | 0.6939 | 0.6557 | 0.4423 | 0.5680 | 0.7172 | 0.6910 | | 53 | `vn_EXPL_reg` | 3544 | 0.1151 | 0.3516 | 0.2977 | 0.1599 | 0.1088 | 0.3323 | 0.2825 | 0.1475 | | 54 | `vn_FOCS_reg` | 3544 | 0.4337 | 0.3907 | 0.6086 | 0.4914 | 0.4098 | 0.3692 | 0.6395 | 0.5212 | | 55 | `vn_FULL_reg` | 3544 | 0.3593 | 0.4066 | 0.5248 | 0.5048 | 0.3466 | 0.3922 | 0.5457 | 0.5310 | | 56 | `vn_GEND_reg` | 3544 | 0.5459 | 0.3428 | 0.8868 | 0.8254 | 0.5735 | 0.3601 | 0.8836 | 0.8256 | | 57 | `vn_HARM_reg` | 3544 | 0.4697 | 0.5674 | 0.6904 | 0.6492 | 0.4512 | 0.5450 | 0.7242 | 0.6856 | | 58 | `vn_METL_reg` | 3544 | 0.4922 | 0.5140 | 0.6026 | 0.5499 | 0.4832 | 0.5046 | 0.6160 | 0.5692 | | 59 | `vn_RANG_reg` | 3544 | 0.5298 | 0.5943 | 0.7133 | 0.7109 | 0.5125 | 0.5749 | 0.7316 | 0.7340 | | 60 | `vn_RCQL_reg` | 3544 | 0.4049 | 0.4498 | 0.6185 | 0.6144 | 0.4084 | 0.4536 | 0.6119 | 0.6139 | | 61 | `vn_REGS_reg` | 3544 | 0.4636 | 0.3690 | 0.8743 | 0.8353 | 0.4724 | 0.3760 | 0.8706 | 0.8324 | | 62 | `vn_RESP_reg` | 3544 | 0.4666 | 0.4362 | 0.8088 | 0.7550 | 0.4444 | 0.4155 | 0.8261 | 0.7708 | | 63 | `vn_ROUG_reg` | 3544 | 0.5316 | 0.6029 | 0.7408 | 0.7184 | 0.5216 | 0.5916 | 0.7509 | 0.7339 | | 64 | `vn_R_CHST_reg` | 3544 | 0.4628 | 0.5450 | 0.7755 | 0.7775 | 0.4618 | 0.5439 | 0.7789 | 0.7825 | | 65 | `vn_R_HEAD_reg` | 3544 | 0.5097 | 0.6189 | 0.7728 | 0.7756 | 0.5088 | 0.6178 | 0.7719 | 0.7709 | | 66 | `vn_R_MASK_reg` | 3544 | 0.4401 | 0.4301 | 0.7060 | 0.6774 | 0.4369 | 0.4269 | 0.7118 | 0.6882 | | 67 | `vn_R_MIXD_reg` | 3544 | 0.4149 | 0.4957 | 0.6177 | 0.5642 | 0.4046 | 0.4835 | 0.6333 | 0.5872 | | 68 | `vn_R_NASL_reg` | 3544 | 0.4205 | 0.3982 | 0.4078 | 0.3681 | 0.4063 | 0.3847 | 0.4375 | 0.4148 | | 69 | `vn_R_ORAL_reg` | 3544 | 0.3597 | 0.4834 | 0.6342 | 0.6177 | 0.3516 | 0.4725 | 0.6439 | 0.6331 | | 70 | `vn_R_THRT_reg` | 3544 | 0.4400 | 0.5464 | 0.6952 | 0.6673 | 0.4269 | 0.5301 | 0.7085 | 0.6869 | | 71 | `vn_SMTH_reg` | 3544 | 0.4650 | 0.4184 | 0.7178 | 0.7130 | 0.4497 | 0.4046 | 0.7341 | 0.7314 | | 72 | `vn_STNC_reg` | 3544 | 0.6337 | 0.5565 | 0.6729 | 0.6393 | 0.6053 | 0.5316 | 0.7046 | 0.6692 | | 73 | `vn_STRU_reg` | 3544 | 0.4699 | 0.4766 | 0.7898 | 0.7386 | 0.4659 | 0.4724 | 0.7923 | 0.7465 | | 74 | `vn_S_ASMR_reg` | 3544 | 0.5717 | 0.4147 | 0.7268 | 0.6562 | 0.5479 | 0.3974 | 0.7467 | 0.6773 | | 75 | `vn_S_AUTH_reg` | 3544 | 0.7103 | 0.6837 | 0.7108 | 0.6953 | 0.6791 | 0.6537 | 0.7404 | 0.7287 | | 76 | `vn_S_CART_reg` | 3544 | 0.6246 | 0.4562 | 0.6659 | 0.5904 | 0.6244 | 0.4560 | 0.6549 | 0.5974 | | 77 | `vn_S_CASU_reg` | 3544 | 0.7998 | 0.6269 | 0.6985 | 0.7130 | 0.7619 | 0.5971 | 0.7284 | 0.7413 | | 78 | `vn_S_CONV_reg` | 3544 | 0.6987 | 0.5353 | 0.7534 | 0.7393 | 0.6797 | 0.5207 | 0.7696 | 0.7600 | | 79 | `vn_S_DRAM_reg` | 3544 | 0.6893 | 0.6050 | 0.7871 | 0.7833 | 0.6613 | 0.5804 | 0.8062 | 0.8037 | | 80 | `vn_S_FORM_reg` | 3544 | 0.6598 | 0.6543 | 0.7273 | 0.7316 | 0.6426 | 0.6373 | 0.7447 | 0.7507 | | 81 | `vn_S_MONO_reg` | 3544 | 0.8582 | 0.6745 | 0.6484 | 0.6408 | 0.8287 | 0.6514 | 0.6732 | 0.6646 | | 82 | `vn_S_NARR_reg` | 3544 | 0.7113 | 0.4956 | 0.7613 | 0.7441 | 0.6963 | 0.4851 | 0.7702 | 0.7539 | | 83 | `vn_S_NEWS_reg` | 3544 | 0.4990 | 0.9056 | 0.6904 | 0.6742 | 0.4881 | 0.8858 | 0.7103 | 0.6925 | | 84 | `vn_S_PLAY_reg` | 3544 | 0.8496 | 0.6535 | 0.6578 | 0.6327 | 0.7954 | 0.6119 | 0.7052 | 0.6929 | | 85 | `vn_S_RANT_reg` | 3544 | 0.6595 | 0.5168 | 0.5771 | 0.4963 | 0.6214 | 0.4870 | 0.6483 | 0.5575 | | 86 | `vn_S_STRY_reg` | 3544 | 0.7550 | 0.5417 | 0.6284 | 0.6181 | 0.7309 | 0.5245 | 0.6553 | 0.6423 | | 87 | `vn_S_TECH_reg` | 3544 | 0.7180 | 0.5895 | 0.7318 | 0.7560 | 0.6733 | 0.5529 | 0.7636 | 0.7858 | | 88 | `vn_S_WHIS_reg` | 3544 | 0.5282 | 0.3304 | 0.6921 | 0.6367 | 0.4893 | 0.3061 | 0.7177 | 0.6580 | | 89 | `vn_TEMP_reg` | 3544 | 0.4112 | 0.3424 | 0.7711 | 0.7410 | 0.4034 | 0.3360 | 0.7789 | 0.7479 | | 90 | `vn_TENS_reg` | 3544 | 0.5110 | 0.6140 | 0.7012 | 0.6084 | 0.4922 | 0.5914 | 0.7223 | 0.6347 | | 91 | `vn_VALN_reg` | 3544 | 0.7516 | 0.6652 | 0.6443 | 0.6258 | 0.6572 | 0.5816 | 0.7361 | 0.7195 | | 92 | `vn_VALS_reg` | 3544 | 0.4307 | 0.3921 | 0.4358 | 0.3966 | 0.4074 | 0.3709 | 0.5246 | 0.4904 | | 93 | `vn_VFLX_reg` | 3544 | 0.3329 | 0.2758 | 0.1541 | 0.1562 | 0.3150 | 0.2611 | 0.1945 | 0.1881 | | 94 | `vn_VOLT_reg` | 3544 | 0.4732 | 0.5448 | 0.7000 | 0.6775 | 0.4696 | 0.5408 | 0.7030 | 0.6810 | | 95 | `vn_VULN_reg` | 3544 | 0.6739 | 0.6803 | 0.7484 | 0.7045 | 0.6464 | 0.6526 | 0.7704 | 0.7329 | | 96 | `vn_WARM_reg` | 3544 | 0.4646 | 0.5134 | 0.5795 | 0.5277 | 0.4362 | 0.4820 | 0.6354 | 0.5966 | | 97 | `genuineness_0_6` | 3544 | 0.4774 | 0.2971 | 0.5561 | 0.5957 | 0.4704 | 0.2928 | 0.5616 | 0.5979 | | 98 | `blend_0_10` | 2391 | 0.8773 | 0.3228 | 0.2822 | 0.3074 | 0.8417 | 0.3097 | 0.3151 | 0.3216 | | 99 | `R_quality` | 2413 | 0.1025 | 0.3095 | 0.9180 | 0.8908 | 0.0972 | 0.2937 | 0.9246 | 0.8994 | | 100 | `eiv_extra:Age` | 3368 | 0.1391 | 0.4716 | 0.8225 | 0.7135 | 0.1344 | 0.4557 | 0.8399 | 0.7298 | | 101 | `eiv_extra:Arousal` | 3368 | 0.2575 | 0.5906 | 0.8606 | 0.8790 | 0.2399 | 0.5503 | 0.8760 | 0.8948 | | 102 | `eiv_extra:Authenticity` | 3368 | 0.0930 | 0.3718 | 0.9065 | 0.9116 | 0.0872 | 0.3487 | 0.9187 | 0.9240 | | 103 | `eiv_extra:Background_Noise` | 3368 | 0.1414 | 0.5358 | 0.8755 | 0.8651 | 0.1324 | 0.5018 | 0.8925 | 0.8775 | | 104 | `eiv_extra:Confident_vs._Hesitant` | 3368 | 0.1701 | 0.5593 | 0.9002 | 0.9047 | 0.1545 | 0.5081 | 0.9178 | 0.9235 | | 105 | `eiv_extra:Gender` | 3368 | 0.2693 | 0.2700 | 0.9372 | 0.9178 | 0.2534 | 0.2540 | 0.9454 | 0.9256 | | 106 | `eiv_extra:High-Pitched_vs._Low-Pitched` | 3368 | 0.0894 | 0.3577 | 0.9407 | 0.9421 | 0.0802 | 0.3206 | 0.9523 | 0.9529 | | 107 | `eiv_extra:Monotone_vs._Expressive` | 3368 | 0.1635 | 0.3478 | 0.9344 | 0.9346 | 0.1480 | 0.3147 | 0.9471 | 0.9475 | | 108 | `eiv_extra:Recording_Quality` | 3368 | 0.1351 | 0.4062 | 0.9468 | 0.9465 | 0.1256 | 0.3775 | 0.9545 | 0.9551 | | 109 | `eiv_extra:Serious_vs._Humorous` | 3368 | 0.1966 | 0.5766 | 0.8678 | 0.8451 | 0.1877 | 0.5506 | 0.8799 | 0.8581 | | 110 | `eiv_extra:Soft_vs._Harsh` | 3368 | 0.1796 | 0.5623 | 0.7959 | 0.8006 | 0.1717 | 0.5377 | 0.8127 | 0.8101 | | 111 | `eiv_extra:Submissive_vs._Dominant` | 3368 | 0.1440 | 0.4455 | 0.8502 | 0.8475 | 0.1348 | 0.4172 | 0.8704 | 0.8672 | | 112 | `eiv_extra:Valence` | 3368 | 0.3711 | 0.9724 | 0.8634 | 0.8222 | 0.3430 | 0.8989 | 0.8808 | 0.8320 | | 113 | `eiv_extra:Vulnerable_vs._Emotionally_Detached` | 3368 | 0.1589 | 0.3646 | 0.9240 | 0.9231 | 0.1509 | 0.3462 | 0.9328 | 0.9323 | | 114 | `eiv_extra:Warm_vs._Cold` | 3368 | 0.1916 | 0.6387 | 0.8206 | 0.7840 | 0.1808 | 0.6030 | 0.8421 | 0.8045 | | 115 | `eiv_extra:score_background_quality` | 3368 | 0.1026 | 0.1479 | 0.9482 | 0.8992 | 0.0996 | 0.1436 | 0.9518 | 0.9062 | | 116 | `eiv_extra:score_content_enjoyment` | 3368 | 0.0765 | 0.2518 | 0.9534 | 0.9410 | 0.0728 | 0.2398 | 0.9577 | 0.9455 | | 117 | `eiv_extra:score_overall_quality` | 3368 | 0.0819 | 0.1613 | 0.9615 | 0.9363 | 0.0786 | 0.1549 | 0.9645 | 0.9391 | | 118 | `eiv_extra:score_speech_quality` | 3368 | 0.0479 | 0.1915 | 0.8332 | 0.7329 | 0.0445 | 0.1779 | 0.8574 | 0.7713 | | 119 | `audiobox:CE` | 3544 | 0.2039 | 0.2006 | 0.9459 | 0.9448 | 0.1902 | 0.1871 | 0.9546 | 0.9518 | | 120 | `audiobox:CU` | 3544 | 0.2151 | 0.2153 | 0.9356 | 0.9426 | 0.1998 | 0.2000 | 0.9454 | 0.9494 | | 121 | `audiobox:PC` | 3544 | 0.1312 | 0.0812 | 0.9535 | 0.8055 | 0.1204 | 0.0745 | 0.9619 | 0.8117 | | 122 | `audiobox:PQ` | 3544 | 0.2230 | 0.2537 | 0.9322 | 0.9349 | 0.2057 | 0.2340 | 0.9429 | 0.9408 | | 123 | `dnsmos:SIG_raw` | 3368 | 0.1495 | 0.2234 | 0.9061 | 0.8503 | 0.1441 | 0.2153 | 0.9133 | 0.8592 | | 124 | `dnsmos:BAK_raw` | 3368 | 0.2130 | 0.1893 | 0.9205 | 0.8742 | 0.2055 | 0.1826 | 0.9258 | 0.8810 | | 125 | `dnsmos:OVRL_raw` | 3368 | 0.1801 | 0.2215 | 0.9296 | 0.8908 | 0.1743 | 0.2144 | 0.9339 | 0.8969 | | 126 | `dnsmos:SIG` | 3368 | 0.0938 | 0.1872 | 0.9033 | 0.8464 | 0.0901 | 0.1797 | 0.9108 | 0.8563 | | 127 | `dnsmos:BAK` | 3368 | 0.1440 | 0.1551 | 0.9250 | 0.8659 | 0.1395 | 0.1502 | 0.9301 | 0.8699 | | 128 | `dnsmos:OVRL` | 3368 | 0.1200 | 0.1978 | 0.9327 | 0.8902 | 0.1161 | 0.1913 | 0.9371 | 0.8965 | | 129 | `dnsmos:P808_MOS` | 3368 | 0.1283 | 0.2585 | 0.9176 | 0.8696 | 0.1235 | 0.2487 | 0.9230 | 0.8819 | | 130 | `burst_count_log1p` | 3544 | 0.2474 | 0.3815 | 0.8305 | 0.8198 | 0.2492 | 0.3843 | 0.8273 | 0.8164 | | 131 | `voiceclap_attribute:Affection` | 3544 | 0.5078 | 0.9158 | 0.5838 | 0.4973 | 0.4877 | 0.8795 | 0.6392 | 0.5332 | | 132 | `voiceclap_attribute:Age` | 3368 | 0.1578 | 0.2902 | 0.8188 | 0.7483 | 0.1504 | 0.2766 | 0.8313 | 0.7624 | | 133 | `voiceclap_attribute:Amusement` | 3544 | 0.5768 | 1.0050 | 0.7031 | 0.6284 | 0.5497 | 0.9579 | 0.7273 | 0.6620 | | 134 | `voiceclap_attribute:Anger` | 3544 | 0.4138 | 0.8156 | 0.6069 | 0.4560 | 0.3912 | 0.7711 | 0.6250 | 0.4527 | | 135 | `voiceclap_attribute:Arousal` | 3368 | 0.2821 | 0.4532 | 0.8416 | 0.8537 | 0.2863 | 0.4599 | 0.8385 | 0.8489 | | 136 | `voiceclap_attribute:Astonishment_Surprise` | 3544 | 0.4791 | 1.0005 | 0.6094 | 0.5371 | 0.4772 | 0.9966 | 0.6296 | 0.5575 | | 137 | `voiceclap_attribute:Authenticity` | 3368 | 0.0910 | 0.3639 | 0.8870 | 0.8965 | 0.0938 | 0.3751 | 0.8810 | 0.8852 | | 138 | `voiceclap_attribute:Awe` | 3544 | 0.4582 | 1.7662 | 0.6179 | 0.4764 | 0.4115 | 1.5863 | 0.7134 | 0.5330 | | 139 | `voiceclap_attribute:Background_Noise` | 3368 | 0.0962 | 0.3644 | 0.8873 | 0.8881 | 0.1001 | 0.3789 | 0.8754 | 0.8735 | | 140 | `voiceclap_attribute:Bitterness` | 3544 | 0.4419 | 1.7677 | 0.5745 | 0.4753 | 0.4032 | 1.6130 | 0.6572 | 0.5235 | | 141 | `voiceclap_attribute:Concentration` | 3544 | 0.5051 | 0.9521 | 0.6333 | 0.6247 | 0.5024 | 0.9471 | 0.6427 | 0.6434 | | 142 | `voiceclap_attribute:Confident_vs._Hesitant` | 3368 | 0.2211 | 0.4716 | 0.8456 | 0.8455 | 0.2246 | 0.4791 | 0.8385 | 0.8364 | | 143 | `voiceclap_attribute:Confusion` | 3544 | 0.4366 | 1.3024 | 0.4817 | 0.4449 | 0.4303 | 1.2836 | 0.5233 | 0.4768 | | 144 | `voiceclap_attribute:Contemplation` | 3544 | 0.6185 | 1.5260 | 0.6700 | 0.6556 | 0.6125 | 1.5110 | 0.6794 | 0.6710 | | 145 | `voiceclap_attribute:Contempt` | 3544 | 0.3968 | 1.1680 | 0.5738 | 0.3881 | 0.3705 | 1.0907 | 0.6316 | 0.4085 | | 146 | `voiceclap_attribute:Contentment` | 3544 | 0.6851 | 2.2319 | 0.6315 | 0.6411 | 0.6264 | 2.0406 | 0.6953 | 0.6942 | | 147 | `voiceclap_attribute:Disappointment` | 3544 | 0.5357 | 1.8439 | 0.6072 | 0.5267 | 0.4829 | 1.6620 | 0.6753 | 0.5702 | | 148 | `voiceclap_attribute:Disgust` | 3544 | 0.2720 | 1.0882 | 0.5145 | 0.3850 | 0.2585 | 1.0342 | 0.6078 | 0.4214 | | 149 | `voiceclap_attribute:Distress` | 3544 | 0.5497 | 1.5707 | 0.7163 | 0.6190 | 0.4861 | 1.3889 | 0.7655 | 0.6572 | | 150 | `voiceclap_attribute:Doubt` | 3544 | 0.5253 | 2.1012 | 0.5091 | 0.5297 | 0.5016 | 2.0064 | 0.5578 | 0.5718 | | 151 | `voiceclap_attribute:Elation` | 3544 | 0.4601 | 1.1538 | 0.6368 | 0.5550 | 0.4292 | 1.0765 | 0.6666 | 0.5830 | | 152 | `voiceclap_attribute:Embarrassment` | 3544 | 0.1896 | 0.6023 | 0.3048 | 0.2565 | 0.1799 | 0.5715 | 0.2576 | 0.2234 | | 153 | `voiceclap_attribute:Emotional_Numbness` | 3544 | 0.2522 | 0.9473 | 0.3121 | 0.2899 | 0.2507 | 0.9419 | 0.3167 | 0.2929 | | 154 | `voiceclap_attribute:Fatigue_Exhaustion` | 3544 | 0.5180 | 1.9897 | 0.7102 | 0.6597 | 0.4850 | 1.8632 | 0.7385 | 0.6754 | | 155 | `voiceclap_attribute:Fear` | 3544 | 0.2930 | 1.1722 | 0.5463 | 0.4014 | 0.2913 | 1.1654 | 0.5694 | 0.4252 | | 156 | `voiceclap_attribute:Gender` | 3368 | 0.2379 | 0.2368 | 0.9505 | 0.9383 | 0.2396 | 0.2385 | 0.9482 | 0.9338 | | 157 | `voiceclap_attribute:Helplessness` | 3544 | 0.5094 | 1.8990 | 0.7148 | 0.5998 | 0.4535 | 1.6905 | 0.7714 | 0.6370 | | 158 | `voiceclap_attribute:High-Pitched_vs._Low-Pitched` | 3368 | 0.1280 | 0.3427 | 0.8771 | 0.8826 | 0.1272 | 0.3405 | 0.8782 | 0.8872 | | 159 | `voiceclap_attribute:Hope_Enthusiasm_Optimism` | 3544 | 0.6346 | 1.5974 | 0.6166 | 0.6083 | 0.5817 | 1.4642 | 0.6760 | 0.6699 | | 160 | `voiceclap_attribute:Impatience_and_Irritability` | 3544 | 0.5706 | 1.0528 | 0.5800 | 0.4786 | 0.5425 | 1.0009 | 0.6083 | 0.5011 | | 161 | `voiceclap_attribute:Infatuation` | 3544 | 0.1870 | 0.5623 | 0.3724 | 0.2935 | 0.1843 | 0.5543 | 0.4069 | 0.3121 | | 162 | `voiceclap_attribute:Interest` | 3544 | 0.5735 | 1.3054 | 0.6784 | 0.6338 | 0.5473 | 1.2458 | 0.7189 | 0.6805 | | 163 | `voiceclap_attribute:Intoxication_Altered_States_of_Consciousness` | 3544 | 0.1527 | 0.3292 | 0.2480 | 0.1951 | 0.1551 | 0.3344 | 0.2136 | 0.1762 | | 164 | `voiceclap_attribute:Jealousy_&_Envy` | 3544 | 0.1094 | 0.4377 | 0.1123 | 0.1094 | 0.0991 | 0.3965 | 0.1322 | 0.0996 | | 165 | `voiceclap_attribute:Longing` | 3544 | 0.4261 | 1.7046 | 0.5481 | 0.5207 | 0.4193 | 1.6770 | 0.5829 | 0.5493 | | 166 | `voiceclap_attribute:Malevolence_Malice` | 3544 | 0.2214 | 0.6399 | 0.3811 | 0.2645 | 0.2264 | 0.6543 | 0.4078 | 0.3097 | | 167 | `voiceclap_attribute:Monotone_vs._Expressive` | 3368 | 0.2390 | 0.4006 | 0.8655 | 0.8564 | 0.2346 | 0.3934 | 0.8682 | 0.8571 | | 168 | `voiceclap_attribute:Pain` | 3544 | 0.3618 | 1.4471 | 0.6585 | 0.5520 | 0.3395 | 1.3579 | 0.6994 | 0.5805 | | 169 | `voiceclap_attribute:Pleasure_Ecstasy` | 3544 | 0.3606 | 1.1972 | 0.6002 | 0.4927 | 0.3280 | 1.0889 | 0.6716 | 0.5413 | | 170 | `voiceclap_attribute:Pride` | 3544 | 0.5591 | 1.6019 | 0.5473 | 0.5229 | 0.5450 | 1.5617 | 0.5676 | 0.5412 | | 171 | `voiceclap_attribute:Recording_Quality` | 3368 | 0.1340 | 0.4083 | 0.9180 | 0.9200 | 0.1342 | 0.4090 | 0.9179 | 0.9209 | | 172 | `voiceclap_attribute:Relief` | 3544 | 0.5277 | 1.2686 | 0.5372 | 0.4394 | 0.5009 | 1.2041 | 0.5875 | 0.4669 | | 173 | `voiceclap_attribute:Sadness` | 3544 | 0.4434 | 1.4169 | 0.6306 | 0.5121 | 0.3948 | 1.2618 | 0.7219 | 0.5584 | | 174 | `voiceclap_attribute:Serious_vs._Humorous` | 3368 | 0.2265 | 0.3971 | 0.8688 | 0.8286 | 0.2183 | 0.3827 | 0.8774 | 0.8265 | | 175 | `voiceclap_attribute:Sexual_Lust` | 3544 | 0.1490 | 0.3565 | 0.2604 | 0.2024 | 0.1441 | 0.3448 | 0.2510 | 0.2158 | | 176 | `voiceclap_attribute:Shame` | 3544 | 0.1468 | 0.5211 | 0.2995 | 0.1960 | 0.1480 | 0.5254 | 0.2968 | 0.1973 | | 177 | `voiceclap_attribute:Soft_vs._Harsh` | 3368 | 0.1702 | 0.4727 | 0.8318 | 0.8489 | 0.1759 | 0.4883 | 0.8189 | 0.8327 | | 178 | `voiceclap_attribute:Sourness` | 3544 | 0.4401 | 1.5150 | 0.5642 | 0.4561 | 0.4040 | 1.3909 | 0.6341 | 0.5007 | | 179 | `voiceclap_attribute:Submissive_vs._Dominant` | 3368 | 0.1595 | 0.4444 | 0.8189 | 0.8267 | 0.1583 | 0.4408 | 0.8211 | 0.8262 | | 180 | `voiceclap_attribute:Teasing` | 3544 | 0.4834 | 1.2742 | 0.6079 | 0.5211 | 0.4720 | 1.2443 | 0.6292 | 0.5347 | | 181 | `voiceclap_attribute:Thankfulness_Gratitude` | 3544 | 0.3695 | 0.6980 | 0.5228 | 0.4199 | 0.3478 | 0.6570 | 0.6076 | 0.4657 | | 182 | `voiceclap_attribute:Triumph` | 3544 | 0.4139 | 1.1535 | 0.5471 | 0.4353 | 0.4051 | 1.1293 | 0.5561 | 0.4452 | | 183 | `voiceclap_attribute:Valence` | 3368 | 0.4702 | 0.7416 | 0.7155 | 0.7111 | 0.4603 | 0.7259 | 0.7213 | 0.7128 | | 184 | `voiceclap_attribute:Vulnerable_vs._Emotionally_Detached` | 3368 | 0.2181 | 0.3924 | 0.8508 | 0.8581 | 0.2110 | 0.3796 | 0.8608 | 0.8564 | | 185 | `voiceclap_attribute:Warm_vs._Cold` | 3368 | 0.2497 | 0.6254 | 0.7692 | 0.7600 | 0.2392 | 0.5989 | 0.7818 | 0.7665 | | 186 | `voiceclap_attribute:duration` | 3368 | 0.9696 | 0.1531 | 0.9725 | 0.9677 | 1.0755 | 0.1698 | 0.9654 | 0.9629 | | 187 | `voiceclap_attribute:score_background_quality` | 3368 | 0.1229 | 0.2464 | 0.8736 | 0.8523 | 0.1219 | 0.2446 | 0.8726 | 0.8543 | | 188 | `voiceclap_attribute:score_content_enjoyment` | 3368 | 0.0622 | 0.2489 | 0.9033 | 0.9136 | 0.0648 | 0.2593 | 0.8934 | 0.9053 | | 189 | `voiceclap_attribute:score_overall_quality` | 3368 | 0.0981 | 0.2378 | 0.9158 | 0.9063 | 0.0979 | 0.2374 | 0.9155 | 0.9048 | | 190 | `voiceclap_attribute:score_speech_quality` | 3368 | 0.0272 | 0.1087 | 0.8570 | 0.8576 | 0.0273 | 0.1091 | 0.8536 | 0.8539 | | 191 | `voiceclap_attribute:talking_speed` | 3368 | 1.4987 | 0.3994 | 0.9193 | 0.9197 | 1.3978 | 0.3725 | 0.9310 | 0.9277 | ## Canonical 53-class vocal-burst vocabulary The classifier predicts these categories. Fine labels outside this vocabulary use the explicit aliases/fallbacks in `gemini_burst_mapping.json`; the fine-to-group23 map is also preserved. | ID | Canonical vocal-burst class | | ---: | --- | | 0 | Other / rare burst | | 1 | Affirmative Grunt | | 2 | Ahem | | 3 | Breath | | 4 | Breathy Giggle | | 5 | Cackle | | 6 | Childlike Giggle | | 7 | Chuckle | | 8 | Contented Sigh | | 9 | Cough | | 10 | Crying | | 11 | Displeased Grunt | | 12 | Effort Grunt | | 13 | Exasperated Sigh | | 14 | Exhausted Groan | | 15 | Fearful Gasp | | 16 | Frustrated Groan | | 17 | Growl | | 18 | Guffaw | | 19 | Gulp | | 20 | Hiccup | | 21 | Hum | | 22 | Kissing Sound | | 23 | Laughter | | 24 | Lip Smack | | 25 | Low Mumble | | 26 | Mournful Wail | | 27 | Nervous Giggle | | 28 | Nervous Gulp | | 29 | Purr | | 30 | Quiet Sob | | 31 | Relief Sigh | | 32 | Scream | | 33 | Sharp Inhale | | 34 | Sharp Whistle | | 35 | Shriek | | 36 | Sigh | | 37 | Sneeze | | 38 | Snicker | | 39 | Sniff | | 40 | Snort | | 41 | Snorting Giggle | | 42 | Soft Whistle | | 43 | Spitting | | 44 | Surprised Gasp | | 45 | Swallows | | 46 | Throat Clearing | | 47 | Tongue Click | | 48 | Trembling Whimper | | 49 | Tsk | | 50 | Whispered Mumble | | 51 | Wistful Sigh | | 52 | Yawn |