- Humaneness Ears: Base and Medium audio understanding
- Results and related releases
- What the models output
- Quick start: local CPU or GPU inference
- Architecture
- Training data, annotation priority and splits
- Exact fine-tuning procedure
- Normalization
- How the data were mounted and how to reproduce
- Held-out Flash target agreement
- Public human-label benchmarks
- P3 retention and limitations
- Annotation models, taxonomies and packages
- Repository contents and export verification
- License and attribution
- All 192 scalar targets: exact held-out results
- Canonical 53-class vocal-burst vocabulary
- Results and related releases
Humaneness Ears: Base and Medium audio understanding
By Christoph Schuhmann · LAION · 7 October 2026 · CC BY 4.0
This repository releases the best completed two-epoch fine-tunes
of LAION's encoder-only Whisper Base and Whisper Small multitask models.
The release names are Humaneness Ears Base (Whisper Base backbone) and
Humaneness Ears Medium (Whisper Small backbone). Given a recording,
one forward pass predicts emotions, speaking style, sound quality, speaker-characteristic
vectors and the location and type of non-speech vocal events such as laughter or crying.
Both models are included: base/model.safetensors and medium/model.safetensors.
The older training logs, evaluation JSON and the machine-readable listening manifest retain the
small key as the historical Whisper backbone label; it refers to Humaneness Ears Medium.
These are audio encoders with task heads. There is no Whisper text decoder, ASR transcript, generated caption or language model. Gemini transcripts and captions helped prepare annotations; captions were not model inputs or an autoregressive training objective. CPS is a scalar prediction rather than a transcript. The requested separate MOSS caption LoRAs are outside this release and were not trained as part of these two runs.
Explore 1,000 real-audio examples in the Humaneness Ears atlas: 50 timed vocal bursts, 20-query Orange timbre and identity neighbor pages, 40 EmoNet and 57 VoiceNet top-ten rankings, and five-tier quality pages. Download the Gemini annotation dataset.
Open the full benchmark: Humaneness Ears Base/Medium versus CLAPv2 XS/M — human-label results, all 192 targets, Orange speaker-vector cosine and vocal-burst timing. Jump to the CLAP comparison.
Results and related releases
- Comprehensive English benchmark report: 36 frozen embedding probes, original and Humaneness Ears Base/Medium after Gemini tuning, and the two full CLAPv2 fine-tunes; human-label emotion benchmarks, all 192 score targets, speaker vectors, validation, burst timing and limitations.
- Which model is best for each task? includes genuineness and vocal-burst blend.
- Every target and validation result.
- Original layered Whisper weights: the Base S4 and original Whisper Small S3 checkpoints are the actual starting states for this release.
Medium has better Orange speaker-vector cosine and lower AudioBox/DNSMOS target error than Base, and a higher EmoNet intensity correlation. Base has slightly better CPS error, Flash frame F1 and VoiceNet-Emo mean rank correlation. A frozen probe can outperform both on other tasks: see the comprehensive report for all winners. This release does not claim one universal best model or statistical significance from small numerical differences.
What the models output
| Output | Shape / scale | Meaning |
|---|---|---|
| EmoNet emotions | 40 scalars, reference 0–4 | Fine-grained emotion intensity; simultaneous regression, not a mutually exclusive class. |
| VoiceNet dimensions | 57 scalars, mostly 0–6 | Delivery, register, timbre and style; Flash BKGN is 0–4 and EXPL 0–2. |
| Genuineness | 1 scalar, 0–6 | The annotation rubric's perceived genuineness. |
| Vocal-burst blend | 1 scalar, 0–10 | How a vocal event blends with surrounding speech; meaningful only when a burst is present. |
| R_quality | 1 legacy scalar | Its distinct original teacher target and units; not overwritten with another quality construct. |
| Empathic Insight Plus extras | 19 scalars | Voice traits, arousal/valence, recording/background quality and content enjoyment. |
| AudioBox Aesthetics | 4 scalars | CE: content enjoyment; CU: content usefulness; PC: production complexity; PQ: production quality. |
| DNSMOS | 7 scalars | SIG, BAK, OVRL, their raw versions and P808 MOS. |
| Burst count | 1 scalar trained in log1p space | Number of annotated events; inference additionally reports expm1 of the nonnegative prediction. |
| VoiceCLAP attributes | 61 scalars | Additional emotion, voice and quality targets; some emotion targets overlap the first 40. |
| Characters per second (CPS) | 1 separate scalar | Unicode transcript characters, including spaces/punctuation, divided by total recording duration; bracketed nonlexical placeholders removed. |
| Orange timbre | 128D unit vector | Regression to the Orange timbre teacher embedding. |
| Orange identity | 250D unit vector | Regression to the Orange identity teacher embedding; not a person's name or an identification guarantee. |
| Burst occupancy | 20 ms frame probabilities | Whether each encoder frame belongs to at least one vocal burst. |
| Independent burst events | Start and end seconds | Onset plus duration proposals can overlap; maximum 32 at the default evaluation setting. |
| Burst type | 53-way event-local class probabilities | The top three labels and uncalibrated softmax probabilities for each predicted interval. |
There are 192 jointly predicted scalar outputs plus the separate CPS head. A complete, ordered list of exact output keys and all their held-out results appears at the end of this card. The models return raw units, training-standardized values and display z-scores. Regression outputs are continuous and are not automatically clipped to the annotation's ordinal range.
Quick start: local CPU or GPU inference
Use Python 3.10 or newer. The measured production environment used Python 3.13.5,
PyTorch 2.9.1, Transformers 5.14.1, SoundFile 0.14 and Safetensors. Install the appropriate
PyTorch CPU/CUDA wheel for your machine; see requirements.txt for the other dependencies.
Inference uses a custom model.py, not pipeline('automatic-speech-recognition') or
WhisperForConditionalGeneration.from_pretrained.
Download inference assets without the large optimizer checkpoints:
from huggingface_hub import snapshot_download
snapshot_download(
repo_id="laion/humaneness-ears-base-medium",
local_dir="humaneness-ears",
allow_patterns=["*.py", "requirements.txt", "*.json", "base/*.json",
"medium/*.json", "base/model.safetensors", "medium/model.safetensors"],
)
cd humaneness-ears
pip install -r requirements.txt
python inference.py clip.wav --model medium --device cpu --threads 4 > medium_prediction.json
python inference.py clip.wav --model base --device cpu --threads 1 > base_prediction.json
python inference.py clip.wav --model medium --device cuda:0 --threads 4 \
--include-frame-probabilities > gpu_prediction.json
Input recordings must be 0.1–30 seconds. SoundFile decodes the supplied audio, stereo is averaged to mono, and SciPy resamples to 16 kHz. The bundled Whisper frontend creates 80-bin log-mel features. Longer recordings need an explicit external chunking policy. No original OpenAI weights, teacher checkpoints or Gemini API access are required at inference.
Reuse the model when processing multiple files:
from inference import Predictor
predictor = Predictor(".", model_size="medium", device="cpu", threads=4)
result = predictor.predict("clip.wav", include_frames=True)
print(result["scores_raw"]["genuineness_0_6"])
print(result["scores_z"]["emo_Amusement"])
print(result["burst_event_proposals"]) # start_s, end_s, top3_classes
--frame-threshold and --event-threshold default to 0.5; --max-events defaults to 32.
This matches the new report's proposal operating point. Lowering the event threshold can
increase recall and false positives. Frame regions merge overlapping events, whereas the
independent proposal head can represent them separately. Neither class softmax values nor
sigmoid values have been calibrated as real-world confidence estimates.
Architecture
Base uses six encoder layers, hidden width 512, and 26,441,735 total parameters. Medium uses 12 encoder layers, hidden width 768, and 97,102,031 total parameters. These counts include our heads and exclude the absent Whisper decoder. The encoder alone contains 20,590,592 / 88,154,112 parameters respectively.
From the final encoder sequence we concatenate masked mean, minimum, maximum and standard deviation. Separately, each transformer layer contributes its masked mean vector. Shared per-family projection MLPs reduce these means to 32 or 64 dimensions, and a sample-specific softmax gate combines layer information. These features add residuals to the 192-score head. Speaker heads use 128D layer-mixture routes; CPS uses 32D. The frame head predicts occupancy; the proposal head predicts independent onset logits and log durations. A 64D layer-mixture route for event classification averages only frames inside that event from every layer, then combines them with the final-layer local mean and a 53-way classifier.
model.py is the exact architecture used in training. Exported weights preserve every saved
tensor without quantization or conversion. The inherited Whisper configuration still contains
some unused decoder fields; architectures and is_encoder_decoder are set for this encoder-only
release. No decoder tensors are present. The custom inference code must be used.
Training data, annotation priority and splits
The selected pool contains 66,199 exact-audio unique clips / 197.36 hours after deduplication across the four selections. Valid Flash annotations exist for 65,163 clips / 194.17 hours. The 1,036 invalid or provider-blocked responses are excluded from this tuning stage. They were not converted into zero-emotion or no-burst targets.
| Split actually used | Clips | Purpose |
|---|---|---|
| Training | 58,380 | Gradient updates for both full fine-tunes. |
| Validation | 3,239 | Full validation after each epoch; selects best checkpoint. |
| Test | 3,544 | Matched before/after evaluation; not used for gradient updates or epoch selection. |
Sources are the balanced S1 100-hour, 400-bucket selection from TTS Scaling Ladder DE/EN, LAION's Got Talent raw, balanced audio snippets 40×3k and voice annotation POC. Final selections preserve all source/task memberships while annotating each exact compressed-audio SHA256 once. Source membership counts are nonexclusive. Deterministic 90/5/5 hashing groups explicit speaker IDs, otherwise known synthetic audio families, otherwise exact audio hashes; observed counts differ from exact percentages because groups are kept intact.
training/provenance/valid_split_summary.json records valid counts, hours, language strings,
memberships and the hash of the exact prepared target snapshot. The older TRAINING_READY.json
also contains split counts including invalid responses; those are not the training counts above.
Source revisions and original pool details are in training/provenance/selected_pool.json.
The data, Gemini responses and teacher TARs are not redistributed by this model repository.
The recorded annotation model identifier is gemini-3.8-flash. Schema-valid Flash outputs
take priority for 40 emotions, 57 VoiceNet dimensions, genuineness, blend, event counts,
event start/end times and transcript-derived CPS. Corresponding VoiceCLAP emotion attributes
also take the Flash value. Distinct constructs such as DNSMOS BAK or the original authenticity
axis keep their own teachers. A null blend on a no-burst clip is masked, not replaced by an old
blend score. Missing values remain masked. Speech-only objectives use domain masks; Orange
speaker losses require valid known single-speaker clips. Teacher vectors are preserved even
when their training loss is masked.
The prompt permits 78 vocal fine labels. Explicit aliases map compatible labels into the
existing 53-class checkpoint vocabulary; unrepresented labels map to Other / rare burst.
See gemini_burst_mapping.json for every mapping. Original fine labels and event captions remain
in the annotation sidecars. End times may be clamped to the decoded duration within the accepted
0.3-second annotation tolerance, with original timing and the clamp flag retained in the data.
Exact fine-tuning procedure
Starting models were the lowest-validation-loss Base S4 and Whisper Small S3 snapshots from
the original S1→S10 layered curriculum campaign. Each snapshot contains only the progress up
to its own stage, rather than the final S10 weights: Base's original update was 266,238;
Medium's was 237,434. The original campaign ran one epoch per nested stage with 10% P3 synthetic
mixture exposures per stage and a continuous cosine schedule. Its full code and recorded run
configs are included under training/ and training/original_curriculum/.
Gemini tuning strictly restored all encoder and head weights, started a fresh AdamW optimizer and cosine schedule, and unfroze the entire encoder, including its positional embedding table. This is full fine-tuning, not LoRA or a frozen-encoder probe.
| Setting | Base | Medium |
|---|---|---|
| Epochs | 2 | 2 |
| Actual optimizer updates | 914 | 914 |
| Best checkpoint | Epoch 2 | Epoch 2 |
| Peak encoder learning rate | 1e-5 | 5e-6 |
| Peak head learning rate | 1e-4 | 1e-4 |
| Warmup | 5% (46 updates) | 5% (46 updates) |
| Cosine minimum / peak | 10% | 10% |
| Weight decay | 0.01 | 0.01 |
| Global effective batch | 128 | 128 |
| Microbatch / GPU | 32 | 16 |
| Accumulation | 1 | 2 |
| Data workers / GPU | 4 | 4 |
| Hardware | 1 node, 4 GH200 GPUs | 1 node, 4 GH200 GPUs |
| Precision | BF16 autocast; FP32 saved weights | BF16 autocast; FP32 saved weights |
| Gradient clipping | Global norm 1.0 | Global norm 1.0 |
| Seed | 20261005 | 20261005 |
| Selected class-weighted validation loss | 1.7683769 | 1.6989245 |
| Training Slurm job | 2199189 | 2199190 |
| Recorded allocation elapsed time | 10m17s | 10m31s |
There is no P3 replay in this extra two-epoch stage. A DDP sampler may pad a few examples
to distribute batches equally; validation is evaluated completely on rank zero without
duplicating its samples. The full config, class counts and epoch loss logs are under training/.
Losses
The implementation is training/train_layered_curriculum.py::loss_terms, reused unchanged
by training/gemini_finetune/train_whisper.py:
- Masked Huber regression, delta 1, on standardized scalar targets; per-family weights
are the exact
GROUPSconstants intraining/multitask_p3_smoke.py. - CPS: standardized Huber delta 1, weight 0.2.
- Speaker vectors: 0.2 × (cosine distance + 0.1 × Huber delta 0.1), only valid single speakers.
- Frames: weighted BCE (positive weight 5), weight 0.5, plus soft Dice weight 0.25.
- Onsets: weighted BCE (positive weight 100), weight 0.2.
- Event log duration: Huber delta 0.5, weight 0.1, evaluated at labeled starts.
- Event class: categorical cross entropy, weight 0.2, label smoothing 0.02; training-only class weights sqrt(median event count / class count), clipped to 0.25–4.
Every epoch saves latest and, when improved, best states. Published base/checkpoint.pt
and medium/checkpoint.pt are the exact best files: model, AdamW, scheduler, four-rank RNG,
completed epoch, update, normalization, taxonomy and original config. They support epoch-boundary
continuation with the same DDP world size. Both runs already finished epoch 2; extending training
requires an explicit new phase/schedule, not blindly resuming the completed two-epoch plan.
Inference uses only Safetensors and never loads the .pt training pickle.
Normalization
raw_score = standardized_network_output × training_std + training_mean
display_z = (raw_score − original_training_teacher_median) / original_training_teacher_std
CPS_raw = CPS_network_output × CPS_training_std + CPS_training_mean
training_normalization.json is the original frozen checkpoint normalization, preserved
through Gemini tuning. It is required for correct raw units. display_normalization.json
is the train-only median/std reference from the original release, retained for compatible
visualization and rankings. gemini_train_statistics.json separately stores train-only Flash
mean/std/counts; its inherited “provisional” description refers to the incremental parser,
but the file here is the completed final export. It was not substituted into network decoding.
Display z-scores are relative context, not class probabilities or confidence intervals.
How the data were mounted and how to reproduce
Training ran on JUPITER Booster. /e/home, /e/project1, /e/scratch and /e/software
were existing shared HPC filesystems visible to every rank; no Docker bind mount or network
download occurred inside the GPU training. Source audio and same-audio teacher outputs stayed
in WebDataset TARs. TarReader caches a bounded set of open archives and reads named members;
it does not extract millions of individual files. Teachers are cached annotations, not called
inside the gradient loop.
| Original absolute path | Contents / use |
|---|---|
/e/home/jusers/schuhmann1/jupiter/whisper_score_regression |
Production training and inference source. |
/e/scratch/reformo/schuhmann1_moss/whisper_score_regression/hf_layered_release/model |
Original Base S4 / Whisper Small S3 states, frontend, normalization and classes. |
/e/scratch/reformo/schuhmann1_moss/whisper_score_regression/gemini_multisource_100h_20261004 |
Combined exact-audio selection and Gemini response sidecars. |
/e/scratch/reformo/schuhmann1_moss/whisper_score_regression/gemini_finetune_preparation_20261005/prepared |
targets.jsonl, final training gate, normalization and burst mapping. |
Same preparation root, teacher_sidecars/ |
SHA-bound TAR companions with Orange vectors and all additional teacher outputs. |
Same preparation root, training/whisper_{base,small} |
Exact best/latest states and per-epoch metrics. |
/e/scratch/reformo/schuhmann1_moss/code/fastgen.sh |
Python/software paths for the original offline HPC environment. |
targets.jsonl contains audio_tar / audio_member and embedding tar / member references.
On another machine, obtain the same complete prepared data, copy TARs and rewrite the absolute
TAR prefixes to your local mounts while preserving member names and audio hashes. Provide
checkpoint_normalization.json, gemini_burst_mapping.json and TRAINING_READY.json beside it.
The repository includes the parser, schema prompt and teacher backfill code; recreating its
data requires the upstream datasets, teacher weights and your own annotation service access.
There is no credential embedded here. Source availability and terms still apply.
To reproduce the two-epoch experiment, download the original Base S4 or Whisper Small S3 .pt
initializer from the previous repository, not this release's already-tuned state. Then:
pip install -r requirements-training.txt
python training/make_finetune_config.py --model base \
--data-root /your/prepared --initial-checkpoint /your/original/base/S4/checkpoint.pt \
--output-dir /your/new-base-run --config /your/base-run.json
python -m torch.distributed.run --standalone --nproc_per_node=4 \
training/gemini_finetune/train_whisper.py --config /your/base-run.json
For the historical Whisper Small training run, use --model small and the original small/S3/checkpoint.pt. The config helper
changes paths while preserving recorded hyperparameters. training/gemini_finetune/train_whisper.sbatch
is the original submitted Slurm launcher, with original mount paths for reference. Evaluator
scripts likewise preserve original study paths; adapt study_paths.py and benchmark data paths
if reproducing the full comparison. Safe-weight export tools and verification receipts are included.
Held-out Flash target agreement
Both models are tested on the same 3,544 valid Flash test clips. These targets are machine annotations, not a fresh human panel. Speaker cosine uses 1,718 valid single-speaker clips; blend uses 2,391 clips with non-null blend. Family MAE is the unweighted mean over valid axes.
| Flash test target / metric | Base | Medium |
|---|---|---|
| Burst frame F1 ↑ | 0.6929 | 0.6892 |
| Burst event F1, IoU ≥0.5 ↑ | 0.1945 | 0.1949 |
| Burst start MAE on IoU ≥0.1 matches, seconds ↓ | 0.1201 | 0.1150 |
| Burst end MAE on IoU ≥0.1 matches, seconds ↓ | 0.1973 | 0.2006 |
| 53-class accuracy with reference spans ↑ | 0.4280 | 0.4483 |
| 53-class accuracy on matched predicted spans ↑ | 0.4189 | 0.4370 |
| Orange timbre cosine ↑ (1,718 valid single-speaker clips) | 0.9038 | 0.9223 |
| Orange identity cosine ↑ (1,718 valid single-speaker clips) | 0.8541 | 0.8741 |
| CPS raw MAE, characters/second ↓ | 0.9012 | 0.9309 |
| Genuineness raw MAE on 0–6 scale ↓ (3,544 clips) | 0.4774 | 0.4704 |
| Genuineness normalized MAE ↓ | 0.2971 | 0.2928 |
| Blend raw MAE on 0–10 scale ↓ (2,391 valid clips) | 0.8773 | 0.8417 |
| Blend normalized MAE ↓ | 0.3228 | 0.3097 |
| 40 emotions, mean normalized MAE ↓ | 0.5873 | 0.5605 |
| 57 VoiceNet dimensions, mean normalized MAE ↓ | 0.4967 | 0.4813 |
| AudioBox four axes, mean normalized MAE ↓ | 0.1877 | 0.1739 |
| DNSMOS seven outputs, mean normalized MAE ↓ | 0.2047 | 0.1975 |
Normalized MAE is measured in frozen training standard deviations. Burst interval metrics use Hungarian matching of independent onset/duration proposals, not binary contiguous-frame regions. Boundary errors and predicted-span class accuracy are conditional on matches at IoU ≥0.1. Reference-span class accuracy gives the classifier the true interval; it is not end-to-end localization accuracy. Maximum proposals is 32, with a fixed 0.5 onset threshold. Event precision is low despite high recall; the report includes all IoU cutoffs and precision/recall.
The validation tables in evaluation/{base,small}/validation_metrics.json (historical backbone keys) use the existing
3,239 checkpoint-selection clips, not an untouched test. Supplementary evaluator loss uses
unit class weights, while the training selection loss above used the configured class weights.
Different loss values under those two policies are expected.
Public human-label benchmarks
| Human-label benchmark / metric | Base | Medium |
|---|---|---|
| EmoNet-Voice intensity, Pearson r ↑ | 0.4340 | 0.4512 |
| EmoNet-Voice intensity, Spearman ρ ↑ | 0.4577 | 0.4679 |
| VoiceNet-Emo, mean prompt Spearman ρ ↑ | 0.4288 | 0.4282 |
| VoiceNet-Ext / emolia-dim, mean prompt Spearman ρ ↑ | 0.1566 | 0.1502 |
| CREMA-D matched actor-CV score-MLP accuracy ↑ | 0.6310 | 0.6423 |
| RAVDESS matched actor-CV score-MLP accuracy ↑ | 0.6597 | 0.6639 |
EmoNet uses 12,000 mapped clips and fixed 0–4→0–10 endpoint scaling; two unsupported labels from the published 12,600-clip benchmark are omitted. VoiceNet-Emo uses 7,986 current-repository questions with at least two raters; VoiceNet-Ext uses 13,917. EXT and DIM refer to the same benchmark. Current snapshot coverage differs from paper coverage. The report includes unflagged, untruncated and alternate agreement cuts and explains optimistic oracle threshold numbers.
CREMA-D and RAVDESS figures are supervised matched nested actor-cross-validation score adapters,
not zero-shot paper numbers. Each adapter sees the model's 192 raw predictions, a train-fold
StandardScaler and a 64-unit GELU MLP: 12,742 / 12,872 parameters respectively. Five outer actor
folds and three inner folds choose learning rate 0.001/0.003 and 20/50 epochs. Outer test actors
never fit that fold's scaler or adapter. These benchmark-only adapters are not the model's burst
classifier and are not automatically applied by inference.py. Full fold records, metrics and
confusions are included; the report compares the same adapter budget across all backbones.
Upstream training/source overlap with the public benchmarks was not exhaustively audited.
P3 retention and limitations
On the identical fixed 2,000-clip P3 test audit, frame F1 changes from 0.9622 to 0.8051 for Base and 0.9693 to 0.8392 for Medium after Gemini tuning. Flash test frame F1 rises from 0.3072/0.3171 to 0.6929/0.6892 respectively. This is a target/domain tradeoff: there was no P3 replay in this phase. All matched pre/post P3 validation/test JSONs are included. It is not valid to claim universal burst improvement from Flash agreement alone.
The selected pool is not a representative random sample of all natural audio. Quality scores imitate AudioBox/DNSMOS/Empathic teachers rather than measuring new human MOS. Speaker cosine measures regression agreement with Orange, not verification accuracy or human identity. The burst classifier groups unsupported fine labels, event proposals have many false positives, and emotional or expressive clips can differ from the synthetic P3 construction labels. Recordings longer than 30 seconds and multi-speaker embeddings require additional policies.
Annotation models, taxonomies and packages
| Supervision / package | Primary reference |
|---|---|
| Original Whisper encoders | OpenAI Whisper Base, OpenAI Whisper Small, Whisper code. |
| Gemini annotation schema | The exact 108 KB master prompt is training/gemini_s10_smoke/gemini_voice_annotation_master_prompt_detailed.txt; parser and validation code are included. |
| 40 emotions | EmoNet taxonomy, emotion annotations toolkit. |
| 57 voice dimensions | VoiceNet taxonomy, commercial dimension predictors. |
| Original genuineness teacher | VoiceCLAP commercial genuineness; matching targets superseded by valid Flash labels. |
| Original blend teacher | VoiceCLAP commercial vocalburst blend; matching targets superseded by valid Flash labels. |
| VoiceCLAP attributes / intermediate embedding | Commercial VoiceCLAP, attribute heads; 768D intermediate embeddings are cached in data, not an output head. |
| Empathic extras / original emotion teacher | Empathic Insight Voice Plus, BUD-E Whisper; 3072D teacher features are cached, not an output head. |
| AudioBox | Meta AudioBox Aesthetics. |
| DNSMOS | Microsoft DNS Challenge / DNSMOS. |
| Orange timbre, 128D | Orange Speaker-wavLM-tbr, pinned annotation revision b8d2608d56f18e2b6e27ba566cf69132e2c2ad6c. |
| Orange identity, 250D | Orange Speaker-wavLM-id, pinned revision abbb3c7b8d220ceebc33d9bd3bb5aedb580342c7. |
| Timbre generation implementation | LAION generation script. |
| Vocal event taxonomy / locator | LAION voice taxonomies, vocalburst locator, canonical classes and Flash mapping bundled here. |
| Human emotion/style benchmark | LAION emolia-bench, EmoNet-Voice Bench. |
Teacher inference runs on the exact selected waveform. New scores and float16 target vectors are persisted in SHA-bound companion TARs; the gradient loop reuses those sidecars. Their model revisions and target-level provenance remain in the prepared annotations.
Repository contents and export verification
| Path | Contents |
|---|---|
base/model.safetensors, medium/model.safetensors |
Complete best Gemini-tuned encoder and head weights; no decoder. |
base/checkpoint.pt, medium/checkpoint.pt |
Exact best resumable training states including optimizer, scheduler and per-rank RNG. |
{base,medium}/config.json, preprocessor_config.json |
Corresponding encoder architecture and Whisper log-mel frontend. |
model.py, inference.py, requirements.txt |
Complete standalone inference, cached Predictor API and CLI. |
training/, vocal_burst_pool/p3/ |
Actual training source and local dependency closure, target parser, prompt, backfill and evaluation code, configs, logs and reproduction helper. |
training_normalization.json, display_normalization.json, gemini_train_statistics.json |
Frozen decoding, contextual display and separate final Flash train-only statistics. |
classes.json, gemini_burst_mapping.json |
53 canonical class names and all Flash aliases/fallbacks. |
evaluation/{base,small}/ (historical keys) |
Flash validation/test, pre/post P3 retention, public benchmarks and actor-adapter folds. |
verification/{base,small}.json (historical keys) |
Bitwise safe-weight and model-output parity; real held-out audio API checks, dimensions and timestamp validity. |
release_manifest.json |
File sizes and SHA256 hashes, original checkpoint hashes and provenance. |
release_tools/ |
Export, verification and publication source. |
LICENSE, NOTICE.md |
CC BY 4.0 and upstream attribution. |
No audio, source transcripts, raw Gemini responses or service credentials are published by this model repository. Model verification uses a real held-out waveform internally and saves only its hash and output-contract evidence. Evaluation JSONs contain aggregate benchmark results, prompts/classes and fold protocols; no listening audio is copied into this release.
License and attribution
The new fine-tuned weights, new code and documentation are released by LAION under
Creative Commons Attribution 4.0 International. Credit Christoph Schuhmann
and LAION, link this repository and indicate changes. This license does not relicense
upstream source audio, teacher models or external packages. The original OpenAI Whisper
components remain MIT; Orange's identity teacher retains its upstream CC BY-SA 3.0 terms.
See NOTICE.md and upstream model/data cards for their notices.
@misc{schuhmann_humaneness_ears_2026,
title = {Humaneness Ears: Base and Medium Encoder-Only Audio Understanding},
author = {Schuhmann, Christoph and {LAION}},
year = {2026},
url = {https://huggingface.co/laion/humaneness-ears-base-medium}
}
All 192 scalar targets: exact held-out results
These are Flash test targets, with the remaining AudioBox/DNSMOS/Orange/Empathic families
provided by their own teachers. Raw MAE uses each target's units. Norm. MAE uses its frozen
training standard deviation. r is Pearson and ρ is Spearman. Missing metrics remain
unavailable. The source JSON has separate counts for every model/target.
| Index | Exact target key | Valid N | Base raw MAE | Base norm. MAE | Base r | Base ρ | Medium raw MAE | Medium norm. MAE | Medium r | Medium ρ |
|---|---|---|---|---|---|---|---|---|---|---|
| 0 | emo_Affection |
3544 | 0.4034 | 0.7418 | 0.7163 | 0.5685 | 0.3911 | 0.7192 | 0.7345 | 0.5839 |
| 1 | emo_Amusement |
3544 | 0.5187 | 0.5978 | 0.7413 | 0.6742 | 0.4879 | 0.5623 | 0.7606 | 0.6992 |
| 2 | emo_Anger |
3544 | 0.3462 | 0.5968 | 0.6770 | 0.4799 | 0.3376 | 0.5819 | 0.6729 | 0.5020 |
| 3 | emo_Astonishment_Surprise |
3544 | 0.4007 | 0.7257 | 0.6906 | 0.5951 | 0.3838 | 0.6950 | 0.7075 | 0.6081 |
| 4 | emo_Awe |
3544 | 0.3259 | 0.7764 | 0.7652 | 0.5528 | 0.3002 | 0.7152 | 0.7946 | 0.5699 |
| 5 | emo_Bitterness |
3544 | 0.3718 | 0.9136 | 0.6692 | 0.5244 | 0.3488 | 0.8571 | 0.6932 | 0.5548 |
| 6 | emo_Concentration |
3544 | 0.4721 | 0.6463 | 0.6937 | 0.6930 | 0.4634 | 0.6343 | 0.7024 | 0.7005 |
| 7 | emo_Confusion |
3544 | 0.3659 | 0.6860 | 0.6370 | 0.5124 | 0.3462 | 0.6490 | 0.6617 | 0.5392 |
| 8 | emo_Contemplation |
3544 | 0.5538 | 0.9403 | 0.7380 | 0.7293 | 0.5367 | 0.9113 | 0.7552 | 0.7449 |
| 9 | emo_Contempt |
3544 | 0.3103 | 0.6878 | 0.6823 | 0.4370 | 0.2999 | 0.6648 | 0.6994 | 0.4587 |
| 10 | emo_Contentment |
3544 | 0.5967 | 1.1843 | 0.6973 | 0.6890 | 0.5675 | 1.1264 | 0.7270 | 0.7214 |
| 11 | emo_Disappointment |
3544 | 0.4548 | 0.6485 | 0.7058 | 0.5985 | 0.4328 | 0.6171 | 0.7290 | 0.6164 |
| 12 | emo_Disgust |
3544 | 0.2244 | 0.4576 | 0.6700 | 0.4164 | 0.2140 | 0.4365 | 0.7002 | 0.4420 |
| 13 | emo_Distress |
3544 | 0.4677 | 0.5722 | 0.7701 | 0.6569 | 0.4445 | 0.5438 | 0.7917 | 0.6834 |
| 14 | emo_Doubt |
3544 | 0.4587 | 0.9491 | 0.6336 | 0.6068 | 0.4360 | 0.9020 | 0.6652 | 0.6390 |
| 15 | emo_Elation |
3544 | 0.3643 | 0.5071 | 0.7254 | 0.5958 | 0.3513 | 0.4890 | 0.7449 | 0.6097 |
| 16 | emo_Embarrassment |
3544 | 0.1472 | 0.3367 | 0.5970 | 0.2711 | 0.1434 | 0.3282 | 0.6350 | 0.2784 |
| 17 | emo_Emotional_Numbness |
3544 | 0.2516 | 0.2617 | 0.5849 | 0.4138 | 0.2434 | 0.2532 | 0.6081 | 0.4254 |
| 18 | emo_Fatigue_Exhaustion |
3544 | 0.4472 | 0.8269 | 0.7700 | 0.6846 | 0.4385 | 0.8108 | 0.7746 | 0.6890 |
| 19 | emo_Fear |
3544 | 0.2303 | 0.3764 | 0.7266 | 0.4491 | 0.2201 | 0.3597 | 0.7331 | 0.4586 |
| 20 | emo_Helplessness |
3544 | 0.4093 | 0.5781 | 0.7839 | 0.6445 | 0.3889 | 0.5492 | 0.8065 | 0.6608 |
| 21 | emo_Hope_Enthusiasm_Optimism |
3544 | 0.5715 | 0.7535 | 0.6933 | 0.6714 | 0.5412 | 0.7135 | 0.7238 | 0.7046 |
| 22 | emo_Impatience_and_Irritability |
3544 | 0.5024 | 0.7343 | 0.6582 | 0.5313 | 0.4835 | 0.7068 | 0.6653 | 0.5517 |
| 23 | emo_Infatuation |
3544 | 0.1289 | 0.2382 | 0.7034 | 0.3314 | 0.1205 | 0.2226 | 0.7270 | 0.3399 |
| 24 | emo_Interest |
3544 | 0.5271 | 0.5766 | 0.7397 | 0.7085 | 0.5056 | 0.5531 | 0.7626 | 0.7317 |
| 25 | emo_Intoxication_Altered_States_of_Consciousness |
3544 | 0.1044 | 0.1877 | 0.4826 | 0.2741 | 0.1046 | 0.1881 | 0.4474 | 0.2786 |
| 26 | emo_Jealousy_and_Envy |
3544 | 0.0624 | 0.1400 | 0.6992 | 0.1091 | 0.0559 | 0.1254 | 0.7301 | 0.1062 |
| 27 | emo_Longing |
3544 | 0.3539 | 0.5566 | 0.7124 | 0.5817 | 0.3355 | 0.5277 | 0.7392 | 0.5894 |
| 28 | emo_Malevolence_Malice |
3544 | 0.1676 | 0.3200 | 0.6364 | 0.3145 | 0.1636 | 0.3124 | 0.6531 | 0.3299 |
| 29 | emo_Pain |
3544 | 0.2795 | 0.4859 | 0.7490 | 0.5796 | 0.2580 | 0.4485 | 0.7785 | 0.5889 |
| 30 | emo_Pleasure_Ecstasy |
3544 | 0.2679 | 0.4388 | 0.7381 | 0.5662 | 0.2471 | 0.4048 | 0.7743 | 0.5881 |
| 31 | emo_Pride |
3544 | 0.5015 | 0.9645 | 0.6547 | 0.5704 | 0.4816 | 0.9263 | 0.6812 | 0.5928 |
| 32 | emo_Relief |
3544 | 0.4468 | 0.6403 | 0.6572 | 0.5138 | 0.4189 | 0.6002 | 0.6932 | 0.5388 |
| 33 | emo_Sadness |
3544 | 0.3507 | 0.4589 | 0.7617 | 0.5775 | 0.3287 | 0.4301 | 0.7960 | 0.5914 |
| 34 | emo_Sexual_Lust |
3544 | 0.0958 | 0.1924 | 0.6895 | 0.2417 | 0.0930 | 0.1868 | 0.6770 | 0.2413 |
| 35 | emo_Shame |
3544 | 0.0955 | 0.1822 | 0.6960 | 0.2302 | 0.0896 | 0.1710 | 0.7326 | 0.2385 |
| 36 | emo_Sourness |
3544 | 0.3766 | 0.8710 | 0.6533 | 0.5006 | 0.3576 | 0.8271 | 0.6761 | 0.5236 |
| 37 | emo_Teasing |
3544 | 0.4089 | 0.7599 | 0.6783 | 0.5625 | 0.4021 | 0.7473 | 0.6848 | 0.5687 |
| 38 | emo_Thankfulness_Gratitude |
3544 | 0.2490 | 0.3602 | 0.7406 | 0.4788 | 0.2260 | 0.3269 | 0.7811 | 0.5085 |
| 39 | emo_Triumph |
3544 | 0.3263 | 0.6180 | 0.6941 | 0.5031 | 0.3151 | 0.5967 | 0.7145 | 0.5123 |
| 40 | vn_AGEV_reg |
3544 | 0.3185 | 0.2527 | 0.4460 | 0.4312 | 0.3093 | 0.2455 | 0.4414 | 0.4270 |
| 41 | vn_AROU_reg |
3544 | 0.5437 | 0.4826 | 0.7243 | 0.6929 | 0.5284 | 0.4690 | 0.7376 | 0.7127 |
| 42 | vn_ARSH_reg |
3544 | 0.4911 | 0.5604 | 0.3468 | 0.3323 | 0.4776 | 0.5450 | 0.3800 | 0.3731 |
| 43 | vn_ATCK_reg |
3544 | 0.4661 | 0.3540 | 0.6870 | 0.6544 | 0.4635 | 0.3520 | 0.6932 | 0.6690 |
| 44 | vn_BKGN_reg |
3544 | 0.3384 | 0.4472 | 0.6167 | 0.6001 | 0.3359 | 0.4439 | 0.6118 | 0.5989 |
| 45 | vn_BRGT_reg |
3544 | 0.3844 | 0.4743 | 0.6861 | 0.6711 | 0.3833 | 0.4729 | 0.6862 | 0.6676 |
| 46 | vn_CHNK_reg |
3544 | 0.4343 | 0.3484 | 0.7160 | 0.7076 | 0.4312 | 0.3459 | 0.7241 | 0.7166 |
| 47 | vn_CLRT_reg |
3544 | 0.4416 | 0.4595 | 0.8017 | 0.6867 | 0.4334 | 0.4510 | 0.8035 | 0.6948 |
| 48 | vn_COGL_reg |
3544 | 0.5413 | 0.5135 | 0.7345 | 0.7335 | 0.5232 | 0.4963 | 0.7542 | 0.7556 |
| 49 | vn_DARC_reg |
3544 | 0.4769 | 0.3501 | 0.4287 | 0.3938 | 0.4663 | 0.3423 | 0.4556 | 0.4118 |
| 50 | vn_DFLU_reg |
3544 | 0.5694 | 0.4820 | 0.8008 | 0.7930 | 0.5619 | 0.4757 | 0.8075 | 0.8025 |
| 51 | vn_EMPH_reg |
3544 | 0.4545 | 0.2893 | 0.6617 | 0.6401 | 0.4437 | 0.2824 | 0.6713 | 0.6575 |
| 52 | vn_ESTH_reg |
3544 | 0.4608 | 0.5917 | 0.6939 | 0.6557 | 0.4423 | 0.5680 | 0.7172 | 0.6910 |
| 53 | vn_EXPL_reg |
3544 | 0.1151 | 0.3516 | 0.2977 | 0.1599 | 0.1088 | 0.3323 | 0.2825 | 0.1475 |
| 54 | vn_FOCS_reg |
3544 | 0.4337 | 0.3907 | 0.6086 | 0.4914 | 0.4098 | 0.3692 | 0.6395 | 0.5212 |
| 55 | vn_FULL_reg |
3544 | 0.3593 | 0.4066 | 0.5248 | 0.5048 | 0.3466 | 0.3922 | 0.5457 | 0.5310 |
| 56 | vn_GEND_reg |
3544 | 0.5459 | 0.3428 | 0.8868 | 0.8254 | 0.5735 | 0.3601 | 0.8836 | 0.8256 |
| 57 | vn_HARM_reg |
3544 | 0.4697 | 0.5674 | 0.6904 | 0.6492 | 0.4512 | 0.5450 | 0.7242 | 0.6856 |
| 58 | vn_METL_reg |
3544 | 0.4922 | 0.5140 | 0.6026 | 0.5499 | 0.4832 | 0.5046 | 0.6160 | 0.5692 |
| 59 | vn_RANG_reg |
3544 | 0.5298 | 0.5943 | 0.7133 | 0.7109 | 0.5125 | 0.5749 | 0.7316 | 0.7340 |
| 60 | vn_RCQL_reg |
3544 | 0.4049 | 0.4498 | 0.6185 | 0.6144 | 0.4084 | 0.4536 | 0.6119 | 0.6139 |
| 61 | vn_REGS_reg |
3544 | 0.4636 | 0.3690 | 0.8743 | 0.8353 | 0.4724 | 0.3760 | 0.8706 | 0.8324 |
| 62 | vn_RESP_reg |
3544 | 0.4666 | 0.4362 | 0.8088 | 0.7550 | 0.4444 | 0.4155 | 0.8261 | 0.7708 |
| 63 | vn_ROUG_reg |
3544 | 0.5316 | 0.6029 | 0.7408 | 0.7184 | 0.5216 | 0.5916 | 0.7509 | 0.7339 |
| 64 | vn_R_CHST_reg |
3544 | 0.4628 | 0.5450 | 0.7755 | 0.7775 | 0.4618 | 0.5439 | 0.7789 | 0.7825 |
| 65 | vn_R_HEAD_reg |
3544 | 0.5097 | 0.6189 | 0.7728 | 0.7756 | 0.5088 | 0.6178 | 0.7719 | 0.7709 |
| 66 | vn_R_MASK_reg |
3544 | 0.4401 | 0.4301 | 0.7060 | 0.6774 | 0.4369 | 0.4269 | 0.7118 | 0.6882 |
| 67 | vn_R_MIXD_reg |
3544 | 0.4149 | 0.4957 | 0.6177 | 0.5642 | 0.4046 | 0.4835 | 0.6333 | 0.5872 |
| 68 | vn_R_NASL_reg |
3544 | 0.4205 | 0.3982 | 0.4078 | 0.3681 | 0.4063 | 0.3847 | 0.4375 | 0.4148 |
| 69 | vn_R_ORAL_reg |
3544 | 0.3597 | 0.4834 | 0.6342 | 0.6177 | 0.3516 | 0.4725 | 0.6439 | 0.6331 |
| 70 | vn_R_THRT_reg |
3544 | 0.4400 | 0.5464 | 0.6952 | 0.6673 | 0.4269 | 0.5301 | 0.7085 | 0.6869 |
| 71 | vn_SMTH_reg |
3544 | 0.4650 | 0.4184 | 0.7178 | 0.7130 | 0.4497 | 0.4046 | 0.7341 | 0.7314 |
| 72 | vn_STNC_reg |
3544 | 0.6337 | 0.5565 | 0.6729 | 0.6393 | 0.6053 | 0.5316 | 0.7046 | 0.6692 |
| 73 | vn_STRU_reg |
3544 | 0.4699 | 0.4766 | 0.7898 | 0.7386 | 0.4659 | 0.4724 | 0.7923 | 0.7465 |
| 74 | vn_S_ASMR_reg |
3544 | 0.5717 | 0.4147 | 0.7268 | 0.6562 | 0.5479 | 0.3974 | 0.7467 | 0.6773 |
| 75 | vn_S_AUTH_reg |
3544 | 0.7103 | 0.6837 | 0.7108 | 0.6953 | 0.6791 | 0.6537 | 0.7404 | 0.7287 |
| 76 | vn_S_CART_reg |
3544 | 0.6246 | 0.4562 | 0.6659 | 0.5904 | 0.6244 | 0.4560 | 0.6549 | 0.5974 |
| 77 | vn_S_CASU_reg |
3544 | 0.7998 | 0.6269 | 0.6985 | 0.7130 | 0.7619 | 0.5971 | 0.7284 | 0.7413 |
| 78 | vn_S_CONV_reg |
3544 | 0.6987 | 0.5353 | 0.7534 | 0.7393 | 0.6797 | 0.5207 | 0.7696 | 0.7600 |
| 79 | vn_S_DRAM_reg |
3544 | 0.6893 | 0.6050 | 0.7871 | 0.7833 | 0.6613 | 0.5804 | 0.8062 | 0.8037 |
| 80 | vn_S_FORM_reg |
3544 | 0.6598 | 0.6543 | 0.7273 | 0.7316 | 0.6426 | 0.6373 | 0.7447 | 0.7507 |
| 81 | vn_S_MONO_reg |
3544 | 0.8582 | 0.6745 | 0.6484 | 0.6408 | 0.8287 | 0.6514 | 0.6732 | 0.6646 |
| 82 | vn_S_NARR_reg |
3544 | 0.7113 | 0.4956 | 0.7613 | 0.7441 | 0.6963 | 0.4851 | 0.7702 | 0.7539 |
| 83 | vn_S_NEWS_reg |
3544 | 0.4990 | 0.9056 | 0.6904 | 0.6742 | 0.4881 | 0.8858 | 0.7103 | 0.6925 |
| 84 | vn_S_PLAY_reg |
3544 | 0.8496 | 0.6535 | 0.6578 | 0.6327 | 0.7954 | 0.6119 | 0.7052 | 0.6929 |
| 85 | vn_S_RANT_reg |
3544 | 0.6595 | 0.5168 | 0.5771 | 0.4963 | 0.6214 | 0.4870 | 0.6483 | 0.5575 |
| 86 | vn_S_STRY_reg |
3544 | 0.7550 | 0.5417 | 0.6284 | 0.6181 | 0.7309 | 0.5245 | 0.6553 | 0.6423 |
| 87 | vn_S_TECH_reg |
3544 | 0.7180 | 0.5895 | 0.7318 | 0.7560 | 0.6733 | 0.5529 | 0.7636 | 0.7858 |
| 88 | vn_S_WHIS_reg |
3544 | 0.5282 | 0.3304 | 0.6921 | 0.6367 | 0.4893 | 0.3061 | 0.7177 | 0.6580 |
| 89 | vn_TEMP_reg |
3544 | 0.4112 | 0.3424 | 0.7711 | 0.7410 | 0.4034 | 0.3360 | 0.7789 | 0.7479 |
| 90 | vn_TENS_reg |
3544 | 0.5110 | 0.6140 | 0.7012 | 0.6084 | 0.4922 | 0.5914 | 0.7223 | 0.6347 |
| 91 | vn_VALN_reg |
3544 | 0.7516 | 0.6652 | 0.6443 | 0.6258 | 0.6572 | 0.5816 | 0.7361 | 0.7195 |
| 92 | vn_VALS_reg |
3544 | 0.4307 | 0.3921 | 0.4358 | 0.3966 | 0.4074 | 0.3709 | 0.5246 | 0.4904 |
| 93 | vn_VFLX_reg |
3544 | 0.3329 | 0.2758 | 0.1541 | 0.1562 | 0.3150 | 0.2611 | 0.1945 | 0.1881 |
| 94 | vn_VOLT_reg |
3544 | 0.4732 | 0.5448 | 0.7000 | 0.6775 | 0.4696 | 0.5408 | 0.7030 | 0.6810 |
| 95 | vn_VULN_reg |
3544 | 0.6739 | 0.6803 | 0.7484 | 0.7045 | 0.6464 | 0.6526 | 0.7704 | 0.7329 |
| 96 | vn_WARM_reg |
3544 | 0.4646 | 0.5134 | 0.5795 | 0.5277 | 0.4362 | 0.4820 | 0.6354 | 0.5966 |
| 97 | genuineness_0_6 |
3544 | 0.4774 | 0.2971 | 0.5561 | 0.5957 | 0.4704 | 0.2928 | 0.5616 | 0.5979 |
| 98 | blend_0_10 |
2391 | 0.8773 | 0.3228 | 0.2822 | 0.3074 | 0.8417 | 0.3097 | 0.3151 | 0.3216 |
| 99 | R_quality |
2413 | 0.1025 | 0.3095 | 0.9180 | 0.8908 | 0.0972 | 0.2937 | 0.9246 | 0.8994 |
| 100 | eiv_extra:Age |
3368 | 0.1391 | 0.4716 | 0.8225 | 0.7135 | 0.1344 | 0.4557 | 0.8399 | 0.7298 |
| 101 | eiv_extra:Arousal |
3368 | 0.2575 | 0.5906 | 0.8606 | 0.8790 | 0.2399 | 0.5503 | 0.8760 | 0.8948 |
| 102 | eiv_extra:Authenticity |
3368 | 0.0930 | 0.3718 | 0.9065 | 0.9116 | 0.0872 | 0.3487 | 0.9187 | 0.9240 |
| 103 | eiv_extra:Background_Noise |
3368 | 0.1414 | 0.5358 | 0.8755 | 0.8651 | 0.1324 | 0.5018 | 0.8925 | 0.8775 |
| 104 | eiv_extra:Confident_vs._Hesitant |
3368 | 0.1701 | 0.5593 | 0.9002 | 0.9047 | 0.1545 | 0.5081 | 0.9178 | 0.9235 |
| 105 | eiv_extra:Gender |
3368 | 0.2693 | 0.2700 | 0.9372 | 0.9178 | 0.2534 | 0.2540 | 0.9454 | 0.9256 |
| 106 | eiv_extra:High-Pitched_vs._Low-Pitched |
3368 | 0.0894 | 0.3577 | 0.9407 | 0.9421 | 0.0802 | 0.3206 | 0.9523 | 0.9529 |
| 107 | eiv_extra:Monotone_vs._Expressive |
3368 | 0.1635 | 0.3478 | 0.9344 | 0.9346 | 0.1480 | 0.3147 | 0.9471 | 0.9475 |
| 108 | eiv_extra:Recording_Quality |
3368 | 0.1351 | 0.4062 | 0.9468 | 0.9465 | 0.1256 | 0.3775 | 0.9545 | 0.9551 |
| 109 | eiv_extra:Serious_vs._Humorous |
3368 | 0.1966 | 0.5766 | 0.8678 | 0.8451 | 0.1877 | 0.5506 | 0.8799 | 0.8581 |
| 110 | eiv_extra:Soft_vs._Harsh |
3368 | 0.1796 | 0.5623 | 0.7959 | 0.8006 | 0.1717 | 0.5377 | 0.8127 | 0.8101 |
| 111 | eiv_extra:Submissive_vs._Dominant |
3368 | 0.1440 | 0.4455 | 0.8502 | 0.8475 | 0.1348 | 0.4172 | 0.8704 | 0.8672 |
| 112 | eiv_extra:Valence |
3368 | 0.3711 | 0.9724 | 0.8634 | 0.8222 | 0.3430 | 0.8989 | 0.8808 | 0.8320 |
| 113 | eiv_extra:Vulnerable_vs._Emotionally_Detached |
3368 | 0.1589 | 0.3646 | 0.9240 | 0.9231 | 0.1509 | 0.3462 | 0.9328 | 0.9323 |
| 114 | eiv_extra:Warm_vs._Cold |
3368 | 0.1916 | 0.6387 | 0.8206 | 0.7840 | 0.1808 | 0.6030 | 0.8421 | 0.8045 |
| 115 | eiv_extra:score_background_quality |
3368 | 0.1026 | 0.1479 | 0.9482 | 0.8992 | 0.0996 | 0.1436 | 0.9518 | 0.9062 |
| 116 | eiv_extra:score_content_enjoyment |
3368 | 0.0765 | 0.2518 | 0.9534 | 0.9410 | 0.0728 | 0.2398 | 0.9577 | 0.9455 |
| 117 | eiv_extra:score_overall_quality |
3368 | 0.0819 | 0.1613 | 0.9615 | 0.9363 | 0.0786 | 0.1549 | 0.9645 | 0.9391 |
| 118 | eiv_extra:score_speech_quality |
3368 | 0.0479 | 0.1915 | 0.8332 | 0.7329 | 0.0445 | 0.1779 | 0.8574 | 0.7713 |
| 119 | audiobox:CE |
3544 | 0.2039 | 0.2006 | 0.9459 | 0.9448 | 0.1902 | 0.1871 | 0.9546 | 0.9518 |
| 120 | audiobox:CU |
3544 | 0.2151 | 0.2153 | 0.9356 | 0.9426 | 0.1998 | 0.2000 | 0.9454 | 0.9494 |
| 121 | audiobox:PC |
3544 | 0.1312 | 0.0812 | 0.9535 | 0.8055 | 0.1204 | 0.0745 | 0.9619 | 0.8117 |
| 122 | audiobox:PQ |
3544 | 0.2230 | 0.2537 | 0.9322 | 0.9349 | 0.2057 | 0.2340 | 0.9429 | 0.9408 |
| 123 | dnsmos:SIG_raw |
3368 | 0.1495 | 0.2234 | 0.9061 | 0.8503 | 0.1441 | 0.2153 | 0.9133 | 0.8592 |
| 124 | dnsmos:BAK_raw |
3368 | 0.2130 | 0.1893 | 0.9205 | 0.8742 | 0.2055 | 0.1826 | 0.9258 | 0.8810 |
| 125 | dnsmos:OVRL_raw |
3368 | 0.1801 | 0.2215 | 0.9296 | 0.8908 | 0.1743 | 0.2144 | 0.9339 | 0.8969 |
| 126 | dnsmos:SIG |
3368 | 0.0938 | 0.1872 | 0.9033 | 0.8464 | 0.0901 | 0.1797 | 0.9108 | 0.8563 |
| 127 | dnsmos:BAK |
3368 | 0.1440 | 0.1551 | 0.9250 | 0.8659 | 0.1395 | 0.1502 | 0.9301 | 0.8699 |
| 128 | dnsmos:OVRL |
3368 | 0.1200 | 0.1978 | 0.9327 | 0.8902 | 0.1161 | 0.1913 | 0.9371 | 0.8965 |
| 129 | dnsmos:P808_MOS |
3368 | 0.1283 | 0.2585 | 0.9176 | 0.8696 | 0.1235 | 0.2487 | 0.9230 | 0.8819 |
| 130 | burst_count_log1p |
3544 | 0.2474 | 0.3815 | 0.8305 | 0.8198 | 0.2492 | 0.3843 | 0.8273 | 0.8164 |
| 131 | voiceclap_attribute:Affection |
3544 | 0.5078 | 0.9158 | 0.5838 | 0.4973 | 0.4877 | 0.8795 | 0.6392 | 0.5332 |
| 132 | voiceclap_attribute:Age |
3368 | 0.1578 | 0.2902 | 0.8188 | 0.7483 | 0.1504 | 0.2766 | 0.8313 | 0.7624 |
| 133 | voiceclap_attribute:Amusement |
3544 | 0.5768 | 1.0050 | 0.7031 | 0.6284 | 0.5497 | 0.9579 | 0.7273 | 0.6620 |
| 134 | voiceclap_attribute:Anger |
3544 | 0.4138 | 0.8156 | 0.6069 | 0.4560 | 0.3912 | 0.7711 | 0.6250 | 0.4527 |
| 135 | voiceclap_attribute:Arousal |
3368 | 0.2821 | 0.4532 | 0.8416 | 0.8537 | 0.2863 | 0.4599 | 0.8385 | 0.8489 |
| 136 | voiceclap_attribute:Astonishment_Surprise |
3544 | 0.4791 | 1.0005 | 0.6094 | 0.5371 | 0.4772 | 0.9966 | 0.6296 | 0.5575 |
| 137 | voiceclap_attribute:Authenticity |
3368 | 0.0910 | 0.3639 | 0.8870 | 0.8965 | 0.0938 | 0.3751 | 0.8810 | 0.8852 |
| 138 | voiceclap_attribute:Awe |
3544 | 0.4582 | 1.7662 | 0.6179 | 0.4764 | 0.4115 | 1.5863 | 0.7134 | 0.5330 |
| 139 | voiceclap_attribute:Background_Noise |
3368 | 0.0962 | 0.3644 | 0.8873 | 0.8881 | 0.1001 | 0.3789 | 0.8754 | 0.8735 |
| 140 | voiceclap_attribute:Bitterness |
3544 | 0.4419 | 1.7677 | 0.5745 | 0.4753 | 0.4032 | 1.6130 | 0.6572 | 0.5235 |
| 141 | voiceclap_attribute:Concentration |
3544 | 0.5051 | 0.9521 | 0.6333 | 0.6247 | 0.5024 | 0.9471 | 0.6427 | 0.6434 |
| 142 | voiceclap_attribute:Confident_vs._Hesitant |
3368 | 0.2211 | 0.4716 | 0.8456 | 0.8455 | 0.2246 | 0.4791 | 0.8385 | 0.8364 |
| 143 | voiceclap_attribute:Confusion |
3544 | 0.4366 | 1.3024 | 0.4817 | 0.4449 | 0.4303 | 1.2836 | 0.5233 | 0.4768 |
| 144 | voiceclap_attribute:Contemplation |
3544 | 0.6185 | 1.5260 | 0.6700 | 0.6556 | 0.6125 | 1.5110 | 0.6794 | 0.6710 |
| 145 | voiceclap_attribute:Contempt |
3544 | 0.3968 | 1.1680 | 0.5738 | 0.3881 | 0.3705 | 1.0907 | 0.6316 | 0.4085 |
| 146 | voiceclap_attribute:Contentment |
3544 | 0.6851 | 2.2319 | 0.6315 | 0.6411 | 0.6264 | 2.0406 | 0.6953 | 0.6942 |
| 147 | voiceclap_attribute:Disappointment |
3544 | 0.5357 | 1.8439 | 0.6072 | 0.5267 | 0.4829 | 1.6620 | 0.6753 | 0.5702 |
| 148 | voiceclap_attribute:Disgust |
3544 | 0.2720 | 1.0882 | 0.5145 | 0.3850 | 0.2585 | 1.0342 | 0.6078 | 0.4214 |
| 149 | voiceclap_attribute:Distress |
3544 | 0.5497 | 1.5707 | 0.7163 | 0.6190 | 0.4861 | 1.3889 | 0.7655 | 0.6572 |
| 150 | voiceclap_attribute:Doubt |
3544 | 0.5253 | 2.1012 | 0.5091 | 0.5297 | 0.5016 | 2.0064 | 0.5578 | 0.5718 |
| 151 | voiceclap_attribute:Elation |
3544 | 0.4601 | 1.1538 | 0.6368 | 0.5550 | 0.4292 | 1.0765 | 0.6666 | 0.5830 |
| 152 | voiceclap_attribute:Embarrassment |
3544 | 0.1896 | 0.6023 | 0.3048 | 0.2565 | 0.1799 | 0.5715 | 0.2576 | 0.2234 |
| 153 | voiceclap_attribute:Emotional_Numbness |
3544 | 0.2522 | 0.9473 | 0.3121 | 0.2899 | 0.2507 | 0.9419 | 0.3167 | 0.2929 |
| 154 | voiceclap_attribute:Fatigue_Exhaustion |
3544 | 0.5180 | 1.9897 | 0.7102 | 0.6597 | 0.4850 | 1.8632 | 0.7385 | 0.6754 |
| 155 | voiceclap_attribute:Fear |
3544 | 0.2930 | 1.1722 | 0.5463 | 0.4014 | 0.2913 | 1.1654 | 0.5694 | 0.4252 |
| 156 | voiceclap_attribute:Gender |
3368 | 0.2379 | 0.2368 | 0.9505 | 0.9383 | 0.2396 | 0.2385 | 0.9482 | 0.9338 |
| 157 | voiceclap_attribute:Helplessness |
3544 | 0.5094 | 1.8990 | 0.7148 | 0.5998 | 0.4535 | 1.6905 | 0.7714 | 0.6370 |
| 158 | voiceclap_attribute:High-Pitched_vs._Low-Pitched |
3368 | 0.1280 | 0.3427 | 0.8771 | 0.8826 | 0.1272 | 0.3405 | 0.8782 | 0.8872 |
| 159 | voiceclap_attribute:Hope_Enthusiasm_Optimism |
3544 | 0.6346 | 1.5974 | 0.6166 | 0.6083 | 0.5817 | 1.4642 | 0.6760 | 0.6699 |
| 160 | voiceclap_attribute:Impatience_and_Irritability |
3544 | 0.5706 | 1.0528 | 0.5800 | 0.4786 | 0.5425 | 1.0009 | 0.6083 | 0.5011 |
| 161 | voiceclap_attribute:Infatuation |
3544 | 0.1870 | 0.5623 | 0.3724 | 0.2935 | 0.1843 | 0.5543 | 0.4069 | 0.3121 |
| 162 | voiceclap_attribute:Interest |
3544 | 0.5735 | 1.3054 | 0.6784 | 0.6338 | 0.5473 | 1.2458 | 0.7189 | 0.6805 |
| 163 | voiceclap_attribute:Intoxication_Altered_States_of_Consciousness |
3544 | 0.1527 | 0.3292 | 0.2480 | 0.1951 | 0.1551 | 0.3344 | 0.2136 | 0.1762 |
| 164 | voiceclap_attribute:Jealousy_&_Envy |
3544 | 0.1094 | 0.4377 | 0.1123 | 0.1094 | 0.0991 | 0.3965 | 0.1322 | 0.0996 |
| 165 | voiceclap_attribute:Longing |
3544 | 0.4261 | 1.7046 | 0.5481 | 0.5207 | 0.4193 | 1.6770 | 0.5829 | 0.5493 |
| 166 | voiceclap_attribute:Malevolence_Malice |
3544 | 0.2214 | 0.6399 | 0.3811 | 0.2645 | 0.2264 | 0.6543 | 0.4078 | 0.3097 |
| 167 | voiceclap_attribute:Monotone_vs._Expressive |
3368 | 0.2390 | 0.4006 | 0.8655 | 0.8564 | 0.2346 | 0.3934 | 0.8682 | 0.8571 |
| 168 | voiceclap_attribute:Pain |
3544 | 0.3618 | 1.4471 | 0.6585 | 0.5520 | 0.3395 | 1.3579 | 0.6994 | 0.5805 |
| 169 | voiceclap_attribute:Pleasure_Ecstasy |
3544 | 0.3606 | 1.1972 | 0.6002 | 0.4927 | 0.3280 | 1.0889 | 0.6716 | 0.5413 |
| 170 | voiceclap_attribute:Pride |
3544 | 0.5591 | 1.6019 | 0.5473 | 0.5229 | 0.5450 | 1.5617 | 0.5676 | 0.5412 |
| 171 | voiceclap_attribute:Recording_Quality |
3368 | 0.1340 | 0.4083 | 0.9180 | 0.9200 | 0.1342 | 0.4090 | 0.9179 | 0.9209 |
| 172 | voiceclap_attribute:Relief |
3544 | 0.5277 | 1.2686 | 0.5372 | 0.4394 | 0.5009 | 1.2041 | 0.5875 | 0.4669 |
| 173 | voiceclap_attribute:Sadness |
3544 | 0.4434 | 1.4169 | 0.6306 | 0.5121 | 0.3948 | 1.2618 | 0.7219 | 0.5584 |
| 174 | voiceclap_attribute:Serious_vs._Humorous |
3368 | 0.2265 | 0.3971 | 0.8688 | 0.8286 | 0.2183 | 0.3827 | 0.8774 | 0.8265 |
| 175 | voiceclap_attribute:Sexual_Lust |
3544 | 0.1490 | 0.3565 | 0.2604 | 0.2024 | 0.1441 | 0.3448 | 0.2510 | 0.2158 |
| 176 | voiceclap_attribute:Shame |
3544 | 0.1468 | 0.5211 | 0.2995 | 0.1960 | 0.1480 | 0.5254 | 0.2968 | 0.1973 |
| 177 | voiceclap_attribute:Soft_vs._Harsh |
3368 | 0.1702 | 0.4727 | 0.8318 | 0.8489 | 0.1759 | 0.4883 | 0.8189 | 0.8327 |
| 178 | voiceclap_attribute:Sourness |
3544 | 0.4401 | 1.5150 | 0.5642 | 0.4561 | 0.4040 | 1.3909 | 0.6341 | 0.5007 |
| 179 | voiceclap_attribute:Submissive_vs._Dominant |
3368 | 0.1595 | 0.4444 | 0.8189 | 0.8267 | 0.1583 | 0.4408 | 0.8211 | 0.8262 |
| 180 | voiceclap_attribute:Teasing |
3544 | 0.4834 | 1.2742 | 0.6079 | 0.5211 | 0.4720 | 1.2443 | 0.6292 | 0.5347 |
| 181 | voiceclap_attribute:Thankfulness_Gratitude |
3544 | 0.3695 | 0.6980 | 0.5228 | 0.4199 | 0.3478 | 0.6570 | 0.6076 | 0.4657 |
| 182 | voiceclap_attribute:Triumph |
3544 | 0.4139 | 1.1535 | 0.5471 | 0.4353 | 0.4051 | 1.1293 | 0.5561 | 0.4452 |
| 183 | voiceclap_attribute:Valence |
3368 | 0.4702 | 0.7416 | 0.7155 | 0.7111 | 0.4603 | 0.7259 | 0.7213 | 0.7128 |
| 184 | voiceclap_attribute:Vulnerable_vs._Emotionally_Detached |
3368 | 0.2181 | 0.3924 | 0.8508 | 0.8581 | 0.2110 | 0.3796 | 0.8608 | 0.8564 |
| 185 | voiceclap_attribute:Warm_vs._Cold |
3368 | 0.2497 | 0.6254 | 0.7692 | 0.7600 | 0.2392 | 0.5989 | 0.7818 | 0.7665 |
| 186 | voiceclap_attribute:duration |
3368 | 0.9696 | 0.1531 | 0.9725 | 0.9677 | 1.0755 | 0.1698 | 0.9654 | 0.9629 |
| 187 | voiceclap_attribute:score_background_quality |
3368 | 0.1229 | 0.2464 | 0.8736 | 0.8523 | 0.1219 | 0.2446 | 0.8726 | 0.8543 |
| 188 | voiceclap_attribute:score_content_enjoyment |
3368 | 0.0622 | 0.2489 | 0.9033 | 0.9136 | 0.0648 | 0.2593 | 0.8934 | 0.9053 |
| 189 | voiceclap_attribute:score_overall_quality |
3368 | 0.0981 | 0.2378 | 0.9158 | 0.9063 | 0.0979 | 0.2374 | 0.9155 | 0.9048 |
| 190 | voiceclap_attribute:score_speech_quality |
3368 | 0.0272 | 0.1087 | 0.8570 | 0.8576 | 0.0273 | 0.1091 | 0.8536 | 0.8539 |
| 191 | voiceclap_attribute:talking_speed |
3368 | 1.4987 | 0.3994 | 0.9193 | 0.9197 | 1.3978 | 0.3725 | 0.9310 | 0.9277 |
Canonical 53-class vocal-burst vocabulary
The classifier predicts these categories. Fine labels outside this vocabulary use the explicit
aliases/fallbacks in gemini_burst_mapping.json; the fine-to-group23 map is also preserved.
| ID | Canonical vocal-burst class |
|---|---|
| 0 | Other / rare burst |
| 1 | Affirmative Grunt |
| 2 | Ahem |
| 3 | Breath |
| 4 | Breathy Giggle |
| 5 | Cackle |
| 6 | Childlike Giggle |
| 7 | Chuckle |
| 8 | Contented Sigh |
| 9 | Cough |
| 10 | Crying |
| 11 | Displeased Grunt |
| 12 | Effort Grunt |
| 13 | Exasperated Sigh |
| 14 | Exhausted Groan |
| 15 | Fearful Gasp |
| 16 | Frustrated Groan |
| 17 | Growl |
| 18 | Guffaw |
| 19 | Gulp |
| 20 | Hiccup |
| 21 | Hum |
| 22 | Kissing Sound |
| 23 | Laughter |
| 24 | Lip Smack |
| 25 | Low Mumble |
| 26 | Mournful Wail |
| 27 | Nervous Giggle |
| 28 | Nervous Gulp |
| 29 | Purr |
| 30 | Quiet Sob |
| 31 | Relief Sigh |
| 32 | Scream |
| 33 | Sharp Inhale |
| 34 | Sharp Whistle |
| 35 | Shriek |
| 36 | Sigh |
| 37 | Sneeze |
| 38 | Snicker |
| 39 | Sniff |
| 40 | Snort |
| 41 | Snorting Giggle |
| 42 | Soft Whistle |
| 43 | Spitting |
| 44 | Surprised Gasp |
| 45 | Swallows |
| 46 | Throat Clearing |
| 47 | Tongue Click |
| 48 | Trembling Whimper |
| 49 | Tsk |
| 50 | Whispered Mumble |
| 51 | Wistful Sigh |
| 52 | Yawn |
Model tree for laion/humaneness-ears-base-medium
Base model
openai/whisper-base