Humaneness Ears: Base and Medium audio understanding

By Christoph Schuhmann · LAION · 7 October 2026 · CC BY 4.0

This repository releases the best completed two-epoch fine-tunes of LAION's encoder-only Whisper Base and Whisper Small multitask models. The release names are Humaneness Ears Base (Whisper Base backbone) and Humaneness Ears Medium (Whisper Small backbone). Given a recording, one forward pass predicts emotions, speaking style, sound quality, speaker-characteristic vectors and the location and type of non-speech vocal events such as laughter or crying. Both models are included: base/model.safetensors and medium/model.safetensors. The older training logs, evaluation JSON and the machine-readable listening manifest retain the small key as the historical Whisper backbone label; it refers to Humaneness Ears Medium.

These are audio encoders with task heads. There is no Whisper text decoder, ASR transcript, generated caption or language model. Gemini transcripts and captions helped prepare annotations; captions were not model inputs or an autoregressive training objective. CPS is a scalar prediction rather than a transcript. The requested separate MOSS caption LoRAs are outside this release and were not trained as part of these two runs.

Explore 1,000 real-audio examples in the Humaneness Ears atlas: 50 timed vocal bursts, 20-query Orange timbre and identity neighbor pages, 40 EmoNet and 57 VoiceNet top-ten rankings, and five-tier quality pages. Download the Gemini annotation dataset.

Open the full benchmark: Humaneness Ears Base/Medium versus CLAPv2 XS/M — human-label results, all 192 targets, Orange speaker-vector cosine and vocal-burst timing. Jump to the CLAP comparison.

Results and related releases

Medium has better Orange speaker-vector cosine and lower AudioBox/DNSMOS target error than Base, and a higher EmoNet intensity correlation. Base has slightly better CPS error, Flash frame F1 and VoiceNet-Emo mean rank correlation. A frozen probe can outperform both on other tasks: see the comprehensive report for all winners. This release does not claim one universal best model or statistical significance from small numerical differences.

What the models output

Output Shape / scale Meaning
EmoNet emotions 40 scalars, reference 0–4 Fine-grained emotion intensity; simultaneous regression, not a mutually exclusive class.
VoiceNet dimensions 57 scalars, mostly 0–6 Delivery, register, timbre and style; Flash BKGN is 0–4 and EXPL 0–2.
Genuineness 1 scalar, 0–6 The annotation rubric's perceived genuineness.
Vocal-burst blend 1 scalar, 0–10 How a vocal event blends with surrounding speech; meaningful only when a burst is present.
R_quality 1 legacy scalar Its distinct original teacher target and units; not overwritten with another quality construct.
Empathic Insight Plus extras 19 scalars Voice traits, arousal/valence, recording/background quality and content enjoyment.
AudioBox Aesthetics 4 scalars CE: content enjoyment; CU: content usefulness; PC: production complexity; PQ: production quality.
DNSMOS 7 scalars SIG, BAK, OVRL, their raw versions and P808 MOS.
Burst count 1 scalar trained in log1p space Number of annotated events; inference additionally reports expm1 of the nonnegative prediction.
VoiceCLAP attributes 61 scalars Additional emotion, voice and quality targets; some emotion targets overlap the first 40.
Characters per second (CPS) 1 separate scalar Unicode transcript characters, including spaces/punctuation, divided by total recording duration; bracketed nonlexical placeholders removed.
Orange timbre 128D unit vector Regression to the Orange timbre teacher embedding.
Orange identity 250D unit vector Regression to the Orange identity teacher embedding; not a person's name or an identification guarantee.
Burst occupancy 20 ms frame probabilities Whether each encoder frame belongs to at least one vocal burst.
Independent burst events Start and end seconds Onset plus duration proposals can overlap; maximum 32 at the default evaluation setting.
Burst type 53-way event-local class probabilities The top three labels and uncalibrated softmax probabilities for each predicted interval.

There are 192 jointly predicted scalar outputs plus the separate CPS head. A complete, ordered list of exact output keys and all their held-out results appears at the end of this card. The models return raw units, training-standardized values and display z-scores. Regression outputs are continuous and are not automatically clipped to the annotation's ordinal range.

Quick start: local CPU or GPU inference

Use Python 3.10 or newer. The measured production environment used Python 3.13.5, PyTorch 2.9.1, Transformers 5.14.1, SoundFile 0.14 and Safetensors. Install the appropriate PyTorch CPU/CUDA wheel for your machine; see requirements.txt for the other dependencies. Inference uses a custom model.py, not pipeline('automatic-speech-recognition') or WhisperForConditionalGeneration.from_pretrained.

Download inference assets without the large optimizer checkpoints:

from huggingface_hub import snapshot_download

snapshot_download(
    repo_id="laion/humaneness-ears-base-medium",
    local_dir="humaneness-ears",
    allow_patterns=["*.py", "requirements.txt", "*.json", "base/*.json",
                    "medium/*.json", "base/model.safetensors", "medium/model.safetensors"],
)
cd humaneness-ears
pip install -r requirements.txt
python inference.py clip.wav --model medium --device cpu --threads 4 > medium_prediction.json
python inference.py clip.wav --model base --device cpu --threads 1 > base_prediction.json
python inference.py clip.wav --model medium --device cuda:0 --threads 4 \
  --include-frame-probabilities > gpu_prediction.json

Input recordings must be 0.1–30 seconds. SoundFile decodes the supplied audio, stereo is averaged to mono, and SciPy resamples to 16 kHz. The bundled Whisper frontend creates 80-bin log-mel features. Longer recordings need an explicit external chunking policy. No original OpenAI weights, teacher checkpoints or Gemini API access are required at inference.

Reuse the model when processing multiple files:

from inference import Predictor

predictor = Predictor(".", model_size="medium", device="cpu", threads=4)
result = predictor.predict("clip.wav", include_frames=True)
print(result["scores_raw"]["genuineness_0_6"])
print(result["scores_z"]["emo_Amusement"])
print(result["burst_event_proposals"])  # start_s, end_s, top3_classes

--frame-threshold and --event-threshold default to 0.5; --max-events defaults to 32. This matches the new report's proposal operating point. Lowering the event threshold can increase recall and false positives. Frame regions merge overlapping events, whereas the independent proposal head can represent them separately. Neither class softmax values nor sigmoid values have been calibrated as real-world confidence estimates.

Architecture

Base uses six encoder layers, hidden width 512, and 26,441,735 total parameters. Medium uses 12 encoder layers, hidden width 768, and 97,102,031 total parameters. These counts include our heads and exclude the absent Whisper decoder. The encoder alone contains 20,590,592 / 88,154,112 parameters respectively.

From the final encoder sequence we concatenate masked mean, minimum, maximum and standard deviation. Separately, each transformer layer contributes its masked mean vector. Shared per-family projection MLPs reduce these means to 32 or 64 dimensions, and a sample-specific softmax gate combines layer information. These features add residuals to the 192-score head. Speaker heads use 128D layer-mixture routes; CPS uses 32D. The frame head predicts occupancy; the proposal head predicts independent onset logits and log durations. A 64D layer-mixture route for event classification averages only frames inside that event from every layer, then combines them with the final-layer local mean and a 53-way classifier.

model.py is the exact architecture used in training. Exported weights preserve every saved tensor without quantization or conversion. The inherited Whisper configuration still contains some unused decoder fields; architectures and is_encoder_decoder are set for this encoder-only release. No decoder tensors are present. The custom inference code must be used.

Training data, annotation priority and splits

The selected pool contains 66,199 exact-audio unique clips / 197.36 hours after deduplication across the four selections. Valid Flash annotations exist for 65,163 clips / 194.17 hours. The 1,036 invalid or provider-blocked responses are excluded from this tuning stage. They were not converted into zero-emotion or no-burst targets.

Split actually used Clips Purpose
Training 58,380 Gradient updates for both full fine-tunes.
Validation 3,239 Full validation after each epoch; selects best checkpoint.
Test 3,544 Matched before/after evaluation; not used for gradient updates or epoch selection.

Sources are the balanced S1 100-hour, 400-bucket selection from TTS Scaling Ladder DE/EN, LAION's Got Talent raw, balanced audio snippets 40×3k and voice annotation POC. Final selections preserve all source/task memberships while annotating each exact compressed-audio SHA256 once. Source membership counts are nonexclusive. Deterministic 90/5/5 hashing groups explicit speaker IDs, otherwise known synthetic audio families, otherwise exact audio hashes; observed counts differ from exact percentages because groups are kept intact.

training/provenance/valid_split_summary.json records valid counts, hours, language strings, memberships and the hash of the exact prepared target snapshot. The older TRAINING_READY.json also contains split counts including invalid responses; those are not the training counts above. Source revisions and original pool details are in training/provenance/selected_pool.json. The data, Gemini responses and teacher TARs are not redistributed by this model repository.

The recorded annotation model identifier is gemini-3.8-flash. Schema-valid Flash outputs take priority for 40 emotions, 57 VoiceNet dimensions, genuineness, blend, event counts, event start/end times and transcript-derived CPS. Corresponding VoiceCLAP emotion attributes also take the Flash value. Distinct constructs such as DNSMOS BAK or the original authenticity axis keep their own teachers. A null blend on a no-burst clip is masked, not replaced by an old blend score. Missing values remain masked. Speech-only objectives use domain masks; Orange speaker losses require valid known single-speaker clips. Teacher vectors are preserved even when their training loss is masked.

The prompt permits 78 vocal fine labels. Explicit aliases map compatible labels into the existing 53-class checkpoint vocabulary; unrepresented labels map to Other / rare burst. See gemini_burst_mapping.json for every mapping. Original fine labels and event captions remain in the annotation sidecars. End times may be clamped to the decoded duration within the accepted 0.3-second annotation tolerance, with original timing and the clamp flag retained in the data.

Exact fine-tuning procedure

Starting models were the lowest-validation-loss Base S4 and Whisper Small S3 snapshots from the original S1→S10 layered curriculum campaign. Each snapshot contains only the progress up to its own stage, rather than the final S10 weights: Base's original update was 266,238; Medium's was 237,434. The original campaign ran one epoch per nested stage with 10% P3 synthetic mixture exposures per stage and a continuous cosine schedule. Its full code and recorded run configs are included under training/ and training/original_curriculum/.

Gemini tuning strictly restored all encoder and head weights, started a fresh AdamW optimizer and cosine schedule, and unfroze the entire encoder, including its positional embedding table. This is full fine-tuning, not LoRA or a frozen-encoder probe.

Setting Base Medium
Epochs 2 2
Actual optimizer updates 914 914
Best checkpoint Epoch 2 Epoch 2
Peak encoder learning rate 1e-5 5e-6
Peak head learning rate 1e-4 1e-4
Warmup 5% (46 updates) 5% (46 updates)
Cosine minimum / peak 10% 10%
Weight decay 0.01 0.01
Global effective batch 128 128
Microbatch / GPU 32 16
Accumulation 1 2
Data workers / GPU 4 4
Hardware 1 node, 4 GH200 GPUs 1 node, 4 GH200 GPUs
Precision BF16 autocast; FP32 saved weights BF16 autocast; FP32 saved weights
Gradient clipping Global norm 1.0 Global norm 1.0
Seed 20261005 20261005
Selected class-weighted validation loss 1.7683769 1.6989245
Training Slurm job 2199189 2199190
Recorded allocation elapsed time 10m17s 10m31s

There is no P3 replay in this extra two-epoch stage. A DDP sampler may pad a few examples to distribute batches equally; validation is evaluated completely on rank zero without duplicating its samples. The full config, class counts and epoch loss logs are under training/.

Losses

The implementation is training/train_layered_curriculum.py::loss_terms, reused unchanged by training/gemini_finetune/train_whisper.py:

  • Masked Huber regression, delta 1, on standardized scalar targets; per-family weights are the exact GROUPS constants in training/multitask_p3_smoke.py.
  • CPS: standardized Huber delta 1, weight 0.2.
  • Speaker vectors: 0.2 × (cosine distance + 0.1 × Huber delta 0.1), only valid single speakers.
  • Frames: weighted BCE (positive weight 5), weight 0.5, plus soft Dice weight 0.25.
  • Onsets: weighted BCE (positive weight 100), weight 0.2.
  • Event log duration: Huber delta 0.5, weight 0.1, evaluated at labeled starts.
  • Event class: categorical cross entropy, weight 0.2, label smoothing 0.02; training-only class weights sqrt(median event count / class count), clipped to 0.25–4.

Every epoch saves latest and, when improved, best states. Published base/checkpoint.pt and medium/checkpoint.pt are the exact best files: model, AdamW, scheduler, four-rank RNG, completed epoch, update, normalization, taxonomy and original config. They support epoch-boundary continuation with the same DDP world size. Both runs already finished epoch 2; extending training requires an explicit new phase/schedule, not blindly resuming the completed two-epoch plan. Inference uses only Safetensors and never loads the .pt training pickle.

Normalization

raw_score = standardized_network_output × training_std + training_mean
display_z = (raw_score − original_training_teacher_median) / original_training_teacher_std
CPS_raw = CPS_network_output × CPS_training_std + CPS_training_mean

training_normalization.json is the original frozen checkpoint normalization, preserved through Gemini tuning. It is required for correct raw units. display_normalization.json is the train-only median/std reference from the original release, retained for compatible visualization and rankings. gemini_train_statistics.json separately stores train-only Flash mean/std/counts; its inherited “provisional” description refers to the incremental parser, but the file here is the completed final export. It was not substituted into network decoding. Display z-scores are relative context, not class probabilities or confidence intervals.

How the data were mounted and how to reproduce

Training ran on JUPITER Booster. /e/home, /e/project1, /e/scratch and /e/software were existing shared HPC filesystems visible to every rank; no Docker bind mount or network download occurred inside the GPU training. Source audio and same-audio teacher outputs stayed in WebDataset TARs. TarReader caches a bounded set of open archives and reads named members; it does not extract millions of individual files. Teachers are cached annotations, not called inside the gradient loop.

Original absolute path Contents / use
/e/home/jusers/schuhmann1/jupiter/whisper_score_regression Production training and inference source.
/e/scratch/reformo/schuhmann1_moss/whisper_score_regression/hf_layered_release/model Original Base S4 / Whisper Small S3 states, frontend, normalization and classes.
/e/scratch/reformo/schuhmann1_moss/whisper_score_regression/gemini_multisource_100h_20261004 Combined exact-audio selection and Gemini response sidecars.
/e/scratch/reformo/schuhmann1_moss/whisper_score_regression/gemini_finetune_preparation_20261005/prepared targets.jsonl, final training gate, normalization and burst mapping.
Same preparation root, teacher_sidecars/ SHA-bound TAR companions with Orange vectors and all additional teacher outputs.
Same preparation root, training/whisper_{base,small} Exact best/latest states and per-epoch metrics.
/e/scratch/reformo/schuhmann1_moss/code/fastgen.sh Python/software paths for the original offline HPC environment.

targets.jsonl contains audio_tar / audio_member and embedding tar / member references. On another machine, obtain the same complete prepared data, copy TARs and rewrite the absolute TAR prefixes to your local mounts while preserving member names and audio hashes. Provide checkpoint_normalization.json, gemini_burst_mapping.json and TRAINING_READY.json beside it. The repository includes the parser, schema prompt and teacher backfill code; recreating its data requires the upstream datasets, teacher weights and your own annotation service access. There is no credential embedded here. Source availability and terms still apply.

To reproduce the two-epoch experiment, download the original Base S4 or Whisper Small S3 .pt initializer from the previous repository, not this release's already-tuned state. Then:

pip install -r requirements-training.txt
python training/make_finetune_config.py --model base \
  --data-root /your/prepared --initial-checkpoint /your/original/base/S4/checkpoint.pt \
  --output-dir /your/new-base-run --config /your/base-run.json
python -m torch.distributed.run --standalone --nproc_per_node=4 \
  training/gemini_finetune/train_whisper.py --config /your/base-run.json

For the historical Whisper Small training run, use --model small and the original small/S3/checkpoint.pt. The config helper changes paths while preserving recorded hyperparameters. training/gemini_finetune/train_whisper.sbatch is the original submitted Slurm launcher, with original mount paths for reference. Evaluator scripts likewise preserve original study paths; adapt study_paths.py and benchmark data paths if reproducing the full comparison. Safe-weight export tools and verification receipts are included.

Held-out Flash target agreement

Both models are tested on the same 3,544 valid Flash test clips. These targets are machine annotations, not a fresh human panel. Speaker cosine uses 1,718 valid single-speaker clips; blend uses 2,391 clips with non-null blend. Family MAE is the unweighted mean over valid axes.

Flash test target / metric Base Medium
Burst frame F1 ↑ 0.6929 0.6892
Burst event F1, IoU ≥0.5 ↑ 0.1945 0.1949
Burst start MAE on IoU ≥0.1 matches, seconds ↓ 0.1201 0.1150
Burst end MAE on IoU ≥0.1 matches, seconds ↓ 0.1973 0.2006
53-class accuracy with reference spans ↑ 0.4280 0.4483
53-class accuracy on matched predicted spans ↑ 0.4189 0.4370
Orange timbre cosine ↑ (1,718 valid single-speaker clips) 0.9038 0.9223
Orange identity cosine ↑ (1,718 valid single-speaker clips) 0.8541 0.8741
CPS raw MAE, characters/second ↓ 0.9012 0.9309
Genuineness raw MAE on 0–6 scale ↓ (3,544 clips) 0.4774 0.4704
Genuineness normalized MAE ↓ 0.2971 0.2928
Blend raw MAE on 0–10 scale ↓ (2,391 valid clips) 0.8773 0.8417
Blend normalized MAE ↓ 0.3228 0.3097
40 emotions, mean normalized MAE ↓ 0.5873 0.5605
57 VoiceNet dimensions, mean normalized MAE ↓ 0.4967 0.4813
AudioBox four axes, mean normalized MAE ↓ 0.1877 0.1739
DNSMOS seven outputs, mean normalized MAE ↓ 0.2047 0.1975

Normalized MAE is measured in frozen training standard deviations. Burst interval metrics use Hungarian matching of independent onset/duration proposals, not binary contiguous-frame regions. Boundary errors and predicted-span class accuracy are conditional on matches at IoU ≥0.1. Reference-span class accuracy gives the classifier the true interval; it is not end-to-end localization accuracy. Maximum proposals is 32, with a fixed 0.5 onset threshold. Event precision is low despite high recall; the report includes all IoU cutoffs and precision/recall.

The validation tables in evaluation/{base,small}/validation_metrics.json (historical backbone keys) use the existing 3,239 checkpoint-selection clips, not an untouched test. Supplementary evaluator loss uses unit class weights, while the training selection loss above used the configured class weights. Different loss values under those two policies are expected.

Public human-label benchmarks

Human-label benchmark / metric Base Medium
EmoNet-Voice intensity, Pearson r ↑ 0.4340 0.4512
EmoNet-Voice intensity, Spearman ρ ↑ 0.4577 0.4679
VoiceNet-Emo, mean prompt Spearman ρ ↑ 0.4288 0.4282
VoiceNet-Ext / emolia-dim, mean prompt Spearman ρ ↑ 0.1566 0.1502
CREMA-D matched actor-CV score-MLP accuracy ↑ 0.6310 0.6423
RAVDESS matched actor-CV score-MLP accuracy ↑ 0.6597 0.6639

EmoNet uses 12,000 mapped clips and fixed 0–4→0–10 endpoint scaling; two unsupported labels from the published 12,600-clip benchmark are omitted. VoiceNet-Emo uses 7,986 current-repository questions with at least two raters; VoiceNet-Ext uses 13,917. EXT and DIM refer to the same benchmark. Current snapshot coverage differs from paper coverage. The report includes unflagged, untruncated and alternate agreement cuts and explains optimistic oracle threshold numbers.

CREMA-D and RAVDESS figures are supervised matched nested actor-cross-validation score adapters, not zero-shot paper numbers. Each adapter sees the model's 192 raw predictions, a train-fold StandardScaler and a 64-unit GELU MLP: 12,742 / 12,872 parameters respectively. Five outer actor folds and three inner folds choose learning rate 0.001/0.003 and 20/50 epochs. Outer test actors never fit that fold's scaler or adapter. These benchmark-only adapters are not the model's burst classifier and are not automatically applied by inference.py. Full fold records, metrics and confusions are included; the report compares the same adapter budget across all backbones. Upstream training/source overlap with the public benchmarks was not exhaustively audited.

P3 retention and limitations

On the identical fixed 2,000-clip P3 test audit, frame F1 changes from 0.9622 to 0.8051 for Base and 0.9693 to 0.8392 for Medium after Gemini tuning. Flash test frame F1 rises from 0.3072/0.3171 to 0.6929/0.6892 respectively. This is a target/domain tradeoff: there was no P3 replay in this phase. All matched pre/post P3 validation/test JSONs are included. It is not valid to claim universal burst improvement from Flash agreement alone.

The selected pool is not a representative random sample of all natural audio. Quality scores imitate AudioBox/DNSMOS/Empathic teachers rather than measuring new human MOS. Speaker cosine measures regression agreement with Orange, not verification accuracy or human identity. The burst classifier groups unsupported fine labels, event proposals have many false positives, and emotional or expressive clips can differ from the synthetic P3 construction labels. Recordings longer than 30 seconds and multi-speaker embeddings require additional policies.

Annotation models, taxonomies and packages

Supervision / package Primary reference
Original Whisper encoders OpenAI Whisper Base, OpenAI Whisper Small, Whisper code.
Gemini annotation schema The exact 108 KB master prompt is training/gemini_s10_smoke/gemini_voice_annotation_master_prompt_detailed.txt; parser and validation code are included.
40 emotions EmoNet taxonomy, emotion annotations toolkit.
57 voice dimensions VoiceNet taxonomy, commercial dimension predictors.
Original genuineness teacher VoiceCLAP commercial genuineness; matching targets superseded by valid Flash labels.
Original blend teacher VoiceCLAP commercial vocalburst blend; matching targets superseded by valid Flash labels.
VoiceCLAP attributes / intermediate embedding Commercial VoiceCLAP, attribute heads; 768D intermediate embeddings are cached in data, not an output head.
Empathic extras / original emotion teacher Empathic Insight Voice Plus, BUD-E Whisper; 3072D teacher features are cached, not an output head.
AudioBox Meta AudioBox Aesthetics.
DNSMOS Microsoft DNS Challenge / DNSMOS.
Orange timbre, 128D Orange Speaker-wavLM-tbr, pinned annotation revision b8d2608d56f18e2b6e27ba566cf69132e2c2ad6c.
Orange identity, 250D Orange Speaker-wavLM-id, pinned revision abbb3c7b8d220ceebc33d9bd3bb5aedb580342c7.
Timbre generation implementation LAION generation script.
Vocal event taxonomy / locator LAION voice taxonomies, vocalburst locator, canonical classes and Flash mapping bundled here.
Human emotion/style benchmark LAION emolia-bench, EmoNet-Voice Bench.

Teacher inference runs on the exact selected waveform. New scores and float16 target vectors are persisted in SHA-bound companion TARs; the gradient loop reuses those sidecars. Their model revisions and target-level provenance remain in the prepared annotations.

Repository contents and export verification

Path Contents
base/model.safetensors, medium/model.safetensors Complete best Gemini-tuned encoder and head weights; no decoder.
base/checkpoint.pt, medium/checkpoint.pt Exact best resumable training states including optimizer, scheduler and per-rank RNG.
{base,medium}/config.json, preprocessor_config.json Corresponding encoder architecture and Whisper log-mel frontend.
model.py, inference.py, requirements.txt Complete standalone inference, cached Predictor API and CLI.
training/, vocal_burst_pool/p3/ Actual training source and local dependency closure, target parser, prompt, backfill and evaluation code, configs, logs and reproduction helper.
training_normalization.json, display_normalization.json, gemini_train_statistics.json Frozen decoding, contextual display and separate final Flash train-only statistics.
classes.json, gemini_burst_mapping.json 53 canonical class names and all Flash aliases/fallbacks.
evaluation/{base,small}/ (historical keys) Flash validation/test, pre/post P3 retention, public benchmarks and actor-adapter folds.
verification/{base,small}.json (historical keys) Bitwise safe-weight and model-output parity; real held-out audio API checks, dimensions and timestamp validity.
release_manifest.json File sizes and SHA256 hashes, original checkpoint hashes and provenance.
release_tools/ Export, verification and publication source.
LICENSE, NOTICE.md CC BY 4.0 and upstream attribution.

No audio, source transcripts, raw Gemini responses or service credentials are published by this model repository. Model verification uses a real held-out waveform internally and saves only its hash and output-contract evidence. Evaluation JSONs contain aggregate benchmark results, prompts/classes and fold protocols; no listening audio is copied into this release.

License and attribution

The new fine-tuned weights, new code and documentation are released by LAION under Creative Commons Attribution 4.0 International. Credit Christoph Schuhmann and LAION, link this repository and indicate changes. This license does not relicense upstream source audio, teacher models or external packages. The original OpenAI Whisper components remain MIT; Orange's identity teacher retains its upstream CC BY-SA 3.0 terms. See NOTICE.md and upstream model/data cards for their notices.

@misc{schuhmann_humaneness_ears_2026,
  title = {Humaneness Ears: Base and Medium Encoder-Only Audio Understanding},
  author = {Schuhmann, Christoph and {LAION}},
  year = {2026},
  url = {https://huggingface.co/laion/humaneness-ears-base-medium}
}

All 192 scalar targets: exact held-out results

These are Flash test targets, with the remaining AudioBox/DNSMOS/Orange/Empathic families provided by their own teachers. Raw MAE uses each target's units. Norm. MAE uses its frozen training standard deviation. r is Pearson and ρ is Spearman. Missing metrics remain unavailable. The source JSON has separate counts for every model/target.

Index Exact target key Valid N Base raw MAE Base norm. MAE Base r Base ρ Medium raw MAE Medium norm. MAE Medium r Medium ρ
0 emo_Affection 3544 0.4034 0.7418 0.7163 0.5685 0.3911 0.7192 0.7345 0.5839
1 emo_Amusement 3544 0.5187 0.5978 0.7413 0.6742 0.4879 0.5623 0.7606 0.6992
2 emo_Anger 3544 0.3462 0.5968 0.6770 0.4799 0.3376 0.5819 0.6729 0.5020
3 emo_Astonishment_Surprise 3544 0.4007 0.7257 0.6906 0.5951 0.3838 0.6950 0.7075 0.6081
4 emo_Awe 3544 0.3259 0.7764 0.7652 0.5528 0.3002 0.7152 0.7946 0.5699
5 emo_Bitterness 3544 0.3718 0.9136 0.6692 0.5244 0.3488 0.8571 0.6932 0.5548
6 emo_Concentration 3544 0.4721 0.6463 0.6937 0.6930 0.4634 0.6343 0.7024 0.7005
7 emo_Confusion 3544 0.3659 0.6860 0.6370 0.5124 0.3462 0.6490 0.6617 0.5392
8 emo_Contemplation 3544 0.5538 0.9403 0.7380 0.7293 0.5367 0.9113 0.7552 0.7449
9 emo_Contempt 3544 0.3103 0.6878 0.6823 0.4370 0.2999 0.6648 0.6994 0.4587
10 emo_Contentment 3544 0.5967 1.1843 0.6973 0.6890 0.5675 1.1264 0.7270 0.7214
11 emo_Disappointment 3544 0.4548 0.6485 0.7058 0.5985 0.4328 0.6171 0.7290 0.6164
12 emo_Disgust 3544 0.2244 0.4576 0.6700 0.4164 0.2140 0.4365 0.7002 0.4420
13 emo_Distress 3544 0.4677 0.5722 0.7701 0.6569 0.4445 0.5438 0.7917 0.6834
14 emo_Doubt 3544 0.4587 0.9491 0.6336 0.6068 0.4360 0.9020 0.6652 0.6390
15 emo_Elation 3544 0.3643 0.5071 0.7254 0.5958 0.3513 0.4890 0.7449 0.6097
16 emo_Embarrassment 3544 0.1472 0.3367 0.5970 0.2711 0.1434 0.3282 0.6350 0.2784
17 emo_Emotional_Numbness 3544 0.2516 0.2617 0.5849 0.4138 0.2434 0.2532 0.6081 0.4254
18 emo_Fatigue_Exhaustion 3544 0.4472 0.8269 0.7700 0.6846 0.4385 0.8108 0.7746 0.6890
19 emo_Fear 3544 0.2303 0.3764 0.7266 0.4491 0.2201 0.3597 0.7331 0.4586
20 emo_Helplessness 3544 0.4093 0.5781 0.7839 0.6445 0.3889 0.5492 0.8065 0.6608
21 emo_Hope_Enthusiasm_Optimism 3544 0.5715 0.7535 0.6933 0.6714 0.5412 0.7135 0.7238 0.7046
22 emo_Impatience_and_Irritability 3544 0.5024 0.7343 0.6582 0.5313 0.4835 0.7068 0.6653 0.5517
23 emo_Infatuation 3544 0.1289 0.2382 0.7034 0.3314 0.1205 0.2226 0.7270 0.3399
24 emo_Interest 3544 0.5271 0.5766 0.7397 0.7085 0.5056 0.5531 0.7626 0.7317
25 emo_Intoxication_Altered_States_of_Consciousness 3544 0.1044 0.1877 0.4826 0.2741 0.1046 0.1881 0.4474 0.2786
26 emo_Jealousy_and_Envy 3544 0.0624 0.1400 0.6992 0.1091 0.0559 0.1254 0.7301 0.1062
27 emo_Longing 3544 0.3539 0.5566 0.7124 0.5817 0.3355 0.5277 0.7392 0.5894
28 emo_Malevolence_Malice 3544 0.1676 0.3200 0.6364 0.3145 0.1636 0.3124 0.6531 0.3299
29 emo_Pain 3544 0.2795 0.4859 0.7490 0.5796 0.2580 0.4485 0.7785 0.5889
30 emo_Pleasure_Ecstasy 3544 0.2679 0.4388 0.7381 0.5662 0.2471 0.4048 0.7743 0.5881
31 emo_Pride 3544 0.5015 0.9645 0.6547 0.5704 0.4816 0.9263 0.6812 0.5928
32 emo_Relief 3544 0.4468 0.6403 0.6572 0.5138 0.4189 0.6002 0.6932 0.5388
33 emo_Sadness 3544 0.3507 0.4589 0.7617 0.5775 0.3287 0.4301 0.7960 0.5914
34 emo_Sexual_Lust 3544 0.0958 0.1924 0.6895 0.2417 0.0930 0.1868 0.6770 0.2413
35 emo_Shame 3544 0.0955 0.1822 0.6960 0.2302 0.0896 0.1710 0.7326 0.2385
36 emo_Sourness 3544 0.3766 0.8710 0.6533 0.5006 0.3576 0.8271 0.6761 0.5236
37 emo_Teasing 3544 0.4089 0.7599 0.6783 0.5625 0.4021 0.7473 0.6848 0.5687
38 emo_Thankfulness_Gratitude 3544 0.2490 0.3602 0.7406 0.4788 0.2260 0.3269 0.7811 0.5085
39 emo_Triumph 3544 0.3263 0.6180 0.6941 0.5031 0.3151 0.5967 0.7145 0.5123
40 vn_AGEV_reg 3544 0.3185 0.2527 0.4460 0.4312 0.3093 0.2455 0.4414 0.4270
41 vn_AROU_reg 3544 0.5437 0.4826 0.7243 0.6929 0.5284 0.4690 0.7376 0.7127
42 vn_ARSH_reg 3544 0.4911 0.5604 0.3468 0.3323 0.4776 0.5450 0.3800 0.3731
43 vn_ATCK_reg 3544 0.4661 0.3540 0.6870 0.6544 0.4635 0.3520 0.6932 0.6690
44 vn_BKGN_reg 3544 0.3384 0.4472 0.6167 0.6001 0.3359 0.4439 0.6118 0.5989
45 vn_BRGT_reg 3544 0.3844 0.4743 0.6861 0.6711 0.3833 0.4729 0.6862 0.6676
46 vn_CHNK_reg 3544 0.4343 0.3484 0.7160 0.7076 0.4312 0.3459 0.7241 0.7166
47 vn_CLRT_reg 3544 0.4416 0.4595 0.8017 0.6867 0.4334 0.4510 0.8035 0.6948
48 vn_COGL_reg 3544 0.5413 0.5135 0.7345 0.7335 0.5232 0.4963 0.7542 0.7556
49 vn_DARC_reg 3544 0.4769 0.3501 0.4287 0.3938 0.4663 0.3423 0.4556 0.4118
50 vn_DFLU_reg 3544 0.5694 0.4820 0.8008 0.7930 0.5619 0.4757 0.8075 0.8025
51 vn_EMPH_reg 3544 0.4545 0.2893 0.6617 0.6401 0.4437 0.2824 0.6713 0.6575
52 vn_ESTH_reg 3544 0.4608 0.5917 0.6939 0.6557 0.4423 0.5680 0.7172 0.6910
53 vn_EXPL_reg 3544 0.1151 0.3516 0.2977 0.1599 0.1088 0.3323 0.2825 0.1475
54 vn_FOCS_reg 3544 0.4337 0.3907 0.6086 0.4914 0.4098 0.3692 0.6395 0.5212
55 vn_FULL_reg 3544 0.3593 0.4066 0.5248 0.5048 0.3466 0.3922 0.5457 0.5310
56 vn_GEND_reg 3544 0.5459 0.3428 0.8868 0.8254 0.5735 0.3601 0.8836 0.8256
57 vn_HARM_reg 3544 0.4697 0.5674 0.6904 0.6492 0.4512 0.5450 0.7242 0.6856
58 vn_METL_reg 3544 0.4922 0.5140 0.6026 0.5499 0.4832 0.5046 0.6160 0.5692
59 vn_RANG_reg 3544 0.5298 0.5943 0.7133 0.7109 0.5125 0.5749 0.7316 0.7340
60 vn_RCQL_reg 3544 0.4049 0.4498 0.6185 0.6144 0.4084 0.4536 0.6119 0.6139
61 vn_REGS_reg 3544 0.4636 0.3690 0.8743 0.8353 0.4724 0.3760 0.8706 0.8324
62 vn_RESP_reg 3544 0.4666 0.4362 0.8088 0.7550 0.4444 0.4155 0.8261 0.7708
63 vn_ROUG_reg 3544 0.5316 0.6029 0.7408 0.7184 0.5216 0.5916 0.7509 0.7339
64 vn_R_CHST_reg 3544 0.4628 0.5450 0.7755 0.7775 0.4618 0.5439 0.7789 0.7825
65 vn_R_HEAD_reg 3544 0.5097 0.6189 0.7728 0.7756 0.5088 0.6178 0.7719 0.7709
66 vn_R_MASK_reg 3544 0.4401 0.4301 0.7060 0.6774 0.4369 0.4269 0.7118 0.6882
67 vn_R_MIXD_reg 3544 0.4149 0.4957 0.6177 0.5642 0.4046 0.4835 0.6333 0.5872
68 vn_R_NASL_reg 3544 0.4205 0.3982 0.4078 0.3681 0.4063 0.3847 0.4375 0.4148
69 vn_R_ORAL_reg 3544 0.3597 0.4834 0.6342 0.6177 0.3516 0.4725 0.6439 0.6331
70 vn_R_THRT_reg 3544 0.4400 0.5464 0.6952 0.6673 0.4269 0.5301 0.7085 0.6869
71 vn_SMTH_reg 3544 0.4650 0.4184 0.7178 0.7130 0.4497 0.4046 0.7341 0.7314
72 vn_STNC_reg 3544 0.6337 0.5565 0.6729 0.6393 0.6053 0.5316 0.7046 0.6692
73 vn_STRU_reg 3544 0.4699 0.4766 0.7898 0.7386 0.4659 0.4724 0.7923 0.7465
74 vn_S_ASMR_reg 3544 0.5717 0.4147 0.7268 0.6562 0.5479 0.3974 0.7467 0.6773
75 vn_S_AUTH_reg 3544 0.7103 0.6837 0.7108 0.6953 0.6791 0.6537 0.7404 0.7287
76 vn_S_CART_reg 3544 0.6246 0.4562 0.6659 0.5904 0.6244 0.4560 0.6549 0.5974
77 vn_S_CASU_reg 3544 0.7998 0.6269 0.6985 0.7130 0.7619 0.5971 0.7284 0.7413
78 vn_S_CONV_reg 3544 0.6987 0.5353 0.7534 0.7393 0.6797 0.5207 0.7696 0.7600
79 vn_S_DRAM_reg 3544 0.6893 0.6050 0.7871 0.7833 0.6613 0.5804 0.8062 0.8037
80 vn_S_FORM_reg 3544 0.6598 0.6543 0.7273 0.7316 0.6426 0.6373 0.7447 0.7507
81 vn_S_MONO_reg 3544 0.8582 0.6745 0.6484 0.6408 0.8287 0.6514 0.6732 0.6646
82 vn_S_NARR_reg 3544 0.7113 0.4956 0.7613 0.7441 0.6963 0.4851 0.7702 0.7539
83 vn_S_NEWS_reg 3544 0.4990 0.9056 0.6904 0.6742 0.4881 0.8858 0.7103 0.6925
84 vn_S_PLAY_reg 3544 0.8496 0.6535 0.6578 0.6327 0.7954 0.6119 0.7052 0.6929
85 vn_S_RANT_reg 3544 0.6595 0.5168 0.5771 0.4963 0.6214 0.4870 0.6483 0.5575
86 vn_S_STRY_reg 3544 0.7550 0.5417 0.6284 0.6181 0.7309 0.5245 0.6553 0.6423
87 vn_S_TECH_reg 3544 0.7180 0.5895 0.7318 0.7560 0.6733 0.5529 0.7636 0.7858
88 vn_S_WHIS_reg 3544 0.5282 0.3304 0.6921 0.6367 0.4893 0.3061 0.7177 0.6580
89 vn_TEMP_reg 3544 0.4112 0.3424 0.7711 0.7410 0.4034 0.3360 0.7789 0.7479
90 vn_TENS_reg 3544 0.5110 0.6140 0.7012 0.6084 0.4922 0.5914 0.7223 0.6347
91 vn_VALN_reg 3544 0.7516 0.6652 0.6443 0.6258 0.6572 0.5816 0.7361 0.7195
92 vn_VALS_reg 3544 0.4307 0.3921 0.4358 0.3966 0.4074 0.3709 0.5246 0.4904
93 vn_VFLX_reg 3544 0.3329 0.2758 0.1541 0.1562 0.3150 0.2611 0.1945 0.1881
94 vn_VOLT_reg 3544 0.4732 0.5448 0.7000 0.6775 0.4696 0.5408 0.7030 0.6810
95 vn_VULN_reg 3544 0.6739 0.6803 0.7484 0.7045 0.6464 0.6526 0.7704 0.7329
96 vn_WARM_reg 3544 0.4646 0.5134 0.5795 0.5277 0.4362 0.4820 0.6354 0.5966
97 genuineness_0_6 3544 0.4774 0.2971 0.5561 0.5957 0.4704 0.2928 0.5616 0.5979
98 blend_0_10 2391 0.8773 0.3228 0.2822 0.3074 0.8417 0.3097 0.3151 0.3216
99 R_quality 2413 0.1025 0.3095 0.9180 0.8908 0.0972 0.2937 0.9246 0.8994
100 eiv_extra:Age 3368 0.1391 0.4716 0.8225 0.7135 0.1344 0.4557 0.8399 0.7298
101 eiv_extra:Arousal 3368 0.2575 0.5906 0.8606 0.8790 0.2399 0.5503 0.8760 0.8948
102 eiv_extra:Authenticity 3368 0.0930 0.3718 0.9065 0.9116 0.0872 0.3487 0.9187 0.9240
103 eiv_extra:Background_Noise 3368 0.1414 0.5358 0.8755 0.8651 0.1324 0.5018 0.8925 0.8775
104 eiv_extra:Confident_vs._Hesitant 3368 0.1701 0.5593 0.9002 0.9047 0.1545 0.5081 0.9178 0.9235
105 eiv_extra:Gender 3368 0.2693 0.2700 0.9372 0.9178 0.2534 0.2540 0.9454 0.9256
106 eiv_extra:High-Pitched_vs._Low-Pitched 3368 0.0894 0.3577 0.9407 0.9421 0.0802 0.3206 0.9523 0.9529
107 eiv_extra:Monotone_vs._Expressive 3368 0.1635 0.3478 0.9344 0.9346 0.1480 0.3147 0.9471 0.9475
108 eiv_extra:Recording_Quality 3368 0.1351 0.4062 0.9468 0.9465 0.1256 0.3775 0.9545 0.9551
109 eiv_extra:Serious_vs._Humorous 3368 0.1966 0.5766 0.8678 0.8451 0.1877 0.5506 0.8799 0.8581
110 eiv_extra:Soft_vs._Harsh 3368 0.1796 0.5623 0.7959 0.8006 0.1717 0.5377 0.8127 0.8101
111 eiv_extra:Submissive_vs._Dominant 3368 0.1440 0.4455 0.8502 0.8475 0.1348 0.4172 0.8704 0.8672
112 eiv_extra:Valence 3368 0.3711 0.9724 0.8634 0.8222 0.3430 0.8989 0.8808 0.8320
113 eiv_extra:Vulnerable_vs._Emotionally_Detached 3368 0.1589 0.3646 0.9240 0.9231 0.1509 0.3462 0.9328 0.9323
114 eiv_extra:Warm_vs._Cold 3368 0.1916 0.6387 0.8206 0.7840 0.1808 0.6030 0.8421 0.8045
115 eiv_extra:score_background_quality 3368 0.1026 0.1479 0.9482 0.8992 0.0996 0.1436 0.9518 0.9062
116 eiv_extra:score_content_enjoyment 3368 0.0765 0.2518 0.9534 0.9410 0.0728 0.2398 0.9577 0.9455
117 eiv_extra:score_overall_quality 3368 0.0819 0.1613 0.9615 0.9363 0.0786 0.1549 0.9645 0.9391
118 eiv_extra:score_speech_quality 3368 0.0479 0.1915 0.8332 0.7329 0.0445 0.1779 0.8574 0.7713
119 audiobox:CE 3544 0.2039 0.2006 0.9459 0.9448 0.1902 0.1871 0.9546 0.9518
120 audiobox:CU 3544 0.2151 0.2153 0.9356 0.9426 0.1998 0.2000 0.9454 0.9494
121 audiobox:PC 3544 0.1312 0.0812 0.9535 0.8055 0.1204 0.0745 0.9619 0.8117
122 audiobox:PQ 3544 0.2230 0.2537 0.9322 0.9349 0.2057 0.2340 0.9429 0.9408
123 dnsmos:SIG_raw 3368 0.1495 0.2234 0.9061 0.8503 0.1441 0.2153 0.9133 0.8592
124 dnsmos:BAK_raw 3368 0.2130 0.1893 0.9205 0.8742 0.2055 0.1826 0.9258 0.8810
125 dnsmos:OVRL_raw 3368 0.1801 0.2215 0.9296 0.8908 0.1743 0.2144 0.9339 0.8969
126 dnsmos:SIG 3368 0.0938 0.1872 0.9033 0.8464 0.0901 0.1797 0.9108 0.8563
127 dnsmos:BAK 3368 0.1440 0.1551 0.9250 0.8659 0.1395 0.1502 0.9301 0.8699
128 dnsmos:OVRL 3368 0.1200 0.1978 0.9327 0.8902 0.1161 0.1913 0.9371 0.8965
129 dnsmos:P808_MOS 3368 0.1283 0.2585 0.9176 0.8696 0.1235 0.2487 0.9230 0.8819
130 burst_count_log1p 3544 0.2474 0.3815 0.8305 0.8198 0.2492 0.3843 0.8273 0.8164
131 voiceclap_attribute:Affection 3544 0.5078 0.9158 0.5838 0.4973 0.4877 0.8795 0.6392 0.5332
132 voiceclap_attribute:Age 3368 0.1578 0.2902 0.8188 0.7483 0.1504 0.2766 0.8313 0.7624
133 voiceclap_attribute:Amusement 3544 0.5768 1.0050 0.7031 0.6284 0.5497 0.9579 0.7273 0.6620
134 voiceclap_attribute:Anger 3544 0.4138 0.8156 0.6069 0.4560 0.3912 0.7711 0.6250 0.4527
135 voiceclap_attribute:Arousal 3368 0.2821 0.4532 0.8416 0.8537 0.2863 0.4599 0.8385 0.8489
136 voiceclap_attribute:Astonishment_Surprise 3544 0.4791 1.0005 0.6094 0.5371 0.4772 0.9966 0.6296 0.5575
137 voiceclap_attribute:Authenticity 3368 0.0910 0.3639 0.8870 0.8965 0.0938 0.3751 0.8810 0.8852
138 voiceclap_attribute:Awe 3544 0.4582 1.7662 0.6179 0.4764 0.4115 1.5863 0.7134 0.5330
139 voiceclap_attribute:Background_Noise 3368 0.0962 0.3644 0.8873 0.8881 0.1001 0.3789 0.8754 0.8735
140 voiceclap_attribute:Bitterness 3544 0.4419 1.7677 0.5745 0.4753 0.4032 1.6130 0.6572 0.5235
141 voiceclap_attribute:Concentration 3544 0.5051 0.9521 0.6333 0.6247 0.5024 0.9471 0.6427 0.6434
142 voiceclap_attribute:Confident_vs._Hesitant 3368 0.2211 0.4716 0.8456 0.8455 0.2246 0.4791 0.8385 0.8364
143 voiceclap_attribute:Confusion 3544 0.4366 1.3024 0.4817 0.4449 0.4303 1.2836 0.5233 0.4768
144 voiceclap_attribute:Contemplation 3544 0.6185 1.5260 0.6700 0.6556 0.6125 1.5110 0.6794 0.6710
145 voiceclap_attribute:Contempt 3544 0.3968 1.1680 0.5738 0.3881 0.3705 1.0907 0.6316 0.4085
146 voiceclap_attribute:Contentment 3544 0.6851 2.2319 0.6315 0.6411 0.6264 2.0406 0.6953 0.6942
147 voiceclap_attribute:Disappointment 3544 0.5357 1.8439 0.6072 0.5267 0.4829 1.6620 0.6753 0.5702
148 voiceclap_attribute:Disgust 3544 0.2720 1.0882 0.5145 0.3850 0.2585 1.0342 0.6078 0.4214
149 voiceclap_attribute:Distress 3544 0.5497 1.5707 0.7163 0.6190 0.4861 1.3889 0.7655 0.6572
150 voiceclap_attribute:Doubt 3544 0.5253 2.1012 0.5091 0.5297 0.5016 2.0064 0.5578 0.5718
151 voiceclap_attribute:Elation 3544 0.4601 1.1538 0.6368 0.5550 0.4292 1.0765 0.6666 0.5830
152 voiceclap_attribute:Embarrassment 3544 0.1896 0.6023 0.3048 0.2565 0.1799 0.5715 0.2576 0.2234
153 voiceclap_attribute:Emotional_Numbness 3544 0.2522 0.9473 0.3121 0.2899 0.2507 0.9419 0.3167 0.2929
154 voiceclap_attribute:Fatigue_Exhaustion 3544 0.5180 1.9897 0.7102 0.6597 0.4850 1.8632 0.7385 0.6754
155 voiceclap_attribute:Fear 3544 0.2930 1.1722 0.5463 0.4014 0.2913 1.1654 0.5694 0.4252
156 voiceclap_attribute:Gender 3368 0.2379 0.2368 0.9505 0.9383 0.2396 0.2385 0.9482 0.9338
157 voiceclap_attribute:Helplessness 3544 0.5094 1.8990 0.7148 0.5998 0.4535 1.6905 0.7714 0.6370
158 voiceclap_attribute:High-Pitched_vs._Low-Pitched 3368 0.1280 0.3427 0.8771 0.8826 0.1272 0.3405 0.8782 0.8872
159 voiceclap_attribute:Hope_Enthusiasm_Optimism 3544 0.6346 1.5974 0.6166 0.6083 0.5817 1.4642 0.6760 0.6699
160 voiceclap_attribute:Impatience_and_Irritability 3544 0.5706 1.0528 0.5800 0.4786 0.5425 1.0009 0.6083 0.5011
161 voiceclap_attribute:Infatuation 3544 0.1870 0.5623 0.3724 0.2935 0.1843 0.5543 0.4069 0.3121
162 voiceclap_attribute:Interest 3544 0.5735 1.3054 0.6784 0.6338 0.5473 1.2458 0.7189 0.6805
163 voiceclap_attribute:Intoxication_Altered_States_of_Consciousness 3544 0.1527 0.3292 0.2480 0.1951 0.1551 0.3344 0.2136 0.1762
164 voiceclap_attribute:Jealousy_&_Envy 3544 0.1094 0.4377 0.1123 0.1094 0.0991 0.3965 0.1322 0.0996
165 voiceclap_attribute:Longing 3544 0.4261 1.7046 0.5481 0.5207 0.4193 1.6770 0.5829 0.5493
166 voiceclap_attribute:Malevolence_Malice 3544 0.2214 0.6399 0.3811 0.2645 0.2264 0.6543 0.4078 0.3097
167 voiceclap_attribute:Monotone_vs._Expressive 3368 0.2390 0.4006 0.8655 0.8564 0.2346 0.3934 0.8682 0.8571
168 voiceclap_attribute:Pain 3544 0.3618 1.4471 0.6585 0.5520 0.3395 1.3579 0.6994 0.5805
169 voiceclap_attribute:Pleasure_Ecstasy 3544 0.3606 1.1972 0.6002 0.4927 0.3280 1.0889 0.6716 0.5413
170 voiceclap_attribute:Pride 3544 0.5591 1.6019 0.5473 0.5229 0.5450 1.5617 0.5676 0.5412
171 voiceclap_attribute:Recording_Quality 3368 0.1340 0.4083 0.9180 0.9200 0.1342 0.4090 0.9179 0.9209
172 voiceclap_attribute:Relief 3544 0.5277 1.2686 0.5372 0.4394 0.5009 1.2041 0.5875 0.4669
173 voiceclap_attribute:Sadness 3544 0.4434 1.4169 0.6306 0.5121 0.3948 1.2618 0.7219 0.5584
174 voiceclap_attribute:Serious_vs._Humorous 3368 0.2265 0.3971 0.8688 0.8286 0.2183 0.3827 0.8774 0.8265
175 voiceclap_attribute:Sexual_Lust 3544 0.1490 0.3565 0.2604 0.2024 0.1441 0.3448 0.2510 0.2158
176 voiceclap_attribute:Shame 3544 0.1468 0.5211 0.2995 0.1960 0.1480 0.5254 0.2968 0.1973
177 voiceclap_attribute:Soft_vs._Harsh 3368 0.1702 0.4727 0.8318 0.8489 0.1759 0.4883 0.8189 0.8327
178 voiceclap_attribute:Sourness 3544 0.4401 1.5150 0.5642 0.4561 0.4040 1.3909 0.6341 0.5007
179 voiceclap_attribute:Submissive_vs._Dominant 3368 0.1595 0.4444 0.8189 0.8267 0.1583 0.4408 0.8211 0.8262
180 voiceclap_attribute:Teasing 3544 0.4834 1.2742 0.6079 0.5211 0.4720 1.2443 0.6292 0.5347
181 voiceclap_attribute:Thankfulness_Gratitude 3544 0.3695 0.6980 0.5228 0.4199 0.3478 0.6570 0.6076 0.4657
182 voiceclap_attribute:Triumph 3544 0.4139 1.1535 0.5471 0.4353 0.4051 1.1293 0.5561 0.4452
183 voiceclap_attribute:Valence 3368 0.4702 0.7416 0.7155 0.7111 0.4603 0.7259 0.7213 0.7128
184 voiceclap_attribute:Vulnerable_vs._Emotionally_Detached 3368 0.2181 0.3924 0.8508 0.8581 0.2110 0.3796 0.8608 0.8564
185 voiceclap_attribute:Warm_vs._Cold 3368 0.2497 0.6254 0.7692 0.7600 0.2392 0.5989 0.7818 0.7665
186 voiceclap_attribute:duration 3368 0.9696 0.1531 0.9725 0.9677 1.0755 0.1698 0.9654 0.9629
187 voiceclap_attribute:score_background_quality 3368 0.1229 0.2464 0.8736 0.8523 0.1219 0.2446 0.8726 0.8543
188 voiceclap_attribute:score_content_enjoyment 3368 0.0622 0.2489 0.9033 0.9136 0.0648 0.2593 0.8934 0.9053
189 voiceclap_attribute:score_overall_quality 3368 0.0981 0.2378 0.9158 0.9063 0.0979 0.2374 0.9155 0.9048
190 voiceclap_attribute:score_speech_quality 3368 0.0272 0.1087 0.8570 0.8576 0.0273 0.1091 0.8536 0.8539
191 voiceclap_attribute:talking_speed 3368 1.4987 0.3994 0.9193 0.9197 1.3978 0.3725 0.9310 0.9277

Canonical 53-class vocal-burst vocabulary

The classifier predicts these categories. Fine labels outside this vocabulary use the explicit aliases/fallbacks in gemini_burst_mapping.json; the fine-to-group23 map is also preserved.

ID Canonical vocal-burst class
0 Other / rare burst
1 Affirmative Grunt
2 Ahem
3 Breath
4 Breathy Giggle
5 Cackle
6 Childlike Giggle
7 Chuckle
8 Contented Sigh
9 Cough
10 Crying
11 Displeased Grunt
12 Effort Grunt
13 Exasperated Sigh
14 Exhausted Groan
15 Fearful Gasp
16 Frustrated Groan
17 Growl
18 Guffaw
19 Gulp
20 Hiccup
21 Hum
22 Kissing Sound
23 Laughter
24 Lip Smack
25 Low Mumble
26 Mournful Wail
27 Nervous Giggle
28 Nervous Gulp
29 Purr
30 Quiet Sob
31 Relief Sigh
32 Scream
33 Sharp Inhale
34 Sharp Whistle
35 Shriek
36 Sigh
37 Sneeze
38 Snicker
39 Sniff
40 Snort
41 Snorting Giggle
42 Soft Whistle
43 Spitting
44 Surprised Gasp
45 Swallows
46 Throat Clearing
47 Tongue Click
48 Trembling Whimper
49 Tsk
50 Whispered Mumble
51 Wistful Sigh
52 Yawn
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for laion/humaneness-ears-base-medium

Finetuned
(1)
this model

Datasets used to train laion/humaneness-ears-base-medium

Space using laion/humaneness-ears-base-medium 1