emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
Paper • 2312.15185 • Published
ONNX export of emotion2vec/emotion2vec_plus_base for on-device inference with ONNX Runtime.
| File | Description |
|---|---|
emotion2vec_plus_base.onnx |
Backbone. Input is raw 16 kHz mono float32 of shape (1, N), values in [-1, 1]. Output is per-frame features of dimension 768. |
emotion2vec_head.json |
Linear classification head: labels, weight (9 x 768), bias (9). |
Labels: angry, disgusted, fearful, happy, neutral, other, sad,
surprised, unknown
The backbone alone does not classify. Mean-pool its output over the frame axis, apply the linear head, then softmax:
import json
import numpy as np
import onnxruntime as ort
head = json.load(open("emotion2vec_head.json"))
W = np.array(head["weight"], dtype=np.float32)
B = np.array(head["bias"], dtype=np.float32)
labels = head["labels"]
sess = ort.InferenceSession("emotion2vec_plus_base.onnx",
providers=["CPUExecutionProvider"])
# audio: 16 kHz mono float32 in [-1, 1]
feats = sess.run(None, {sess.get_inputs()[0].name: audio.reshape(1, -1)})[0]
pooled = feats[0].mean(axis=0)
logits = W @ pooled + B
probs = np.exp(logits - logits.max())
probs /= probs.sum()
print(labels[int(probs.argmax())])
FunASR Model Open Source License, inherited from the original model. See https://github.com/modelscope/FunASR/blob/main/MODEL_LICENSE.
Commercial use and redistribution are permitted. Attribution and retention of model names are required.
Base model
emotion2vec/emotion2vec_plus_base