Firebot Voice Intent Classifier & Speaker Verification Models

Official model checkpoints and speaker verification voiceprints for Firebot, an autonomous tactical firefighting robot platform with dual-tier offline voice intent recognition and personalized acoustic adaptation.

Overview

The Firebot voice system is engineered for zero-latency, high-reliability command recognition in noisy operational environments. It combines:

  1. Frozen Whisper Encoder Feature Extraction: 384-dimensional pooled audio embeddings from OpenAI's Whisper model (tiny.en / base.en).
  2. L2-SP Regularized MLP Intent Classifier: A lightweight Multi-Layer Perceptron (LayerNorm $\to$ Linear 384$\times$128 $\to$ GELU $\to$ Dropout $\to$ Linear 128$\times$15) fine-tuned with an L2 penalty pulling weights toward the base anchor.
  3. ECAPA-TDNN Multi-Clip Acoustic Voiceprints: 192-dimensional speaker embeddings for tactical operator authentication and biometric voice gating.
  4. Adaptive Online Calibration: Few-shot per-operator head adaptation enabling rapid personalization from 1–5 takes per command with held-out generalization validation.

Model Architecture & Directory Layout

firebot-voice-intent/
β”œβ”€β”€ intent_head.pt           # Base 15-class Whisper intent classification head
β”œβ”€β”€ intent_head.json         # Architecture configuration, pooling, and class map
β”œβ”€β”€ intent_prototypes.pt     # Latent prototype embeddings for few-shot similarity
β”œβ”€β”€ users/
β”‚   β”œβ”€β”€ ananya.pt            # Calibrated personal Whisper head for Operator Ananya
β”‚   β”œβ”€β”€ ananya.json          # Empirical training metrics and per-class take breakdown
β”‚   β”œβ”€β”€ ananya.history.jsonl # Training iteration telemetry log
β”‚   β”œβ”€β”€ avinandan.pt         # Calibrated personal Whisper head for Operator Avinandan
β”‚   β”œβ”€β”€ avinandan.json       # Calibration metrics
β”‚   └── avinandan.history.jsonl
└── voiceprints/
    β”œβ”€β”€ ananya.npy           # ECAPA-TDNN biometric speaker voiceprint embedding
    └── avinandan.npy        # ECAPA-TDNN biometric speaker voiceprint embedding

Command Vocabulary (15 Closed-Set Classes)

Class ID Canonical Label Voice Phrases / Synonyms Target Coordinate / Action
0 STOP "stop", "halt", "freeze", "abort mission", "e-stop" Emergency brake & valve cutoff
1 EXTINGUISH "put out the fire", "extinguish", "spray the fire", "suppress" Engage high-pressure water pump
2 RETURN_HOME "return home", "go to base", "back to dock", "retreat" Autonomous navigation to dock (1.2, 1.0)
3 STATUS "status report", "give me a report", "tank level", "sitrep" Telemetry & battery check
4 UNKNOWN "hello", "what's the weather", "testing" Out-of-domain conversational filter
5 GOTO_HOME "go to home", "head to charging station" Waypoint navigation: (1.2, 1.0)
6 GOTO_CENTER "go to center", "move to the middle" Waypoint navigation: (6.0, 4.0)
7 GOTO_NORTH "go to north", "move to the north side" Waypoint navigation: (6.0, 7.0)
8 GOTO_SOUTH "go to south", "head south" Waypoint navigation: (6.0, 1.0)
9 GOTO_EAST "go to east", "move right" Waypoint navigation: (10.5, 4.0)
10 GOTO_WEST "go to west", "move left" Waypoint navigation: (2.0, 4.0)
11 GOTO_NORTHEAST "go to northeast", "upper right corner" Waypoint navigation: (10.5, 7.0)
12 GOTO_NORTHWEST "go to northwest", "upper left corner" Waypoint navigation: (1.5, 7.0)
13 GOTO_SOUTHEAST "go to southeast", "bottom right corner" Waypoint navigation: (10.5, 1.0)
14 GOTO_SOUTHWEST "go to southwest", "bottom left corner" Waypoint navigation: (1.5, 1.2)

Quickstart: Python Inference

import soundfile as sf
import torch
from huggingface_hub import hf_hub_download
from firebot.voice_intent.infer import IntentClassifier

# 1. Download model from Hugging Face Hub
ckpt_path = hf_hub_download(repo_id="anabaena/firebot-voice-intent", filename="intent_head.pt")

# 2. Instantiate Intent Classifier (loads frozen Whisper + lightweight head)
classifier = IntentClassifier(checkpoint_path=ckpt_path)

# 3. Classify raw audio (16kHz mono WAV)
audio, sr = sf.read("command_sample.wav", dtype="float32")
prediction = classifier.predict(audio, sample_rate=sr)

print(f"Predicted Intent: {prediction['label']}")
print(f"Confidence: {prediction['confidence']:.2%}")
print(f"Action: {prediction['canonical_phrase']}")

Biometric Speaker Verification (ECAPA-TDNN)

import numpy as np

# Load enrolled speaker voiceprint
enrolled = np.load("voiceprints/ananya.npy")

# Compare with live speaker embedding
cosine_sim = np.dot(enrolled, live_embedding) / (np.linalg.norm(enrolled) * np.linalg.norm(live_embedding))
is_authorized = cosine_sim >= 0.72  # calibrated acceptance threshold
print(f"Speaker Verified: {is_authorized} (Similarity: {cosine_sim:.3f})")

Performance & Evaluation

  • Inference Latency: ~35ms on CPU (Apple Silicon / modern x86_64)
  • Base Intent Accuracy: 94.2% on diverse acoustic validation test set
  • Personalized Operator Accuracy: 98.8% – 100.0% with 3–5 takes per command
  • Memory Footprint: ~210 KB for the MLP head, ~75 MB for frozen Whisper tiny encoder

License

MIT License. Designed and developed as part of the Firebot Autonomous Robotic Platform.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Evaluation results