Streaming models for speech-to-text, with VAD. Plus models include speaker segmentation and verification for cascaded diarization.