Developed in collaboration with Mirelo (AI https://www.mirelo.ai/), MuScriptor is an open-weight model for multi-instrument automatic transcription
AI & ML interests
Our mission is to build and democratize artificial general intelligence through open science
Recent Activity
Papers
MuScriptor: An Open Model for Multi-Instrument Music Transcription
The FID Lottery: Quantifying Hidden Randomness in Generative-Model Evaluation
Full-duplex speech models post-trained with reinforcement learning for improved conversational interactivity.
-
Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models
Paper • 2606.11167 • Published • 5 -
kyutai/moshika-rl-seamless
Audio-to-Audio • 8B • Updated • 172 • 30 -
kyutai/personaplex-rl-seamless
Audio-to-Audio • 8B • Updated • 8.87k • 31 -
kyutai/interactivity-alignment-samples
Viewer • Updated • 5.56k • 643 • 8
Streaming speech translation without the need for word-level alignments
-
Hibiki Zero Samples
🏆13Demo samples of the speech translation model Hibiki-Zero.
-
Simultaneous Speech-to-Speech Translation Without Aligned Data
Paper • 2602.11072 • Published • 2 -
kyutai/Audio-NTREX-4L
Viewer • Updated • 3.6k • 462 • 6 -
kyutai/hibiki-zero-3b-pytorch-bf16
Audio-to-Audio • Updated • 2.05k • 56
CASA: Cross-Attention over Self-Attention for Efficient Vision-Language Fusion on long-context streaming inputs
-
CASA Gallery
🏠4Video Gallery for CASA: Cross-Attention over Self-Attention
-
CASA: Cross-Attention via Self-Attention for Efficient Vision-Language Fusion
Paper • 2512.19535 • Published • 13 -
kyutai/CASA-Helium1-VL-2B
Image-Text-to-Text • 3B • Updated • 24 • 9 -
kyutai/CASA-Qwen2_5-VL-3B
Image-Text-to-Text • 4B • Updated • 19 • 3
https://kyutai.org/next/tts
Helium 1: a modular and multilingual LLM
Hibiki is a model for streaming speech translation , which can run on device! See https://github.com/kyutai-labs/hibiki.
Developed with General Intuition (https://www.generalintuition.com/), in collaboration with Epic Games. MIRA is a Multiplayer Interactive World Model.
Candle & PyTorch model checkpoints released as part of the MoshiRAG release from Kyutai. Run inference via: https://github.com/kyutai-labs/moshi-rag
Temporal pretraining checkpoints and KairosQA evaluation dataset
Pretrained ARC-Encoders and a fine-tuning dataset: context compression for unmodified LLMs.
-
ARC-Encoder: learning compressed text representations for large language models
Paper • 2510.20535 • Published • 10 -
kyutai/ARC8_Encoder_Llama
Feature Extraction • 3B • Updated • 14 • 3 -
kyutai/ARC_finetuning
Preview • Updated • 29 • 1 -
kyutai/ARC8_Encoder_multi
Feature Extraction • 3B • Updated • 7 • 7
https://kyutai.org/next/stt
MoshiVis is a Vision Speech Model built as a perceptually-augmented version of Moshi v0.1 for conversing about image inputs
MLX, Candle & PyTorch model checkpoints released as part of the Moshi release from Kyutai. Run inference via: https://github.com/kyutai-labs/moshi
Developed in collaboration with Mirelo (AI https://www.mirelo.ai/), MuScriptor is an open-weight model for multi-instrument automatic transcription
Developed with General Intuition (https://www.generalintuition.com/), in collaboration with Epic Games. MIRA is a Multiplayer Interactive World Model.
Full-duplex speech models post-trained with reinforcement learning for improved conversational interactivity.
-
Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models
Paper • 2606.11167 • Published • 5 -
kyutai/moshika-rl-seamless
Audio-to-Audio • 8B • Updated • 172 • 30 -
kyutai/personaplex-rl-seamless
Audio-to-Audio • 8B • Updated • 8.87k • 31 -
kyutai/interactivity-alignment-samples
Viewer • Updated • 5.56k • 643 • 8
Candle & PyTorch model checkpoints released as part of the MoshiRAG release from Kyutai. Run inference via: https://github.com/kyutai-labs/moshi-rag
Streaming speech translation without the need for word-level alignments
-
Hibiki Zero Samples
🏆13Demo samples of the speech translation model Hibiki-Zero.
-
Simultaneous Speech-to-Speech Translation Without Aligned Data
Paper • 2602.11072 • Published • 2 -
kyutai/Audio-NTREX-4L
Viewer • Updated • 3.6k • 462 • 6 -
kyutai/hibiki-zero-3b-pytorch-bf16
Audio-to-Audio • Updated • 2.05k • 56
Temporal pretraining checkpoints and KairosQA evaluation dataset
CASA: Cross-Attention over Self-Attention for Efficient Vision-Language Fusion on long-context streaming inputs
-
CASA Gallery
🏠4Video Gallery for CASA: Cross-Attention over Self-Attention
-
CASA: Cross-Attention via Self-Attention for Efficient Vision-Language Fusion
Paper • 2512.19535 • Published • 13 -
kyutai/CASA-Helium1-VL-2B
Image-Text-to-Text • 3B • Updated • 24 • 9 -
kyutai/CASA-Qwen2_5-VL-3B
Image-Text-to-Text • 4B • Updated • 19 • 3
Pretrained ARC-Encoders and a fine-tuning dataset: context compression for unmodified LLMs.
-
ARC-Encoder: learning compressed text representations for large language models
Paper • 2510.20535 • Published • 10 -
kyutai/ARC8_Encoder_Llama
Feature Extraction • 3B • Updated • 14 • 3 -
kyutai/ARC_finetuning
Preview • Updated • 29 • 1 -
kyutai/ARC8_Encoder_multi
Feature Extraction • 3B • Updated • 7 • 7
https://kyutai.org/next/tts
https://kyutai.org/next/stt
Helium 1: a modular and multilingual LLM
MoshiVis is a Vision Speech Model built as a perceptually-augmented version of Moshi v0.1 for conversing about image inputs
Hibiki is a model for streaming speech translation , which can run on device! See https://github.com/kyutai-labs/hibiki.
MLX, Candle & PyTorch model checkpoints released as part of the Moshi release from Kyutai. Run inference via: https://github.com/kyutai-labs/moshi