Impressive scale and dataset effort!
Out of curiosity, do you have any plans to add speaker turn labels or alignment confidence scores in future versions?
When we built OleSpeech-IV from similar web streams, we added human-sourced speaker labels and word-level confidence scores so researchers can easily handle multi-speaker turns and filter noisy web captions out of the box: https://huggingface.co/datasets/olewave/OleSpeech-IV-2025-EN-AR-100
Great work expanding open multilingual speech resources for the community!