AI & ML interests

Accelerate the frontier of AI development with enterprise-grade, deeply curated datasets engineered to enhance pre-training, alignment, and real-world performance.

Recent Activity

RohitManglik  updated a collection about 13 hours ago
STEM & Non-STEM Q&A Datasets for LLM Training
RohitManglik  updated a collection about 15 hours ago
STEM & Non-STEM Q&A Datasets for LLM Training
RohitManglik  updated a dataset about 15 hours ago
InfoBayAI/Multimodel_QA_Dataset
View all activity

InfoBayAI 's collections 14

Dual Channel Global Customer-Agent Interaction Datasets
Sample Datasets of dual-channel call center audio with separate agent and customer channels for ASR, diarization, and conversational AI training.
Healthcare AI Datasets for Clinical & LLM Training
Sample dataset from an enterprise-grade medical corpus built for clinical AI, diagnosis support, and healthcare LLM training.
Single-channel Podcast Speech Audio Datasets
Sample from a podcast audio dataset, designed for ASR, speech recognition, and conversational AI training using diverse, real-world spoken content.
Academic Textbook Corpora for LLM Training
Sample of a 2.6+ word textbook corpus across 39K+ books, 5K+ subjects, and 15 languages for LLM training and multilingual knowledge modeling.
Healthcare AI Datasets for Clinical & LLM Training
Sample dataset from an enterprise-grade medical corpus built for clinical AI, diagnosis support, and healthcare LLM training.
Dual Channel Global Customer-Agent Interaction Datasets
Sample Datasets of dual-channel call center audio with separate agent and customer channels for ASR, diarization, and conversational AI training.
Single-channel Podcast Speech Audio Datasets
Sample from a podcast audio dataset, designed for ASR, speech recognition, and conversational AI training using diverse, real-world spoken content.
Academic Textbook Corpora for LLM Training
Sample of a 2.6+ word textbook corpus across 39K+ books, 5K+ subjects, and 15 languages for LLM training and multilingual knowledge modeling.