Personalization LLM User-LLM: Efficient LLM Contextualization with User Embeddings Paper • 2402.13598 • Published Feb 21, 2024 • 22
User-LLM: Efficient LLM Contextualization with User Embeddings Paper • 2402.13598 • Published Feb 21, 2024 • 22
Indic Datasets List of text and voice datasets to train and finetune Indic LLMs ai4bharat/sangraha Viewer • Updated Mar 5, 2025 • 268M • 39.3k • 83 uonlp/CulturaX Viewer • Updated Dec 16, 2024 • 7.18B • 29.2k • 688 pary/hind_encorp Updated Jan 18, 2024 • 192 • 2 PleIAs/YouTube-Commons Updated Jun 26, 2024 • 6.64k • 396
Papers SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training Paper • 2501.17161 • Published Jan 28, 2025 • 127
SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training Paper • 2501.17161 • Published Jan 28, 2025 • 127
Alignment Dataset English and other model alignment datasets. H-D-T/Buzz-8b-Large-v0.5 Text Generation • 8B • Updated May 14, 2024 • 51 • 29 allenai/WildChat-1M Viewer • Updated Oct 17, 2024 • 838k • 27.6k • 464 nvidia/ChatQA-Training-Data Viewer • Updated Jun 4, 2024 • 442k • 914 • 177 nvidia/ChatRAG-Bench Viewer • Updated May 24, 2024 • 34.6k • 1.89k • 118
Papers SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training Paper • 2501.17161 • Published Jan 28, 2025 • 127
SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training Paper • 2501.17161 • Published Jan 28, 2025 • 127
Personalization LLM User-LLM: Efficient LLM Contextualization with User Embeddings Paper • 2402.13598 • Published Feb 21, 2024 • 22
User-LLM: Efficient LLM Contextualization with User Embeddings Paper • 2402.13598 • Published Feb 21, 2024 • 22
Indic Datasets List of text and voice datasets to train and finetune Indic LLMs ai4bharat/sangraha Viewer • Updated Mar 5, 2025 • 268M • 39.3k • 83 uonlp/CulturaX Viewer • Updated Dec 16, 2024 • 7.18B • 29.2k • 688 pary/hind_encorp Updated Jan 18, 2024 • 192 • 2 PleIAs/YouTube-Commons Updated Jun 26, 2024 • 6.64k • 396
Alignment Dataset English and other model alignment datasets. H-D-T/Buzz-8b-Large-v0.5 Text Generation • 8B • Updated May 14, 2024 • 51 • 29 allenai/WildChat-1M Viewer • Updated Oct 17, 2024 • 838k • 27.6k • 464 nvidia/ChatQA-Training-Data Viewer • Updated Jun 4, 2024 • 442k • 914 • 177 nvidia/ChatRAG-Bench Viewer • Updated May 24, 2024 • 34.6k • 1.89k • 118