JetBrains/Mellum2-12B-A2.5B-Thinking-SFT Text Generation • 12B • Updated about 13 hours ago • 324 • 25
JetBrains/Mellum2-12B-A2.5B-Instruct-SFT Text Generation • 12B • Updated about 13 hours ago • 158 • 14
JetBrains/Mellum2-12B-A2.5B-Thinking Text Generation • 12B • Updated about 13 hours ago • 4.9k • 331
JetBrains/Mellum2-12B-A2.5B-Instruct Text Generation • 12B • Updated about 13 hours ago • 4.3k • 82
The Best of N Worlds: Aligning Reinforcement Learning with Best-of-N Sampling via max@k Optimisation Paper • 2510.23393 • Published Oct 27, 2025 • 21
On Pretraining for Project-Level Code Completion Paper • 2510.13697 • Published Oct 15, 2025 • 7
Diff-XYZ: A Benchmark for Evaluating Diff Understanding Paper • 2510.12487 • Published Oct 14, 2025 • 9
The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management Paper • 2508.21433 • Published Aug 29, 2025 • 9
CrowdSpeech and VoxDIY: Benchmark Datasets for Crowdsourced Audio Transcription Paper • 2107.01091 • Published Jul 2, 2021
Best Prompts for Text-to-Image Models and How to Find Them Paper • 2209.11711 • Published Sep 23, 2022 • 3
Spherical convolutions on molecular graphs for protein model quality assessment Paper • 2011.07980 • Published Nov 16, 2020
IMDB-WIKI-SbS: An Evaluation Dataset for Crowdsourced Pairwise Comparisons Paper • 2110.14990 • Published Oct 28, 2021
Towards Realistic Evaluation of Commit Message Generation by Matching Online and Offline Settings Paper • 2410.12046 • Published Oct 15, 2024
Towards Realistic Evaluation of Commit Message Generation by Matching Online and Offline Settings Paper • 2410.12046 • Published Oct 15, 2024
Long Code Arena: a Set of Benchmarks for Long-Context Code Models Paper • 2406.11612 • Published Jun 17, 2024 • 26
Long Code Arena: a Set of Benchmarks for Long-Context Code Models Paper • 2406.11612 • Published Jun 17, 2024 • 26