Hiera: A Hierarchical Vision Transformer without the Bells-and-Whistles Paper • 2306.00989 • Published Jun 1, 2023 • 2
view article Article NVIDIA Nemotron 3 Embed Ranks #1 Overall on RTEB, Advancing Agentic Retrieval nvidia • 19 days ago • 58
Inkling Collection Inkling is a versatile, customizable model that reasons over text, images, audio, with variable and efficient thinking effort. • 4 items • Updated 7 days ago • 45
Olmo 3.1 Collection The latest members of the Olmo 3 family: another 3 weeks of RL for 32B Think, the 32B Instruct model, large post-training research datasets... • 9 items • Updated Dec 23, 2025 • 54
Perception Encoder Collection OpenCLIP (PE Core image + text) and timm PE Core, Spatial, Lang (ViT only) weights. NOTE: These weights do not work with original modeling code. • 19 items • Updated Sep 19, 2025 • 8
Franca: Nested Matryoshka Clustering for Scalable Visual Representation Learning Paper • 2507.14137 • Published Jul 18, 2025 • 36
Cluster and Predict Latents Patches for Improved Masked Image Modeling Paper • 2502.08769 • Published Feb 12, 2025 • 5
INaturalist-2021 Fine-tunes Collection Fine-tune experiments for various `timm` models on the INaturalist 2021 Challenge dataset (https://github.com/visipedia/inat_comp/tree/master/2021) • 5 items • Updated Oct 25, 2023 • 7