ThinkJEPA: Empowering Latent World Models with Large Vision-Language Reasoning Model Paper • 2603.22281 • Published Mar 23 • 20
Ref-Adv: Exploring MLLM Visual Reasoning in Referring Expression Tasks Paper • 2602.23898 • Published Feb 27 • 10
Fine-T2I: An Open, Large-Scale, and Diverse Dataset for High-Quality T2I Fine-Tuning Paper • 2602.09439 • Published Feb 10 • 14
The Quest for the Right Mediator: A History, Survey, and Theoretical Grounding of Causal Interpretability Paper • 2408.01416 • Published Aug 2, 2024 • 1
NNsight and NDIF: Democratizing Access to Foundation Model Internals Paper • 2407.14561 • Published Jul 18, 2024 • 35
ThinkGrasp: A Vision-Language System for Strategic Part Grasping in Clutter Paper • 2407.11298 • Published Jul 16, 2024 • 6
Imagination Policy: Using Generative Point Cloud Models for Learning Manipulation Policies Paper • 2406.11740 • Published Jun 17, 2024 • 1
Token Erasure as a Footprint of Implicit Vocabulary Items in LLMs Paper • 2406.20086 • Published Jun 28, 2024 • 6
Concept Sliders: LoRA Adaptors for Precise Control in Diffusion Models Paper • 2311.12092 • Published Nov 20, 2023 • 22