DeltaV: Thinking with Visual State Updates in Unified Large Multimodal Models Paper • 2607.08434 • Published about 1 month ago • 2
MonkeyOCRv2 Collection The collection of the paper `MonkeyOCRv2: A Visual-Text Foundation Model for Document AI`. • 10 items • Updated 15 days ago • 2
MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Paper • 2607.11562 • Published 26 days ago • 77
MonkeyOCR: Document Parsing with a Structure-Recognition-Relation Triplet Paradigm Paper • 2506.05218 • Published Jun 5, 2025 • 3
TextSquare: Scaling up Text-Centric Visual Instruction Tuning Paper • 2404.12803 • Published Apr 19, 2024 • 30
Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models Paper • 2311.06607 • Published Nov 11, 2023 • 4