⨠Highlights ⢠┠Up to 3.4Ć faster on dense multi-region captioning, with stable per-image latency ⢠š PerceptionDLM-Base beats LLaDA-V on 15/16 multimodal benchmarks (new SOTA among open diffusion VLMs) ⢠š New benchmark: ParaDLC-Bench ā jointly evaluates caption quality AND inference efficiency ⢠š Code, models & benchmark all open-sourced
⨠Highlights ⢠┠Up to 3.4Ć faster on dense multi-region captioning, with stable per-image latency ⢠š PerceptionDLM-Base beats LLaDA-V on 15/16 multimodal benchmarks (new SOTA among open diffusion VLMs) ⢠š New benchmark: ParaDLC-Bench ā jointly evaluates caption quality AND inference efficiency ⢠š Code, models & benchmark all open-sourced