OmniConfess: Eliciting Token Confessions to Mitigate Omni-Modal Hallucination
Abstract
Omni-modal large language models (OmniLLMs) unify text, images, audio, and video, yet hallucinate when generation relies on the wrong evidence. Existing inference-time methods can reduce hallucinations, but rarely reveal which evidence sustains a generated commitment. We introduce OmniConfess, a training-free method for mitigating omni-modal hallucinations. It fixes a candidate response and re-scores it at token resolution under controlled channel-wise evidence interventions, producing a structured token-by-channel confession that reveals the response's evidential dependence. OmniConfess uses this confession to preserve grounded content and correct commitments driven by irrelevant or contradictory evidence. To evaluate OmniConfess, we construct OmniHalluBench, a 3,540-example benchmark built from six datasets spanning text, image, audio, and video settings and both judgment and free-form generation. Experiments show that OmniConfess mitigates hallucinations across heterogeneous modality and task settings. Our code and benchmark are publicly available at https://github.com/RongHuiQiang/OmniConfess.
Community
Hi everyone! We introduce OmniConfess, a training-free method for mitigating omni-modal hallucinations. It anchors a candidate response, examines each token’s dependence on individual evidence channels, and uses the resulting confession to guide correction.
We also introduce OmniHalluBench, a 3,540-example benchmark spanning text, image, audio, and video, covering both judgment and free-form generation.
We welcome your feedback and discussion!
Get this paper in your agent:
hf papers read 2610.02999 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 1
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper