Multimodal OCR model for complex document understanding.
qwen3-vl
glm-ocr
Convert spoken words to text