Instructions to use SeerRay-Lab/Xiaomi-OCR-0 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use SeerRay-Lab/Xiaomi-OCR-0 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="SeerRay-Lab/Xiaomi-OCR-0") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("SeerRay-Lab/Xiaomi-OCR-0") model = AutoModelForMultimodalLM.from_pretrained("SeerRay-Lab/Xiaomi-OCR-0", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use SeerRay-Lab/Xiaomi-OCR-0 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "SeerRay-Lab/Xiaomi-OCR-0" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "SeerRay-Lab/Xiaomi-OCR-0", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/SeerRay-Lab/Xiaomi-OCR-0
- SGLang
How to use SeerRay-Lab/Xiaomi-OCR-0 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "SeerRay-Lab/Xiaomi-OCR-0" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "SeerRay-Lab/Xiaomi-OCR-0", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "SeerRay-Lab/Xiaomi-OCR-0" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "SeerRay-Lab/Xiaomi-OCR-0", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use SeerRay-Lab/Xiaomi-OCR-0 with Docker Model Runner:
docker model run hf.co/SeerRay-Lab/Xiaomi-OCR-0
Xiaomi-OCR-0
A unified 0.8B model for document parsing and OCR-related understanding.
English · 简体中文
Xiaomi-OCR-0 is a unified 0.8B OCR-specialized vision-language model for both document parsing and OCR-centric understanding.
Starting from Qwen3.5-0.8B-Base, Xiaomi-OCR-0 is trained on an approximately 170M-sample OCR-centric corpus with a progressive training recipe:
- 📍 Q-Mask text anchoring: align fine-grained text with its location
- 📚 OCR-centric continued pretraining (CPT): build broad, multi-task capability
- 🎯 Mixed-task reinforcement learning (Mix-RL): use verifiable rewards on hard examples
✨ Highlights
- 🧩 Compact and unified: one 0.8B model covering document parsing, OCR-oriented VQA, and key information extraction (KIE).
- 🧪 Scalable data engine: combines heterogeneous-expert consensus, render-guided verification, multi-factor sample mining, and targeted synthesis.
- 📄 Robust document parsing: stable on standard documents and on scanned, warped, photographed, skewed, and otherwise recaptured ones.
- 💬 OCR-centric understanding: supports question answering and structured field extraction, not just transcription.
📊 Key Performance
↑ means higher is better, ↓ means lower is better; bold only marks Xiaomi-OCR-0 and does not necessarily indicate the best value in a column.
Document parsing: selected comparisons
| Model | Size | OmniDocBench v1.6 ↑ | Real5 ↑ | Wild ↑ |
|---|---|---|---|---|
| Xiaomi-OCR-0 | 0.8B | 96.83 | 95.24 | 87.94 |
| TeleOCR | 1.2B | 96.87 | — | 88.53 |
| OvisOCR2 | 0.8B | 96.58 | 92.29 | 87.91 |
| PaddleOCR-VL-1.6 | 0.9B | 96.33 | 93.19 | 87.36 |
| MinerU2.5-Pro | 1.2B | 95.75 | 88.94 | 87.33 |
| GLM-OCR | 0.9B | 95.22 | 90.32 | 85.08 |
All three columns report Overall score; “—” means the source table does not report it.
OCR-centric visual question answering
| Model | Size | DocVQA | InfoVQA | ChartQA | OCRBench | TextVQA | Mean |
|---|---|---|---|---|---|---|---|
| Xiaomi-OCR-0 | 0.8B | 93.1 | 75.1 | 84.6 | 84.6 | 78.6 | 83.2 |
| Qwen3.5-0.8B | 0.8B | 88.5 | 60.3 | 69.5 | 77.9 | 68.3 | 72.9 |
| Qwen3.5-2B | 2B | 92.4 | 72.4 | 77.0 | 85.9 | 76.9 | 80.9 |
| Qwen3.5-4B | 4B | 94.4 | 80.4 | 82.4 | 86.6 | 80.8 | 84.9 |
| MiniCPM-V-4.5 | 8B | 84.9 | 69.6 | 87.4 | 89.0 | 82.2 | 82.6 |
The mean is the arithmetic average of these five benchmarks on a 0–100 scale.
🧬 Data Engine
Xiaomi-OCR-0 is trained on an approximately 170M-sample corpus covering document parsing and OCR-centric understanding. The parsing data engine provides supervision for text, tables, and formulas through three main components.
Ensemble Triplet Consensus
A heterogeneous expert pool generates candidate annotations for text, tables, and formulas. Xiaomi-OCR-0 uses Ensemble Triplet Consensus (ETC) to rank candidates by pool-wide agreement, select a three-expert consensus subset, and send uncertain samples for further refinement.
The expert pool in the technical report includes PaddleOCR-VL-1.6, MinerU2.5-Pro, GLM-OCR, dots.mocr, HunyuanOCR-1.5, TeleOCR, OvisOCR2, and Qianfan-OCR.
Render-Guided Refine-and-Judge
When experts disagree, candidate annotations are rendered back into images and compared against the source region. This makes structural errors in tables and formulas easier to catch and correct before the samples enter training.
Multi-Factor Mining and Targeted Synthesis
Training samples are selected by a combination of:
- expert agreement
- model self-consistency
- semantic clustering
- verified model failure patterns
Synthetic data is generated in two complementary ways:
- Coverage-driven synthesis: fills in structures, styles, glyphs, and domains that are under-covered
- Failure-driven synthesis: targets errors that recur during training
📸 Key Findings
During training, we studied how OCR-centric understanding data (hereafter “understanding”) affects document parsing metrics (hereafter “parsing”), and compared the gains of Mix-RL against MOPD (one parsing teacher and one understanding teacher), in the hope of offering the community some useful insight ✨
1. Adding understanding data benefits parsing once parsing capability is already strong
The early-stage change is −0.330; the middle-stage gain is +0.639, followed by +0.0749 on the stronger late-stage baseline. The three stages use 9%, 29%, and 50% of all parsing data, and each gain is measured against that stage's parsing baseline.
2. Understanding needs more parameters and is harder than parsing
The 4B and 0.8B models below go through the same CPT process to keep their data aligned, and are then trained on domain-specific data to produce two expert models.
(a) Document parsing: OmniDocBench v1.6
| Model | Overall ↑ | Text Edit ↓ | Formula CDM ↑ | Table TEDS ↑ | Table TEDS-S ↑ | Reading Order Edit ↓ |
|---|---|---|---|---|---|---|
| 4B parsing teacher | 96.9745 | 0.0308 | 98.4102 | 95.5932 | 97.5050 | 0.1228 |
| 0.8B CPT | 96.4739 | 0.0332 | 98.3444 | 94.3973 | 96.6835 | 0.1221 |
| 0.8B MOPD | 96.7417 | 0.0314 | 98.4677 | 94.8975 | 97.1579 | 0.1223 |
| 0.8B Mix-RL | 96.8277 | 0.0313 | 98.5044 | 95.1087 | 97.1929 | 0.1221 |
(b) OCR-centric understanding: five OCR-VQA benchmarks
| Model | Overall ↑ | DocVQA ↑ | InfoVQA ↑ | ChartQA ↑ | OCRBench ↑ | TextVQA ↑ |
|---|---|---|---|---|---|---|
| 4B VQA teacher | 88.1 | 95.9 | 84.4 | 87.1 | 88.3 | 84.8 |
| 0.8B CPT | 78.1 | 92.2 | 72.5 | 83.4 | 80.6 | 62.0 |
| 0.8B MOPD | 82.7 | 93.1 | 74.8 | 84.2 | 83.9 | 77.5 |
| 0.8B Mix-RL | 83.2 | 93.1 | 75.1 | 84.6 | 84.6 | 78.6 |
Scaling from 0.8B to 4B brings a clearly larger gain on understanding (about 4.9 points, 88.1 vs 83.2) than on parsing (about 0.15 points, 96.9745 vs 96.8277). Within each task, the teacher and the student use the same amount of task data; the difference is that each teacher specializes in a single domain while the student trains on both. VQA Overall is the arithmetic average of the five benchmarks on a 0–100 scale.
3. MOPD converges fast, while longer-trained Mix-RL reaches a higher peak
For more detail, including the full analysis behind these findings, see the technical report.
🚀 Usage
Install the verified combination with uv, then run whole-page parsing on any document image:
uv venv --python 3.13
source .venv/bin/activate # Windows: .venv\Scripts\activate
uv pip install "transformers==5.17.0" "torch==2.14.0" "torchvision==0.29.0" pillow
import torch
from huggingface_hub import hf_hub_download
from transformers import AutoModelForImageTextToText, AutoProcessor
MODEL = "SeerRay-Lab/Xiaomi-OCR-0" # Or a local model directory
# Choose "qianziwen.jpeg" or "dense-equations.jpg".
IMAGE = hf_hub_download(
repo_id="SeerRay-Lab/Xiaomi-OCR-0",
filename="example_pics/qianziwen.jpeg",
)
PROMPT = (
"Extract all information from the main body of the document image and represent it "
"in markdown format, ignoring headers and footers. Tables should be expressed in OTSL "
"format, formulas in the document should be represented using LATEX format, and the "
"parsing should be organized according to the reading order."
)
device = (
"cuda" if torch.cuda.is_available()
else "mps" if torch.backends.mps.is_available()
else "cpu"
)
dtype = torch.float32 if device == "cpu" else torch.float16
processor = AutoProcessor.from_pretrained(MODEL)
model = AutoModelForImageTextToText.from_pretrained(
MODEL, dtype=dtype,
).to(device).eval()
messages = [{
"role": "user",
"content": [
{"type": "image", "url": IMAGE},
{"type": "text", "text": PROMPT},
],
}]
inputs = processor.apply_chat_template(
messages, tokenize=True, add_generation_prompt=True,
return_dict=True, return_tensors="pt",
)
inputs = inputs.to(device)
with torch.inference_mode():
output = model.generate(**inputs, max_new_tokens=4096, do_sample=False)
new_tokens = output[:, inputs["input_ids"].shape[1]:]
print(processor.batch_decode(new_tokens, skip_special_tokens=True)[0])
Task prompts
| Task | Prompt | Output |
|---|---|---|
| Document | Extract all information from the main body of the document image and represent it in markdown format, ignoring headers and footers. Tables should be expressed in OTSL format, formulas in the document should be represented using LATEX format, and the parsing should be organized according to the reading order. |
Markdown; tables as OTSL, formulas as LaTeX, in reading order |
| Text region | Extract the text in the image. |
Plain text |
| Table region | Parse the table in the image into OTSL. |
Table output in OTSL format |
| Formula region | Identify the formula in the image and represent it using LATEX format. |
LaTeX |
| Key information extraction | Extract key information in the image, followed by your JSON schema |
JSON object |
Use Examples
The examples below cover whole-page parsing and one KIE case. Inputs and outputs are in example_pics/; the two clips below are screen recordings of the model parsing each document.
Document parsing
Dense mathematical page · dense-equations.jpg dense-equations.gif
Parsed output
We then define the following variables: \(\mathbf{P}_i^j = \prod_{t=i}^j (\mathbf{I} - \beta_t \boldsymbol{k}_t \boldsymbol{k}_t^\top) \in \mathbb{R}^{d \times d}\), \(\mathbf{H}_i^j = \sum_{t=i}^j \beta_t (\boldsymbol{v}_t \boldsymbol{k}_t^\top) \mathbf{P}_{t+1}^j \in \mathbb{R}^{d \times d}\), where we let \(\mathbf{P}_i^j = \mathbf{I}\) whenever \(i > j\). Intuitively, \(\mathbf{P}_i^j\) is the "decay factor" to be applied to \(\mathbf{S}_i\) for obtaining \(\mathbf{S}_j\), and \(\mathbf{H}_i^j\) represents the contributions to \(\mathbf{S}_j\) starting from token \(i\). (Hence \(\mathbf{S}_t = \mathbf{H}_1^t\)). The chunkwise recurrence can then be written as,
\[\mathbf {S} _ {[ t ]} ^ {r} = \mathbf {S} _ {[ t ]} ^ {0} \mathbf {P} _ {[ t ]} ^ {r} + \mathbf {H} _ {[ t ]} ^ {r} \tag{5}\]
where we define the chunkwise variables \(\mathbf{S}_{[t]}^{i} = \mathbf{S}_{tC + i}\), \(\mathbf{P}_{[t]}^{r} = \mathbf{P}_{tC + 1}^{tC + r}\), \(\mathbf{H}_{[t]}^{r} = \mathbf{H}_{tC + 1}^{tC + r}\). Here we have \(\frac{L}{C}\) chunks of size \(C\). The trick is to now efficiently represent the \(\mathbf{P}_{[t]}^{r}, \mathbf{H}_{[t]}^{r} \in \mathbb{R}^{d \times d}\) matrices using a similar approach described in §3.1, so that these matrices can be stored in \(\mathcal{O}(d)\) memory,
\[\mathbf {P} _ {[ t ]} ^ {r} = \mathbf {I} - \sum_ {i = 1} ^ {r} \boldsymbol {w} _ {[ t ]} ^ {i} \boldsymbol {k} _ {[ t ]} ^ {i ^ {\top}}, \quad \mathbf {H} _ {[ t ]} ^ {r} = \sum_ {i = 1} ^ {r} \boldsymbol {u} _ {[ t ]} ^ {i} \boldsymbol {k} _ {[ t ]} ^ {i ^ {\top}} \quad \in \mathbb {R} ^ {d \times d} \tag{6}\]
\[\boldsymbol {w} _ {[ t ]} ^ {r} = \beta_ {[ t ]} ^ {r} \left(\boldsymbol {k} _ {[ t ]} ^ {r} - \sum_ {i = 1} ^ {r - 1} \boldsymbol {w} _ {[ t ]} ^ {i} (\boldsymbol {k} _ {[ t ]} ^ {i ^ {\top}} \boldsymbol {k} _ {[ t ]} ^ {r})\right), \quad \boldsymbol {u} _ {[ t ]} ^ {r} = \beta_ {[ t ]} ^ {r} \left(\boldsymbol {v} _ {[ t ]} ^ {r} - \sum_ {i = 1} ^ {r - 1} \boldsymbol {u} _ {[ t ]} ^ {i} (\boldsymbol {k} _ {[ t ]} ^ {i ^ {\top}} \boldsymbol {k} _ {[ t ]} ^ {r})\right) \quad \in \mathbb {R} ^ {d} \tag{7}\]
The derivations for the above can be found in the appendix. Subsequently, based on Eq. 5, we can obtain the chunk-level recurrence for hidden states and outputs as,
\[\mathbf {S} _ {[ t ]} ^ {r} = \mathbf {S} _ {[ t ]} ^ {0} - \left(\mathbf {S} _ {[ t ]} ^ {0} \sum_ {i = 1} ^ {r} \boldsymbol {w} _ {[ t ]} ^ {i} \boldsymbol {k} _ {[ t ]} ^ {i ^ {\top}}\right) + \sum_ {i = 1} ^ {r} \boldsymbol {u} _ {[ t ]} ^ {i} \boldsymbol {k} _ {[ t ]} ^ {i ^ {\top}} = \mathbf {S} _ {[ t ]} ^ {0} + \sum_ {i = 1} ^ {r} \left(\boldsymbol {u} _ {[ t ]} ^ {i} - \mathbf {S} _ {[ t ]} ^ {0} \boldsymbol {w} _ {[ t ]} ^ {i}\right) \boldsymbol {k} _ {[ t ]} ^ {i ^ {\top}},\]
\[\boldsymbol {o} _ {[ t ]} ^ {r} = \mathbf {S} _ {[ t ]} ^ {r} \boldsymbol {q} _ {[ t ]} ^ {r} = \mathbf {S} _ {[ t ]} ^ {0} \boldsymbol {q} _ {[ t ]} ^ {r} + \sum_ {i = 1} ^ {r} \left(\boldsymbol {u} _ {[ t ]} ^ {i} - \mathbf {S} _ {[ t ]} ^ {0} \boldsymbol {w} _ {[ t ]} ^ {i}\right) \left(\boldsymbol {k} _ {[ t ]} ^ {i ^ {\top}} \boldsymbol {q} _ {[ t ]} ^ {i}\right).\]
Letting \(\mathbf{S}_{[t]} = \mathbf{S}_{[t]}^{0}\), the above can be simplified to matrix notations similarly to Eq.1-2,
\[\mathbf {S} _ {[ t + 1 ]} = \mathbf {S} _ {[ t ]} + \left(\mathbf {U} _ {[ t ]} - \mathbf {W} _ {[ t ]} \mathbf {S} _ {[ t ]} ^ {\top}\right) ^ {\top} \mathbf {K} _ {[ t ]}, \tag{8}\]
\[\mathbf {O} _ {[ t ]} = \mathbf {Q} _ {[ t ]} \mathbf {S} _ {[ t ]} ^ {\top} + (\mathbf {Q} _ {[ t ]} \mathbf {K} _ {[ t ]} ^ {\top} \odot \mathbf {M}) \left(\mathbf {U} _ {[ t ]} - \mathbf {W} _ {[ t ]} \mathbf {S} _ {[ t ]} ^ {\top}\right) \tag{9}\]
where \(\square_{[t]} = \square_{[t]}^{1:C} \in \mathbb{R}^{C \times d}\) for \(\square \in \{\mathbf{Q}, \mathbf{K}, \mathbf{V}, \mathbf{O}, \mathbf{U}, \mathbf{W}\}\) defines the chunkwise matrices that are formed from stacking the \(q_t, k_t, v_t, o_t, u_t, w_t\) vectors.
Practical considerations. In the above, Eq. 7 is fully recurrent and thus cannot use tensor cores written as is. To solve this, we further leverage the UT transform [44, 23] (see §B.2 for derivations):
\[\mathbf {T} _ {[ t ]} = \left(\mathbf {I} + \operatorname{tril} (\operatorname{diag} (\beta_ {[ t ]}) \mathbf {K} _ {[ t ]} \mathbf {K} _ {[ t ]} ^ {\top}, - 1)\right) ^ {- 1} \operatorname{diag} \left(\beta_ {[ t ]}\right) \tag{10}\]
\[\mathbf {W} _ {[ t ]} = \mathbf {T} _ {[ t ]} \mathbf {K} _ {[ t ]}, \quad \mathbf {U} _ {[ t ]} = \mathbf {T} _ {[ t ]} \mathbf {V} _ {[ t ]} \tag{11}
Handwritten derivation · handwrite-formula.jpg handwrite-formula.gif
Parsed output
(b) we get Friedmann equation:
\[
\left(\frac{\dot{a}}{a}\right)^{2} = \frac{8\pi G}{3c^{2}}\rho - \frac{kc^{2}}{a^{2}}
\]
and Raychaudhuri equation:
\[
2\frac{\ddot{a}}{a} + \left(\frac{\dot{a}}{a}\right)^{2} + \frac{kc^{2}}{a^{2}} = -\frac{8\pi G P}{c^{2}}
\]
Substituting the Friedmann equation and \(P = w\rho\) into the Raychaudhuri equation:
\[
2\frac{\ddot{a}}{a} + \frac{8\pi G}{3c^{2}}\rho - \frac{kc^{2}}{a^{2}} + \frac{kc^{2}}{a^{2}} = -\frac{8\pi G}{c^{2}}\rho
\]
\[
\frac{\ddot{a}}{a} + \frac{4\pi G}{3c^{2}}\rho = -\frac{4\pi G}{c^{2}}\rho
\]
\[
\frac{\ddot{a}}{a} = -\frac{4\pi G}{3c^{2}}(\rho + 3P) = -\frac{4\pi G}{3c^{2}}(1 + 3w)\rho
\]
Differentiating the Friedmann equation with respect to time; We get:
\[
\begin{array}{r l}
{\left[\left(\frac{\dot{a}}{a}\right)^{2}\right]^{\prime}} & = 2\left(\frac{\dot{a}}{a}\right)\left(\frac{\dot{a}}{a}\right)^{\prime} \\
& = 2\frac{\dot{a}}{a}\left(\frac{\ddot{a}}{a} + \left(-\frac{\dot{a}}{a^{2}}\right)\dot{a}\right) \\
& = \frac{2\dot{a}\ddot{a}}{a^{2}} + (-2)\frac{\dot{a}^{3}}{a^{3}}
\end{array}
\]
\[
\left[\left(\frac{\dot{a}}{a}\right)^{2}\right]^{\prime} = \frac{8\pi G}{3c^{2}}\dot{\rho} - kc^{2}\left(\frac{1}{a^{2}}\right)^{\prime}
\]
\[
\frac{2\dot{a}\ddot{a}}{a^{2}} + \frac{\dot{a}^{3}}{a^{3}}(-2) = \frac{8\pi G}{3c^{2}}\dot{\rho} + \frac{2kc^{2}}{a^{3}}\dot{a}
\]
\[
\frac{\dot{a}\ddot{a}}{a^{2}} - \frac{\dot{a}^{3}}{a^{3}} = \frac{4\pi G}{3c^{2}}\dot{\rho} + \frac{kc^{2}}{a^{3}}\dot{a}
\]
substituting \(\frac{\ddot{a}}{a} = -\frac{4\pi G}{3c^{2}}(1 + 3w)\rho\)
\[
\frac{\dot{a}}{a}\left(-\frac{4\pi G}{3c^{2}}\right)(1 + 3w)\rho = \frac{4\pi G}{3c^{2}}\dot{\rho} + \frac{\dot{a}}{a}\left(\frac{kc^{2}}{a^{2}} + \frac{\dot{a}^{2}}{a^{2}}\right)
\]
and Friedmann equation into it :
\[
\frac{\dot{a}}{a}\left(-\frac{4\pi G}{3c^{2}}\right)(1 + 3w)\rho = \frac{4\pi G}{3c^{2}}\dot{\rho} + \frac{\dot{a}}{a}\cdot\frac{4\pi G}{3c^{2}}\cdot 2\cdot\rho
\]
\[
\begin{array}{l}
(-1)\dot{a}(1 + 3w)\rho = a\dot{\rho} + 2\rho\dot{a} \\
-(3 + 3w)\rho\dot{a} = a\dot{\rho}
\end{array}
\]
\[
\therefore \quad \frac{\dot{\rho}}{\rho} = -3(1 + w)\frac{\dot{a}}{a}.
\]
Running-script Thousand Character Classic · qianziwen.jpeg
行书千文
梁貟外散騎侍郎周興嗣次韻
天地玄黄宇宙洪荒日月
Key information extraction
Receipt · receipt.jpg
Extract key information in the image
Please output the key information in JSON format according to the following schema:
{
"date": "",
"seller_address": "",
"seller_gst_id": "",
"seller_name": "",
"total_amount": "",
"total_tax": ""
}
{
"date": "2017-10-18",
"seller_address": "LOT 1851-A & 1851-B, JALAN KPB 6, KAWASAN PERINDUSTRIAN BALAKONG, 43300 SERI KEMBANGAN, SELANGOR",
"seller_gst_id": "001092886528",
"seller_name": "MR. D.I.Y. SDN BHD",
"total_amount": "5.90",
"total_tax": "0.33"
}
More code
Batch processing, two-stage parsing (PP-DocLayoutV3 → region recognition → page assembly), the Agent Skill, and the evaluation and post-processing code are all in the GitHub repository.
Acknowledgments
Xiaomi-OCR-0 builds on the open OCR and multimodal LLM community. We thank everyone in the community for their tireless effort and outstanding contributions along the way, and especially the authors of Qwen3.5, PaddleOCR-VL-1.6, MinerU2.5-Pro, GLM-OCR, dots.mocr, HunyuanOCR-1.5, TeleOCR, OvisOCR2, Qianfan-OCR, and PP-DocLayoutV3 for the inspiration and contributions that shaped Xiaomi-OCR-0.
Citation and License
- License: Apache-2.0.
- Paper: Xiaomi-OCR-0 Technical Report (arXiv:2609.36136)
@misc{chen2026xiaomiocr0,
title = {Xiaomi-OCR-0 Technical Report},
author = {Chen, Xin and Du, Anan and Feng, Feng and Fu, Pei and Luan, Jian and
Xu, Longwei and Zhang, Shaojie and Li, Hang and Qu, Heng and Tan, Cheng},
year = {2026},
eprint = {2609.36136},
archivePrefix= {arXiv},
primaryClass = {cs.CV},
url = {https://arxiv.org/abs/2609.36136}
}
- Downloads last month
- -
Model tree for SeerRay-Lab/Xiaomi-OCR-0
Base model
Qwen/Qwen3.5-0.8B-Base







