obj_v1 / README.md
objectai's picture
Update README.md
60dba78 verified
|
Raw History Blame Contribute Delete
4.03 kB
---
license: apache-2.0
base_model: Qwen/Qwen3-VL-4B-Instruct
pipeline_tag: image-text-to-text
tags:
- document-understanding
- information-extraction
- vision-language
- qwen3-vl
- structured-data-extraction
- multimodal
- document-ai
library_name: transformers
language:
- en
---
# obj_v1
Vision-language model fine-tuned for structured data extraction from Indian
financial documents. Give it a page image and a JSON schema; it returns the
schema filled in from what is on the page.
A 4B vision-language model, LoRA fine-tuned and merged
## Authors
<p align="left">
<!-- <a href="https://www.linkedin.com/in/ahmedzaweel/">
<img src="https://img.shields.io/badge/LinkedIn-Ahmed%20Zaweel-0A66C2?style=for-the-badge&logo=linkedin&logoColor=white" alt="Ahmed Zaweel on LinkedIn" />
</a> -->
&nbsp;&nbsp;
<a href="https://www.linkedin.com/in/rachit-kumar-b41299228/">
<img src="https://img.shields.io/badge/LinkedIn-Rachit%20Kumar-0A66C2?style=for-the-badge&logo=linkedin&logoColor=white" alt="Rachit Kumar on LinkedIn" />
</a>
&nbsp;&nbsp;
<a href="https://www.linkedin.com/in/ahmedzaweel/">
<img src="https://img.shields.io/badge/LinkedIn-Ahmed%20Zaweel-0A66C2?style=for-the-badge&logo=linkedin&logoColor=white" alt="Ahmed Zaweel on LinkedIn" />
</a>
</p>
## Serving with vLLM
```bash
vllm serve objectai/obj_v1 \
--served-model-name obj_v1 \
--max-model-len 16384 \
--limit-mm-per-prompt '{"image":1}' \
--mm-processor-kwargs '{"max_pixels":1003520}' \
--trust-remote-code
```
`max_pixels` is 1280x28x28, the resolution the model was trained at. Raising it
wastes KV cache; lowering it makes small print unreadable.
## Calling it
The server is OpenAI-compatible, so an ordinary chat completion works:
```python
import base64, json, openai
client = openai.OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
image = base64.b64encode(open("cheque.jpg", "rb").read()).decode()
schema = {"cheque_details": {"amount": "number", "payee": "string",
"date": "string", "cheque_number": "string"}}
response = client.chat.completions.create(
model="obj_v1",
temperature=0.0,
max_tokens=8192,
messages=[
{"role": "system", "content":
"You are a document data extraction model. "
"Extract only values present in the document. "
"Use null for fields that are absent or illegible. "
"Output a single compact JSON object matching the requested schema. "
"No prose, no markdown, no explanation."},
{"role": "user", "content": [
{"type": "image_url",
"image_url": {"url": f"data:image/jpeg;base64,{image}"}},
{"type": "text",
"text": f"document_type: cheque\nschema: {json.dumps(schema)}"},
]},
],
)
print(response.choices[0].message.content)
```
## Prompt format
Match training or accuracy drops. The system prompt above is verbatim, and the
user turn is the image followed by exactly two lines:
```
document_type: <type>
schema: <compact json>
```
Set `temperature=0.0` so the same page yields the same answer.
## Requirements
| | |
| --- | --- |
| Weights | 8.9 GB (bf16) |
| VRAM | 16 GB minimum, 24 GB comfortable |
| Precision | bf16 (Ampere or newer; use fp16 below that) |
| Context | 16384 covers the longest documents |
Runs on an L4, A10G, L40S, A100 or RTX 4090. On a T4 add `--dtype float16`.
## Output
Compact JSON matching the requested schema. Fields absent from the page come
back `null` rather than guessed. Values found on the page that the schema did
not ask for are placed under `extras` when that key is included in the schema.
## Limitations
- Trained on Indian financial documents; other domains and layouts are untested.
- Handwriting is the weakest case, particularly digits at low resolution.
- The model does not verify its own arithmetic. Totals that must reconcile
should be checked by the caller.
## License
Apache 2.0. Fine-tuned from Qwen3-VL-4B-Instruct, which is Apache 2.0.