Image-to-Text
Transformers
Safetensors
molparser_vision_encoder_decoder
image-text-to-text
chemistry
ocsr
markush
e-smiles2.0
custom_code
Instructions to use UniParser/MolParser-Mobile-V2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use UniParser/MolParser-Mobile-V2 with Transformers:
# Use a pipeline as a high-level helper # Warning: Pipeline type "image-to-text" is no longer supported in transformers v5. # You must load the model directly (see below) or downgrade to v4.x with: # pip install "transformers<5.0.0" from transformers import pipeline pipe = pipeline("image-to-text", model="UniParser/MolParser-Mobile-V2", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForImageTextToText model = AutoModelForImageTextToText.from_pretrained("UniParser/MolParser-Mobile-V2", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
File size: 6,284 Bytes
67ee3f1 11d114d 67ee3f1 2c3d3b2 67ee3f1 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 | ---
library_name: transformers
pipeline_tag: image-to-text
tags:
- chemistry
- image-to-text
- ocsr
- markush
- e-smiles2.0
datasets:
- UniParser/MolParser-7M
- UniParser/MolGallery
license: cc-by-nc-sa-4.0
---
# MolParser Mobile V2
<p align="center">
π» <a href="https://github.com/dptech-corp/MolParser">GitHub</a> |
π <a href="https://github.com/dptech-corp/MolParser/blob/main/skills/molparser-extended-smiles/extended-smiles-spec.md">E-SMILES 2.0 Spec</a> |
π <a href="https://arxiv.org/abs/2609.05807">Report</a> |
π <a href="https://huggingface.co/spaces/AI4Industry/molparser-mobile">Demo</a>
</p>
**MolParser-Mobile-V2** is a lightweight Optical Chemical Structure Recognition (OCSR) model that converts molecular structure images directly into **E-SMILES 2.0**. It upgrades MolParser-Mobile for broader recognition of structures found in chemical literature, especially complex Markush structures, while retaining a compact 10M parameter architecture.
## π What's New
* **E-SMILES 2.0 output** with substantially broader coverage of literature molecules and Markush structures.
* **Richer Markush type coverage** for literature molecules, including atom- and ring-indexed substituents, explicit dummy attachments, nested substructures, structural repeating units and polymers, virtual arcs, colored endpoint balls, and axial-chirality annotations.
* **384 Γ 384 input resolution**, increased from 224 Γ 224 in MolParser-Mobile.
* **384-token maximum output length**, increased from 256 tokens.
* **Improved recognition accuracy**, particularly for complex and stereochemical structures.
For notation details, examples, validation, normalization, substitution, and rendering utilities, see the [MolParser Repo](https://github.com/dptech-corp/MolParser) and the [E-SMILES specification](https://github.com/dptech-corp/MolParser/blob/main/skills/molparser-extended-smiles/extended-smiles-spec.md).
## π Performance
Accuracy for MolParser-Mobile-V2 was measured with FP16 inference, greedy decoding, and batch size 512. Deltas are relative to MolParser-Mobile.
| Model | Parameters | Throughput (RTX 4090D) | Uni-Parser Bench | BioVista | WildMol-10k | USPTO |
| ------------------------ | ---------: | ---------------------: | -------------------: | -------------------: | -------------------: | ------------------: |
| MolParser-Mobile | 9.98M | 1,520 Mol/s | 0.823 | 0.801 | 0.734 | 0.836 |
| **MolParser-Mobile-V2** | 10.00M | 1,296 Mol/s | **0.850** (+0.027) | **0.820** (+0.019) | **0.762** (+0.028) | **0.909** (+0.073) |
## β‘ Usage
### Option 1. MolParser Library (Recommended)
The [MolParser library](https://github.com/dptech-corp/MolParser) provides a convenient interface for molecule detection, recognition, E-SMILES 2.0 post-processing, and rendering.
Clone the repository and install the package:
```bash
git clone https://github.com/dptech-corp/MolParser.git
cd MolParser
pip install -e .
```
Then run:
```python
from molparser import MolParser
parser = MolParser(molparser_hf_repo="UniParser/MolParser-Mobile-V2", max_length=384)
result = parser.parse("mol.png", rec_only=True)
```
To render the predicted E-SMILES as SVG or PNG, see [Render E-SMILES](https://github.com/dptech-corp/MolParser/tree/main#render-e-smiles):
```python
from pathlib import Path
from molparser import utils as mutils
raw = "*C(O)c1cc(C(=O)N(*)*)cc(-c2*ccc*2)c1<sep><a>0:CF3</a><a>9:R[3]</a><a>10:R[2]</a><a>14:X</a><a>18:Y</a><r>1:R[1]?1-3</r>"
svg_text = mutils.draw(raw, output_format="svg")
Path("molecule.svg").write_text(svg_text, encoding="utf-8")
png_bytes = mutils.draw(raw, output_format="png")
Path("molecule.png").write_bytes(png_bytes)
```
### Option 2. π€ Transformers
Load MolParser-Mobile-V2 directly with the Hugging Face transformers library.
```python
import torch
from PIL import Image
from transformers import AutoModelForImageTextToText, AutoProcessor
repo_id = "UniParser/MolParser-Mobile-V2"
device = "cuda" if torch.cuda.is_available() else "cpu"
dtype = torch.float16 if device == "cuda" else torch.float32
processor = AutoProcessor.from_pretrained(repo_id, trust_remote_code=True)
model = AutoModelForImageTextToText.from_pretrained(
repo_id,
dtype=dtype,
trust_remote_code=True,
).to(device).eval()
image = Image.open("mol.png").convert("RGB")
inputs = processor(images=image, return_tensors="pt")
inputs = {k: v.to(device, dtype=dtype) for k, v in inputs.items()}
output_ids = model.generate(**inputs, max_length=384, num_beams=1, do_sample=False)
caption = processor.batch_decode(output_ids, skip_special_tokens=True)[0]
print(caption)
```
## π License
### MolParser-Mobile-V2 Weight
The **MolParser-Mobile-V2 model weights** are provided for **non-commercial use only** under CC BY-NC-SA 4.0.
For commercial licensing, please contact **fangxi@dp.tech** or open a discussion on Hugging Face.
### MolParser Github Repo
The **MolParser** library (including E-SMILES post-processing and rendering) is available at https://github.com/dptech-corp/MolParser and is licensed under the **Apache License 2.0**, which permits commercial use, modification, and distribution, provided that the license and copyright notices are retained.
**Note:** Model weights, datasets, and third-party dependencies are subject to their respective licenses.
## π Citation
If you use this model, please cite:
```
@article{fang2026molparserm,
title={MolParser-Mobile: Ultrafast OCSR System for Large-Scale Chemical Literature Mining},
author={Fang, Xi and Lu, Haocheng and Lyu, Han and Luo, Chengxiang and Zhang, Linfeng and Ke, Guolin},
journal={arXiv preprint arXiv:2609.05807},
year={2026}
}
```
```
@inproceedings{fang2025molparser,
title={Molparser: End-to-end visual recognition of molecule structures in the wild},
author={Fang, Xi and Wang, Jiankun and Cai, Xiaochen and Chen, Shangqian and Yang, Shuwen and Tao, Haoyi and Wang, Nan and Yao, Lin and Zhang, Linfeng and Ke, Guolin},
booktitle={Proceedings of the IEEE/CVF International Conference on Computer Vision},
pages={24528--24538},
year={2025}
}
```
|