Image-to-Text
Transformers
Safetensors
molparser_vision_encoder_decoder
image-text-to-text
chemistry
ocsr
markush
e-smiles2.0
custom_code
Instructions to use UniParser/MolParser-Mobile-V2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use UniParser/MolParser-Mobile-V2 with Transformers:
# Use a pipeline as a high-level helper # Warning: Pipeline type "image-to-text" is no longer supported in transformers v5. # You must load the model directly (see below) or downgrade to v4.x with: # pip install "transformers<5.0.0" from transformers import pipeline pipe = pipeline("image-to-text", model="UniParser/MolParser-Mobile-V2", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForImageTextToText model = AutoModelForImageTextToText.from_pretrained("UniParser/MolParser-Mobile-V2", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
|
Download README.md from UniParser/MolParser-Mobile-V2: direct link, hf CLI and curl.
- Browser
- Download file 6.28 kB
-
https://huggingface.co/UniParser/MolParser-Mobile-V2/resolve/main/README.md
- Command line
-
hf download hf://UniParser/MolParser-Mobile-V2/README.md
-
curl -L -o README.md https://huggingface.co/UniParser/MolParser-Mobile-V2/resolve/main/README.md
6.28 kB
| library_name: transformers | |
| pipeline_tag: image-to-text | |
| tags: | |
| - chemistry | |
| - image-to-text | |
| - ocsr | |
| - markush | |
| - e-smiles2.0 | |
| datasets: | |
| - UniParser/MolParser-7M | |
| - UniParser/MolGallery | |
| license: cc-by-nc-sa-4.0 | |
| # MolParser Mobile V2 | |
| <p align="center"> | |
| π» <a href="https://github.com/dptech-corp/MolParser">GitHub</a> | | |
| π <a href="https://github.com/dptech-corp/MolParser/blob/main/skills/molparser-extended-smiles/extended-smiles-spec.md">E-SMILES 2.0 Spec</a> | | |
| π <a href="https://arxiv.org/abs/2609.05807">Report</a> | | |
| π <a href="https://huggingface.co/spaces/AI4Industry/molparser-mobile">Demo</a> | |
| </p> | |
| **MolParser-Mobile-V2** is a lightweight Optical Chemical Structure Recognition (OCSR) model that converts molecular structure images directly into **E-SMILES 2.0**. It upgrades MolParser-Mobile for broader recognition of structures found in chemical literature, especially complex Markush structures, while retaining a compact 10M parameter architecture. | |
| ## π What's New | |
| * **E-SMILES 2.0 output** with substantially broader coverage of literature molecules and Markush structures. | |
| * **Richer Markush type coverage** for literature molecules, including atom- and ring-indexed substituents, explicit dummy attachments, nested substructures, structural repeating units and polymers, virtual arcs, colored endpoint balls, and axial-chirality annotations. | |
| * **384 Γ 384 input resolution**, increased from 224 Γ 224 in MolParser-Mobile. | |
| * **384-token maximum output length**, increased from 256 tokens. | |
| * **Improved recognition accuracy**, particularly for complex and stereochemical structures. | |
| For notation details, examples, validation, normalization, substitution, and rendering utilities, see the [MolParser Repo](https://github.com/dptech-corp/MolParser) and the [E-SMILES specification](https://github.com/dptech-corp/MolParser/blob/main/skills/molparser-extended-smiles/extended-smiles-spec.md). | |
| ## π Performance | |
| Accuracy for MolParser-Mobile-V2 was measured with FP16 inference, greedy decoding, and batch size 512. Deltas are relative to MolParser-Mobile. | |
| | Model | Parameters | Throughput (RTX 4090D) | Uni-Parser Bench | BioVista | WildMol-10k | USPTO | | |
| | ------------------------ | ---------: | ---------------------: | -------------------: | -------------------: | -------------------: | ------------------: | | |
| | MolParser-Mobile | 9.98M | 1,520 Mol/s | 0.823 | 0.801 | 0.734 | 0.836 | | |
| | **MolParser-Mobile-V2** | 10.00M | 1,296 Mol/s | **0.850** (+0.027) | **0.820** (+0.019) | **0.762** (+0.028) | **0.909** (+0.073) | | |
| ## β‘ Usage | |
| ### Option 1. MolParser Library (Recommended) | |
| The [MolParser library](https://github.com/dptech-corp/MolParser) provides a convenient interface for molecule detection, recognition, E-SMILES 2.0 post-processing, and rendering. | |
| Clone the repository and install the package: | |
| ```bash | |
| git clone https://github.com/dptech-corp/MolParser.git | |
| cd MolParser | |
| pip install -e . | |
| ``` | |
| Then run: | |
| ```python | |
| from molparser import MolParser | |
| parser = MolParser(molparser_hf_repo="UniParser/MolParser-Mobile-V2", max_length=384) | |
| result = parser.parse("mol.png", rec_only=True) | |
| ``` | |
| To render the predicted E-SMILES as SVG or PNG, see [Render E-SMILES](https://github.com/dptech-corp/MolParser/tree/main#render-e-smiles): | |
| ```python | |
| from pathlib import Path | |
| from molparser import utils as mutils | |
| raw = "*C(O)c1cc(C(=O)N(*)*)cc(-c2*ccc*2)c1<sep><a>0:CF3</a><a>9:R[3]</a><a>10:R[2]</a><a>14:X</a><a>18:Y</a><r>1:R[1]?1-3</r>" | |
| svg_text = mutils.draw(raw, output_format="svg") | |
| Path("molecule.svg").write_text(svg_text, encoding="utf-8") | |
| png_bytes = mutils.draw(raw, output_format="png") | |
| Path("molecule.png").write_bytes(png_bytes) | |
| ``` | |
| ### Option 2. π€ Transformers | |
| Load MolParser-Mobile-V2 directly with the Hugging Face transformers library. | |
| ```python | |
| import torch | |
| from PIL import Image | |
| from transformers import AutoModelForImageTextToText, AutoProcessor | |
| repo_id = "UniParser/MolParser-Mobile-V2" | |
| device = "cuda" if torch.cuda.is_available() else "cpu" | |
| dtype = torch.float16 if device == "cuda" else torch.float32 | |
| processor = AutoProcessor.from_pretrained(repo_id, trust_remote_code=True) | |
| model = AutoModelForImageTextToText.from_pretrained( | |
| repo_id, | |
| dtype=dtype, | |
| trust_remote_code=True, | |
| ).to(device).eval() | |
| image = Image.open("mol.png").convert("RGB") | |
| inputs = processor(images=image, return_tensors="pt") | |
| inputs = {k: v.to(device, dtype=dtype) for k, v in inputs.items()} | |
| output_ids = model.generate(**inputs, max_length=384, num_beams=1, do_sample=False) | |
| caption = processor.batch_decode(output_ids, skip_special_tokens=True)[0] | |
| print(caption) | |
| ``` | |
| ## π License | |
| ### MolParser-Mobile-V2 Weight | |
| The **MolParser-Mobile-V2 model weights** are provided for **non-commercial use only** under CC BY-NC-SA 4.0. | |
| For commercial licensing, please contact **fangxi@dp.tech** or open a discussion on Hugging Face. | |
| ### MolParser Github Repo | |
| The **MolParser** library (including E-SMILES post-processing and rendering) is available at https://github.com/dptech-corp/MolParser and is licensed under the **Apache License 2.0**, which permits commercial use, modification, and distribution, provided that the license and copyright notices are retained. | |
| **Note:** Model weights, datasets, and third-party dependencies are subject to their respective licenses. | |
| ## π Citation | |
| If you use this model, please cite: | |
| ``` | |
| @article{fang2026molparserm, | |
| title={MolParser-Mobile: Ultrafast OCSR System for Large-Scale Chemical Literature Mining}, | |
| author={Fang, Xi and Lu, Haocheng and Lyu, Han and Luo, Chengxiang and Zhang, Linfeng and Ke, Guolin}, | |
| journal={arXiv preprint arXiv:2609.05807}, | |
| year={2026} | |
| } | |
| ``` | |
| ``` | |
| @inproceedings{fang2025molparser, | |
| title={Molparser: End-to-end visual recognition of molecule structures in the wild}, | |
| author={Fang, Xi and Wang, Jiankun and Cai, Xiaochen and Chen, Shangqian and Yang, Shuwen and Tao, Haoyi and Wang, Nan and Yao, Lin and Zhang, Linfeng and Ke, Guolin}, | |
| booktitle={Proceedings of the IEEE/CVF International Conference on Computer Vision}, | |
| pages={24528--24538}, | |
| year={2025} | |
| } | |
| ``` | |