--- library_name: transformers pipeline_tag: image-to-text tags: - chemistry - image-to-text - ocsr - markush - e-smiles2.0 datasets: - UniParser/MolParser-7M - UniParser/MolGallery license: cc-by-nc-sa-4.0 --- # MolParser Mobile V2
💻 GitHub | 📘 E-SMILES 2.0 Spec | 📄 Report | 🚀 Demo
**MolParser-Mobile-V2** is a lightweight Optical Chemical Structure Recognition (OCSR) model that converts molecular structure images directly into **E-SMILES 2.0**. It upgrades MolParser-Mobile for broader recognition of structures found in chemical literature, especially complex Markush structures, while retaining a compact 10M parameter architecture. ## 🚀 What's New * **E-SMILES 2.0 output** with substantially broader coverage of literature molecules and Markush structures. * **Richer Markush type coverage** for literature molecules, including atom- and ring-indexed substituents, explicit dummy attachments, nested substructures, structural repeating units and polymers, virtual arcs, colored endpoint balls, and axial-chirality annotations. * **384 × 384 input resolution**, increased from 224 × 224 in MolParser-Mobile. * **384-token maximum output length**, increased from 256 tokens. * **Improved recognition accuracy**, particularly for complex and stereochemical structures. For notation details, examples, validation, normalization, substitution, and rendering utilities, see the [MolParser Repo](https://github.com/dptech-corp/MolParser) and the [E-SMILES specification](https://github.com/dptech-corp/MolParser/blob/main/skills/molparser-extended-smiles/extended-smiles-spec.md). ## 📊 Performance Accuracy for MolParser-Mobile-V2 was measured with FP16 inference, greedy decoding, and batch size 512. Deltas are relative to MolParser-Mobile. | Model | Parameters | Throughput (RTX 4090D) | Uni-Parser Bench | BioVista | WildMol-10k | USPTO | | ------------------------ | ---------: | ---------------------: | -------------------: | -------------------: | -------------------: | ------------------: | | MolParser-Mobile | 9.98M | 1,520 Mol/s | 0.823 | 0.801 | 0.734 | 0.836 | | **MolParser-Mobile-V2** | 10.00M | 1,296 Mol/s | **0.850** (+0.027) | **0.820** (+0.019) | **0.762** (+0.028) | **0.909** (+0.073) | ## ⚡ Usage ### Option 1. MolParser Library (Recommended) The [MolParser library](https://github.com/dptech-corp/MolParser) provides a convenient interface for molecule detection, recognition, E-SMILES 2.0 post-processing, and rendering. Clone the repository and install the package: ```bash git clone https://github.com/dptech-corp/MolParser.git cd MolParser pip install -e . ``` Then run: ```python from molparser import MolParser parser = MolParser(molparser_hf_repo="UniParser/MolParser-Mobile-V2", max_length=384) result = parser.parse("mol.png", rec_only=True) ``` To render the predicted E-SMILES as SVG or PNG, see [Render E-SMILES](https://github.com/dptech-corp/MolParser/tree/main#render-e-smiles): ```python from pathlib import Path from molparser import utils as mutils raw = "*C(O)c1cc(C(=O)N(*)*)cc(-c2*ccc*2)c1