AI4Industry commited on
Commit
67ee3f1
·
verified ·
1 Parent(s): 8097e82

Upload MolParser-Mobile-V2 (private)

Browse files
README.md ADDED
@@ -0,0 +1,153 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ library_name: transformers
3
+ pipeline_tag: image-to-text
4
+ tags:
5
+ - chemistry
6
+ - image-to-text
7
+ - ocsr
8
+ - markush
9
+ - e-smiles2.0
10
+ datasets:
11
+ - UniParser/MolParser-7M
12
+ - UniParser/MolGallery
13
+ license: cc-by-nc-sa-4.0
14
+ ---
15
+
16
+ # MolParser Mobile V2
17
+
18
+ <p align="center">
19
+ 💻 <a href="https://github.com/dptech-corp/MolParser">GitHub</a> |
20
+ 📘 <a href="https://github.com/dptech-corp/MolParser/blob/main/skills/molparser-extended-smiles/extended-smiles-spec.md">E-SMILES 2.0 Spec</a> |
21
+ 📄 <a href="https://arxiv.org/abs/2609.05807">Report</a> |
22
+ 🚀 <a href="https://ocsr.dp.tech/">Demo</a>
23
+ </p>
24
+
25
+ **MolParser-Mobile-V2** is a lightweight Optical Chemical Structure Recognition (OCSR) model that converts molecular structure images directly into **E-SMILES 2.0**. It upgrades MolParser-Mobile for broader recognition of structures found in chemical literature, especially complex Markush structures, while retaining a compact 10M parameter architecture.
26
+
27
+
28
+ ## 🚀 What's New
29
+
30
+ * **E-SMILES 2.0 output** with substantially broader coverage of literature molecules and Markush structures.
31
+ * **Richer Markush type coverage** for literature molecules, including atom- and ring-indexed substituents, explicit dummy attachments, nested substructures, structural repeating units and polymers, virtual arcs, colored endpoint balls, and axial-chirality annotations.
32
+ * **384 × 384 input resolution**, increased from 224 × 224 in MolParser-Mobile.
33
+ * **384-token maximum output length**, increased from 256 tokens.
34
+ * **Improved recognition accuracy**, particularly for complex and stereochemical structures.
35
+
36
+ For notation details, examples, validation, normalization, substitution, and rendering utilities, see the [MolParser Repo](https://github.com/dptech-corp/MolParser) and the [E-SMILES specification](https://github.com/dptech-corp/MolParser/blob/main/skills/molparser-extended-smiles/extended-smiles-spec.md).
37
+
38
+
39
+ ## 📊 Performance
40
+
41
+ Accuracy for MolParser-Mobile-V2 was measured with FP16 inference, greedy decoding, and batch size 512. Deltas are relative to MolParser-Mobile.
42
+
43
+ | Model | Parameters | Throughput (RTX 4090D) | Uni-Parser Bench | BioVista | WildMol-10k | USPTO |
44
+ | ------------------------ | ---------: | ---------------------: | -------------------: | -------------------: | -------------------: | ------------------: |
45
+ | MolParser-Mobile | 9.98M | 1,520 Mol/s | 0.823 | 0.801 | 0.734 | 0.836 |
46
+ | **MolParser-Mobile-V2** | 10.00M | 1,296 Mol/s | **0.850** (+0.027) | **0.820** (+0.019) | **0.762** (+0.028) | **0.909** (+0.073) |
47
+
48
+
49
+ ## ⚡ Usage
50
+
51
+ ### Option 1. MolParser Library (Recommended)
52
+
53
+ The [MolParser library](https://github.com/dptech-corp/MolParser) provides a convenient interface for molecule detection, recognition, E-SMILES 2.0 post-processing, and rendering.
54
+
55
+ Clone the repository and install the package:
56
+
57
+ ```bash
58
+ git clone https://github.com/dptech-corp/MolParser.git
59
+ cd MolParser
60
+ pip install -e .
61
+ ```
62
+
63
+ Then run:
64
+
65
+ ```python
66
+ from molparser import MolParser
67
+
68
+ parser = MolParser()
69
+ # default setting (updated to latest main): molparser_hf_repo="UniParser/MolParser-Mobile-V2", max_length=384,
70
+
71
+ result = parser.parse("mol.png", rec_only=True)
72
+ ```
73
+
74
+ To render the predicted E-SMILES as SVG or PNG, see [Render E-SMILES](https://github.com/dptech-corp/MolParser/tree/main#render-e-smiles):
75
+
76
+ ```python
77
+ from pathlib import Path
78
+ from molparser import utils as mutils
79
+
80
+ raw = "*C(O)c1cc(C(=O)N(*)*)cc(-c2*ccc*2)c1<sep><a>0:CF3</a><a>9:R[3]</a><a>10:R[2]</a><a>14:X</a><a>18:Y</a><r>1:R[1]?1-3</r>"
81
+ svg_text = mutils.draw(raw, output_format="svg")
82
+ Path("molecule.svg").write_text(svg_text, encoding="utf-8")
83
+
84
+ png_bytes = mutils.draw(raw, output_format="png")
85
+ Path("molecule.png").write_bytes(png_bytes)
86
+ ```
87
+
88
+ ### Option 2. 🤗 Transformers
89
+
90
+ Load MolParser-Mobile-V2 directly with the Hugging Face transformers library.
91
+
92
+ ```python
93
+ import torch
94
+ from PIL import Image
95
+ from transformers import AutoModelForImageTextToText, AutoProcessor
96
+
97
+ repo_id = "UniParser/MolParser-Mobile-V2"
98
+ device = "cuda" if torch.cuda.is_available() else "cpu"
99
+ dtype = torch.float16 if device == "cuda" else torch.float32
100
+
101
+ processor = AutoProcessor.from_pretrained(repo_id, trust_remote_code=True)
102
+ model = AutoModelForImageTextToText.from_pretrained(
103
+ repo_id,
104
+ dtype=dtype,
105
+ trust_remote_code=True,
106
+ ).to(device).eval()
107
+
108
+ image = Image.open("mol.png").convert("RGB")
109
+ inputs = processor(images=image, return_tensors="pt")
110
+ inputs = {k: v.to(device, dtype=dtype) for k, v in inputs.items()}
111
+
112
+ output_ids = model.generate(**inputs, max_length=384, num_beams=1, do_sample=False)
113
+ caption = processor.batch_decode(output_ids, skip_special_tokens=True)[0]
114
+ print(caption)
115
+ ```
116
+
117
+
118
+ ## 📜 License
119
+
120
+ ### MolParser-Mobile-V2 Weight
121
+
122
+ The **MolParser-Mobile-V2 model weights** are provided for **non-commercial use only** under CC BY-NC-SA 4.0.
123
+
124
+ For commercial licensing, please contact **fangxi@dp.tech** or open a discussion on Hugging Face.
125
+
126
+ ### MolParser Github Repo
127
+
128
+ The **MolParser** library (including E-SMILES post-processing and rendering) is available at https://github.com/dptech-corp/MolParser and is licensed under the **Apache License 2.0**, which permits commercial use, modification, and distribution, provided that the license and copyright notices are retained.
129
+
130
+ **Note:** Model weights, datasets, and third-party dependencies are subject to their respective licenses.
131
+
132
+ ## 📖 Citation
133
+
134
+ If you use this model, please cite:
135
+
136
+ ```
137
+ @article{fang2026molparserm,
138
+ title={MolParser-Mobile: Ultrafast OCSR System for Large-Scale Chemical Literature Mining},
139
+ author={Fang, Xi and Lu, Haocheng and Lyu, Han and Luo, Chengxiang and Zhang, Linfeng and Ke, Guolin},
140
+ journal={arXiv preprint arXiv:2609.05807},
141
+ year={2026}
142
+ }
143
+ ```
144
+
145
+ ```
146
+ @inproceedings{fang2025molparser,
147
+ title={Molparser: End-to-end visual recognition of molecule structures in the wild},
148
+ author={Fang, Xi and Wang, Jiankun and Cai, Xiaochen and Chen, Shangqian and Yang, Shuwen and Tao, Haoyi and Wang, Nan and Yao, Lin and Zhang, Linfeng and Ke, Guolin},
149
+ booktitle={Proceedings of the IEEE/CVF International Conference on Computer Vision},
150
+ pages={24528--24538},
151
+ year={2025}
152
+ }
153
+ ```
config.json ADDED
@@ -0,0 +1,101 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "architectures": [
3
+ "MolParserVisionEncoderDecoderModel"
4
+ ],
5
+ "auto_map": {
6
+ "AutoConfig": "modeling_molparser_mobile.MolParserVisionEncoderDecoderConfig",
7
+ "AutoModelForImageTextToText": "modeling_molparser_mobile.MolParserVisionEncoderDecoderModel"
8
+ },
9
+ "decoder": {
10
+ "_name_or_path": "",
11
+ "activation_dropout": 0.0,
12
+ "activation_function": "gelu",
13
+ "add_cross_attention": true,
14
+ "architectures": null,
15
+ "attention_dropout": 0.0,
16
+ "bos_token_id": 0,
17
+ "chunk_size_feed_forward": 0,
18
+ "classifier_dropout": 0.0,
19
+ "d_model": 192,
20
+ "decoder_attention_heads": 4,
21
+ "decoder_ffn_dim": 768,
22
+ "decoder_layerdrop": 0.0,
23
+ "decoder_layers": 6,
24
+ "decoder_start_token_id": 2,
25
+ "dropout": 0.1,
26
+ "dtype": "float16",
27
+ "encoder_attention_heads": 4,
28
+ "encoder_ffn_dim": 768,
29
+ "encoder_hidden_size": 192,
30
+ "encoder_layerdrop": 0.0,
31
+ "encoder_layers": 0,
32
+ "eos_token_id": 2,
33
+ "forced_eos_token_id": 2,
34
+ "id2label": {
35
+ "0": "LABEL_0",
36
+ "1": "LABEL_1",
37
+ "2": "LABEL_2"
38
+ },
39
+ "init_std": 0.02,
40
+ "is_decoder": true,
41
+ "is_encoder_decoder": false,
42
+ "label2id": {
43
+ "LABEL_0": 0,
44
+ "LABEL_1": 1,
45
+ "LABEL_2": 2
46
+ },
47
+ "max_position_embeddings": 1024,
48
+ "model_type": "bart",
49
+ "output_attentions": false,
50
+ "output_hidden_states": false,
51
+ "pad_token_id": 1,
52
+ "problem_type": null,
53
+ "return_dict": true,
54
+ "scale_embedding": false,
55
+ "tie_word_embeddings": true,
56
+ "use_cache": true,
57
+ "vocab_size": 414
58
+ },
59
+ "decoder_start_token_id": 0,
60
+ "dtype": "float16",
61
+ "encoder": {
62
+ "_name_or_path": "",
63
+ "architectures": null,
64
+ "auto_map": {
65
+ "AutoConfig": "modeling_molparser_mobile.CustomEncoderConfig",
66
+ "AutoModel": "modeling_molparser_mobile.CustomTimmEncoder"
67
+ },
68
+ "chunk_size_feed_forward": 0,
69
+ "dtype": "float16",
70
+ "hidden_size": 192,
71
+ "id2label": {
72
+ "0": "LABEL_0",
73
+ "1": "LABEL_1"
74
+ },
75
+ "initializer_range": 0.02,
76
+ "is_encoder_decoder": false,
77
+ "label2id": {
78
+ "LABEL_0": 0,
79
+ "LABEL_1": 1
80
+ },
81
+ "model_input_size": 384,
82
+ "model_type": "custom_timm_encoder",
83
+ "output_attentions": false,
84
+ "output_hidden_states": false,
85
+ "pixel_unshuffle": 1,
86
+ "problem_type": null,
87
+ "return_dict": true,
88
+ "timm_model_name": "vit_tiny_r_s16_p8_384.augreg_in21k_ft_in1k",
89
+ "timm_output_dim": 192,
90
+ "timm_pretrained": false
91
+ },
92
+ "is_encoder_decoder": true,
93
+ "model_type": "molparser_vision_encoder_decoder",
94
+ "pad_token_id": 1,
95
+ "tie_word_embeddings": false,
96
+ "transformers_version": "5.4.0",
97
+ "use_cache": false,
98
+ "vocab_size": 414,
99
+ "torch_dtype": "float16",
100
+ "model_name": "MolParser Mobile V2"
101
+ }
generation_config.json ADDED
@@ -0,0 +1,39 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "_from_model_config": false,
3
+ "assistant_confidence_threshold": 0.4,
4
+ "assistant_lookbehind": 10,
5
+ "bos_token_id": 0,
6
+ "decoder_start_token_id": 0,
7
+ "diversity_penalty": 0.0,
8
+ "do_sample": false,
9
+ "early_stopping": false,
10
+ "encoder_no_repeat_ngram_size": 0,
11
+ "encoder_repetition_penalty": 1.0,
12
+ "eos_token_id": 2,
13
+ "epsilon_cutoff": 0.0,
14
+ "eta_cutoff": 0.0,
15
+ "forced_eos_token_id": 2,
16
+ "length_penalty": 1.0,
17
+ "max_length": 384,
18
+ "min_length": 0,
19
+ "no_repeat_ngram_size": 0,
20
+ "num_assistant_tokens": 20,
21
+ "num_assistant_tokens_schedule": "constant",
22
+ "num_beam_groups": 1,
23
+ "num_beams": 1,
24
+ "num_return_sequences": 1,
25
+ "output_attentions": false,
26
+ "output_hidden_states": false,
27
+ "output_scores": false,
28
+ "pad_token_id": 1,
29
+ "remove_invalid_values": false,
30
+ "repetition_penalty": 1.0,
31
+ "return_dict_in_generate": false,
32
+ "target_lookbehind": 10,
33
+ "temperature": 1.0,
34
+ "top_k": 50,
35
+ "top_p": 1.0,
36
+ "transformers_version": "5.4.0",
37
+ "typical_p": 1.0,
38
+ "use_cache": false
39
+ }
image_processing_molparser_mobile.py ADDED
@@ -0,0 +1,95 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Image processor for MolParser Mobile."""
2
+
3
+ from __future__ import annotations
4
+
5
+ from typing import List, Sequence, Union
6
+
7
+ import cv2
8
+ import numpy as np
9
+ import torch
10
+ from PIL import Image
11
+ from transformers import BaseImageProcessor
12
+ from transformers.feature_extraction_utils import BatchFeature
13
+
14
+
15
+ ImageInput = Union[str, Image.Image, np.ndarray, torch.Tensor]
16
+
17
+
18
+ class MolParserImageProcessor(BaseImageProcessor):
19
+ model_input_names = ["pixel_values"]
20
+
21
+ def __init__(
22
+ self,
23
+ image_size: int = 384,
24
+ do_resize: bool = True,
25
+ do_normalize: bool = True,
26
+ image_mean: Sequence[float] = (0.485, 0.456, 0.406),
27
+ image_std: Sequence[float] = (0.229, 0.224, 0.225),
28
+ **kwargs,
29
+ ):
30
+ super().__init__(**kwargs)
31
+ self.image_size = int(image_size)
32
+ self.do_resize = bool(do_resize)
33
+ self.do_normalize = bool(do_normalize)
34
+ self.image_mean = list(image_mean)
35
+ self.image_std = list(image_std)
36
+
37
+ @property
38
+ def size(self):
39
+ return {"height": self.image_size, "width": self.image_size}
40
+
41
+ def _to_pil(self, image: ImageInput) -> Image.Image:
42
+ if isinstance(image, Image.Image):
43
+ return image.convert("RGB")
44
+ if isinstance(image, str):
45
+ return Image.open(image).convert("RGB")
46
+ if isinstance(image, torch.Tensor):
47
+ tensor = image.detach().cpu()
48
+ if tensor.ndim == 3 and tensor.shape[0] in {1, 3, 4}:
49
+ tensor = tensor.permute(1, 2, 0)
50
+ array = tensor.numpy()
51
+ else:
52
+ array = np.asarray(image)
53
+ if array.dtype != np.uint8:
54
+ if array.max() <= 1.0:
55
+ array = array * 255.0
56
+ array = np.clip(np.rint(array), 0, 255).astype(np.uint8)
57
+ if array.ndim == 2:
58
+ return Image.fromarray(array, mode="L").convert("RGB")
59
+ if array.shape[-1] == 4:
60
+ return Image.fromarray(array, mode="RGBA").convert("RGB")
61
+ return Image.fromarray(array).convert("RGB")
62
+
63
+ def _preprocess_one(self, image: ImageInput) -> np.ndarray:
64
+ pil_image = self._to_pil(image)
65
+ array = np.asarray(pil_image).astype(np.uint8)
66
+ if self.do_resize:
67
+ # Match deploy/transform.py: albumentations.Resize defaults to OpenCV INTER_LINEAR.
68
+ array = cv2.resize(array, (self.image_size, self.image_size), interpolation=cv2.INTER_LINEAR)
69
+ array = array.astype(np.float32) / 255.0
70
+ if self.do_normalize:
71
+ mean = np.asarray(self.image_mean, dtype=np.float32).reshape(1, 1, 3)
72
+ std = np.asarray(self.image_std, dtype=np.float32).reshape(1, 1, 3)
73
+ array = (array - mean) / std
74
+ return np.transpose(array, (2, 0, 1))
75
+
76
+ def preprocess(
77
+ self,
78
+ images: Union[ImageInput, Sequence[ImageInput]],
79
+ return_tensors: str | None = None,
80
+ **kwargs,
81
+ ) -> BatchFeature:
82
+ if not isinstance(images, (list, tuple)):
83
+ images = [images]
84
+ pixel_values: List[np.ndarray] = [self._preprocess_one(image) for image in images]
85
+ data = {"pixel_values": np.stack(pixel_values, axis=0)}
86
+ encoded = BatchFeature(data=data)
87
+ if return_tensors is not None:
88
+ encoded = encoded.convert_to_tensors(return_tensors)
89
+ return encoded
90
+
91
+ def __call__(self, images: Union[ImageInput, Sequence[ImageInput]], return_tensors: str | None = None, **kwargs):
92
+ return self.preprocess(images=images, return_tensors=return_tensors, **kwargs)
93
+
94
+
95
+ __all__ = ["MolParserImageProcessor"]
model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:2204a871f7975780ead6eee8ac354de11f859c4eccc8863c3462e6f9eb110f02
3
+ size 20037304
modeling_molparser_mobile.py ADDED
@@ -0,0 +1,204 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """MolParser Mobile model code for Hugging Face Hub remote loading."""
2
+
3
+ from __future__ import annotations
4
+
5
+ import torch
6
+ import torch.nn as nn
7
+ import timm
8
+ from transformers import (
9
+ AutoConfig,
10
+ AutoModel,
11
+ PretrainedConfig,
12
+ PreTrainedModel,
13
+ VisionEncoderDecoderConfig,
14
+ VisionEncoderDecoderModel,
15
+ )
16
+ from transformers.modeling_outputs import BaseModelOutput
17
+
18
+
19
+ MOBILE_TIMM_MODEL_NAME = "vit_tiny_r_s16_p8_384.augreg_in21k_ft_in1k"
20
+ MOBILE_IMAGE_SIZE = 384
21
+ MOBILE_HIDDEN_SIZE = 192
22
+ MOBILE_PIXEL_UNSHUFFLE = 1
23
+
24
+
25
+ class CustomEncoderConfig(PretrainedConfig):
26
+ model_type = "custom_timm_encoder"
27
+
28
+ def __init__(
29
+ self,
30
+ timm_model_name: str = MOBILE_TIMM_MODEL_NAME,
31
+ timm_pretrained: bool = False,
32
+ timm_output_dim: int = MOBILE_HIDDEN_SIZE,
33
+ pixel_unshuffle: int = MOBILE_PIXEL_UNSHUFFLE,
34
+ hidden_size: int = MOBILE_HIDDEN_SIZE,
35
+ initializer_range: float = 0.02,
36
+ model_input_size: int = MOBILE_IMAGE_SIZE,
37
+ **kwargs,
38
+ ):
39
+ self.timm_model_name = timm_model_name
40
+ self.timm_pretrained = timm_pretrained
41
+ self.timm_output_dim = timm_output_dim
42
+ self.pixel_unshuffle = pixel_unshuffle
43
+ self.hidden_size = hidden_size
44
+ self.initializer_range = initializer_range
45
+ self.model_input_size = model_input_size
46
+ super().__init__(**kwargs)
47
+ if getattr(self, "auto_map", None) is None:
48
+ self.auto_map = {
49
+ "AutoConfig": "modeling_molparser_mobile.CustomEncoderConfig",
50
+ "AutoModel": "modeling_molparser_mobile.CustomTimmEncoder",
51
+ }
52
+
53
+
54
+ class CustomTimmEncoder(PreTrainedModel):
55
+ config_class = CustomEncoderConfig
56
+ main_input_name = "pixel_values"
57
+
58
+ def __init__(self, config: CustomEncoderConfig):
59
+ super().__init__(config)
60
+ self.config = config
61
+ self.pixel_unshuffle_factor = int(config.pixel_unshuffle)
62
+ self.unshuffle = (
63
+ nn.PixelUnshuffle(self.pixel_unshuffle_factor)
64
+ if self.pixel_unshuffle_factor > 1
65
+ else nn.Identity()
66
+ )
67
+
68
+ timm_kwargs = {
69
+ "pretrained": bool(config.timm_pretrained),
70
+ "features_only": True,
71
+ "num_classes": 0,
72
+ "global_pool": "",
73
+ }
74
+ if getattr(config, "model_input_size", None):
75
+ timm_kwargs["img_size"] = int(config.model_input_size)
76
+ self.timm_model = timm.create_model(config.timm_model_name, **timm_kwargs)
77
+
78
+ in_channels = int(config.timm_output_dim) * self.pixel_unshuffle_factor**2
79
+ self.use_projection = not (
80
+ self.pixel_unshuffle_factor == 1 and in_channels == int(config.hidden_size)
81
+ )
82
+ if self.use_projection:
83
+ self.projection = nn.Sequential(
84
+ nn.Conv2d(in_channels=in_channels, out_channels=config.hidden_size, kernel_size=1),
85
+ nn.GELU(),
86
+ )
87
+ else:
88
+ self.projection = nn.Identity()
89
+
90
+ def forward(self, pixel_values: torch.Tensor, **kwargs):
91
+ encoder_features = self.timm_model(pixel_values)[-1]
92
+ if encoder_features.ndim != 4:
93
+ raise ValueError(f"Expected 4D feature map, got shape={tuple(encoder_features.shape)}")
94
+ if encoder_features.shape[1] == self.config.timm_output_dim:
95
+ pass
96
+ elif encoder_features.shape[-1] == self.config.timm_output_dim:
97
+ encoder_features = encoder_features.permute(0, 3, 1, 2).contiguous()
98
+ else:
99
+ raise ValueError(
100
+ "Unexpected timm feature shape "
101
+ f"{tuple(encoder_features.shape)} for timm_output_dim={self.config.timm_output_dim}"
102
+ )
103
+
104
+ encoder_features = self.unshuffle(encoder_features)
105
+ encoder_features = self.projection(encoder_features)
106
+ _, channels, _, _ = encoder_features.shape
107
+ if channels != self.config.hidden_size:
108
+ raise ValueError(
109
+ f"Unexpected encoder channels={channels}, expected hidden_size={self.config.hidden_size}."
110
+ )
111
+ encoder_hidden_states = encoder_features.flatten(2).transpose(1, 2)
112
+ return BaseModelOutput(last_hidden_state=encoder_hidden_states)
113
+
114
+
115
+ try:
116
+ AutoConfig.register("custom_timm_encoder", CustomEncoderConfig)
117
+ except ValueError:
118
+ pass
119
+ try:
120
+ AutoModel.register(CustomEncoderConfig, CustomTimmEncoder)
121
+ except ValueError:
122
+ pass
123
+
124
+ CustomEncoderConfig.register_for_auto_class()
125
+ CustomTimmEncoder.register_for_auto_class("AutoModel")
126
+
127
+
128
+ class MolParserVisionEncoderDecoderConfig(VisionEncoderDecoderConfig):
129
+ model_type = "molparser_vision_encoder_decoder"
130
+
131
+
132
+ class MolParserVisionEncoderDecoderModel(VisionEncoderDecoderModel):
133
+ config_class = MolParserVisionEncoderDecoderConfig
134
+
135
+ def __init__(self, config=None, encoder=None, decoder=None):
136
+ if config is not None and not isinstance(config, self.config_class):
137
+ config = self.config_class.from_dict(config.to_dict())
138
+ super().__init__(config=config, encoder=encoder, decoder=decoder)
139
+
140
+ @classmethod
141
+ def get_init_context(cls, dtype, is_quantized, _is_ds_init_called, allow_all_kernels):
142
+ contexts = super().get_init_context(dtype, is_quantized, _is_ds_init_called, allow_all_kernels)
143
+ if is_quantized:
144
+ return contexts
145
+ # timm ViT initialization calls tensor.item(), which cannot run on meta tensors.
146
+ return [ctx for ctx in contexts if not (isinstance(ctx, torch.device) and ctx.type == "meta")]
147
+
148
+
149
+ try:
150
+ AutoConfig.register("molparser_vision_encoder_decoder", MolParserVisionEncoderDecoderConfig)
151
+ except ValueError:
152
+ pass
153
+
154
+ MolParserVisionEncoderDecoderConfig.register_for_auto_class()
155
+ MolParserVisionEncoderDecoderModel.register_for_auto_class("AutoModelForImageTextToText")
156
+
157
+
158
+ def assert_mobile_config(config: MolParserVisionEncoderDecoderConfig) -> None:
159
+ encoder = getattr(config, "encoder", None)
160
+ decoder = getattr(config, "decoder", None)
161
+ if encoder is None or decoder is None:
162
+ raise ValueError("MolParser Mobile config must contain encoder and decoder sub-configs.")
163
+
164
+ checks = {
165
+ "encoder.timm_model_name": getattr(encoder, "timm_model_name", None) == MOBILE_TIMM_MODEL_NAME,
166
+ "encoder.model_input_size": int(getattr(encoder, "model_input_size", 0) or 0) == MOBILE_IMAGE_SIZE,
167
+ "encoder.hidden_size": int(getattr(encoder, "hidden_size", 0) or 0) == MOBILE_HIDDEN_SIZE,
168
+ "encoder.pixel_unshuffle": int(getattr(encoder, "pixel_unshuffle", 0) or 0) == MOBILE_PIXEL_UNSHUFFLE,
169
+ "decoder.d_model": int(getattr(decoder, "d_model", 0) or 0) == MOBILE_HIDDEN_SIZE,
170
+ "decoder.decoder_layers": int(getattr(decoder, "decoder_layers", 0) or 0) == 6,
171
+ "decoder.decoder_attention_heads": int(getattr(decoder, "decoder_attention_heads", 0) or 0) == 4,
172
+ }
173
+ failed = [name for name, ok in checks.items() if not ok]
174
+ if failed:
175
+ raise ValueError(
176
+ "This Hub package is mobile-only; non-mobile config values found: "
177
+ + ", ".join(failed)
178
+ )
179
+
180
+
181
+ def load_molparser_mobile_model(checkpoint_path: str) -> MolParserVisionEncoderDecoderModel:
182
+ config = MolParserVisionEncoderDecoderConfig.from_pretrained(
183
+ checkpoint_path,
184
+ trust_remote_code=True,
185
+ )
186
+ assert_mobile_config(config)
187
+ if getattr(config, "encoder", None) is not None:
188
+ config.encoder.timm_pretrained = False
189
+ model = MolParserVisionEncoderDecoderModel.from_pretrained(
190
+ checkpoint_path,
191
+ config=config,
192
+ trust_remote_code=True,
193
+ )
194
+ return model
195
+
196
+
197
+ __all__ = [
198
+ "CustomEncoderConfig",
199
+ "CustomTimmEncoder",
200
+ "MolParserVisionEncoderDecoderConfig",
201
+ "MolParserVisionEncoderDecoderModel",
202
+ "assert_mobile_config",
203
+ "load_molparser_mobile_model",
204
+ ]
preprocessor_config.json ADDED
@@ -0,0 +1,20 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "image_processor_type": "MolParserImageProcessor",
3
+ "model_name": "MolParser Mobile V2",
4
+ "auto_map": {
5
+ "AutoImageProcessor": "image_processing_molparser_mobile.MolParserImageProcessor"
6
+ },
7
+ "image_size": 384,
8
+ "do_resize": true,
9
+ "do_normalize": true,
10
+ "image_mean": [
11
+ 0.485,
12
+ 0.456,
13
+ 0.406
14
+ ],
15
+ "image_std": [
16
+ 0.229,
17
+ 0.224,
18
+ 0.225
19
+ ]
20
+ }
processing_molparser_mobile.py ADDED
@@ -0,0 +1,71 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Processor that combines MolParser Mobile image preprocessing and tokenizer."""
2
+
3
+ from __future__ import annotations
4
+
5
+ from pathlib import Path
6
+ from typing import Sequence
7
+
8
+ from .image_processing_molparser_mobile import MolParserImageProcessor
9
+ from .tokenization_molparser_mobile import MolParserTokenizer
10
+
11
+
12
+ class MolParserProcessor:
13
+ attributes = ["image_processor", "tokenizer"]
14
+ image_processor_class = "MolParserImageProcessor"
15
+ tokenizer_class = "MolParserTokenizer"
16
+
17
+ def __init__(
18
+ self,
19
+ image_processor: MolParserImageProcessor | None = None,
20
+ tokenizer: MolParserTokenizer | None = None,
21
+ ):
22
+ self.image_processor = image_processor or MolParserImageProcessor()
23
+ self.tokenizer = tokenizer
24
+
25
+ @classmethod
26
+ def register_for_auto_class(cls, auto_class: str = "AutoProcessor"):
27
+ cls._auto_class = auto_class
28
+
29
+ @classmethod
30
+ def from_pretrained(cls, pretrained_model_name_or_path: str, **kwargs) -> "MolParserProcessor":
31
+ path = str(pretrained_model_name_or_path)
32
+ image_processor = MolParserImageProcessor.from_pretrained(path, **kwargs)
33
+ tokenizer = MolParserTokenizer.from_pretrained(path, **kwargs)
34
+ return cls(image_processor=image_processor, tokenizer=tokenizer)
35
+
36
+ def save_pretrained(self, save_directory: str, **kwargs):
37
+ Path(save_directory).mkdir(parents=True, exist_ok=True)
38
+ image_files = self.image_processor.save_pretrained(save_directory, **kwargs)
39
+ tokenizer_files = ()
40
+ if self.tokenizer is not None:
41
+ tokenizer_files = self.tokenizer.save_pretrained(save_directory, **kwargs)
42
+ return tuple(image_files) + tuple(tokenizer_files)
43
+
44
+ def __call__(
45
+ self,
46
+ images=None,
47
+ text: str | Sequence[str] | None = None,
48
+ return_tensors: str | None = None,
49
+ **kwargs,
50
+ ):
51
+ encoded = {}
52
+ if images is not None:
53
+ encoded.update(self.image_processor(images=images, return_tensors=return_tensors, **kwargs))
54
+ if text is not None:
55
+ if self.tokenizer is None:
56
+ raise ValueError("MolParserProcessor was created without a tokenizer.")
57
+ encoded.update(self.tokenizer(text, return_tensors=return_tensors, **kwargs))
58
+ return encoded
59
+
60
+ def decode(self, *args, **kwargs):
61
+ if self.tokenizer is None:
62
+ raise ValueError("MolParserProcessor was created without a tokenizer.")
63
+ return self.tokenizer.decode(*args, **kwargs)
64
+
65
+ def batch_decode(self, *args, **kwargs):
66
+ if self.tokenizer is None:
67
+ raise ValueError("MolParserProcessor was created without a tokenizer.")
68
+ return self.tokenizer.batch_decode(*args, **kwargs)
69
+
70
+
71
+ __all__ = ["MolParserProcessor"]
processor_config.json ADDED
@@ -0,0 +1,7 @@
 
 
 
 
 
 
 
 
1
+ {
2
+ "processor_class": "MolParserProcessor",
3
+ "model_name": "MolParser Mobile V2",
4
+ "auto_map": {
5
+ "AutoProcessor": "processing_molparser_mobile.MolParserProcessor"
6
+ }
7
+ }
tokenization_molparser_mobile.py ADDED
@@ -0,0 +1,191 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """MolParser Mobile tokenizer for Hugging Face Hub remote loading."""
2
+
3
+ from __future__ import annotations
4
+
5
+ import json
6
+ import os
7
+ import re
8
+ from pathlib import Path
9
+ from typing import Dict, Iterable, List, Optional, Sequence, Union
10
+
11
+ from huggingface_hub import hf_hub_download
12
+ from transformers import PreTrainedTokenizer
13
+
14
+
15
+ TOKENIZER_CONFIG_NAME = "tokenizer_config.json"
16
+ VOCAB_NAME = "vocab.txt"
17
+
18
+
19
+ class MolParserTokenizer(PreTrainedTokenizer):
20
+ model_input_names = ["input_ids", "attention_mask"]
21
+ padding_side = "right"
22
+
23
+ def __init__(
24
+ self,
25
+ vocab_list: Optional[List[str]] = None,
26
+ special_tokens: Optional[Dict[str, str]] = None,
27
+ additional_special_tokens: Optional[Sequence[str]] = None,
28
+ **kwargs,
29
+ ):
30
+ if vocab_list is None:
31
+ vocab_list = []
32
+ if special_tokens is None:
33
+ special_tokens = {}
34
+ if additional_special_tokens is None:
35
+ additional_special_tokens = []
36
+
37
+ self.special_tokens = {
38
+ "cls_token": special_tokens.get("cls_token", "[CLS]"),
39
+ "pad_token": special_tokens.get("pad_token", "[PAD]"),
40
+ "sep_token": special_tokens.get("sep_token", "[SEP]"),
41
+ "unk_token": special_tokens.get("unk_token", "[UNK]"),
42
+ }
43
+ self.additional_special_tokens = list(dict.fromkeys(additional_special_tokens))
44
+ self.vocab_list = list(vocab_list)
45
+ all_tokens = self._build_full_vocab(self.vocab_list)
46
+ self.vocab = {token: idx for idx, token in enumerate(all_tokens)}
47
+ self.ids_to_tokens = {idx: token for token, idx in self.vocab.items()}
48
+ self._decode_skip_tokens = set(self.special_tokens.values())
49
+ self._compile_pattern()
50
+
51
+ super().__init__(
52
+ cls_token=self.special_tokens["cls_token"],
53
+ pad_token=self.special_tokens["pad_token"],
54
+ sep_token=self.special_tokens["sep_token"],
55
+ unk_token=self.special_tokens["unk_token"],
56
+ bos_token=self.special_tokens["cls_token"],
57
+ eos_token=self.special_tokens["sep_token"],
58
+ additional_special_tokens=self.additional_special_tokens,
59
+ **kwargs,
60
+ )
61
+
62
+ def _build_full_vocab(self, vocab_list: Sequence[str]) -> List[str]:
63
+ ordered_tokens: List[str] = []
64
+ for token in list(self.special_tokens.values()) + list(vocab_list) + list(self.additional_special_tokens):
65
+ if token not in ordered_tokens:
66
+ ordered_tokens.append(token)
67
+ return ordered_tokens
68
+
69
+ def _compile_pattern(self) -> None:
70
+ multi_char_tokens = sorted(self.vocab.keys(), key=len, reverse=True)
71
+ pattern = "(" + "|".join(re.escape(token) for token in multi_char_tokens) + "|.)"
72
+ self.pattern = re.compile(pattern)
73
+
74
+ @property
75
+ def vocab_size(self) -> int:
76
+ return len(self.vocab)
77
+
78
+ def __len__(self) -> int:
79
+ return len(self.vocab)
80
+
81
+ def get_vocab(self) -> Dict[str, int]:
82
+ return dict(self.vocab)
83
+
84
+ def _tokenize(self, text: str) -> List[str]:
85
+ return [token for token in self.pattern.findall(str(text)) if token]
86
+
87
+ def tokenize(self, text: str, **kwargs) -> List[str]:
88
+ return self._tokenize(text)
89
+
90
+ def _convert_token_to_id(self, token: str) -> int:
91
+ return self.vocab.get(token, self.unk_token_id)
92
+
93
+ def _convert_id_to_token(self, index: int) -> str:
94
+ return self.ids_to_tokens.get(int(index), self.unk_token)
95
+
96
+ def convert_tokens_to_string(self, tokens: Sequence[str]) -> str:
97
+ return "".join(tokens)
98
+
99
+ def build_inputs_with_special_tokens(self, token_ids_0, token_ids_1=None):
100
+ if token_ids_1 is None:
101
+ return list(token_ids_0)
102
+ return list(token_ids_0) + list(token_ids_1)
103
+
104
+ def encode(self, text: str, add_special_tokens: bool = False, **kwargs) -> List[int]:
105
+ token_ids = [self._convert_token_to_id(token) for token in self._tokenize(text)]
106
+ if add_special_tokens:
107
+ return [self.bos_token_id] + token_ids + [self.eos_token_id]
108
+ return token_ids
109
+
110
+ def decode(self, token_ids: Iterable[int], skip_special_tokens: bool = False, **kwargs) -> str:
111
+ tokens = [self._convert_id_to_token(idx) for idx in token_ids]
112
+ if skip_special_tokens:
113
+ # Match deploy/tokenizer.py: keep MolParser business tokens such as
114
+ # <sep>, <a>, </a>, <r>, </r>, <c>, </c>, and |Sg:n|.
115
+ tokens = [token for token in tokens if token not in self._decode_skip_tokens]
116
+ return "".join(tokens)
117
+
118
+ def batch_encode(self, texts: Sequence[str], add_special_tokens: bool = False) -> List[List[int]]:
119
+ return [self.encode(text, add_special_tokens=add_special_tokens) for text in texts]
120
+
121
+ def batch_decode(
122
+ self,
123
+ sequences: Sequence[Sequence[int]],
124
+ skip_special_tokens: bool = False,
125
+ **kwargs,
126
+ ) -> List[str]:
127
+ return [self.decode(ids, skip_special_tokens=skip_special_tokens, **kwargs) for ids in sequences]
128
+
129
+ def to_dict(self) -> Dict[str, object]:
130
+ return {
131
+ "vocab_list": self.vocab_list,
132
+ "special_tokens": self.special_tokens,
133
+ "additional_special_tokens": self.additional_special_tokens,
134
+ "tokenizer_class": self.__class__.__name__,
135
+ "auto_map": {
136
+ "AutoTokenizer": [
137
+ "tokenization_molparser_mobile.MolParserTokenizer",
138
+ None,
139
+ ]
140
+ },
141
+ }
142
+
143
+ @classmethod
144
+ def from_dict(cls, config: Dict[str, object]) -> "MolParserTokenizer":
145
+ return cls(
146
+ vocab_list=list(config["vocab_list"]),
147
+ special_tokens=dict(config["special_tokens"]),
148
+ additional_special_tokens=list(config.get("additional_special_tokens", [])),
149
+ )
150
+
151
+ def save_vocabulary(self, save_directory: str, filename_prefix: Optional[str] = None):
152
+ path = Path(save_directory)
153
+ path.mkdir(parents=True, exist_ok=True)
154
+ name = f"{filename_prefix}-{VOCAB_NAME}" if filename_prefix else VOCAB_NAME
155
+ vocab_path = path / name
156
+ vocab_path.write_text("\n".join(self.vocab_list) + "\n", encoding="utf-8")
157
+ return (str(vocab_path),)
158
+
159
+ def save_pretrained(self, save_directory: str, **kwargs):
160
+ os.makedirs(save_directory, exist_ok=True)
161
+ config_path = os.path.join(save_directory, TOKENIZER_CONFIG_NAME)
162
+ with open(config_path, "w", encoding="utf-8") as f:
163
+ json.dump(self.to_dict(), f, ensure_ascii=False, indent=2)
164
+ vocab_files = self.save_vocabulary(save_directory)
165
+ return (config_path, *vocab_files)
166
+
167
+ @classmethod
168
+ def from_pretrained(cls, pretrained_model_name_or_path: str, *args, **kwargs) -> "MolParserTokenizer":
169
+ config_path = Path(pretrained_model_name_or_path)
170
+ if config_path.is_dir():
171
+ config_path = config_path / TOKENIZER_CONFIG_NAME
172
+ elif config_path.is_file():
173
+ pass
174
+ else:
175
+ config_path = Path(
176
+ hf_hub_download(
177
+ repo_id=str(pretrained_model_name_or_path),
178
+ filename=TOKENIZER_CONFIG_NAME,
179
+ repo_type=kwargs.get("repo_type"),
180
+ revision=kwargs.get("revision"),
181
+ cache_dir=kwargs.get("cache_dir"),
182
+ token=kwargs.get("token"),
183
+ local_files_only=kwargs.get("local_files_only", False),
184
+ )
185
+ )
186
+ with open(config_path, "r", encoding="utf-8") as f:
187
+ config = json.load(f)
188
+ return cls.from_dict(config)
189
+
190
+
191
+ __all__ = ["MolParserTokenizer"]
tokenizer_config.json ADDED
@@ -0,0 +1,429 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "vocab_list": [
3
+ "!",
4
+ "\"",
5
+ "#",
6
+ "$",
7
+ "%",
8
+ "&",
9
+ "'",
10
+ "(",
11
+ ")",
12
+ "*",
13
+ "+",
14
+ ",",
15
+ "-",
16
+ ".",
17
+ "/",
18
+ "0",
19
+ "1",
20
+ "2",
21
+ "3",
22
+ "4",
23
+ "5",
24
+ "6",
25
+ "7",
26
+ "8",
27
+ "9",
28
+ "10",
29
+ "11",
30
+ "12",
31
+ "13",
32
+ "14",
33
+ "15",
34
+ "16",
35
+ "17",
36
+ "18",
37
+ "19",
38
+ "20",
39
+ "21",
40
+ "22",
41
+ "23",
42
+ "24",
43
+ "25",
44
+ "26",
45
+ "27",
46
+ "28",
47
+ "29",
48
+ "30",
49
+ "31",
50
+ "32",
51
+ "33",
52
+ "34",
53
+ "35",
54
+ "36",
55
+ "37",
56
+ "38",
57
+ "39",
58
+ "40",
59
+ "41",
60
+ "42",
61
+ "43",
62
+ "44",
63
+ "45",
64
+ "46",
65
+ "47",
66
+ "48",
67
+ "49",
68
+ "50",
69
+ "51",
70
+ "52",
71
+ "53",
72
+ "54",
73
+ "55",
74
+ "56",
75
+ "57",
76
+ "58",
77
+ "59",
78
+ "60",
79
+ "61",
80
+ "62",
81
+ "63",
82
+ "64",
83
+ "65",
84
+ "66",
85
+ "67",
86
+ "68",
87
+ "69",
88
+ "70",
89
+ "71",
90
+ "72",
91
+ "73",
92
+ "74",
93
+ "75",
94
+ "76",
95
+ "77",
96
+ "78",
97
+ "79",
98
+ "80",
99
+ "81",
100
+ "82",
101
+ "83",
102
+ "84",
103
+ "85",
104
+ "86",
105
+ "87",
106
+ "88",
107
+ "89",
108
+ "90",
109
+ "91",
110
+ "92",
111
+ "93",
112
+ "94",
113
+ "95",
114
+ "96",
115
+ "97",
116
+ "98",
117
+ "99",
118
+ "100",
119
+ "101",
120
+ "102",
121
+ "103",
122
+ "104",
123
+ "105",
124
+ "106",
125
+ "107",
126
+ "108",
127
+ "109",
128
+ "110",
129
+ "111",
130
+ "112",
131
+ "113",
132
+ "114",
133
+ "115",
134
+ "116",
135
+ "117",
136
+ "118",
137
+ "119",
138
+ "120",
139
+ "121",
140
+ "122",
141
+ "123",
142
+ "124",
143
+ "125",
144
+ "126",
145
+ "127",
146
+ "128",
147
+ "129",
148
+ "130",
149
+ "131",
150
+ "132",
151
+ "133",
152
+ "134",
153
+ "135",
154
+ "136",
155
+ "137",
156
+ "138",
157
+ "139",
158
+ "140",
159
+ "141",
160
+ "142",
161
+ "143",
162
+ "144",
163
+ "145",
164
+ "146",
165
+ "147",
166
+ "148",
167
+ "149",
168
+ "150",
169
+ "151",
170
+ "152",
171
+ "153",
172
+ "154",
173
+ "155",
174
+ "156",
175
+ "157",
176
+ "158",
177
+ "159",
178
+ "160",
179
+ "161",
180
+ "162",
181
+ "163",
182
+ "164",
183
+ "165",
184
+ "166",
185
+ "167",
186
+ "168",
187
+ "169",
188
+ "170",
189
+ "171",
190
+ "172",
191
+ "173",
192
+ "174",
193
+ "175",
194
+ "176",
195
+ "177",
196
+ "178",
197
+ "179",
198
+ "180",
199
+ "181",
200
+ "182",
201
+ "183",
202
+ "184",
203
+ "185",
204
+ "186",
205
+ "187",
206
+ "188",
207
+ "189",
208
+ "190",
209
+ "191",
210
+ "192",
211
+ "193",
212
+ "194",
213
+ "195",
214
+ "196",
215
+ "197",
216
+ "198",
217
+ "199",
218
+ "200",
219
+ "201",
220
+ "202",
221
+ "203",
222
+ "204",
223
+ "205",
224
+ "206",
225
+ "207",
226
+ "208",
227
+ "209",
228
+ "210",
229
+ "211",
230
+ "212",
231
+ "213",
232
+ "214",
233
+ "215",
234
+ "216",
235
+ "217",
236
+ "218",
237
+ "219",
238
+ "220",
239
+ "221",
240
+ "222",
241
+ "223",
242
+ "224",
243
+ "225",
244
+ "226",
245
+ "227",
246
+ "228",
247
+ "229",
248
+ "230",
249
+ "231",
250
+ "232",
251
+ "233",
252
+ "234",
253
+ "235",
254
+ "236",
255
+ "237",
256
+ "238",
257
+ "239",
258
+ "240",
259
+ "241",
260
+ "242",
261
+ "243",
262
+ "244",
263
+ "245",
264
+ "246",
265
+ "247",
266
+ "248",
267
+ "249",
268
+ "250",
269
+ "251",
270
+ "252",
271
+ "253",
272
+ "254",
273
+ "255",
274
+ ":",
275
+ ";",
276
+ "<",
277
+ "=",
278
+ ">",
279
+ "?",
280
+ "@@",
281
+ "@",
282
+ "A",
283
+ "B",
284
+ "C",
285
+ "D",
286
+ "E",
287
+ "F",
288
+ "G",
289
+ "H",
290
+ "I",
291
+ "J",
292
+ "K",
293
+ "L",
294
+ "M",
295
+ "N",
296
+ "O",
297
+ "P",
298
+ "Q",
299
+ "R",
300
+ "S",
301
+ "T",
302
+ "U",
303
+ "V",
304
+ "W",
305
+ "X",
306
+ "Y",
307
+ "Z",
308
+ "[",
309
+ "\\",
310
+ "]",
311
+ "^",
312
+ "_",
313
+ "`",
314
+ "a",
315
+ "b",
316
+ "c",
317
+ "d",
318
+ "e",
319
+ "f",
320
+ "g",
321
+ "h",
322
+ "i",
323
+ "j",
324
+ "k",
325
+ "l",
326
+ "m",
327
+ "n",
328
+ "o",
329
+ "p",
330
+ "q",
331
+ "r",
332
+ "s",
333
+ "t",
334
+ "u",
335
+ "v",
336
+ "w",
337
+ "x",
338
+ "y",
339
+ "z",
340
+ "{",
341
+ "|",
342
+ "}",
343
+ "~",
344
+ "Ag",
345
+ "Al",
346
+ "As",
347
+ "Au",
348
+ "Br",
349
+ "Ca",
350
+ "Cl",
351
+ "Cr",
352
+ "Cu",
353
+ "Fe",
354
+ "Gd",
355
+ "Hg",
356
+ "Li",
357
+ "Mg",
358
+ "Mn",
359
+ "Na",
360
+ "Ni",
361
+ "Pb",
362
+ "Pt",
363
+ "Sb",
364
+ "Se",
365
+ "Si",
366
+ "Sn",
367
+ "Ti",
368
+ "Zn",
369
+ "Zr",
370
+ "2H",
371
+ "3H",
372
+ "->",
373
+ "<-",
374
+ "ball",
375
+ "grey",
376
+ "black",
377
+ "red",
378
+ "green",
379
+ "blue",
380
+ "yellow",
381
+ "purple",
382
+ "orange",
383
+ "pink",
384
+ "brown",
385
+ "DNA",
386
+ "RNA",
387
+ "star",
388
+ "capsule",
389
+ "other",
390
+ "radioactive",
391
+ "Sg:",
392
+ "Ra",
393
+ "Sa"
394
+ ],
395
+ "special_tokens": {
396
+ "cls_token": "[CLS]",
397
+ "pad_token": "[PAD]",
398
+ "sep_token": "[SEP]",
399
+ "unk_token": "[UNK]"
400
+ },
401
+ "additional_special_tokens": [
402
+ "<sep>",
403
+ "<dum>",
404
+ "<id>",
405
+ "<a>",
406
+ "</a>",
407
+ "<r>",
408
+ "</r>",
409
+ "<c>",
410
+ "</c>",
411
+ "<d>",
412
+ "</d>",
413
+ "<s>",
414
+ "</s>",
415
+ "<g>",
416
+ "</g>",
417
+ "<v>",
418
+ "</v>",
419
+ "<x>",
420
+ "</x>"
421
+ ],
422
+ "tokenizer_class": "MolParserTokenizer",
423
+ "auto_map": {
424
+ "AutoTokenizer": [
425
+ "tokenization_molparser_mobile.MolParserTokenizer",
426
+ null
427
+ ]
428
+ }
429
+ }
vocab.txt ADDED
@@ -0,0 +1,391 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ !
2
+ "
3
+ #
4
+ $
5
+ %
6
+ &
7
+ '
8
+ (
9
+ )
10
+ *
11
+ +
12
+ ,
13
+ -
14
+ .
15
+ /
16
+ 0
17
+ 1
18
+ 2
19
+ 3
20
+ 4
21
+ 5
22
+ 6
23
+ 7
24
+ 8
25
+ 9
26
+ 10
27
+ 11
28
+ 12
29
+ 13
30
+ 14
31
+ 15
32
+ 16
33
+ 17
34
+ 18
35
+ 19
36
+ 20
37
+ 21
38
+ 22
39
+ 23
40
+ 24
41
+ 25
42
+ 26
43
+ 27
44
+ 28
45
+ 29
46
+ 30
47
+ 31
48
+ 32
49
+ 33
50
+ 34
51
+ 35
52
+ 36
53
+ 37
54
+ 38
55
+ 39
56
+ 40
57
+ 41
58
+ 42
59
+ 43
60
+ 44
61
+ 45
62
+ 46
63
+ 47
64
+ 48
65
+ 49
66
+ 50
67
+ 51
68
+ 52
69
+ 53
70
+ 54
71
+ 55
72
+ 56
73
+ 57
74
+ 58
75
+ 59
76
+ 60
77
+ 61
78
+ 62
79
+ 63
80
+ 64
81
+ 65
82
+ 66
83
+ 67
84
+ 68
85
+ 69
86
+ 70
87
+ 71
88
+ 72
89
+ 73
90
+ 74
91
+ 75
92
+ 76
93
+ 77
94
+ 78
95
+ 79
96
+ 80
97
+ 81
98
+ 82
99
+ 83
100
+ 84
101
+ 85
102
+ 86
103
+ 87
104
+ 88
105
+ 89
106
+ 90
107
+ 91
108
+ 92
109
+ 93
110
+ 94
111
+ 95
112
+ 96
113
+ 97
114
+ 98
115
+ 99
116
+ 100
117
+ 101
118
+ 102
119
+ 103
120
+ 104
121
+ 105
122
+ 106
123
+ 107
124
+ 108
125
+ 109
126
+ 110
127
+ 111
128
+ 112
129
+ 113
130
+ 114
131
+ 115
132
+ 116
133
+ 117
134
+ 118
135
+ 119
136
+ 120
137
+ 121
138
+ 122
139
+ 123
140
+ 124
141
+ 125
142
+ 126
143
+ 127
144
+ 128
145
+ 129
146
+ 130
147
+ 131
148
+ 132
149
+ 133
150
+ 134
151
+ 135
152
+ 136
153
+ 137
154
+ 138
155
+ 139
156
+ 140
157
+ 141
158
+ 142
159
+ 143
160
+ 144
161
+ 145
162
+ 146
163
+ 147
164
+ 148
165
+ 149
166
+ 150
167
+ 151
168
+ 152
169
+ 153
170
+ 154
171
+ 155
172
+ 156
173
+ 157
174
+ 158
175
+ 159
176
+ 160
177
+ 161
178
+ 162
179
+ 163
180
+ 164
181
+ 165
182
+ 166
183
+ 167
184
+ 168
185
+ 169
186
+ 170
187
+ 171
188
+ 172
189
+ 173
190
+ 174
191
+ 175
192
+ 176
193
+ 177
194
+ 178
195
+ 179
196
+ 180
197
+ 181
198
+ 182
199
+ 183
200
+ 184
201
+ 185
202
+ 186
203
+ 187
204
+ 188
205
+ 189
206
+ 190
207
+ 191
208
+ 192
209
+ 193
210
+ 194
211
+ 195
212
+ 196
213
+ 197
214
+ 198
215
+ 199
216
+ 200
217
+ 201
218
+ 202
219
+ 203
220
+ 204
221
+ 205
222
+ 206
223
+ 207
224
+ 208
225
+ 209
226
+ 210
227
+ 211
228
+ 212
229
+ 213
230
+ 214
231
+ 215
232
+ 216
233
+ 217
234
+ 218
235
+ 219
236
+ 220
237
+ 221
238
+ 222
239
+ 223
240
+ 224
241
+ 225
242
+ 226
243
+ 227
244
+ 228
245
+ 229
246
+ 230
247
+ 231
248
+ 232
249
+ 233
250
+ 234
251
+ 235
252
+ 236
253
+ 237
254
+ 238
255
+ 239
256
+ 240
257
+ 241
258
+ 242
259
+ 243
260
+ 244
261
+ 245
262
+ 246
263
+ 247
264
+ 248
265
+ 249
266
+ 250
267
+ 251
268
+ 252
269
+ 253
270
+ 254
271
+ 255
272
+ :
273
+ ;
274
+ <
275
+ =
276
+ >
277
+ ?
278
+ @@
279
+ @
280
+ A
281
+ B
282
+ C
283
+ D
284
+ E
285
+ F
286
+ G
287
+ H
288
+ I
289
+ J
290
+ K
291
+ L
292
+ M
293
+ N
294
+ O
295
+ P
296
+ Q
297
+ R
298
+ S
299
+ T
300
+ U
301
+ V
302
+ W
303
+ X
304
+ Y
305
+ Z
306
+ [
307
+ \
308
+ ]
309
+ ^
310
+ _
311
+ `
312
+ a
313
+ b
314
+ c
315
+ d
316
+ e
317
+ f
318
+ g
319
+ h
320
+ i
321
+ j
322
+ k
323
+ l
324
+ m
325
+ n
326
+ o
327
+ p
328
+ q
329
+ r
330
+ s
331
+ t
332
+ u
333
+ v
334
+ w
335
+ x
336
+ y
337
+ z
338
+ {
339
+ |
340
+ }
341
+ ~
342
+ Ag
343
+ Al
344
+ As
345
+ Au
346
+ Br
347
+ Ca
348
+ Cl
349
+ Cr
350
+ Cu
351
+ Fe
352
+ Gd
353
+ Hg
354
+ Li
355
+ Mg
356
+ Mn
357
+ Na
358
+ Ni
359
+ Pb
360
+ Pt
361
+ Sb
362
+ Se
363
+ Si
364
+ Sn
365
+ Ti
366
+ Zn
367
+ Zr
368
+ 2H
369
+ 3H
370
+ ->
371
+ <-
372
+ ball
373
+ grey
374
+ black
375
+ red
376
+ green
377
+ blue
378
+ yellow
379
+ purple
380
+ orange
381
+ pink
382
+ brown
383
+ DNA
384
+ RNA
385
+ star
386
+ capsule
387
+ other
388
+ radioactive
389
+ Sg:
390
+ Ra
391
+ Sa