ChartGalaxyPP commited on
Commit
815d6c5
·
verified ·
1 Parent(s): 1f05665

Improve README with curated examples and usage guidance

Browse files
Files changed (4) hide show
  1. .gitattributes +1 -0
  2. README.md +58 -25
  3. SHA256SUMS +2 -1
  4. assets/image2json-results.png +3 -0
.gitattributes CHANGED
@@ -1,2 +1,3 @@
1
  *.safetensors filter=lfs diff=lfs merge=lfs -text
2
  tokenizer.json filter=lfs diff=lfs merge=lfs -text
 
 
1
  *.safetensors filter=lfs diff=lfs merge=lfs -text
2
  tokenizer.json filter=lfs diff=lfs merge=lfs -text
3
+ assets/image2json-results.png filter=lfs diff=lfs merge=lfs -text
README.md CHANGED
@@ -14,44 +14,77 @@ tags:
14
  - qwen3_5
15
  ---
16
 
17
- # ChartGalaxy++ Image2JSON
18
 
19
- Image2JSON predicts a structured description of an infographic chart, including visual elements, semantic groups, parent links, text, bounding boxes, and appearance descriptions. It is the Qwen3.5-4B-family model fine-tuned for the image-to-scene-graph task in ChartGalaxy++.
20
 
21
- This release contains the full **checkpoint-100000** model in two Safetensors shards, together with its tokenizer, processor, chat template, and inference specification. The nine model files total **9,339,891,123 bytes**. Their hashes match the export used for the paper's evaluation.
22
 
23
- [Dataset](https://huggingface.co/datasets/ChartGalaxyPP/ChartGalaxyPlusPlus) · [Project](https://github.com/ChartGalaxyPP/ChartGalaxyPlusPlus) · [Inference guide](INFERENCE.md) · [Output format](OUTPUT_FORMAT.md)
24
 
25
- ## Model and evaluation
26
 
27
- | Property | Value |
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
28
  | --- | --- |
29
- | Architecture | `Qwen3_5ForConditionalGeneration` |
30
- | Fine-tuning | Supervised fine-tuning for chart scene graph prediction |
31
- | Checkpoint step | 100,000 |
32
- | Weight dtype | BF16 |
33
- | Artifact | Full model, tokenizer, image processor, and chat template |
34
- | Paper test set | 1,000 charts: 500 real and 500 synthetic |
35
- | Element F1 | 89.4% |
36
- | Hierarchy F1 | 85.1% |
37
- | Spatial F1 | 88.0% |
38
 
39
- Scores are those of the paper's evaluation protocol, including its output recovery and graph evaluation procedures. They are not a fresh evaluation of the short inference example below. Stored weight tensors include the vision and auxiliary model components; the architecture's “4B” family name is not the count of all scalars in this export.
40
 
41
- ## Use
 
 
 
 
 
 
 
 
42
 
43
- Download all model files into one directory. Follow [INFERENCE.md](INFERENCE.md) for the exact prompt and local vLLM settings. The standalone example has now been run with Python 3.11.14, vLLM 0.20.2, Transformers 5.12.1, PyTorch 2.11.0 / CUDA 13.0, and xgrammar 0.1.32. The exact tested environment is recorded in `requirements-smoke.txt`.
 
44
 
45
- The native output is **`elements.layout` with `[x0, y0, x1, y1]` bounding boxes on a 0–1000 grid**. The dataset uses `compositional_deconstruction.nodes` and `[y0, x0, y1, x1]`. Apply the mapping in [OUTPUT_FORMAT.md](OUTPUT_FORMAT.md) before comparing outputs or drawing dataset-format boxes. Spatial relationship records are not directly generated by the model.
46
 
47
- ## Limits and validation
48
 
49
- Predictions can miss elements, assign incorrect parents, misread text, or produce invalid JSON. Long charts may exceed the 20,000-token output limit. Closing truncated JSON does not recover omitted nodes or establish annotation correctness.
50
 
51
- Preparation checks verified every source file checksum, all 738 Safetensors entries and their shard mapping, and the identity of the paper checkpoint. A standalone GPU inference completed on the recorded synthetic training example: 11,485 output tokens, 128 layout items, normal stop, and valid JSON without recovery. The output also passed the native evaluator schema and graph-reference checks. See `smoke/validation.json` and the saved output; its input hash identifies the historical version used for that smoke test. This checks usability and structure; it is not an accuracy evaluation.
52
 
53
- ## License and attribution
54
 
55
- The ChartGalaxy++ fine-tuning contributions are licensed under **CC BY-NC 4.0**, as selected by the model owner. See [LICENSE](LICENSE) and [NOTICE.md](NOTICE.md). Retained Qwen materials remain subject to their upstream Apache-2.0 license, included in [LICENSE-QWEN-APACHE-2.0](LICENSE-QWEN-APACHE-2.0).
56
 
57
- `release_metadata.json` lists the verified model identities; `SHA256SUMS` covers the released files. This package contains model artifacts and usage documentation. Annotation, training, evaluation, and review pipeline implementations are outside its scope.
 
14
  - qwen3_5
15
  ---
16
 
17
+ <h1 align="center">ChartGalaxy++ · Image2JSON</h1>
18
 
19
+ <p align="center"><strong>From infographic images to structured chart representations</strong></p>
20
 
21
+ <p align="center"><a href="https://huggingface.co/datasets/ChartGalaxyPP/ChartGalaxyPlusPlus">Dataset</a> &nbsp; · &nbsp; <a href="https://github.com/ChartGalaxyPP/ChartGalaxyPlusPlus">Project</a> &nbsp; · &nbsp; <a href="INFERENCE.md">Run inference</a> &nbsp; · &nbsp; <a href="OUTPUT_FORMAT.md">Output format</a></p>
22
 
23
+ **Image2JSON identifies chart elements and how they belong together.** Fine-tuned from the **Qwen3.5-4B family** on ChartGalaxy++, it predicts visual elements, recognized text, bounding boxes, appearance attributes, semantic groups, and parent links as structured JSON.
24
 
25
+ ![Qualitative comparisons from the paper showing element localization and grouping by Image2JSON and GPT-6 Astra](assets/image2json-results.png)
26
 
27
+ *Paper examples: separating nearby visual elements and placing marks in the correct semantic group. [Enlarge](assets/image2json-results.png)*
28
+
29
+ ## Results
30
+
31
+ Evaluation on **1,000 infographic charts: 500 real and 500 synthetic**.
32
+
33
+ | Model | Node F1 | Hierarchy F1 | Spatial F1 |
34
+ | --- | ---: | ---: | ---: |
35
+ | GPT-6 Astra | 76.0% | 50.3% | 72.5% |
36
+ | **Image2JSON (ours)** | **89.4%** | **85.1%** | **88.0%** |
37
+
38
+ [Full comparison with 10 baselines](https://github.com/ChartGalaxyPP/ChartGalaxyPlusPlus/blob/main/applications/image_to_scene_graph/paper_table.json) · [Metric definitions](https://github.com/ChartGalaxyPP/ChartGalaxyPlusPlus/blob/main/applications/image_to_scene_graph/metric_definitions.json) · [Saved predictions and scores](https://github.com/ChartGalaxyPP/ChartGalaxyPlusPlus/releases/download/v1.0/image-to-scene-graph.tar.gz)
39
+
40
+ These are the paper's scores, including its output-recovery and graph-evaluation procedures. Spatial relationships are derived from the predicted structure and geometry; they are **not directly generated** by this model.
41
+
42
+ ## Download and run
43
+
44
+ ```python
45
+ from huggingface_hub import HfApi, snapshot_download
46
+
47
+ repo = "ChartGalaxyPP/ChartGalaxyPlusPlus-Image2JSON"
48
+ revision = HfApi().model_info(repo).sha
49
+ snapshot_download(repo, revision=revision, local_dir="image2json-model")
50
+ ```
51
+
52
+ Follow **[the complete single-image inference example](INFERENCE.md)** for the prompt, image preprocessing, and generation settings. Keep the weight shards, tokenizer, processor, and chat template together. The tested runtime uses Python 3.11, vLLM 0.20.2, Transformers 5.12.1, and PyTorch 2.11.0 / CUDA 13.0 on one RTX PRO 6000 GPU; exact dependencies are in [requirements-smoke.txt](requirements-smoke.txt).
53
+
54
+ ## Model output
55
+
56
+ | Output | Contents |
57
  | --- | --- |
58
+ | **Elements** | Text, image, and shape items with bounding boxes and appearance descriptions |
59
+ | **Text** | Recognized strings associated with their visual elements |
60
+ | **Groups** | Semantic units and parent links organizing the chart |
61
+ | **Coordinates** | Native `elements.layout` boxes use **`[x0, y0, x1, y1]`** on a 0–1000 grid |
 
 
 
 
 
62
 
63
+ **The dataset uses a different serialization:** `compositional_deconstruction.nodes` with **`[y0, x0, y1, x1]`** boxes. Use the mapping in [OUTPUT_FORMAT.md](OUTPUT_FORMAT.md) when comparing outputs with dataset annotations.
64
 
65
+ ## Checkpoint
66
+
67
+ | Property | Released artifact |
68
+ | --- | --- |
69
+ | Architecture | `Qwen3_5ForConditionalGeneration` |
70
+ | Fine-tuning | Supervised fine-tuning for chart scene graph prediction |
71
+ | Checkpoint | Step 100,000; full model, not an adapter |
72
+ | Precision | BF16 |
73
+ | Weights | Two Safetensors shards; tokenizer, processor, and chat template included |
74
 
75
+ <details>
76
+ <summary>Validation and limitations</summary>
77
 
78
+ The export matches the checkpoint used for the paper's evaluation. File identities are recorded in [release_metadata.json](release_metadata.json) and [SHA256SUMS](SHA256SUMS).
79
 
80
+ A standalone GPU smoke test produced 128 layout items and 11,485 output tokens, stopping normally with valid JSON without recovery. The native schema and graph-reference checks also passed. See [the validation receipt](smoke/validation.json) and [saved prediction](smoke/prediction.json). This checks loading and output structure, not prediction accuracy; its input hash identifies the historical image version.
81
 
82
+ Predictions may omit elements, misread text, assign incorrect parents, or produce invalid JSON. Dense charts may exceed the 20,000-token output limit. Valid JSON alone does not establish completeness or correctness. The family name “4B” is not the count of all scalars in this export, which also contains vision and auxiliary components.
83
 
84
+ </details>
85
 
86
+ ## License
87
 
88
+ ChartGalaxy++ fine-tuning contributions use **CC BY-NC 4.0**. Retained Qwen materials remain subject to their upstream **Apache-2.0** license. See [LICENSE](LICENSE), [NOTICE.md](NOTICE.md), and [the upstream license](LICENSE-QWEN-APACHE-2.0).
89
 
90
+ This repository contains model artifacts and usage documentation. Annotation, training, evaluation, and review pipeline implementations are not included.
SHA256SUMS CHANGED
@@ -4,7 +4,8 @@ cc97a2c4ff7723a1f25179dc7c18f676fa84191bbb36ea7d0061acaf3c102559 LICENSE
4
  50cbab8a892c5f2993b8c7351a99182507472def3b1374558308605d99b86b32 LICENSE-QWEN-APACHE-2.0
5
  806ea75508761d8cb9515bac61c05be851967b5db755f2223f51a6306d9a1d0d NOTICE.md
6
  7eedb3f5f62834ed0fcdaa889101085b24e5864e13d25c9aecdf7c513d05f951 OUTPUT_FORMAT.md
7
- 5b6f27f2427e3bc8dfee1e03be48162b139d853a17a1f8e42116785201c95b8d README.md
 
8
  a4aee8afcf2e0711942cf848899be66016f8d14a889ff9ede07bca099c28f715 chat_template.jinja
9
  881208c0f3f48d8fe920b4caf1bac23b4e2d351557bf2c1314b2040ceb4c6bba config.json
10
  a50023c63e657eff70e2cfdd42f384eedd085ffa34898aa6d7e34cb16618bbc7 inference_spec.json
 
4
  50cbab8a892c5f2993b8c7351a99182507472def3b1374558308605d99b86b32 LICENSE-QWEN-APACHE-2.0
5
  806ea75508761d8cb9515bac61c05be851967b5db755f2223f51a6306d9a1d0d NOTICE.md
6
  7eedb3f5f62834ed0fcdaa889101085b24e5864e13d25c9aecdf7c513d05f951 OUTPUT_FORMAT.md
7
+ 8e42f268a3afdf273ab566134439d6f9444a376a2d4984bbe8e83d90e1599164 README.md
8
+ 30f4be461845fd9c11ac640b5e1ea04b3c72f37523fd5888b353104b5a77cd4a assets/image2json-results.png
9
  a4aee8afcf2e0711942cf848899be66016f8d14a889ff9ede07bca099c28f715 chat_template.jinja
10
  881208c0f3f48d8fe920b4caf1bac23b4e2d351557bf2c1314b2040ceb4c6bba config.json
11
  a50023c63e657eff70e2cfdd42f384eedd085ffa34898aa6d7e34cb16618bbc7 inference_spec.json
assets/image2json-results.png ADDED

Git LFS Details

  • SHA256: 30f4be461845fd9c11ac640b5e1ea04b3c72f37523fd5888b353104b5a77cd4a
  • Pointer size: 132 Bytes
  • Size of remote file: 1.15 MB