Image-to-Text
PEFT
Safetensors
Portuguese
English
vision-language
table-extraction
scientific-figures
markdown-table
qwen2.5-vl
lora
icdar-metric-loss
Instructions to use lucasoc/sci-image-models with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use lucasoc/sci-image-models with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-VL-3B-Instruct") model = PeftModel.from_pretrained(base_model, "lucasoc/sci-image-models") - Notebooks
- Google Colab
- Kaggle
Download src/utils/io.py from lucasoc/sci-image-models: direct link, hf CLI and curl.
- Browser
- Download file 1.14 kB
-
https://huggingface.co/lucasoc/sci-image-models/resolve/main/src/utils/io.py
- Command line
-
hf download hf://lucasoc/sci-image-models/src/utils/io.py
-
curl -L -o io.py https://huggingface.co/lucasoc/sci-image-models/resolve/main/src/utils/io.py
1.14 kB
| """ | |
| I/O utilities for reading and writing dataset files (JSONL, JSON, CSV). | |
| """ | |
| import json | |
| import os | |
| from typing import Any, Dict, List, Generator | |
| def read_jsonl(file_path: str) -> List[Dict[str, Any]]: | |
| """Reads a JSONL file into a list of dictionaries.""" | |
| records = [] | |
| with open(file_path, "r", encoding="utf-8") as f: | |
| for line in f: | |
| line = line.strip() | |
| if line: | |
| records.append(json.loads(line)) | |
| return records | |
| def iter_jsonl(file_path: str) -> Generator[Dict[str, Any], None, None]: | |
| """Yields records one by one from a JSONL file.""" | |
| with open(file_path, "r", encoding="utf-8") as f: | |
| for line in f: | |
| line = line.strip() | |
| if line: | |
| yield json.loads(line) | |
| def write_jsonl(records: List[Dict[str, Any]], file_path: str) -> None: | |
| """Writes a list of dictionaries to a JSONL file.""" | |
| os.makedirs(os.path.dirname(os.path.abspath(file_path)), exist_ok=True) | |
| with open(file_path, "w", encoding="utf-8") as f: | |
| for record in records: | |
| f.write(json.dumps(record, ensure_ascii=False) + "\n") | |