|
Download README.md from OpenExplorer/InternVL2_5-1B: direct link, hf CLI and curl.
- Browser
- Download file 1.51 kB
-
https://huggingface.co/OpenExplorer/InternVL2_5-1B/resolve/main/README.md
- Command line
-
hf download hf://OpenExplorer/InternVL2_5-1B/README.md
-
curl -L -o README.md https://huggingface.co/OpenExplorer/InternVL2_5-1B/resolve/main/README.md
1.51 kB
| license: other | |
| tags: | |
| - heal | |
| - horizon | |
| # InternVL2.5-1B | |
| **Original model repository:** [OpenGVLab/InternVL2_5-1B](https://huggingface.co/OpenGVLab/InternVL2_5-1B) | |
| ## Model Introduction | |
| InternVL2.5-1B is an instruction-tuned Vision-Language Model (VLM) built on the InternVL2.5 architecture. It combines an InternViT-300M vision encoder, an MLP projector, and Qwen2.5-0.5B-Instruct as its language model. Compared with InternVL2, InternVL2.5 improves multimodal training and data quality for tasks such as OCR, document and chart understanding, visual question answering, visual grounding, and multi-image understanding. | |
| ## Deployment Metrics | |
| ### Model Parameters | |
| | Metric | Value | | |
| |---|---:| | |
| | Total model parameters | 938.2M | | |
| | Vision model (ViT) parameters | 308.5M | | |
| | Language model (LM) parameters | 629.7M | | |
| Parameter counts are calculated from the tensors stored in the upstream checkpoint. | |
| ### Performance Metrics | |
| | Chips | Data Type | ViT Image Size | Sequence Length (tokens) | Maximum Context Length (tokens) | BPU Cores (ViT / Prefill / Decode) | ViT Latency (ms) | TTFT (ms) | Prefill TPS (token/s) | Decode TPS (token/s) | BPU Memory (GB) | CPU Memory (GB) | | |
| |---|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:| | |
| | J6P | W8A8 | 448 × 448 | 512 | 1024 | 4 / 4 / 4 | 29.323 | 58.925 | 19,982.194 | 157.002 | 1.01 | 0.68 | | |
| > **Note:** TTFT includes preprocessing and ViT latency. Memory values represent the peak memory usage measured during the specified performance test. | |