InternVL2-2B

Original model repository: OpenGVLab/InternVL2-2B

Model Introduction

InternVL2-2B is an instruction-tuned Vision-Language Model (VLM) for image understanding and text generation. It combines an InternViT-300M vision encoder, an MLP projector, and InternLM2-Chat-1.8B as its language model. Typical applications include visual question answering, OCR, image description, document and chart understanding, and multimodal dialogue.

Deployment Metrics

Model Parameters

Metric Value
Total model parameters 2.206B
Vision model (ViT) parameters 316.6M
Language model (LM) parameters 1.889B

Parameter counts are calculated from the tensors stored in the upstream checkpoint.

Performance Metrics

Chips Data Type ViT Image Size Sequence Length (tokens) Maximum Context Length (tokens) BPU Cores (ViT / Prefill / Decode) ViT Latency (ms) TTFT (ms) Prefill TPS (token/s) Decode TPS (token/s) BPU Memory (GB) CPU Memory (GB)
J6P W8A8 448 × 448 512 1024 4 / 4 / 4 41.535 84.448 13,265.273 70.473 2.4 0.79

Note: TTFT includes preprocessing and ViT latency. Memory values represent the peak memory usage measured during the specified performance test.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including OpenExplorer/InternVL2-2B