yintao02.hong commited on
Commit
b3365a1
·
1 Parent(s): 13e0e76

add model card

Browse files
Files changed (1) hide show
  1. README.md +50 -0
README.md ADDED
@@ -0,0 +1,50 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # InternVL2.5-2B
2
+
3
+ **Original model repository:** [OpenGVLab/InternVL2_5-2B](https://huggingface.co/OpenGVLab/InternVL2_5-2B)
4
+
5
+ ## Model Introduction
6
+
7
+ InternVL2.5-2B is an instruction-tuned Vision-Language Model (VLM) built on the InternVL2.5 architecture. It combines an InternViT-300M vision encoder, an MLP projector, and InternLM2.5-Chat-1.8B as its language model. It is designed for visual question answering, OCR, document and chart understanding, visual grounding, image description, and general multimodal dialogue.
8
+
9
+ ## Deployment Metrics
10
+
11
+ ### Model Parameters
12
+
13
+ | Metric | Value |
14
+ |---|---:|
15
+ | Total model parameters | 2.206B |
16
+ | Vision model (ViT) parameters | 316.6M |
17
+ | Language model (LM) parameters | 1.889B |
18
+
19
+ Parameter counts are calculated from the tensors stored in the upstream checkpoint.
20
+
21
+ ### Performance Metrics
22
+
23
+ #### Test Configuration
24
+
25
+ | Metric | Value |
26
+ |---|---|
27
+ | Platform | Matrix6P |
28
+ | Data type | W8A8 |
29
+ | ViT image size | 448 × 448 |
30
+ | Sequence length | 512 |
31
+ | Maximum context length | 1024 |
32
+ | BPU cores (ViT / Prefill / Decode) | 4 / 4 / 4 |
33
+
34
+ #### Performance Results
35
+
36
+ | Metric | Value |
37
+ |---|---:|
38
+ | ViT latency | 41.984 ms |
39
+ | Time to first token (TTFT) | 82.650 ms |
40
+ | Prefill throughput | 14,105.694 tokens/s |
41
+ | Decode throughput | 71.620 tokens/s |
42
+
43
+ #### Memory Usage
44
+
45
+ | Metric | Value |
46
+ |---|---:|
47
+ | BPU memory | 2.4 GB |
48
+ | CPU memory | 0.79 GB |
49
+
50
+ > **Note:** TTFT includes preprocessing and ViT latency. Memory values represent the peak memory usage measured during the specified performance test.