yintao02.hong commited on
Commit
307be18
·
1 Parent(s): d82b20a

add model card

Browse files
Files changed (1) hide show
  1. README.md +50 -0
README.md ADDED
@@ -0,0 +1,50 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # InternVL2-1B
2
+
3
+ **Original model repository:** [OpenGVLab/InternVL2-1B](https://huggingface.co/OpenGVLab/InternVL2-1B)
4
+
5
+ ## Model Introduction
6
+
7
+ InternVL2-1B is an instruction-tuned Vision-Language Model (VLM) for understanding images and generating text responses. It uses an InternViT-300M vision encoder, an MLP projector, and Qwen2-0.5B-Instruct as its language model. The model can be used for visual question answering, image description, OCR, document and chart understanding, and general multimodal dialogue.
8
+
9
+ ## Deployment Metrics
10
+
11
+ ### Model Parameters
12
+
13
+ | Metric | Value |
14
+ |---|---:|
15
+ | Total model parameters | 938.2M |
16
+ | Vision model (ViT) parameters | 308.5M |
17
+ | Language model (LM) parameters | 629.7M |
18
+
19
+ Parameter counts are calculated from the tensors stored in the upstream checkpoint.
20
+
21
+ ### Performance Metrics
22
+
23
+ #### Test Configuration
24
+
25
+ | Metric | Value |
26
+ |---|---|
27
+ | Platform | Matrix6P |
28
+ | Data type | W8A8 |
29
+ | ViT image size | 448 × 448 |
30
+ | Sequence length | 512 |
31
+ | Maximum context length | 1024 |
32
+ | BPU cores (ViT / Prefill / Decode) | 4 / 4 / 4 |
33
+
34
+ #### Performance Results
35
+
36
+ | Metric | Value |
37
+ |---|---:|
38
+ | ViT latency | 41.451 ms |
39
+ | Time to first token (TTFT) | 65.766 ms |
40
+ | Prefill throughput | 24,955.227 tokens/s |
41
+ | Decode throughput | 159.481 tokens/s |
42
+
43
+ #### Memory Usage
44
+
45
+ | Metric | Value |
46
+ |---|---:|
47
+ | BPU memory | 1.01 GB |
48
+ | CPU memory | 0.68 GB |
49
+
50
+ > **Note:** TTFT includes preprocessing and ViT latency. Memory values represent the peak memory usage measured during the specified performance test.