File size: 1,523 Bytes
d25f18d
 
 
 
 
 
 
9c31a95
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
138192b
 
567e6e8
9c31a95
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
---
license: other
tags:
- heal
- horizon
---

# InternVL3.5-1B-Instruct

**Original model repository:** [OpenGVLab/InternVL3_5-1B-Instruct](https://huggingface.co/OpenGVLab/InternVL3_5-1B-Instruct)

## Model Introduction

InternVL3.5-1B-Instruct is an instruction-tuned Vision-Language Model (VLM) for multimodal perception and reasoning. It uses a vision encoder, an MLP projector, and an autoregressive language model to understand images and generate text. The model is designed for tasks such as OCR, document and chart understanding, visual question answering, multimodal reasoning, spatial understanding, and visual-agent applications.

## Deployment Metrics

### Model Parameters

| Metric | Value |
|---|---:|
| Total model parameters | 1.061B |
| Vision model (ViT) parameters | 309.3M |
| Language model (LM) parameters | 751.6M |

Parameter counts are calculated from the tensors stored in the upstream checkpoint.

### Performance Metrics

| Chips | Data Type | ViT Image Size | Sequence Length (tokens) | Maximum Context Length (tokens) | BPU Cores (ViT / Prefill / Decode) | ViT Latency (ms) | TTFT (ms) | Prefill TPS (token/s) | Decode TPS (token/s) | BPU Memory (GB) | CPU Memory (GB) |
|---|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
| J6P | W8A8 | 448 × 448 | 512 | 1024 | 4 / 4 / 4 | 28.201 | 64.765 | 15,652.844 | 104.543 | 1.6 | 0.76 |

> **Note:** TTFT includes preprocessing and ViT latency. Memory values represent the peak memory usage measured during the specified performance test.