File size: 906 Bytes
8d16bb8
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
---
language:
  - en
library_name: transformers
pipeline_tag: text-generation
license: apache-2.0
base_model: KU-DFI/TelecomGPT-R1
base_model_relation: quantized
tags:
  - fp8
  - compressed-tensors
  - vllm
  - telecom
  - reasoning
---

# TelecomGPT-R1-27B-FP8-Dynamic

FP8 Dynamic quantization of `KU-DFI/TelecomGPT-R1`.

## Quantization

- Scheme: `FP8_DYNAMIC`
- Serialization: `compressed-tensors`
- Target modules: `Linear`
- `lm_head`: unquantized
- Calibration dataset: none
- Source precision: BF16

## Intended use

Telecom reasoning, alarm analysis, root-cause analysis, protocol reasoning,
and evaluation against operator-specific incident datasets.

## A100 note

NVIDIA A100 is an Ampere GPU and does not provide native Hopper-style FP8
Tensor Core execution. The FP8 checkpoint still reduces model-weight memory,
and vLLM can use its supported Ampere execution path when loading the model.