LLMLingua-2 (XLM-RoBERTa Large) - INT8 ONNX
This repository provides an INT8 dynamically quantized ONNX version of microsoft/llmlingua-2-xlm-roberta-large-meetingbank, exported and quantized via Hugging Face Optimum.
Overview
- Base Model:
microsoft/llmlingua-2-xlm-roberta-large-meetingbank - Task: Token classification (prompt compression)
- Format: ONNX Runtime (Dynamic Quantization, INT8)
- Size: ~116 MB (reduced from ~2.2 GB FP32)
- Target: CPU inference (Node.js, Python, edge proxies)
Purpose
LLMLingua-2 classifies prompt tokens into binary retention labels:
LABEL_1: Preserve token.LABEL_0: Discard token.
Filtering tokens based on these predictions allows reducing prompt length and token usage before calling downstream LLMs, with minimal loss in semantic context.
Usage
Python (Optimum)
from optimum.onnxruntime import ORTModelForTokenClassification
from transformers import AutoTokenizer
model_id = "danielfcramos/llmlingua-2-onnx-quantized"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = ORTModelForTokenClassification.from_pretrained(model_id)
text = "This is a long prompt containing redundant words and phrases."
inputs = tokenizer(text, return_tensors="pt")
outputs = model(**inputs)
predictions = outputs.logits.argmax(dim=-1).squeeze().tolist()
tokens = tokenizer.convert_ids_to_tokens(inputs["input_ids"][0])
compressed = [token for token, label in zip(tokens, predictions) if label == 1]
- Downloads last month
- 12