LLMLingua-2 (XLM-RoBERTa Large) - INT8 ONNX

This repository provides an INT8 dynamically quantized ONNX version of microsoft/llmlingua-2-xlm-roberta-large-meetingbank, exported and quantized via Hugging Face Optimum.

Overview

  • Base Model: microsoft/llmlingua-2-xlm-roberta-large-meetingbank
  • Task: Token classification (prompt compression)
  • Format: ONNX Runtime (Dynamic Quantization, INT8)
  • Size: ~116 MB (reduced from ~2.2 GB FP32)
  • Target: CPU inference (Node.js, Python, edge proxies)

Purpose

LLMLingua-2 classifies prompt tokens into binary retention labels:

  • LABEL_1: Preserve token.
  • LABEL_0: Discard token.

Filtering tokens based on these predictions allows reducing prompt length and token usage before calling downstream LLMs, with minimal loss in semantic context.

Usage

Python (Optimum)

from optimum.onnxruntime import ORTModelForTokenClassification
from transformers import AutoTokenizer

model_id = "danielfcramos/llmlingua-2-onnx-quantized"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = ORTModelForTokenClassification.from_pretrained(model_id)

text = "This is a long prompt containing redundant words and phrases."
inputs = tokenizer(text, return_tensors="pt")
outputs = model(**inputs)

predictions = outputs.logits.argmax(dim=-1).squeeze().tolist()
tokens = tokenizer.convert_ids_to_tokens(inputs["input_ids"][0])
compressed = [token for token, label in zip(tokens, predictions) if label == 1]
Downloads last month
12
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for danielfcramos/llmlingua-2-onnx-quantized

Quantized
(3)
this model