Text Classification
fastText
English
data-cleaning
quality-filtering
ultrax

Line Classifier

A lightweight model built with FastText designed to evaluate individual lines of text and classify them as keep, delete, or edit to improve overall corpus quality.

This project is inspired by and based on an analysis of the UltraX project (it is not officially affiliated with it).

πŸ‹οΈ Training & Methodology

  1. The model was trained using FastText's supervised learning with autotune enabled, utilizing the train and validation splits from agentlans/openbmb-UltraX-Preview-line-classification.
  2. After training, the model was compressed using FastText's built-in quantization (.ftz format) to significantly reduce file size and memory footprint while maintaining robust classification performance.
  3. Evaluated on the independent test split of agentlans/openbmb-UltraX-Preview-line-classification, achieving the following overall metrics:
Metric Score Sample Size (N)
Precision (@1) 0.88 526238
Recall (@1) 0.88 526238

πŸš€ Quick Start

First, download the line_classifier.ftz file from this repository, then load and run predictions using Python:

import fasttext

# Load the model
model = fasttext.load_model("line_classifier.ftz")

# Test the model on sample text lines
texts = [
    "You are the 100th visitor today! Claim your prize NOW!",
    "This lecture series will examine the role of sweepstakes on consumer behaviour."
]

labels, probabilities = model.predict(texts)
print(labels)
print(probabilities)

πŸ“Š Example Predictions

Line of Text Prediction Confidence
Type: Whitepaper Poster delete 0.4536
We now have FIVE dogs! I never thought I would have five dogs, but at least two are small dogs (the Boston Terriers β€” Guinness and Rollie) and two are calm senior dogs (Dusty and Tansy). [...] keep 0.9354
The Gifu Prefecture on the main island of Honshu and the state of Baden-WΓΌrttemberg have long nurtured their political and economic ties... keep 0.9527
## LIVE TICKET FOR 1 edit 0.4407
Nclex Categories Breakdown delete 0.8325

⚠️ Limitations & Considerations

  • Lack of Context: The model operates strictly one line at a time. It lacks broader contextual awareness, which can occasionally make it difficult to determine the true importance of a line.
  • Repetitive Content: Standalone line evaluation means duplicate or repetitive lines across a document may not always be successfully flagged or filtered out.
  • Structured Content: The model may struggle with or mangle heavily formatted text layouts, such as Markdown tables, ASCII art, or structural diagrams.

πŸ“„ License

Distributed under the MIT License. See the LICENSE file for details.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Dataset used to train agentlans/fasttext-line-classifier