How to use from the
Use from the
Transformers library
# Use a pipeline as a high-level helper
from transformers import pipeline

pipe = pipeline("text-classification", model="Chima207/distilbert_amazon_book_classification")
# pip install -U transformers accelerate
# Load model directly
from transformers import AutoTokenizer, AutoModelForSequenceClassification

tokenizer = AutoTokenizer.from_pretrained("Chima207/distilbert_amazon_book_classification")
model = AutoModelForSequenceClassification.from_pretrained("Chima207/distilbert_amazon_book_classification", device_map="auto")
Quick Links

distilbert_amazon_book_classification

This model is a fine-tuned version of distilbert-base-uncased on an Kaggle Amazon Kindle Books dataset. It achieves the following results on the evaluation set:

  • Loss: 1.4475
  • Accuracy: 0.5871
  • F1 Score: 0.5865
  • Precision: 0.5967
  • Recall: 0.5871

Model description

This model is a fine-tuned version of distilbert-base-uncased trained directly on structured Amazon book metadata across 31 standardized Kindle categories.

Serving as a clean-data benchmark in comparative analysis, this model evaluates genre classification performance under structured, editorial metadata conditions. It achieves an Accuracy of 58.71% and a Macro F1-Score of 58.65%, demonstrating that high inherent data quality and structured category labels significantly improve the upper-bound performance of transformer architectures.

Intended uses & limitations

  • Direct classification of structured book descriptions into standard Amazon Kindle categories.

  • Comparative benchmarking for domain adaptation and cross-domain transfer learning experiments.

  • May underperform or show sensitivity when applied to highly informal, uncurated, or user-generated text inputs without prior domain adaptation.

Datasets

Training and evaluation data

  • Source: Official Amazon Kindle book product listings and metadata.

  • Target Taxonomy: 31 standardized Amazon Kindle categories (e.g., Literature & Fiction, Sci-Fi & Fantasy, Romance, Business & Money).

  • Input Features: Concatenated book title and author.

  • Metadata Quality: High quality, structured, and editorially curated product descriptions with minimal noise compared to community-driven tags.

  • Splits: Partitioned into stratified (80/20) train and test sets across all 31 target classes.

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • learning_rate: 2e-05
  • train_batch_size: 4
  • eval_batch_size: 4
  • seed: 42
  • gradient_accumulation_steps: 4
  • total_train_batch_size: 16
  • optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08
  • lr_scheduler_type: linear
  • num_epochs: 2
  • mixed_precision_training: Native AMP

Training results

Training Loss Epoch Step Validation Loss Accuracy F1 Score Precision Recall
1.6436 0.9999 9679 1.4688 0.5680 0.5624 0.5822 0.5680
1.0845 1.9998 19358 1.4475 0.5871 0.5865 0.5967 0.5871

Framework versions

  • Transformers 4.45.2
  • Pytorch 2.5.1
  • Datasets 4.1.1
  • Tokenizers 0.20.1

Academic Context & Citation / Akademischer Kontext

This repository and model were developed as part of a Bachelor's thesis in 2026.

  • Title: Classification of Goodreads genres: A methodological comparison of Doc2Vec and DistilBERT
  • License: CC BY-NC 4.0 (Free for research, education, and personal use; commercial use prohibited)

Dieses Repository und Modell wurden im Rahmen einer Bachelorarbeit im Jahr 2026 entwickelt.

  • Titel: Klassifikation von Goodreads-Genres: Ein methodischer Vergleich von Doc2Vec und DistilBERT
  • Lizenz: CC BY-NC 4.0 (Frei für Forschung, Lehre und private Nutzung; kommerzielle Nutzung untersagt)
Downloads last month
35
Safetensors
Model size
67M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Chima207/distilbert_amazon_book_classification

Finetuned
(12645)
this model
Finetunes
1 model