• Task: Multi-dialect Japanese dialect identification
  • The weight decay was drawn from the range 1e-6 to 1e-1.
  • Fifteen trials were conducted using Optuna as the search framework.
  • Input truncation was enabled.

The model released here corresponds to the best-performing trial and checkpoint selected based on macro F1.

Training configuration for this model:

  • Optimizer: AdamW
  • Learning rate: 5e - 5
  • Learning rate scheduler: Linear
  • Batch size: 16
  • Epochs: 3
  • Warmup ratio: 0
  • Warmup steps: 0
  • Random seed: 42
  • Weight decay: 0.07422069766185352
Downloads last month
35
Safetensors
Model size
0.2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for kinlas/hogen-identification-googlebert-ver

Finetuned
(1836)
this model