MetaReason-SFT

This model is a fine-tuned version of Qwen/Qwen3-VL-8B-Instruct on the MathVRTrain_multimodal_train, the ZKPG_multimodal_train and the MathCanvasInstruct_multimodal_train datasets. It achieves the following results on the evaluation set:

  • Loss: 0.2481

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • learning_rate: 1e-05
  • train_batch_size: 1
  • eval_batch_size: 1
  • seed: 42
  • distributed_type: multi-GPU
  • num_devices: 8
  • gradient_accumulation_steps: 4
  • total_train_batch_size: 32
  • total_eval_batch_size: 8
  • optimizer: Use adamw_torch with betas=(0.9,0.999) and epsilon=1e-08 and optimizer_args=No additional optimizer arguments
  • lr_scheduler_type: cosine
  • lr_scheduler_warmup_ratio: 0.1
  • num_epochs: 2.0

Training results

Training Loss Epoch Step Validation Loss
0.2922 0.2867 500 0.2858
0.2697 0.5734 1000 0.2753
0.2635 0.8601 1500 0.2656
0.2103 1.1468 2000 0.2603
0.2022 1.4335 2500 0.2535
0.2048 1.7202 3000 0.2490

Framework versions

  • Transformers 4.57.1
  • Pytorch 2.6.0+cu124
  • Datasets 4.0.0
  • Tokenizers 0.22.1
Downloads last month
-
Safetensors
Model size
770k params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for pH202411/MetaReason-SFT

Finetuned
(535)
this model