Benchmark data in "Beyond Ideal Instruction: A Comprehensive Framework for Evaluating LLMs in Realistic Interactions".
AI & ML interests
LLM reasoning
Recent Activity
View all activity
models 7
Miaow-Lab/Qwen3.5-9B-decision-subtrajectory-32k-epoch3
Text Generation • 9B • Updated
Miaow-Lab/Qwen3.5-9B-decision-subtrajectory-32k-epoch2
Text Generation • 9B • Updated
Miaow-Lab/Qwen3.5-9B-decision-subtrajectory-32k-epoch1
Text Generation • 9B • Updated
Miaow-Lab/RLVR-Linearity-Checkpoints
Text Generation • Updated
Miaow-Lab/STT-Agent-RL
196k • Updated • 21 • 1
Miaow-Lab/STT-Agent-SFT
196k • Updated • 9 • 1
Miaow-Lab/SSAE-Checkpoints
Feature Extraction • Updated