Add SkillsBench evaluation results

#49
by SaylorTwift HF Staff - opened

Add SkillsBench (Avg5 = 48.2) eval result to .eval_results/, mapped to benchflow/skillsbench task skillsbench_v1_1. Source: model card Benchmark Results > Language > Coding Agent. Setup: OpenCode scaffold, 78 self-contained tasks, avg of 5 runs. See https://huggingface.co/docs/hub/eval-results.

PR Description: Add SkillsBench Evaluation Results for Qwen/Qwen3.6-27B

Summary

This PR adds a SkillsBench evaluation result for Qwen/Qwen3.6-27B to the .eval_results/ directory, following the Hugging Face Hub evaluation-results specification.

Benchmark Added

Model Card Benchmark Score Hub Dataset Task ID
SkillsBench (Avg5) 48.2 benchflow/skillsbench skillsbench_v1_1

Source

Evaluation Setup (from the model card)

SkillsBench: Evaluated via OpenCode on 78 tasks (self-contained subset, excluding API-dependent tasks); avg of 5 runs.

Files Added

  • .eval_results/Qwen3.6-27B.yaml

Verification

These results were extracted from the model card's published HTML benchmark table (text table, not OCR). No verified token is provided as these were not run via HF Jobs with inspect-ai.

Ready to merge
This branch is ready to get merged automatically.

Sign up or log in to comment