Spaces:
Running
Running
File size: 808 Bytes
e2b5f64 d879501 e2b5f64 d879501 e2b5f64 d879501 e2b5f64 d879501 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 | ---
title: DAREBench
emoji: 🦞
colorFrom: red
colorTo: blue
sdk: static
app_file: index.html
pinned: true
short_description: Deployment-Aware and Reliable Evaluation of Models as Agents
tags:
- benchmark
- agents
- openclaw
- evaluation
- leaderboard
datasets:
- SeerRay-Lab/DAREBench
thumbnail: https://seerray-lab-darebench.static.hf.space/assets/img/social-card.png
---
# DAREBench: Deployment-Aware and Reliable Evaluation of Models as Agents
Project page for DAREBench. A workload- and deployment-aware benchmark for evaluating models as agents.
- 📄 Paper: [arXiv 2609.06059](https://arxiv.org/abs/2609.06059)
- 💻 Code: [github.com/SeerRay-Lab/DareBench](https://github.com/SeerRay-Lab/DareBench)
- 🤗 Dataset: [SeerRay-Lab/DAREBench](https://huggingface.co/datasets/SeerRay-Lab/DAREBench)
|