README / README.md
brady777's picture
Link tau2-bench leaderboard PR #480
ce06151 verified
|
Raw History Blame Contribute Delete
1.58 kB
---
title: README
sdk: static
pinned: false
---
CloudSurf Software builds **[Surfspace](https://surf.space)** β€” the AI workspace that turns chats into apps from any device β€” and trains **CloudSurf** models: small, US-trained, open-weight models for function calling and agentic tool use, released under Apache-2.0.
## Models
- **[CloudSurf-4B-FC](https://huggingface.co/cloudsurf-software/CloudSurf-4B-FC)** β€” Gemma-4 E4B (4B active / ~8.0B total params) fine-tuned for function calling. BFCL V4 FULL, official harness, 3-seed mean: **55.73** vs 34.81 for the stock base β€” above the published small-model class bar (51.40). Self-run numbers with full settings disclosure; leaderboard submission open at [gorilla PR #1357](https://github.com/ShishirPatil/gorilla/pull/1357).
## Evidence, not just claims
Every score we publish ships with the raw artifacts to check it:
- [CloudSurf-4B-FC-bfcl-results](https://huggingface.co/datasets/cloudsurf-software/CloudSurf-4B-FC-bfcl-results) β€” raw BFCL V4 result files for our runs *and* the stock baselines, with per-category scores.
- [tau2-trajectories-cloudsurf-4b-fc](https://huggingface.co/datasets/cloudsurf-software/tau2-trajectories-cloudsurf-4b-fc) β€” full τ²-bench evaluation trajectories backing our leaderboard submission, open at [tau2-bench PR #480](https://github.com/sierra-research/tau2-bench/pull/480).
Model cards disclose both comparison frames, measured eval-noise bands, and known limitations.
## Links
- Product: [surf.space](https://surf.space)
- Company: [cloudsurf.com](https://cloudsurf.com)