Chat, coding and vision
Language and vision models on EcoHash, with our own TTFT, per-token latency and precision published next to every number.
Text Generation • 753B • Updated • 682k • • 5.15kNote $1.00 in / $3.00 out per 1M. The step up when an 8B to 35B is not enough. Tool calling and structured output verified against our API.
deepseek-ai/DeepSeek-V4-Pro
Text Generation • 1.6T • Updated • 436k • • 5.64kNote $0.91 in / $2.72 out per 1M. Tool calling and structured output verified against our API.
Qwen/Qwen3.6-35B-A3B
Image-Text-to-Text • 36B • Updated • 3.52M • • 2.94kNote 35B MoE, 3B active. 170 ms TTFT p95, 4 ms per token, ~213 tok/s, FP8. $0.40 in and out per 1M. Larger and faster than our 8B dense model - sparse MoE means parameter count stopped predicting speed.
Qwen/Qwen3-Coder-30B-A3B-Instruct
Text Generation • 31B • Updated • 412k • • 1.27kNote 30B MoE, 3B active. 30 ms TTFT p95 - the lowest in our catalogue - 9 ms per token, ~111 tok/s, BF16. $0.10 in / $0.30 out per 1M, 32k context. TTFT is what agent loops pay on every step of a plan.
openai/gpt-oss-20b
Text Generation • 21B • Updated • 6.19M • • 5.14kNote 20B MoE, 3.6B active. 170 ms TTFT p95, 8 ms per token, ~125 tok/s, MXFP4 as the vendor ships it. $0.20 in / $0.28 out per 1M, 128k context. Peak 8,900 tok/s under concurrency on one card.
Qwen/Qwen3-VL-8B-Instruct
Image-Text-to-Text • 9B • Updated • 10.4M • • 1.17kNote 8B vision-language. 162 ms TTFT p95, 12 ms per token, ~78 tok/s, BF16. $0.15 in / $0.50 out per 1M. Documents, charts, screenshots and photographs.
Vision Chat
🚀Qwen3-VL-8B and gpt-oss-20b with live latency and cost
Note Live demo. Streaming chat with image input, reporting time to first token, throughput, token counts and the exact cost of each reply.