AI & ML interests
Open-model inference APIs, OpenAI-compatible endpoints, LLM · vision · speech · video, RTX PRO 6000 GPU Cloud
Recent Activity
Articles
EcoHash
RTX PRO 6000 GPU cloud and OpenAI-compatible APIs for open models.
Chat, vision-language, image generation, speech recognition, speech synthesis, embeddings and reranking — one key, one endpoint, billed by the second. Moving an existing OpenAI integration across is one line:
client = OpenAI(api_key="eco_...", base_url="https://api.ecohash.com/v1")
Every model runs on our own NVIDIA RTX PRO 6000 Blackwell fleet — 96 GB per GPU, US region, not resold from a third party. GPU instances from $1.89 per GPU-hour.
The collections here carry our measured latency, throughput and price for each model. The Spaces let you check those numbers against your own input.
Model catalog · Pricing · Documentation · Open benchmarks · Blog
Create an account — new accounts get $1 of free credit.
-
zai-org/GLM-5.2
Text Generation • 753B • Updated • 928k • • 5.13k -
deepseek-ai/DeepSeek-V4-Pro
Text Generation • 1.6T • Updated • 514k • • 5.59k -
Qwen/Qwen3.6-35B-A3B
Image-Text-to-Text • 36B • Updated • 3.14M • • 2.87k -
Qwen/Qwen3-Coder-30B-A3B-Instruct
Text Generation • 31B • Updated • 550k • • 1.25k
-
zai-org/GLM-5.2
Text Generation • 753B • Updated • 928k • • 5.13k -
deepseek-ai/DeepSeek-V4-Pro
Text Generation • 1.6T • Updated • 514k • • 5.59k -
Qwen/Qwen3.6-35B-A3B
Image-Text-to-Text • 36B • Updated • 3.14M • • 2.87k -
Qwen/Qwen3-Coder-30B-A3B-Instruct
Text Generation • 31B • Updated • 550k • • 1.25k
spaces 6
Vision Chat
Qwen3-VL-8B and gpt-oss-20b with live latency and cost
Speech Recognition Benchmark
Whisper v3 Turbo vs Fun-ASR-Nano, measured side by side
Retrieval Pipeline
Embed, rerank and generate on one key, cost per stage
Image Generation
Z-Image-Turbo and Qwen-Image at $0.01 and $0.03 an image
Text-to-Speech Studio
Kokoro-82M and Qwen3-TTS with measured latency and cost