AI & ML interests
Open-model inference APIs, OpenAI-compatible endpoints, LLM · vision · speech · video, RTX PRO 6000 GPU Cloud
Recent Activity
Articles
pinned
Running
Agents
Vision Chat
🚀
Qwen3-VL-8B and gpt-oss-20b with live latency and cost
pinned
Running
Agents
Speech Recognition Benchmark
🚀
Whisper v3 Turbo vs Fun-ASR-Nano, measured side by side
Running
README
🚀
Sleeping
Agents
Retrieval Pipeline
🚀
Embed, rerank and generate on one key, cost per stage
Sleeping
Agents
Image Generation
🚀
Z-Image-Turbo and Qwen-Image at $0.01 and $0.03 an image
Running
Agents
Text-to-Speech Studio
🚀
Kokoro-82M and Qwen3-TTS with measured latency and cost