Garage Inference: self-hosted AI on one 24 GB GPU Collection Models run by the garage-inference stacks: LLM serving, speech-to-text and voice cloning sharing one GPU. Book and code on GitHub. • 7 items • Updated 8 days ago
FastVideo/FastVideo-FastH3-4-step-Preview-v1-VSA-DataFree Text-to-Video • 35B • Updated 20 days ago • 592k • 314