Garage Inference: self-hosted AI on one 24 GB GPU Collection Models run by the garage-inference stacks: LLM serving, speech-to-text and voice cloning sharing one GPU. Book and code on GitHub. • 7 items • Updated 5 days ago
FastVideo/FastVideo-FastH3-4-step-Preview-v1-VSA-DataFree Text-to-Video • 35B • Updated 16 days ago • 356k • 308