Know your VRAM
before you deploy
Real memory requirements for the models everyone is running, and the exact Locally server that fits each one. No marketing math.
Llama 3.1 8B
MetaThe default starting point. Fast, capable, runs everywhere.
Llama 3.1 70B
MetaGPT-4-class answers on your own machine. The reason people rent 4090s.
DeepSeek R1 70B
DeepSeekOpen reasoning model. Chain-of-thought that rivals closed labs.
Qwen 2.5 32B
AlibabaThe sweet spot: near-70B quality at half the memory bill.
Mistral 7B
MistralTiny, quick and surprisingly sharp. Happy even on CPU.
Gemma 2 27B
GoogleGoogle's open weights. Great multilingual chat under 24 GB.
Flux.1 dev
BFLState of the art open image generation. Text rendering that works.
Stable Diffusion XL
StabilityThe workhorse of image gen. Massive LoRA and ControlNet ecosystem.
Whisper large-v3
OpenAIBest-in-class transcription in 99 languages, fully offline.
XTTS v2
CoquiVoice cloning and TTS in 17 languages from 6 seconds of audio.
Qwen 2.5 Coder 32B
AlibabaOpen code model that trades blows with closed assistants.
DeepSeek Coder V2 16B
DeepSeekLight MoE coder. Fast completions without a monster GPU.