NEW

GPU instances are live. Use code LOCAL10 for 10% off your first month.

Deploy now
model library

Know your VRAM
before you deploy

Real memory requirements for the models everyone is running, and the exact Locally server that fits each one. No marketing math.

Llama 3.1 8B

Meta

The default starting point. Fast, capable, runs everywhere.

Q4 quant5 GB VRAM
FP1616 GB VRAM
Fits RTX 3060 at $59/mo · 60+ tok/s
Deploy for this model

Llama 3.1 70B

Meta

GPT-4-class answers on your own machine. The reason people rent 4090s.

Q4 quant20 GB VRAM
FP16140 GB VRAM
Fits RTX 4090 at $129/mo · 40+ tok/s
Deploy for this model

DeepSeek R1 70B

DeepSeek

Open reasoning model. Chain-of-thought that rivals closed labs.

Q4 quant20 GB VRAM
FP16140 GB VRAM
Fits RTX 4090 at $129/mo · full precision on A100
Deploy for this model

Qwen 2.5 32B

Alibaba

The sweet spot: near-70B quality at half the memory bill.

Q4 quant18 GB VRAM
FP1664 GB VRAM
Fits RTX 3090 at $115/mo · best value
Deploy for this model

Mistral 7B

Mistral

Tiny, quick and surprisingly sharp. Happy even on CPU.

Q4 quant4 GB VRAM
CPU RAM8 GB
Runs on VPS Dev at $12.99/mo via Ollama
Deploy for this model

Gemma 2 27B

Google

Google's open weights. Great multilingual chat under 24 GB.

Q4 quant16 GB VRAM
FP1654 GB VRAM
Fits RTX 4060 Ti at $79/mo
Deploy for this model

Flux.1 dev

BFL

State of the art open image generation. Text rendering that works.

FP823 GB VRAM
GGUF Q816 GB VRAM
Fits RTX 4090 at $129/mo · seconds per image
Deploy for this model

Stable Diffusion XL

Stability

The workhorse of image gen. Massive LoRA and ControlNet ecosystem.

Base10 GB VRAM
+ refiner16 GB VRAM
Fits RTX 4060 Ti at $79/mo
Deploy for this model

Whisper large-v3

OpenAI

Best-in-class transcription in 99 languages, fully offline.

GPU5 GB VRAM
CPU RAM10 GB
Fits GTX 1660 at $39/mo · 1h audio in minutes
Deploy for this model

XTTS v2

Coqui

Voice cloning and TTS in 17 languages from 6 seconds of audio.

GPU6 GB VRAM
CPU RAM12 GB
Fits GTX 1660 at $39/mo
Deploy for this model

Qwen 2.5 Coder 32B

Alibaba

Open code model that trades blows with closed assistants.

Q4 quant18 GB VRAM
FP1664 GB VRAM
Fits RTX 3090 at $115/mo · self-hosted Copilot
Deploy for this model

DeepSeek Coder V2 16B

DeepSeek

Light MoE coder. Fast completions without a monster GPU.

Q4 quant10 GB VRAM
FP1632 GB VRAM
Fits RTX 3060 at $59/mo
Deploy for this model
Not sure which one? Use the plan finder and get matched in one click, or ask us, we run these daily.