Roughly 40GB at 4-bit quantization, on a single high-end GPU. That one number decides whether your project is feasible or fantasy, so it is up front. - Full precision wants around 70GB, which is why almost nobody runs a 70B that way. - At 4-bit, quality stays surprisingly close to the original for most tasks. Go lower and the quality cost stops being subtle. - Context length eats memory too. Crank the context up and the KV cache needs headroom past the base number. - Size honestly and round up. Running out of memory mid-inference is the most annoying way to learn any of this. PrivateLLM deploy handles the whole thing: private LLM on AWS, $50 setup plus usage, sized right from day one.