Quantization sounds like scary math. It is not: weights are numbers, numbers can live in fewer bits, fewer bits means less memory. The catch is the rounding costs a little precision. - At 4-bit, good models keep most of their quality while shrinking a lot. That is why 4-bit is the default. Most local setups live there and never look back. - Below 4-bit the deal changes: casual chat fine, code and precise reasoning suffer. Match the quantization to the task, not your ambition. - Quantization is not the only memory consumer. A giant context window can blow past your VRAM on its own. - Headroom is cheap insurance against the weirdest errors you will ever debug. Want it running without the setup? PrivateLLM deploy puts a private LLM on your AWS for $50 plus usage.