Five checks, in the order that actually matters. Run them and 200 models become five. Most eliminate themselves the moment you apply real constraints. - Hardware: can you run it at the quantization you will actually use, with headroom for your context length? - License: if anything might ship commercially, check now. Apache 2.0 and MIT are green lights. Skip this and regret it later. - Context window: how much the model can consider at once. Nobody checks it early enough. Everybody should. - Benchmarks, read skeptically: good for shortlisting, useless as a final verdict. - Maintenance: is the model family alive? Active projects get fixes, better quantizations, and midnight answers. PrivateLLM deploy handles the whole thing for businesses: private LLM on AWS, $50 setup plus usage, your data stays yours.