| 16 GiBentry local AI | 3B–8B models; some 12B/14B at aggressive or balanced quantization and short context | chatlight codinglearning | OS pressure arrives quickly; avoid assuming max advertised context. |
| 24 GiBcompact workstation | 8B–14B with generous context; many ~20B-class Q4 deployments | codingRAGsmall agents | 24B-class Q4 can be feasible but leaves less room for long context. |
| 32 GiBstrong mainstream tier | 24B dense Q4; ~30B total-parameter MoE Q4; 8B/14B with very large context | serious codingmultilinguallocal RAG | “Fits in 32 GB” often means moderate, not maximum, context. |
| 48 GiBupper-mid local tier | 30B-class at higher precision; 70B at Q3/Q4 with constrained context | researchagents | 70B Q4 is possible but operational headroom can be narrow on some stacks. |
| 64 GiBlarge-model sweet spot | 70B Q4 with short-to-moderate context; smaller models with long context or concurrency | 70B-classadvanced coding | Full 128K context can invalidate the fit depending on KV geometry and precision. |
| 96 GiBhigh-headroom local | 70B Q5/Q6; 70B Q4 with long context; 100B-class Q4 in selective cases | long contextmulti-agent | Memory bandwidth increasingly dominates usability. |
| 128 GiBpower workstation | 70B at high quantization fidelity; 100B+ Q4; larger MoE models if total weights fit | researchlarge RAG | MoE total weights still matter; “active B” is not a RAM number. |
| 192–256+ GiBspecialist / server class | 100B–400B-class quantized experiments, depending strongly on precision and architecture | labmulti-modelextreme context | Fit becomes a weak proxy for practicality; bandwidth, interconnect and backend support are decisive. |