Part 4
Models, Compute, And Cost
When to spend $200 a month on a frontier model and when to spend $0 a month on hardware in your closet. The math, the trade-offs, and the vocabulary you need to read a model's spec sheet without nodding along to things you don't understand. Sections 4.2, 4.3, and 4.5 are deep-dives; you can skip them on first pass and come back when you're buying hardware.
Start with section 4.1 →- 4.1
Frontier Vs. Local: When Each Makes Sense
The most important budgeting decision you'll make. Not every problem deserves Claude Max; not every problem can be solved with a Mac mini.
- 4.2
How Models Actually Work (weights, Parameters, Active Parameters)
DEEP-DIVE[Deep-dive] You don't need the math. You need the vocabulary to read a model spec sheet without nodding along.
- 4.3
Ram, Vram, And Unified Memory
DEEP-DIVE[Deep-dive] The number that determines whether you can run a model is "how much memory you have, and what kind." The kinds matter.
- 4.4
Self-hosting On Mac (mini, Studio, Pro)
The easiest on-ramp to running models on hardware you own. Mac mini for solo use, Mac Studio for small teams, Mac Pro for the edge cases.
- 4.5
Self-hosting On Nvidia (spark, Orin, Full Gpus)
DEEP-DIVE[Deep-dive] The NVIDIA path. More flexible, more setup, often cheaper for comparable capability. Jetson Orin to DGX Spark to multi-GPU rigs.
- 4.6
Cost Literacy: Tokens, Context Windows, And The $200 Question
The cheat sheet: tokens, input vs output pricing, context windows, and the actual math on whether Claude Max is worth $200 a month.