Own the Harness, Rent the Model: The Quantified Case for AI's Great Rebalancing
Open-weight models now trail the closed frontier by only ~3 points on average, and with 4-bit quantization costing just 1.5–2% quality, routing 80% of traffic to self-hosted hardware cuts inference bi…