AI
Local models and the end of the API tax
For three years the default assumption was that serious inference happened somewhere else, on hardware you rented by the token. That assumption is now wrong for a widening band of tasks.
Where the line sits
Classification, extraction, rewriting, and most retrieval-augmented question answering run acceptably on a machine you already own. Long-horizon reasoning still does not.
The interesting part is not the benchmark. It is that the tasks below the line no longer generate a log entry on someone else’s server, which changes what you are willing to feed them.