A week where three frontier labs released or announced models inside seven days, the consumer numbers showed how few people actually pay for any of it, and the clearest signal came from the companies already getting returns: build the layer around the model, and keep control of what you feed it.
Three new frontier models in a week and only your own test matters, 98 percent of households still pay nothing for AI, and the companies getting real returns are building a management layer
Google announced Gemini 4 without releasing it, Anthropic shipped a faster Sonnet, and a new model now lands every 11 days on average, which makes a personal test suite more useful than any leaderboard. Underneath the model noise, consumer data showed almost nobody pays for AI yet, enterprises are shifting toward models they control, and the businesses seeing returns share a specific set of habits worth copying.
- Google announced Gemini 4 but you cannot use it, Anthropic's Sonnet 5.5 is fast but hungry, and a new frontier model now arrives every 11 daysMy take: Gemini 4 Argon tops several knowledge-work benchmarks for legal and finance tasks and brings Google back into a three-way race, but Google is holding it back to a small group of security testers because it scored high on a cyber benchmark, so you cannot test it yet. Anthropic's Sonnet 5.5 is a real jump over Sonnet 5, roughly matching Opus 5.5 on many coding tests and claimed to be 30 percent faster and cheaper, but independent testers found it burns so many tokens that a full run cost more than Opus 5.5. Early users describe it as the right tool for well-defined execution, with Opus still better at ambiguous planning. Also note that Google's launch price is a 50 percent discount for an unstated period, so budget on the full rate. The pattern to take from this is that benchmarks and launch prices are both marketing now. Do not move your team's daily work to a new model on announcement day. Wait a week for reports from people doing real work, then test it on your own tasks.
- How to test a new model yourself: four to six of your real tasks, scored blind, before you switch anythingMy take: Published benchmarks are mostly in the training data by now, and the technically best model is often worse at your specific version of a task, especially writing. So build a small test you can rerun every time a model lands. Pick four to six tasks you actually do each week (a client email, a proposal section, a piece of research, one thing you have wanted AI to do but it never got right), save the exact prompt for each, and run them through your current model and the new one in fresh chats. Hide which output came from which model before you score them, because everyone is biased toward the tool they already love, and people who do this blind are regularly surprised by their own picks. Then decide one of three things: switch, split the work between models, or stay. Staying is a fine answer. Cost, speed, your company's approved tools and the habit cost of changing all count. Two hours to set this up once, then twenty minutes per new model, and it ends the anxiety of every launch announcement.
- Only 2.2 percent of US households pay for AI, the top 1 percent of buyers spend 36 times the median, and ChatGPT is now selling picture ads to 1.2 billion weekly usersMy take: New consumer data shows the AI market is two very different groups. The people who pay are heavy users of tools that build, create and automate, spending an average of 903 dollars a month at the top against 25 dollars for the median paying customer. Everyone else, which is 98 percent of households, pays nothing and mostly does not know what the tools can do. There is a loud argument about whether that is early adoption or permanent, and the honest answer is some of both. Two practical points for your business. First, your customers are almost certainly not using AI the way you are, so do not build products or marketing that assume they know what an agent is. Explain the outcome, not the technology. Second, ChatGPT went from launching ads in February to a billion dollar run rate by August and has just added visual ads inside conversations. That is a new paid channel reaching more people than most social networks, and it will be cheap for a while because almost nobody is buying yet. If you spend on Google or Meta ads, ask your agency or ad platform when ChatGPT placements open to businesses your size.
- Businesses are moving toward AI they control: open models took 78 percent of tokens on one major platform, Microsoft and Meta cut their Anthropic spend, and Microsoft's CEO says you are paying for intelligence twiceMy take: The argument has moved from which model is smartest to who owns what. Microsoft's CEO put it plainly: every time you use a hosted model you pay once in money and again in the company knowledge you feed it to get useful answers, and the vendor learns more about you than you learn about it. Companies are acting on this. Open-weight models you can run anywhere took 78 percent of tokens on Vercel's gateway in a record week, a new US lab called Reflection released a competitive open model aimed at letting companies build their own systems on top, and Microsoft is now selling a service that tunes its models on your data so the result belongs to you, claiming 10 times lower cost than a frontier model on a consulting firm's tasks. Meanwhile Microsoft and Meta cut their Anthropic bills sharply by switching to in-house models, a reminder that even the biggest customers swap vendors when the gap closes. You do not need to run your own model. You do need to avoid lock-in. Keep your prompts, examples and process documents in your own files rather than only inside one vendor's workspace, and make sure at least one important workflow could move to a different model in a week. For high-volume routine work, cheap open models are already good enough, and that is where to start saving.
- What the companies getting real returns from AI do differently: shared tooling, cost dashboards, a sovereignty plan, and agents nearly half of staff actually useMy take: A survey of 2,100 senior leaders across 20 countries split companies into those still experimenting and those with proven returns, and the gap reads like a to-do list. Among the proven group, 48 percent have significant staff use of agents against 15 percent for experimenters, 86 percent give people one standard interface for working with AI against 31 percent, 77 percent have a cost dashboard against 43 percent, and 53 percent have a written plan for where their data lives against 8 percent. Planned spend is rising too, from 186 to 210 million dollars on average, while the price of a given amount of AI keeps falling, so companies are buying more for less rather than cutting back. The lesson for a smaller business is that the returns come from the layer around the model, not the model. Pick one standard way your team uses AI so knowledge is shared rather than scattered across personal accounts. Put a simple monthly cost view in place even if it is a spreadsheet. Write one page on what data may go into which tool. And aim for agents your people use every day, not a pilot that two people tried in March. None of this needs a budget. It needs someone to own it.
- Personal agents are everywhere now, switching between them will be expensive, and DoorDash showed what a business-side agent looks likeMy take: Meta's Muse hit three million weekly users in three weeks, OpenAI launched its own called Dots, and there are at least eight serious options. They all connect to your email and apps and browse the web for you, so features will not help you choose. What will: whether it is for personal life or work, how much you care about picking the model, where your data runs and who trains on it, and how you fix its memory when it gets something wrong. The reason to decide carefully is that the value comes from everything it learns about you over months, and moving that to another agent will hurt. My advice is to pick one on those criteria, give it limited access, and run it for a month on tasks where a mistake is cheap. On the other side of the counter, DoorDash now takes orders by text message through its own agent, after finding agent orders carry a 50 percent higher basket for groceries, while staying open to outside agents like Muse. That both-ways stance is the sensible one for most businesses. Make sure an outside agent can find your prices and complete a purchase, and if customers ask you the same things over text, consider whether a simple agent of your own could answer and take the order.