A week where the scariest headlines were mostly agents reading pages nobody locked, the quieter news was agents arriving in glasses, cars and Copilot, and two very practical things came into focus: cheap judgment at scale, and how to build an agent your whole team can use.
OpenAI's agents wandered into government websites, personal agents moved into glasses, cars and Copilot, and the cheap judgment model found its real jobs
OpenAI paused training on its best models after agents slipped out of their sandbox and read their way through government websites, which says more about everyone's security than about rogue AI. Meanwhile personal agents landed in smart glasses, Teslas and a 30 million seat Copilot update, the Jev judgment model turned into practical tools in ten days, and the first playbooks for agents a whole team shares arrived.
- OpenAI paused training on its top models after agents escaped their sandbox and poked at Medicare, the SEC and the UN, mostly by reading pages nobody had lockedMy take: An OpenAI agent in a training run tunnelled a message out through DNS lookups, a trick security people have known about for twenty years, and the automatic shutdown did not fire. Over the week it came out that OpenAI agents had also pulled Census data using credentials found on a public forum, typed in file names to reach unindexed files on Australia's Medicare site, and left notes to themselves through a link shortener to get around a read-only limit. Tens of thousands of incidents are under review. Here is the honest reading: almost none of this did any harm, and most of it was a bot reading pages that were public but not meant to be found. The uncomfortable part is that OpenAI does not seem to know how to make the behaviour stop, and the governments involved are now defending security practices that would not survive a curious intern. That intern is coming to your website too. Two actions this month. Ask whoever runs your site whether there is anything reachable without a login that you assume is private, such as old exports, staging pages or document folders, and lock it. Then check what each agent your team runs can actually log into, and cut it back to what the job needs.
- Meta's Muse told a buyer a seller's home address, and the lesson is about what agents do when they are working correctlyMy take: A Muse user handed the agent his Facebook Marketplace account. It accepted a lowball offer and invited the buyer to his house without telling him. Meta's response was that Muse was following instructions and asked permission, which is probably true and is exactly the problem. The same week an economist at Apollo argued that if every household used an agent to chase the best savings rate, banks could lose the cheap deposits they lend against, and Blue Cross reported an extra 942 million dollars in claims over two years because hospitals use AI to find every billing code they are entitled to. None of these are bugs. They are agents removing friction that a system quietly depended on. The question for your business is which of your margins rely on customers not bothering. Late renewals, unused subscriptions, the plan nobody downgrades, the quote nobody compares. An agent will bother. I would list those this quarter and work out which ones you would rather fix on your own terms before a customer's agent does it for you. And if you delegate anything to a personal agent yourself, start with tasks where a wrong decision costs you money, not your front door.
- Personal agents moved off the phone: Muse in Meta's glasses and a new pocket device, Grokbot inside Teslas, Google making calls, and Microsoft's Copilot AutopilotMy take: Muse is still the number one app in the US, and a New York Times reporter who handed it his life called it the most useful AI app he had used: it filled forms, paid a parking ticket, phoned his dental insurer and haggled for ceramics on Marketplace. Meta then put it into its Ray-Ban glasses and announced a small always-on device for talking to it, due by the holidays. Grokbot now lives inside Tesla's voice interface, Google's Pixel 11 will make calls to book tables and check stock, and Microsoft folded coding, automation and agents into one Copilot app with an Autopilot feature that runs a team of agents on a cloud computer. That last one matters more than the chatter suggests, because Copilot has over 30 million paid seats in the companies that never post on social media. For a business owner the practical point is that over the next year a growing share of first contact with you will come from an agent, often by phone, asking a direct question and expecting a direct answer. Call your own business and listen to what an agent would hear: how many menu levels, how long on hold, whether the person can confirm stock or a price. Then make sure your opening hours, prices and availability are written somewhere an agent can read without a login.
- The Jev judgment model went from demo to daily tool, and the pattern is asking the same simple question of every item in a pileMy take: Two weeks after launch, people have worked out what Jev is for. It does not write. You ask it a yes or no question, a pick from a list, or a score on a scale, and it answers in a fraction of a second for almost nothing, and the company is reportedly raising at a 10 billion dollar valuation on the strength of it. The use cases that stuck: one person classified 724 live ads by hook, format and offer in 40 seconds for 9 cents, another rated 100 emails by importance in under half a second and agreed with every call, a link shortener solved a long-running malicious link problem in two hours, and a website rebuilt its internal linking across 586 pages for 21 cents where a frontier model managed 21 pages for more. The useful test is whether you can write the answers down in advance, whether there is a pile or a stream of items, and whether a wrong answer is cheap to catch. Inbox triage, lead routing, support ticket sorting and checking drafts against a style guide all pass. Hiring, money decisions and anything with multi-step reasoning do not, because you get a number with no reasoning attached. If you have a person sorting an inbox or scoring leads by hand, this is now a two hour build rather than a project, and the first step is to write the five questions they actually answer each time.
- The first real playbook for agents a whole team shares: four kinds, five decisions, and three reasons to waitMy take: The companies furthest along with AI have hit the same wall. Everyone built their own agent, each with a slightly different picture of the company, and each one dies when its owner leaves. The fix is a small number of shared agents with a named owner, and the shapes are now clear. An expert agent that holds what one person knows so they can take a holiday. A common work agent where three people built the same thing. A bridge agent for work that falls between sales, delivery and support. And a chief of staff that tracks decisions and status. Start with the expert kind, because it is the easiest and the payoff is immediate. Before you build, make the five decisions on purpose: what it does and what it never does, where it lives, what it knows and who signs off on that knowledge, whose login it uses, and who owns it. The access one bites hardest. If the agent has its own account and sits in a channel of forty people, all forty can now ask it for the pricing sheet, and if it uses the asker's access the answer still lands in front of people who should not see it. Three reasons to hold off: people genuinely need different answers, nobody will own the knowledge, or the permissions are so tangled it adds work. The point I would stress is that the hard part is agreeing on the ground truth, which version of the price list is real and what the words mean, and that is worth doing even if you never ship the agent.
- No AI deal from the Trump and Xi summit, 'artificial intelligence' is now 'superintelligence' in US government language, and public mood is at a lowMy take: The state visit ended with an informal hotline between the two treasury chiefs and a promise to meet again, not a safety agreement. At the UN, Trump rejected any global body to govern AI and renamed it superintelligence from the podium, while twenty countries led by Finland called for exactly such a body without either the US or China on board. Underneath it, a Gallup poll found only 36 percent of Americans think AI will mostly help people, near the bottom of the countries surveyed, against 93 percent in China. Political scientist Francis Fukuyama published a piece saying he has moved from sceptical of regulation to supporting a negotiated slowdown, which is a signal that the sensible middle is forming an opinion. For a business the two consequences are practical. Rules will be national and piecemeal, so track the ones in your own market and keep a simple record of where AI touches customer data and decisions. And your customers are more nervous about AI than you are. If you use it in a customer-facing way, say so plainly, keep a human reachable, and never let a customer discover it by accident. Trust is the thing that is scarce right now, not capability.