One developer burned 69 billion tokens in July, roughly $87,000 of API spend across twelve extra Claude Max accounts, to run eighteen agents against a single codebase. His conclusion after an earlier attempt collapsed: agent harnesses have to be bonded into the application, not bought as a framework. (Steve Yegge)
An unreleased OpenAI model solved ten math problems that had gone a decade or more without progress, then wrote each argument into a Lean proof a computer can check line by line. The tokens behind all ten would cost roughly $2,000 at current API rates. (OpenAI)
An AI chatbot out-charmed human scammers, talking nearly half of 22 test subjects into installing an app after a week, versus fewer than one in five for the humans. Running the model through the long trust-building phase and handing off to a person for the final ask slips past vendor safeguards entirely. (Ars Technica)
Language models agree with each other more than twice as much as two human readers do, and no model matched readers better than another reader did. Across 18 models from 11 vendors tested against 2,523 reader highlight sets, only the smallest models picked sentences at human-level agreement. (arXiv)
Microsoft cannot patch as fast as AI finds bugs: Anthropic’s Mythos model turned up 90 critical and 141 important SharePoint flaws in April alone, most of them still unfixed weeks later. July 14 brought patches for more than 600 bugs, an all-time record against June’s previous high of about 200. (ProPublica)