Ad
Skip to content
Read full article about: Top mathematicians say LLMs are strong calculators but poor creative thinkers

LLMs can't jump, Part II. Mathematicians Timothy Gowers and Peter Sarnak credit large language models with serious math skills but see hard limits for genuinely new ideas. Gowers argues current models are good at combining known methods and trying many search paths but lack the intuition to pick the few productive routes in a vast search space. Sarnak agrees: AI can derive results from existing theory but fails to develop the abstractions that underpin major proofs when starting from an elementary question.

DeepMind researcher Tom Zahavy reached a similar conclusion. In his paper "LLMs Can't Jump," he pins the bottleneck on "manipulative abduction," the ability to invent new foundational assumptions with no linguistic precedent. World models could offer a way forward. These assessments feed into a broader debate about whether LLMs are actually becoming more versatile or "just" getting better at benchmarks and familiar problem spaces.

Comment Source: AMS | Gowers
Read full article about: OpenAI dissolved the team built to catch catastrophic AI risks, reassigning its work to other groups

OpenAI shut down its "Preparedness" team at the end of July. The team evaluated whether the company's AI models could pose serious or catastrophic risks, the Financial Times reports, citing internal sources. Its work on biological and cyber risks has been parceled out to existing teams.

Former unit lead Dylan Scandinaro now focuses on safety risks from "recursively self-improving" AI, systems that can optimize themselves and train other models. Co-founder Greg Brockman said OpenAI has woven safety work more tightly into model development.

Several safety staffers have left recently, including Chief Ethics Officer Chloe Bakalar and Joshua Achiam. Internally, unease is building. One source described a "burbling sense of responsibility and dread" that OpenAI isn't doing enough on safety. Employees have also spoken up publicly, especially after the autonomous hacking incident involving Hugging Face. One employee said he hoped OpenAI would treat it as a "warning shot."

Ad
Read full article about: Anthropic's bio-weapons filter was down for nearly a year, exposing 133 million requests

From the company whose CEO calls AI-assisted development of chemical and biological weapons a bigger threat than cyberattacks comes a safety report revealing that Anthropic's blocking biological classifiers were inactive from May 2025 through April 2026. These filters are designed to prevent AI models from being used to extract dangerous knowledge about chemical or biological weapons. For almost a year, all traffic from external contractors providing human feedback ran without them.

Image
Screenshot

The gap affected a pool of about 50,000 people who ran roughly 133 million chats with the models. According to Anthropic, these individuals were vetted only by external vendors whose screening processes were often insufficient. Anthropic says its internal investigation turned up no evidence of actual misuse. The company has since tightened contractor requirements.

Anthropic also recently loosened its classifiers on Fable 5 after researchers complained the filters were so aggressive they blocked legitimate research.

Optima tackles AI benchmarking's biggest flaw by letting users test models against their own data

Artificial Analysis has launched Optima, a platform that lets users build custom AI benchmarks from their own data and workflows. Models can be compared not just on quality but also on cost and time per task. For agent-based applications, those metrics often tell you more than raw token pricing.

Ad

Investor pressure forces Nvidia to shrink its OpenAI bet just as Anthropic's numbers defy bubble warnings

Nvidia has cut its guarantee for OpenAI’s planned data center in Ohio nearly in half, from $250 billion to just under $120 billion, after investors pushed back on the risk. Meanwhile, Anthropic is complicating the AI bubble debate with revenue that jumped from $4.7 billion to $11.5 billion in a single quarter.

Plaintiff hid invisible AI instructions in court filings to secretly influence automated review

A plaintiff in Connecticut embedded invisible prompt injections in court filings, formatted in 3-point white text on a white background, to manipulate a potential AI review system. Judge Spader compared the attempt to secretly tampering with a jury and revoked the plaintiff’s electronic filing privileges. The court stressed that Connecticut doesn’t use AI to review filings, but the intent alone was enough to warrant sanctions.

Ad

World Labs turns one real-world robot task into thousands of simulated variations for training

World Labs, the startup founded by AI pioneer Fei-Fei Li, has unveiled a simulation engine that trains robot controllers entirely in virtual environments. From a single real-world task, the system generates thousands of controlled variations. The trained models then ran for one hour each on five different robot platforms without human intervention. How well the results hold up in more complex everyday situations remains to be seen.