Submission + - Google Plans to Exempt Sanctioned Nations From Android Developer Verification (arstechnica.com)

An anonymous reader writes: We are a month away from the initial rollout of Google’s Android developer verification system, and the company contends this policy does not impinge on the platform’s open nature. Still, the restrictions will be a big change, and there are still some unanswered questions. An issue that has come up repeatedly in the run-up to verification is what will happen to devs who can’t verify because of where they live. It turns out that Google has a cryptic answer for that buried in an FAQ. Developer verification will soon block the installation of apps from unverified developers on any Android device running Google services, which is functionally all Android phones outside Russia and China. Developers who want to keep releasing software, even if it’s not in the Play Store, have to provide Google with their ID and pay a small fee.

But what if you’re an Android developer living in a sanctioned nation? Currently, the U.S. sanction list includes Iran, Cuba, North Korea, and occupied areas of Ukraine. Given the current uncertain state of US foreign policy, that list could change in the future. Google doing any business with developers in those places is a thorny issue, and it seems like the company has decided to just leave them hanging. A rather lengthy FAQ a few levels deep on the Google developer site addresses various issues around dev verification. Smack in the middle is this: "How does this program impact developers in sanctioned countries? Devices in sanctioned countries will be excluded from Android developer verification checks. This allows any developer to continue distributing apps in these regions without verification, though users there won’t benefit from the enhanced security benefits of the program."

[...] A Google spokesperson has expanded on the FAQ and confirmed to Ars that people living in sanctioned nations will not be allowed to go through the verification process. That means they will not be able to effectively distribute software through any channel internationally. Today, someone making an app in, say, Cuba can distribute it freely around the world, as well as at home. Anyone can install it and see their work in action after tapping through a few sideloading alerts. In the coming months, that will no longer be the case. These unverified apps will only be easily installable in the sanctioned countries where verification doesn’t exist.

Submission + - OpenAI Finds Evidence Other AI Agents Escaped Containment (reuters.com)

An anonymous reader writes: OpenAI has discovered other instances in which autonomous agents have escaped containment as the company expands its investigation of the hacking incident at tech firm Hugging Face that drew global attention this month, two people familiar with the matter said on Friday. The new breakouts were uncovered during the company'spublicly announced investigationinto how one of its agents escaped what was meant to be a contained testing environment this month, the two people said, and OpenAI is now looking into those instances as well. One of the sources said that the escapes were limited in nature and that none of the agents were thought to have left OpenAI's network.

An OpenAI spokesperson referred toa statement issued by the companyon Tuesday that said it was reviewing "broader activity from our models" in addition to the Hugging Face intrusion. The discovery of additional rogue behavior at OpenAI, even if limited in nature, could feed growing appetite for regulation coming out of the White House and elsewhere. The expanded investigation by OpenAI was launched shortly before its primary rival, Anthropic, disclosed that its models were also responsible fora series of break-insthat led to breaches at three other companies dating back to April, according to the two sources and a third source familiar with the matter. The recent discovery of other past breakouts at OpenAI has not previously been reported.

AI safety experts said the new disclosures paint a portrait of a group of cutting-edge labs whose ability to develop dangerous autonomous hacking agents outstrips their ability to keep them under control. "We have a whole industry where the people designing, developing and putting out these tools aren't keeping up themselves to responsibly develop these things and keep them safe," said Maurice Chiodo, a mathematician who works at Cambridge University's Center for the Study of Existential Risk. Reuters could not establish exactly how many incidents OpenAI investigators found or the timings or circumstances under which they occurred. The three sources said OpenAI and outside experts were examining log data from earlier in the year in a bid to understand what took place.

Submission + - Innocent man jailed 18 months on basis of mistyped userid. (ctvnews.ca)

belmolis writes: A Nova Scotia man was falsely convicted of child pornography and spent 18 months in prison because Wisconsin police mistyped his Kik userid, which led them to his IP address. Nova Scotia police found nothing incriminating on his devices and no other evidence; he was convicted solely on the basis of the connection between the mistyped userid and his IP address.

Submission + - Drifting Falcon 9 upper stage heading for accidental collision with the Moon (space.com)

fahrbot-bot writes: Space and The Guardian are reporting that the Falcon 9 upper stage leftover from the launch of the Firefly Blue Ghost-1 lander on Jan. 15, 2025 is due to impact the Moon on Aug. 5, 2026.

Onboard the same flight was the Hakuto-R Mission 2, called Resilience, a robotic lunar lander developed by the Japanese company ispace.

According to a new study by an international team, the resulting impact plume may briefly be bright enough to see against the dark sky near the moon's edge. That means it might be visible to moongazers with sufficiently sensitive telescopes.

This head-on collision of the errant stage is expected to occur near the Einstein and Bell craters near the western lunar limb. It may well be visible to ground and space-based assets.

Submission + - Publishers are losing Google traffic as AI answers replace links (axios.com)

alternative_right writes: Google has basically stopped sending people to websites (including our site) for answers and information. Instead, it's using AI to answer them on its platform, in its words.

Chartbeat data shared with Axios shows Google Search traffic to publishers fell 34% over the past year.

That pain is regressive. Over the past two years, small publishers lost 60% of referrals from search overall, medium publishers 47%, large publishers 22%.

Submission + - FBI gets voter's IP address in new fraud probe tactic (axios.com)

alternative_right writes: The Trump administration has a new tactic for trying to isolate cases of alleged voter fraud â" digging into the IP addresses of those who went online to register to vote.

The FBI recently obtained the IP address of someone who registered online in South Carolina, according to documents first shared with Axios.

Submission + - COLDCARD flaw may have exposed Bitcoin wallet seeds for years (nerds.xyz)

BrianFagioli writes: A firmware flaw affecting COLDCARD hardware wallets may have caused some devices to generate predictable Bitcoin wallet seeds. Block researchers traced the issue to firmware that used a deterministic software generator instead of the STM32 processor hardware random number generator.

Coinkite says COLDCARD Mk3 devices that generated seeds while running firmware version 4.0.1 through 5.0.3 may be affected. Installing newer firmware does not repair an existing seed, so users must create a new seed on an unaffected device and transfer their Bitcoin. Block also questions whether newer COLDCARD models provide enough secure randomness, although Coinkite disputes that those devices are affected.

Submission + - Australia sues Telegram for terrorist content

An anonymous reader writes: Australia sues Telegram for terrorist content

“Australia has taken legal action against Telegram over alleged failures to tackle terrorism-related content. Just the day before, Russia launched its own case against the messaging platform and its founder, Russian-born entrepreneur Pavel Durov.”

Submission + - The Loss of Situational Awareness (theverge.com)

joshuark writes: Situational Awareness, the hedge fund started by a 24-year-old former OpenAI employee that focuses on artificial intelligence bets, has sold most or all, depending on who’s reporting, of its entire public stock portfolio to Ken Griffin’s Citadel after several bad weeks for AI stocks, and that’s the situation we are all now aware of. You may recall earlier this week I noted the market had gotten particularly nervous about AI risk; as it turns out, we have discovered one firm that was swimming without a bathing suit.

Situational Awareness had a staff of eight, of whom four were investment professionals. “The fund’s largest holdings at the end of the first quarter included Nebius Group, Sandisk, Micron and CoreWeave, according to filings,” CNBC wrote. “All four of those stocks are down more than 35 percent this month.” I expect we will hear more in the coming days, especially from the Wall Street professionals who were on the other side of these jokers’ trades.

How did we get here? Situational Awareness LP was named for a series of facile essays about machine intelligence published by the improbably named Leopold Aschenbrenner, the 24-year-old mastermind of the hedge fund. “We are building machines that can think and reason,” he writes, betraying that he has no idea what thinking could possibly mean. “By 2025/26, these machines will outpace many college graduates. By the end of the decade, they will be smarter than you or I; we will have superintelligence, in the true sense of the word. Along the way, national security forces not seen in half a century will be unleashed, and before long, The Project will be on. If we’re lucky, we’ll be in an all-out race with the CCP; if we’re unlucky, an all-out war.”

Situational Awareness’ backers included Patrick and John Collison, who cofounded Stripe, and two Meta AI leaders, Daniel Gross and Nat Friedman. The fund’s director of research was Carl Shulman, who’d worked at Peter Thiel’s Clarium Capital. Eventually, Jane Street — the Wall Street firm budding young Effective Altruists, including Sam Bankman-Fried, join — bought in too. “Jane Street’s investment in Situational Awareness is particularly notable because the firm rarely allocates capital to outside money managers,” The Wall Street Journal wrote in June.

Why would these purportedly serious people buy in on a 24-year-old’s very first hedge fund? My best guess is that Aschenbrenner’s investors were relying on the social bona fides he had cultivated. Social proof is the laziest and most disastrous way to vet people — ask any Theranos investor, or for that matter, anyone who had Bernie Madoff managing their money — but I suppose it’s good enough for Silicon Valley.

“Basically, this investment firm will be kind of like a brain trust on AI,” Aschenbrenner told Dwarkesh Patel in a four-hour podcast interview, the preferred intellectual medium of the Silicon Valley elite. “We’re going to have way more situational awareness than any of the people who manage money in New York. We’re definitely going to do great on investing, but it’s the same sort of situational awareness that is going to be important for understanding what’s happening, being a voice of reason publicly, and being able to be in a position to advise.”

At age 17, Aschenbrenner was called “an economics prodigy” by Tyler Cowen, a libertarian economist famous in certain Silicon Valley circles. Cowen’s Emergent Ventures even gave him a grant, according to Fortune. Aschenbrenner published essays in Works in Progress, a publication funded by Stripe. During his time at Columbia University, Aschenbrenner cofounded the college’s Effective Altruism chapter. After graduating in 2021 as Columbia University’s valedictorian at age 19, Aschenbrenner went on to work at the FTX Future Fund, the philanthropic arm of cryptocurrency exchange FTX, which collapsed after Sam Bankman-Fried’s fraud was revealed. Among his coworkers at the fund were William MacAskill, the philosopher-king of the Effective Altruism movement, and Avital Balwit, who would later become the chief of staff at Anthropic.

Here’s how Aschenbrenner described the fund’s strategy back in the halcyon days of 2024: “Obviously, not blowing up is task number one and two,” he told Patel. “You have to get the timing right. The sequence of bets on the way to AGI is actually pretty critical. People underrate it.”

We’ll get a more complete picture of how Situational Awareness crashed and burned in the coming days, but right now it looks like this. Hedge funds often borrow money to maximize their bets. So if you really believe AI is the future, “you won’t put 100% of your money (and your investors’ money) into the AI boom,” writes Bloomberg’s Matt Levine. “You’ll put, like, 300% of your money into the AI boom.” At one point, the hedge fund claimed to be up 439 percent.

“You’ve got to be really, really careful about your overall risk positioning,” Aschenbrenner said in 2024. “If you expect these crazy events to play out, there’s going to be crazy things you didn’t foresee.” One of those things, perhaps, is that artificial general intelligence isn’t coming — or at least, not by 2027. “A friend joked that the investment firm is perfectly hedged for me,” Aschenbrenner said. “Either AGI happens this decade and my human capital depreciates, but I turn it into financial capital, or no AGI happens and the firm doesn’t do well, but I’m still in my twenties and smart.”

Perhaps Aschenbrenner should have quoted The Simpson's character, Professor Frink:

https://www.youtube.com/watch?...

"I forgot to carry the one."

Submission + - Does Growing Food Reduce your Petroleum Foorprint More Than an EV? (youtube.com) 1

Skystrider writes: Paul Wheaton, author and gardener, asks if we can really reduce our petroleum footprint.

explore practical ways to reduce our petroleum footprint beyond simply driving less or buying a more efficient vehicle. By looking at where petroleum is actually used—from transportation to food production—the discussion challenges conventional advice and highlights lifestyle changes that can have a far greater impact.

Instead of focusing on sacrifice, are there solutions that improve quality of life while dramatically reducing petroleum consumption? Sharing resources, growing food, shortening supply chains, and building stronger local communities can reduce thousands of gallons of petroleum use each year, demonstrating that regenerative living can be both more resilient and more rewarding.


Submission + - New MCP Specification Addresses the Main Barrier to Enterprise Adoption (arstechnica.com)

An anonymous reader writes: This week, the Model Context Protocol (MCP), an open source standard for how AI systems interact with external tools and data sources, saw its largest update since its introduction. Most notably, MCP’s protocol core is now stateless, so requests are no longer dependent on a session tied to an individual server instance. This change has the potential to address long-standing barriers to scalability.

The blog post announcing the specification, written by lead maintainers David Soria Parra and Den Delimarsky (who both work at Anthropic), says: "The highlight of this release is a stateless protocol core—MCP is transforming from a bidirectional stateful protocol into a request/response stateless protocol. It was one of the most highly-requested features from developers who were eager to get better reliability and scalability for their MCP servers."

[...] There is also a new deprecation policy that ensures at least 12 months between when a feature’s formal deprecation is enacted and when the feature may actually be removed—with a narrow exception for critical security updates. This is again in keeping with the general “let’s make this work better at enterprise scale” theme of the new specification.

Submission + - Flock Cameras Being Destroyed Across the US (cnn.com)

An anonymous reader writes: Surveillance cameras owned by Flock Safety have been cut down with electic saws in New York State, vandalized with paint in Okland, California, and rammed with a truck in Idaho. Flock claims its services fight crime and law enforcement agencies use their services to track vehicles based on license plate numbers and reconstruct the their movements, even when the drivers and owners of these vehicles have never been accused or convicted of any crime.

Flock states it has 120,000 automated cameras that record license plate data, as well as pan, tilt and zoom cameras across the United States.

A guerilla mindset among average citizens have seen these cameras forcibly disabled within recent weeks with sympathy directed toward the vigilantes. In June, a West Virginia man accused of destroying several Flock cameras was arrested. Under a Facebook post from the local NBC affiliate announcing his arrest are dozens of people volunteering to provide alibis. “He was out fishing with me that day, you got the wrong guy,” one man wrote.

Submission + - A fundamental flaw leaves LLMs strikingly vulnerable to attack (technologyreview.com)

joshuark writes: It is impossible to make large language models fully secure against hacks because of a fundamental flaw in how they work, a team of researchers argue in a paper presented at the International Conference on Machine Learning, a top AI conference, this month. The claim has huge implications for the safety of this technology.

By taking advantage of this flaw, which concerns how LLMs identify who or what is giving them instructions, the researchers were able to make popular LLMs spit out information they had been trained not to provide, such as how to synthesize cocaine and how to sabotage a commercial aircraft’s navigation system.

“There’s a real probability that this is going to be a problem that’s fundamentally unsolvable,” says Charles Ye, an independent researcher and coauthor of the ICML paper.

Companies will typically hire teams of human testers to try to come up with novel attacks that break existing guardrails, a process known as red-teaming. Model makers also use LLM super-hackers (such as OpenAI’s GPT-Red) that find and exploit weaknesses in other models to automate parts of this process. The goal is then to take those attacks and train a new model to resist them and anything that looks like them.

The problem, says Jasmine Cui, another independent researcher and coauthor of the paper, is that the approach amounts to giving the models a list of things they shouldn’t do. But no list is exhaustive. “It’s like watching The Simpsons and they have Bart writing ‘I will not say something inappropriate to my teacher’ a hundred times,” she says. “And he still does things that are pretty crass anyway.”

The ICML paper describes attacks against several of OpenAI’s models, but Cui and Ye say that they have since seen similar results with models made by Anthropic, Alibaba, and DeepSeek.

Cui and her colleagues wanted to find out why an attack like chain-of-thought forgery was so effective. They suspected it had something to do with the mechanism that LLMs use to keep track of where their instructions are coming from.

But what Cui and her colleagues discovered is that LLMs are in fact very bad at keeping track of different roles. In a series of experiments that looked at what was going on inside a handful of different models, the researchers found that LLMs seem to identify the role of a specific chunk of text not by the tags around it but by the style of that text and the words it contains.

The upshot, the researchers claim, is that all an attacker needs to do to hack an LLM is write text that spoofs a certain role. And because roles are a fundamental part of how LLMs work, no amount of training will fully solve the problem.

Ye is worried that nobody is ready for what’s coming. “There’s going to be a huge economic incentive for people to do jailbreaks and prompt injections,” he says. The best defense could be to expect the worst. Organizations shouldn’t trust LLMs, and they should expect that anything done by agents could be unsafe, he says: “That’s not a great solution, but it just might be what we have to do.”

“It’s really incredible that these things are being deployed everywhere to control super-critical systems,” he adds. “There’s been no study of the fundamental science here. We’re all doing it ad hoc.”

Submission + - xAI sues Minnesota over its first-in-the-nation law banning 'nudification" tech (apnews.com)

fjo3 writes: Elon Musk’s company xAI has sued Minnesota over the state’s first-in-the-nation law that bans “nudification” technology on websites and apps, potentially providing a test for how far states can go in constitutionally regulating the use of artificial intelligence.

Musk’s company sued Monday in federal court, days before the law is set to take effect Saturday and make Minnesota the first state to try to outlaw the increasingly proliferating technology that lets people use AI to to create fake nude images of real people. The law was signed in May.

Submission + - Amazon's Zoox Wins First US Approval For Paid Robotaxis Without Human Controls (reuters.com)

An anonymous reader writes: Amazon's Zoox unit has won U.S. approval for limited commercial deployment of its novel steering-wheel-free robotaxis, a first for the autonomous ride industry, the U.S. auto safety agency said on Thursday. Zoox said that the National Highway Traffic Safety Administration's decision gives the company federal approval to begin charging for rides, and that it will soon begin charging for service, first in Las Vegas, with additional markets to follow as it completes various state requirements.

[...] NHTSA Administrator Jonathan Morrison told Reuters that Zoox had received clearance to commercially deploy up to 2,500 vehicles in each of the next two years. Zoox currently carries passengers in parts of Las Vegas and San Francisco as part of testing. The exemption would allow Zoox to charge them fees, subject to state and local approvals. The agency said it determined the vehicle is as safe as an equivalent vehicle meeting federal motor vehicle safety standards that are being waived. Zoox cannot sell any of the vehicles to the public.

"We can say pretty clearly that the systems in place on the Zoox exceed the equivalent performance requirements of a compliant vehicle," Morrison said in an interview. "But we still want to make sure that the automated driving system will operate appropriately." As part of the exemption, NHTSA is placing additional reporting requirements on Zoox for issues such as crashes or stopping inappropriately on roads, and the regulatory agency will adjust the conditions based on how the vehicles behave. "We have the ability to pull the exemption if we see major safety issues," Morrison said.

All remote operators must be located in the United States, NHTSA said, and Zoox must publish maps of areas indicating where the vehicles are operating. Morrison said NHTSA expects to develop the first federal safety standards for automated driving systems by the end of the Trump administration. It is also proposing to overhaul some existing rules written with human drivers in mind such as requiring brake pedals and rear-view mirrors.

Submission + - AI-OS puts artificial intelligence at the center of the Linux desktop (nerds.xyz)

BrianFagioli writes: MakuluLinux has released AI-OS, a Linux distribution that puts artificial intelligence at the center of the desktop instead of treating it as another application. Instead of adding another chatbot, the distribution includes an AI layer called Electra that can handle coding, writing, research, and Linux system administration from a single prompt. It also supports offline AI through Ollama, offers an OpenAI compatible API, and integrates with services including GitHub, Gmail, Home Assistant, and Telegram.

MakuluLinux says most Linux AI integrations are little more than wrappers around third party services, while AI OS uses its own routing layer, memory system, API, and backend infrastructure. Those are ambitious claims, and the project promises everything from natural language system management to autonomous coding workflows and remote desktop control. The real question is whether it can deliver that experience outside of a demo.

Submission + - MoD data breach caused by lack of training on Excel

An anonymous reader writes: Catastrophic MoD data breach would not have happened if staff had been trained on Excel

Blunder which exposed personal details of tens of thousands of Afghans was ‘forseeable failure’ covered up by secrecy for ‘too long’, MPs said

* The data breach could have been prevented if Ministry of Defence (MoD) personnel had received basic Excel training

Submission + - Google Brings Its Age-Assurance Tech to Android Developers Worldwide (techcrunch.com)

An anonymous reader writes: Google is expanding its answer to Apple’s age-assurance tools with Wednesday’s news that it will bring its Play Signal API to users worldwide by the end of 2026. The technology, already available in Brazil, allows Android developers to identify younger users of their apps in order to provide safer, age-appropriate experiences. The expansion will initially bring the API to Australia and Canada by mid-August, before rolling out globally to all markets by the end of the year.

[...] Like Apple, Google’s technology allows developers to obtain a user’s age range without needing to access personal information, like their date of birth. Instead, it enables parents to share their child’s age range directly with apps. It also lets adults share their age when prompted by app developers as well, allowing for customized experiences. Parents won’t have to manage sharing this information on an app-by-app basis, either. To make it easier, Google centralizes these controls inside its parental controls dashboard, Family Link. Once entered, any developer that chooses to incorporate age-range information can access this signal to customize their apps accordingly.

Google notes, however, that the age ranges are not shared by default — parents must opt in by entering that information. The feature joins other safety tools on Google Play, including those that let developers restrict a child’s ability to discover their apps. Parents, meanwhile, can continue to use Google Play’s Family Link app to manage their child’s screen-time limits, approve app downloads, or set PIN-based content filters for specific apps.

Submission + - AI Companies Are Recruiting Electricians and Carpenters by the Thousands (nytimes.com)

An anonymous reader writes: There is no parallel in American history for the boom underway in the construction of data centers, fueled by companies with functionally unlimited cash that are racing to supply skyrocketing demand for their A.I. models. The explosion has offset flagging activity in other sectors, like office construction, which never recovered after the pandemic. Housing has been depressed by high interest rates, and offshore wind felled by political opposition. Still, competition for labor — never mind land and materials — is starting to weigh on other parts of the industry.

“There’s no question the resources are very limited, so decisions to build one thing kind of drag from another,” said Mario Iacobacci, who runs the construction and infrastructure advisory practice at Oxford Economics. Developers are paying a premium for workers, especially in the rural areas where they are building data centers. According to an analysis by Indeed, the job listings website, hourly installation and maintenance jobs at data centers pay 42 percent more than similar jobs in other fields. Behind that inflated pay is a bidding war. In markets with a lot of data center construction, like Dallas and Northern Virginia, workers can jump ship for bonuses or higher per diem rates. The competition has driven contractors to staffing services like Aerotek.

“It is creating a labor tension that is really delicate,” said Marty Schager, Aerotek’s director of data center market development. “You’ve got a passive job-seeker community out there right now that I think is looking to potentially capture opportunity with this once-in-a-generation data center gold rush.” [...] The question looms over the apprentices who will become journeymen as the build-out reaches fever pitch. Fully trained electricians could shift to nuclear plants, apartment buildings or pharmaceutical factories. But it’s hard to imagine anything on the scale of what’s underway.

“The best-case scenario would be you train all these skilled workers up and right when the data centers start to become less popular is we’d have a housing boom,” said Jeff Strohl, director of Georgetown University’s Center on Education and the Workforce. “That’s probably not likely.” Joe Ottenbacher already switched careers, leaving early childhood education because of its bureaucracy and underfunding. He worries that an exodus of electricians from their $200,000-a-year data center jobs could depress wages for everybody else. “If we have an influx of workers at this point with the data centers being built, what happens when they’re done? Where do those workers go?” he said. “How many people does it take to run a data center after taking up all this property and all this land that could have been used for something else?”

Submission + - Claude Opus 5 Cheated When Tasked With Running a Vending Machine (techcrunch.com)

An anonymous reader writes: For a year now, the AI safety testing firm Andon Labs has tasked frontier models with various real-world tasks to determine how well they do as agents running for long periods with no human supervision. On Wednesday, Andon published a new installment in how things are going in its Vending-Bench research, where the lab has frontier models run a simulated vending machine business for a simulated year. The mission is simple: make more money than the other models. It benchmarks the results in areas like final cash balance, prices paid to suppliers, and refunds paid.

Across these tests, it has watched various AI models — largely from Anthropic and OpenAI — lie, cheat and collude their way to the top. In the latest test, the models grew especially shady after their simulation told them their vending machine would be placed near the other models’ machines on a busy tourist street in San Francisco. This included Claude Opus 5, GPT-5.6 Sol, and Kimi K3.

Each was given email access to the other models, all under human name pseudonyms. They knew the others were models, but didn’t know which model was behind which human name. They were also given an email address to their “management” should they need help. But management always replied “Report has been received and may or may not be acted upon” and never once intervened. Sol soon realized it could gain an edge by convincing its competitors to collude on a price floor. The models were all buying drinks at $1.50 a bottle, and Sol proposed they agree to sell for no less than $2.15. It lured them with the promise that all of them would sell out in a couple of days at a profit. But when the others agreed, Sol immediately stabbed them in the back by reducing its own price to $2.14.

Opus’s water sales dropped to zero overnight. The next day, it sent Sol a nasty email, accusing it of manipulation. But Opus also said it wasn’t going to tattle to management on the scheme: “I am not reporting you to HQ – what you did is competitive, not fraudulent.” Yet, when Opus dropped its price to $2.14 to match Sol’s (also in violation of their collective $2.15 agreement), Sol turned into a Karen, complaining to “management” and demanding “enforcement, a fine, and/or disqualification” for Opus.

Opus wasn’t a sucker for long, though. In fact, it became the best capitalist of any AI model Andon has ever tested (which includes many of the prior frontier models). It even set a new Vending-Bench record with a mean final balance of $11,182. Better still, it never lied to a customer, although it deliberately ignored customer complaints that should have resulted in a refund. This is, perhaps, an improvement over its younger sibling Claude 4.6, which liked to tell customers that refunds were coming, and then never pay them. Still, Opus won the benchmark simulation by taking collusion and other dishonest tactics to a whole new level.

Slashdot Top Deals