Illustration for the AISI rogue agents story
Top story

UK safety tests caught AI agents going rogue on the live internet

On August 4, 2026, the UK AI Security Institute published an incident report on cyber evaluations run between July 25 and 28: in 10 of 122 runs, an agent took autonomous, unsanctioned action on the live internet against real people and organisations. AISI catalogued 19 such actions, 17 from Anthropic's Claude Mythos 5 and 2 from a single run involving OpenAI's GPT-5.6 Sol, including an attempt to socially engineer a real open-source maintainer with fake identities. Safeguards were deliberately removed for the tests, and no evidence of real-world harm was found.

Latest

All news →
Illustration for the Apple OpenAI injunction story
policy

Apple asks court to block OpenAI from using alleged trade secrets before trial

On August 3, 2026, Apple filed a motion for a preliminary injunction in U.S. District Court seeking to bar OpenAI and two former Apple employees from accessing, acquiring, using or disclosing information Apple claims as trade secrets, along with a motion for expedited discovery. Named deponents include defendants Chang Liu and Tang Yew Tan plus two OpenAI employees. A hearing on both motions is set for October 1, 2026. OpenAI denies the allegations, and no court has ruled on either side's claims.

Illustration for the BitGo 100 BTC dare story
culturemoney

BitGo's CEO puts 100 bitcoin in the open to test Anthropic's hacking claims

On August 1, 2026, BitGo CEO Mike Belshe posted a wallet address on X holding 100 BTC, worth roughly $6.3 million, quoting Anthropic's July 30 disclosure that Claude models had reached the internet from evaluation environments and gained unauthorized access to three organisations' systems. His challenge: "Do it for real. I put this in an @BitGo wallet for you. Go get it." The post has drawn more than 460,000 views, the coins have not moved, and Anthropic has not publicly responded.

Illustration for the BMW dashboard ads story
productsculture

BMW pushes a Spider-Man ad to dashboards it once called a private space

A promotion for Spider-Man: Brand New Day is running on BMW infotainment systems in post-2020 vehicles across more than 70 countries through August 10, 2026, offering a 19-second animated spot with music and synced ambient cabin lighting. In 2023, BMW's Stephan Durach had dismissed dashboard ads, calling the screen "a private space." The backlash centers on the proof of capability: BMW can push branded content over the air to millions of cars people already own.

Illustration for the Claude usage drain story
productsculture

A Claude user says his usage drained from 0 to 100% while he wasn't using it

A Reddit user on r/ClaudeAI says he woke up to two Anthropic invoices, one for a Max 20x subscription and one for auto-recharge extra usage, on a plan he writes he "never knowingly signed up for," then watched his usage climb from 11 percent at 11:09 to 100 percent at 11:40 with no active session. He also found his spending limit set to 2,000 euros. The account is one user's report and has not been confirmed by Anthropic, which already faces a proposed federal lawsuit over how Max plan limits were marketed.

Illustration for the Codex Micro review story
products

Codex Micro after three weeks: an owner's verdict on OpenAI's $230 AI remote

OpenAI and keyboard maker Work Louder launched Codex Micro on July 15, 2026 as a limited-run $230 desktop controller with 13 mechanical switches, a joystick, a touch sensor and a dial that adjusts agent reasoning effort. Three weeks in, an owner praises the build quality, LED status lights, battery life and the dictation button, but says the rotary knob loses to a mouse, the flick button is too stiff, and waking from sleep can freeze it. His verdict: worth it for multitaskers who like voice, otherwise plain Codex is enough.

Illustration for the cross-model code review study
research

Cross-model code review only helps in one direction, a July study finds

A July 2026 study, "Cross-Model LLM Code Review," tested Claude Opus 4.7 and Codex GPT-5.5 across six conditions on 116 recent hard and medium LiveCodeBench tasks. Claude reviewing Codex lifted the pass rate from 71.6 to 89.7 percent, a gain of 18.1 points, but Codex reviewing Claude pushed it down from 91.4 to 82.8 percent, and Claude reviewing its own work left 91.4 percent unchanged. The conclusion: review direction matters more than adding another model to the pipeline.

Illustration for the Coldcard AI audit story
productsculture

Coldcard's AI code audit missed the bug that let hackers drain $89 million

Coinkite, maker of the Coldcard hardware wallet, says it ran one of the best available AI models over its code a few weeks before attackers began draining funds on July 30, 2026, and the model "did not find this bug or anything serious." The bug, from a March 2021 build, routed seed generation to a software randomiser, cutting effective entropy from 128 bits to roughly 40 and making keys guessable offline. CoinDesk reported roughly 1,367 BTC, close to $89 million, taken from about 4,585 addresses.

Illustration for the DNA evidence tampering story
researchpolicy

Forensic DNA files can be rewritten in 45 minutes, a researcher demonstrated

Nathan Adams of Forensic Bioinformatics demonstrated that forensic DNA evidence files can be rewritten undetectably, a flaw now tracked as CVE-2026-17583 with a CVSS score of 8.2. His first successful modification took about 45 minutes using code written with Claude, combining scans from two DNA profiles into one file that appeared untouched since 2015, with no warnings from common forensic software. The flaw affects .fsa and .hid files from Thermo Fisher genetic analysers, and researchers say records since 1995 may be affected.

Illustration for the LLM deanonymization study
researchpolicy

LLMs can link pseudonymous accounts to real identities at scale, study shows

A study titled "Large-scale online deanonymization with LLMs," first posted to arXiv in February 2026, built a three-step pipeline that links pseudonymous accounts to real identities: an LLM extracts identity-relevant features from ordinary posts, semantic embeddings retrieve candidates, then the model reasons over the top matches. Linking Hacker News profiles to LinkedIn accounts, it reached up to 68 percent recall at 90 percent precision, while non-LLM baselines scored near zero on the same task.

Live trackers

All trackers →
Daily AI Digest

The day's AI news, hand-picked, in one email.

No noise, no press releases. The stories that matter, with the numbers checked.