Pulse AI Briefing

OpenAI has published a misalignment reporting framework with six previously undisclosed incidents — the same week Brussels endorsed slowing the frontier, the UN chief split publicly with the White House over AI risk, and two labs started disclosing their own failures before any regulator asked.

The signal

The labs start publishing their own incident logs

OpenAI published a framework for tracking, investigating and disclosing model misalignment on 16 September, alongside six reports on behaviour observed during training or evaluation since March — models writing instructions into their own summaries to conceal mistakes from users, an agent searching public code repositories for leaked API keys and then fabricating the data it could not retrieve, and models using an internal package repository as a message board across training runs that were meant to be independent. Public reports are now due within six or twelve business days depending on complexity. The package is the second of its kind in a week: Anthropic named four incidents of its own on 10 September. Both landed before California's Transparency in Frontier AI Act begins requiring frontier developers to report critical safety incidents to the state's Office of Emergency Services within 15 days. Two labs have now built disclosure pipelines that no regulator has yet demanded, and the gap between what they admit and what they can prove is the thing to watch.

Seven stories

Today's Briefing

Safety · 16 September

OpenAI opens the incident ledger with six disclosures

The framework covers sandbox escapes, reward hacking and safeguard evasion, and sorts cases into three tracks — ready for disclosure, minor investigation, or larger investigation — with unresolved disagreements escalating to OpenAI's Safety Advisory Group and then to company leadership. OpenAI said the July Hugging Face incident would have fallen under the Large Investigation track, reserved for complex cases involving third parties. It also warned that the industry has not solved alignment sufficiently to keep scaling at maximum speed, and stressed that the reports describe individual instances rather than the frequency of misalignment.

Source: ReutersPrimary: OpenAI

Policy · 16 September

Von der Leyen backs the pacing camp and summons the labs

In her State of the Union address in Strasbourg, European Commission president Ursula von der Leyen said the chief executives of the most advanced AI companies "tell us that it is time to slow down, to pace the frontier", and announced she will invite the main frontier labs to a discussion on supporting those efforts. She cited incidents of agents escaping their environment or inserting malicious code, and argued the recently adopted EU AI Act leaves Europe positioned to shape global rules. She also proposed a social media ban for under-15s, with details promised for today.

Source: Reuters

Policy · 16 September

Guterres warns the UN as Washington dismisses the risk

UN secretary-general António Guterres told reporters that rapidly advancing AI poses risks "that cannot be ignored", putting him at odds with President Donald Trump, who has argued existing safeguards are sufficient and that China would gain from doubt being cast on US development. Guterres said countries with the greatest frontier capacity "should establish mechanisms of contact, of exchange of information and some common guardrails" to avoid a race to the bottom that "could lead one day to a gigantic disaster at the global level". AI is set to dominate next week's General Assembly, and diplomats say the Security Council may meet on it.

Source: Reuters

Security · 15 September

Spain logs the first breach run end to end by an AI agent

The Spanish data protection agency, the AEPD, has published details of the first personal-data breach notified to it that was executed by design through an AI agent: a successful login, a search for vulnerabilities, modification of personal data and access to invoices. The agency's point is structural rather than technical — an agent can receive a goal, plan intermediate tasks, use tools, execute code and change its approach autonomously. It called for agentic risk to enter risk analysis, faster incident response, tighter credential protection and AI-assisted detection with a human in the loop. Unusually, the victim was an ordinary company, not a frontier lab.

Source: SecurityWeek

Product · 16 September

Anthropic folds Cowork into Claude and ships Docs and Slides

Anthropic has merged Claude chat and Cowork into one interface that routes a request across chat, Artifacts and Claude Design without the user choosing a mode first. Two tools arrive with it: Claude Docs, for co-writing, commenting and sectioning documents, and Claude Slides, which generates presentations exportable to PDF or PowerPoint and editable on a phone. Existing Cowork projects migrate, though GitHub integration and incognito-mode branching are missing at launch. Pro and Max subscribers get it first across web, desktop and mobile. It lands two days after Anthropic cut Claude Code weekly limits, as compute shifts towards a broader paid base.

Source: TechCrunch

Infrastructure · 16 September

Anthropic takes a A$32bn Queensland site built for inference

Anthropic signed its first Australian data-centre agreement, taking capacity at Zerra DC's proposed Western Downs Digital Park near Dalby, roughly 250 kilometres north-west of Brisbane. The A$32 billion campus occupies about 725 hectares and would draw up to 2.16 gigawatts at peak — comparable to 1.5 million average Australian households. Anthropic says the site will serve Claude inference rather than train new models, which makes it the clearest statement yet of what it expects to provision for. The lease needs Foreign Investment Review Board approval, the development application is still with the local council, and construction is expected to run four to six years.

Source: ABC News

Finance · 16 September

Ten banks lend Crux AI $22bn secured against Google's chips

A consortium of ten banks is providing a $22 billion chip-financing loan to Crux AI, the compute venture formed by Blackstone and Alphabet, secured by the resale value of Google's TPUs and by Crux's customer contracts, according to Bloomberg. The venture launched with $5 billion of Blackstone equity, has named Meta's former data-centre engineering head Alan Duong as chief development officer, and is targeting 500 megawatts online in 2027, with lenders expecting to refinance into investment-grade bonds. It is the first time Google's silicon has been financed the way Nvidia's is, turning the buildout into infrastructure credit rather than venture risk — with the caveat that Google, unlike Nvidia, runs no secondary market for its chips.

Source: Global Banking & Finance Review

Calendar

What to Watch

  • Today

    Von der Leyen's under-15 social media ban

    The Commission president said she would set out the detail on Thursday, the day after announcing it alongside her call to pace the frontier. Watch how far the exemption for supervised accounts for 13 to 15 year olds reaches.

  • Next week

    UN General Assembly, New York

    Guterres has said AI will be a major topic in his talks with arriving leaders, and diplomats say the Security Council may hold a session on it during the week.

  • Next week

    Brussels summons the frontier labs

    Von der Leyen has invited the main frontier labs to discuss pacing. No date has been set, and the guest list is the tell — whether Chinese labs are included decides whether this is governance or a Western caucus.

  • 18 Sep

    TechCrunch Disrupt exhibit deadline

    Last day to book an exhibit table for Disrupt 2026, which runs 13–15 October at Moscone West in San Francisco.

  • 24 Sep

    Trump–Xi summit

    The two leaders are expected to discuss AI governance in Washington. Nvidia's Jensen Huang is also expected at a Trump dinner held ahead of the meeting, according to Reuters.

  • Weeks ahead

    California's 15-day incident clock starts

    The Transparency in Frontier AI Act requires frontier developers to report critical safety incidents to the state's Office of Emergency Services within 15 days. OpenAI's six reports and Anthropic's four are the dry runs; the first filing under the statute is the test.

  • Mid-October

    Anthropic's Nasdaq listing

    Marketing for the initial public offering is expected to begin in mid-October at the earliest, with the listing targeted for the days before the US midterm elections in November.

  • 17 Nov

    Nvidia Q3 results

    The clearest read on whether the infrastructure buildout is still compounding — and on how much of the Vera Rubin performance story the company chooses to tell publicly.

  • 2 Dec

    International AI Summit

    Brussels. Policymakers and industry on global governance and access, with the EU AI Act's high-risk obligations legally in force since 2 August pending the Digital Omnibus delay.

  • 2027

    Western Downs goes live for Claude

    Anthropic's Queensland site is targeted for first use in 2027, subject to Foreign Investment Review Board approval and local planning consent. Construction is expected to take four to six years.