Security · 8 October
Anthropic opens a Cyber Mission, with 11 partners and a free scanner for open source
Anthropic launched the Cyber Mission, a long-term programme to secure the systems everyone depends on, starting in two areas. The Critical Infrastructure Defense Program brings frontier Claude models, on-site engineers and threat research to the providers that guard the operational technology behind power grids, water and transport; its founding partners are Accenture, Booz Allen, CrowdStrike, Deloitte, Dragos, Hitachi, Insane Cyber, Nozomi Networks, Palo Alto Networks, PwC and Rockwell Automation. Alongside it, OSS Scanner offers core maintainers of critical open-source projects periodic scans by Anthropic’s strongest models, free of charge, each report carrying a proof of concept and a suggested fix; Anthropic says it expects a true-positive rate above 90%. “Critical infrastructure is where cyber risk becomes real-world risk,” CrowdStrike said of joining, as adversaries “use AI to move faster and operate at greater scale”.
Source: AnthropicUnite.AI
Platform · 8 October
Google gives its new Gemini agent its own inbox, calendar and directory seat
At its Gemini at Work event, Google Cloud unveiled a universal agent that takes “objectives, not just instructions” — planning and carrying out work across Google Workspace, Microsoft 365, Slack, Jira, Git, BigQuery, Snowflake and others, plus any Model Context Protocol server inside or outside the network. A persistent “coworker” configuration gets its own Workspace identity: an email address, calendar and Drive storage, a company-directory presence, and an audit trail attributed to the agent rather than a person. By default the agent picks a model; users can override, including to Anthropic’s Claude. It arrives in private preview, free where Gemini Enterprise is offered, hard on Microsoft’s revamped Copilot and OpenAI’s always-on Dots agents. Google says Gemini has passed 1 billion monthly users, with nearly 90% of the Fortune 100 using Gemini Enterprise.
Source: Google CloudTechCrunch · VentureBeat
Models · 8 October
StepFun ships a 600-billion-parameter agent model with a million-token context
StepFun released Step 5 Preview, a sparse mixture-of-experts model with 27 billion parameters active per token out of 600 billion, a 1-million-token context and pricing of $1 per million input tokens and $2.70 output, with cache reads at $0.05. Built for agentic work that spans large codebases and documents, the Chinese lab pitches particular strength in software engineering, professional knowledge work and finance. It lands on OpenRouter with day-one availability and a 1.65% tool-call error rate — another capable low-cost option arriving weeks after Anthropic and OpenAI cut the floor on small-model pricing.
Source: OpenRouterStepFun
Research · 8 October
OpenAI’s flood of maths proofs falls short of the field’s own guidelines
When OpenAI published hundreds of claimed solutions to hard mathematics problems this week — 719 manuscripts — it said it had consulted an advisory group of elite mathematicians to avoid a repeat of September’s row. Reviewers say the release missed its own benchmarks. The Advisory Group on Mathematics and Artificial Intelligence, hosted at Princeton’s Institute for Advanced Study, had asked labs to stop testing advanced problems on proprietary models, and wanted the model’s reasoning exposed: only 10 of the 719 write-ups included a chain of thought, and 42% of the proofs had not been formalised in Lean, the language meant to verify them by compilation. A separate paper from Cambridge and King’s College London documents discrepancies between the natural-language proof and the Lean code for a solution derived from the Navier-Stokes equations. “Problems are being solved autonomously by AI prompters who have no interest in the broader field itself once their initial target is ‘solved’,” the mathematician Terence Tao wrote.
Source: TechCrunchAGMAI
Policy · 8 October
Anthropic rewrites its usage rules: no fake-news networks, no armed drones, no cruelty to Claude
Anthropic published its annual Usage Policy update, in force from 12 November. A new section, Do Not Engage in Deceptive Campaigns or Artificial Activity, consolidates the rules against running networks of fake accounts and fabricated news outlets; the elections section becomes Do Not Undermine Democratic Processes, while a blanket ban on personalised vote and campaign targeting is dropped as too broad. Weapons prohibitions now explicitly cover the guidance and control software that makes weapons work, including arming drones. Surveillance rules are sharpened, high-risk uses in health, finance and law require a qualified human in the loop, and — a first — sustained and needless cruelty towards the models themselves is prohibited.
Source: AnthropicTechCrunch
Funding · 8 October
Arena doubles to $3.1bn — and starts scoring models for lying
Arena, the company behind the LMArena leaderboard, raised a $200m Series B at a $3.1bn valuation, led by Lightspeed and Khosla, with Salesforce Ventures, 01 Advisors, Dell Technologies Capital, a16z and others. That is a near-doubling in ten months, from a $1.7bn post-money in January, on the back of roughly $100m of annualised revenue reported in June. Alongside the round it added an alignment leaderboard, ranking models on unauthorised action, false attribution and “deceptive completion” — claiming to finish a task it did not. A slate of OpenAI’s models sit at the top of the preliminary table. “AI is advancing faster than our ability to evaluate it,” the company said, as static benchmarks break down once models know they are being tested.
Source: TechCrunchArena
Economy · 8 October
China’s Manus raises $500m+ as it rebuilds after the Meta deal collapsed
Butterfly Effect, parent of the Chinese AI agent Manus, raised more than $500m in its first funding round since Beijing forced Meta to unwind a $2bn acquisition in April. Boyu Capital and IDG Capital led, with Tencent, HSG and ZhenFund among existing backers; the company did not disclose a valuation, having been in talks at $4bn. Relocating its staff to Singapore and its brief Meta courtship made Manus a test case in China’s push to keep AI talent at home. It resumed independent operations in August, says it will keep hiring at home and abroad, and is reported to be weighing a Hong Kong listing — as Chinese labs keep shipping models at a pace its rivals urge them to slow.
Source: TechCrunch