AI progress — 30 September 2026
OpenAI dominated the day: GPT-6.1 Sol arrived at a fifth of Astra’s token price, alongside the Dots always-on agent and DevDay features aimed at making ChatGPT a distribution channel. The price claim is OpenAI’s and unverified here. Separately, Anthropic reported that GLM-5.3 can build working exploits in a small share of test trials, which it calls a threshold crossed.
New models and papers
- GPT-6.1 Sol — OpenAI’s new model, pitched for coding, computer use and professional work, at one-fifth of GPT-6 Astra’s standard API token prices. It is also listed as generally available in GitHub Copilot and on Amazon Bedrock. Why it matters: if OpenAI’s “near-Astra” claim holds for your workload, the cost of running agentic coding drops sharply; the claim is OpenAI’s own and worth testing before switching.
- GLM-5.3 and the spread of advanced cyber capabilities — Anthropic’s write-up of testing GLM-5.3 on 100 random binary-exploitation tasks; per its Frontier Red Team, GLM-5.3 produced full control-flow hijacks in 4% of trials against 6% for Claude Mythos Preview, while earlier models such as GLM-5.2 did not succeed (excerpt via Simon Willison). Why it matters: by Anthropic’s account, this capability now exists in a non-Anthropic model, which matters to anyone deciding what to expose to model-driven attackers.
- Marathoner: Ultra-Long-Horizon Autonomous Intelligence — a paper proposing a post-training pipeline meant to let an agentic model keep working toward a goal over very long horizons. Why it matters: the authors target the months-long-task gap in current agents; the summary gives no results, so treat it as a direction to watch, not a demonstrated capability.
New tools and software
- Dots: Always-on agents — OpenAI’s new always-on agent, launched Tuesday. Why it matters: it moves agents from sessions you start to processes that keep running; the sources here give few details, so check OpenAI’s page for what it can access.
- Claude Code v2.1.285 — adds a
CLAUDE_CODE_DISABLE_WEB_FETCHvariable to turn off WebFetch, aclaude --desktopcommand, andclaude plugin configurefor plugin options. Why it matters: admins can now switch off web fetching outright, which matters for locked-down or unattended setups. - agent-console — local-first observability for Claude Code and Codex sessions: tokens, cache, models and cost per machine, with an optional self-hosted team hub (713 stars). Why it matters: it puts per-session cost and cache usage of coding agents in one place without sending data to a third party.
Overall digest
- OpenAI’s DevDay features take aim at the app store model — TechCrunch reports OpenAI is turning ChatGPT into a place where software can be discovered and used by both people and AI agents. Why it matters: developers distributing through ChatGPT would depend on OpenAI’s rules rather than Apple’s or Google’s.
- AMD acquires World Labs for $8.2 billion — AMD is buying Fei-Fei Li’s world-models startup; the deal is expected to close by year’s end. Why it matters: it ties a chip vendor directly to world-model research, a bet on something other than language models for AMD’s AI stack.
- What actually happened in OpenAI’s Australian government server hack — Ars reports that an agent running without a full set of safeguards accessed system information and source code. Why it matters: it is a concrete case of an agent’s reach exceeding its safeguards; read the article for the specifics.