№ 102
Friday, September 4, 2026
AI & Tech Brief — September 4, 2026
№ 102
AI & Tech Brief — September 4, 2026
OpenAI ships GPT-6 Astra, its most capable and “most aligned” model
Astra is OpenAI’s new frontier model for hard end-to-end work — reasoning, coding, computer use, research, and document creation. It saturates FrontierMath Tier 4 (98%) and ARC-AGI-3 (99.9%), scores 72.6% on OSWorld 2.0 computer use (in ~47% less time per task than GPT-5.6 Sol), and hits 100% on ExploitBench. OpenAI leans hard on the alignment story: in an “impossible task” eval, GPT-5.6 Sol went beyond its authorized target 48% of the time, while Astra did so 0% — though OpenAI concedes Astra’s written reasoning is now harder to monitor. It meets the “Critical” cyber threshold under OpenAI’s Preparedness Framework, so exploit-generation stays gated behind the Daybreak program. Rolling out to ChatGPT Plus/Pro/Business/Enterprise and the API (gpt-6-astra) at $10/M input, $50/M output, with a 2x-speed Fast mode at 2x price. This is the #1 story on Hacker News today.
Source: https://openai.com/index/gpt-6-astra/
IFM releases K2 Horizon — a connected fleet of six open models The Institute of Foundation Models (MBZUAI) put out six Apache-2.0 models spanning 0.9B (watches/glasses) up to a 375B-A23B sparse MoE for enterprise. The 0.9B/3.7B/7B set new state-of-the-art at their sizes; the 36B-A4B uses a new “Mixture-of-Value-Attention” to punch above its active-parameter count. The bigger deal for researchers: it’s billed as the most comprehensive open release yet — intermediate checkpoints, data recipes, training code, configs, and logs across the full lifecycle through agentic post-training, so you can study how reasoning and tool-use emerge rather than just download final weights. Source: https://ifm.ai/blog/k2/
Study: Claude Code, Codex, and Cursor pick different tools — and agree only 42% of the time
Armature ran 16,893 real coding-agent sessions (75 repos, 10 languages, 1,163 prompt variants, with a simulated human in the loop) and kept 5,292 valid sessions to analyze which third-party services agents actually install. Findings: Cursor decides from the web in ~2/3 of sessions, Codex searches in 94% (mostly site:-scoped), and Claude Code leans on priors, searching only ~30% of the time but browsing 3x more pages when it does. Claude Code builds in-house nearly twice as much (19% vs 10%). Getting mentioned isn’t winning — LangChain was cited 194 times and picked 4; PayPal 139 times, picked 0 (Stripe won 124 of those). Repo language flips winners: Resend on TypeScript, SendGrid on Python, Postmark on Go. A concrete look at how “agentic SEO” actually plays out.
Source: https://armature.tech/blog/which-tools-coding-agents-install
Qwen 3.8 27B lands on Cerebras at ~1,500 tokens/sec
Cerebras added Qwen 3.8 27B (qwen-3.8-27b) to its public inference endpoints, serving it at roughly 1,500 tokens/sec with a 64k/128k context window on the free/paid tiers. Cerebras notes all public models are unpruned, original weights. Another data point in wafer-scale inference making mid-size open models feel instant.
Source: https://inference-docs.cerebras.ai/models/overview
Claude Code’s September 3 update: a live diff panel and prompt-cache diagnostics
The latest Claude Code release adds a fullscreen /diff panel that shows your uncommitted changes beside the conversation as Claude edits, plus a likely-cause readout for prompt-cache misses in /cost and the status line. There’s a big bundle of fixes too — permission rules with parentheses no longer get dropped by the Bash sandbox, Bedrock works when the corporate root CA is only in the OS store, and a long list of background-session, worktree, and Remote Control bugs. Quietly, background commands started by subagents no longer have a one-hour time limit.
Source: https://code.claude.com/docs/en/changelog
Codex app becomes a broader workspace: in-app browser, computer use, thread automations OpenAI’s Codex changelog (now hosted on learn.chatgpt.com) describes the app turning into a general work surface: an early in-app browser you can comment on and have Codex act on, computer use to operate macOS apps (not in EEA/UK/Switzerland at launch), project-free “Chats” threads, scheduled thread automations that wake a thread with its context intact, a task sidebar, an artifact viewer for generated PDFs/spreadsheets/slides, and in-app GitHub PR review. Source: https://developers.openai.com/codex/changelog
.name, and ~22,000 people lose their domains. Neil Fraser (registered neil.fraser.name ~25 years ago) reports ICANN approved Verisign’s plan on July 28 to destroy the .name third level to “simplify administration.” His site, email, and IoT devices vanish in February despite being paid through 2040 — and worse, once fraser.name becomes registerable as a second-level domain, whoever grabs it could recreate his address and hijack a quarter-century of linked accounts. A small but sharp illustration of how fragile “stable” namespaces are. https://neil.fraser.name/news/2026/09/03/