The Sequence Radar - Issue 944 — Last Week in AI: Last Week in AI: OpenAI Connects the Dots, Gemini Levels Up, and Agents Cash In
Next Week in The Sequence:
Another installment of our series about recursive self-improvement.
We dive into OpenAI and Gemini releases of this week.
In the opinion section, we dive into the Jev revolution and the new explosion in decision model releases.
We deep dive into another robotics stack.
Subscribe and don’t miss out:
📝 Editorial: Last Week in AI: OpenAI Connects the Dots, Gemini Levels Up, and Agents Cash In
AI is getting an expense account. That, at least, is the impression left by a week of persistent agents, enterprise ambitions, spatial intelligence, and billion-dollar funding rounds. The industry increasingly wants permission to enter our workflows, operate our software, and complete the tasks we keep postponing. Intelligence is becoming a job application.
OpenAI’s DevDay made that ambition concrete. Dots are persistent agents designed to pursue goals between conversations, while ChatGPT Space gives people and AI a shared workspace. GPT-6.1 Sol brings what OpenAI describes as near-Astra performance at lower cost, and Astra Ultrafast accelerates token generation. Together, these releases address the economics and logistics of delegation. An agent that works repeatedly needs affordable inference, persistent context, useful tools, and somewhere to deliver results. OpenAI is assembling that machinery around the model. The chatbot is acquiring an office.
Google supplied another candidate for employee of the month: Gemini 4 Argon. Its most striking specification is a one-million-token output limit, up from 64,000, providing more room for extended reasoning and complex tasks. Google reports 77.9% on DeepSWE v1.1 and describes internal work on code migrations and data-center memory optimization. Those are company-reported results, and initial access is restricted to trusted cyber defenders through the Fairwind Program. The engineering question is fascinating: how much useful work can a model sustain before errors, latency, and cost overwhelm the benefits of thinking longer?
Meta’s contribution came with an executive appointment. The company announced Meta Enterprise Platform and recruited MongoDB CEO Chirantan “CJ” Desai to lead it, reporting directly to Mark Zuckerberg. The initiative will bring products including Muse, Meta Business Agent, Muse API, and Muse Code to businesses and developers. Desai’s arrival suggests Meta understands the distance between consumer enthusiasm and enterprise adoption. Corporate buyers eventually ask awkward questions about integration, permissions, support, and accountability. A charming agent still has to survive procurement.
AMD approached the opportunity from underneath the software stack, agreeing to acquire Fei-Fei Li’s World Labs for approximately $8.2 billion in stock. World Labs builds spatial-intelligence models that generate and simulate interactive 3D environments. The transaction remains subject to approvals. My reading is that AMD wants model research to inform its infrastructure roadmap directly. Robotics and simulation create computational demands that chip designers need to understand early. Owning expertise in those workloads could help connect tomorrow’s algorithms to tomorrow’s silicon.
Meanwhile, Instinct raised a $1 billion Series C from Sequoia, Benchmark, and Coatue at a $10 billion valuation. Its personal agent, still in early access, tackles travel bookings, groceries, and forgotten subscriptions. Investors are placing an enormous bet on the value of becoming the trusted intermediary for everyday tasks and transactions. Apparently, cancelling the gym membership is now a frontier technology market.
Across these announcements, delegation is the common ambition. Success will depend on completed tasks, sensible permissions, recoverable mistakes, and costs that justify continued use. The winning agent may be the one you finally stop checking every five minutes.
🔎 AI Research
RRSI: Regularized Recursive Self-Improvement of Agent Harnesses
AI Lab: Google Cloud AI Research (with UNC-Chapel Hill, Stanford, Washington University in St. Louis)
Summary: Harness self-evolution often overfits the evolve-set; RRSI keeps every harness component editable but regularizes the search—annealed edit budgets and evidence-aware proposals on one side, a leakage critic, noise-adjusted floor, cost rule, and pruning on the other. Across eight coding, agentic-workspace, and engineering-design benchmarks (frozen Claude Opus 4.8), it averages +4.0 pts on the three suites it evolves against and +3.4 on six held-out benchmarks (all improve), while using ~36% fewer policy tokens per trial than unregularized evolution.
Follow the Entities: A Corpus Map for Agentic Search
AI Lab: Microsoft (with KAIST)
Summary: CorpusMap builds an offline entity–document navigation layer—Entity Pages that aggregate and link every document mentioning a recurring entity—so agents follow shared paths instead of rediscovering cross-document relations per query. On EnterpriseRAG-Bench, WixQA, and HERB with GPT-5.5/5.6 models, it lifts Overall Quality over raw-corpus search (e.g. +6.45 pts with GPT-5.5; up to +11.74 with Luna) while cutting relative input tokens to ~0.43–0.65×, and outperforms four alternative navigation layers.
Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents
AI Lab: NVIDIA (with KAIST)
Summary: Mid-Harness samples candidate terminal actions and verifies them before execution, leaving the generator and harness unchanged. With TMAX-9B on TerminalBench-Lite, a GPT-5.6 Sol verifier raises Pass@1 from 50.00% to 68.03% at N=8; pairwise self-verification reaches 54.76% (57.14% after distilling the frontier verifier), and pairing Mid-Harness with Best-of-T (T=3) hits 66.33%.
AIM: Agentic Idea Management for Automated Research
AI Lab: Google Cloud AI Research (with University of Wisconsin-Madison)
Summary: AIM is an idea-driven research manager that uses an Agentic Surrogate and Acquisition loop (Bayesian-opt style), a Solution Auditor for idea–solution integrity, and a Resource Planner for budgeted parallelism. On 10 AutoLab tasks it scores 67.0% on System Optimization and 55.8% on Model Development & CUDA (+1.6 / +4.9 pts over ScientistOne) and reaches the baseline’s best score up to 3.1× faster in wall-clock time.
Scaling Laws for Looped Mixture of Experts
AI Lab: Meta AI
Summary: Loop Scaling Laws jointly predict held-out loss from model size, data, recurrence, and MoE sparsity via a bounded, sparsity-conditional effective-parameter gain (recovering dense and MoE laws as special cases). Downstream, sparsity yields ~3× active-parameter efficiency and recurrence ~2× total-parameter efficiency on reasoning; at matched compute a 0.3B-active/1.3B-total looped MoE matches a ~2× larger non-looped MoE on reasoning benchmarks while enabling test-time recurrence scaling.
🤖 AI Tech Releases
Open Agent Safety Platform
NVIDIA launched Open Agent Safety Platform, pairing open-source OpenShell (a secure runtime boundary for agent access and policy) with Sentry on BlueField-4 DPUs that monitors agents out-of-band and can quarantine breakouts in milliseconds—backed at launch by Anthropic, Microsoft, SpaceXAI, Figure, JPMorganChase, and 100+ other organizations.
Dots
OpenAI introduced Dots, always-on agents powered by GPT-6 Astra with their own cloud computer that connect to 4,000+ apps and can keep working 24/7 across ChatGPT, Slack, and Teams—rolling out today to Pro and Business Premium users in eligible markets (first dot included at no extra cost), with an Enterprise beta when a workspace admin enables it.
GPT-6.1 Sol
OpenAI introduced GPT-6.1 Sol, an upgrade that nearly matches GPT-6 Astra on agentic coding, computer use, and professional work at about one-fifth of Astra’s standard token prices ($2/$10 per MTok input/output; $0.10/M cached input)—available today in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise, and Edu (API id gpt-6.1-sol
), with GPT-6.1 Sol Ultrafast coming in days at up to 8× faster token generation in Codex.
OpenAI DevDay extras
Alongside Dots and GPT-6.1 Sol, OpenAI’s DevDay also shipped Ultrafast mode (up to 8× faster token gen in Codex; Astra Ultrafast on Pro 500 and Enterprise), ChatGPT Space and Pages for human–agent collaboration, Agents API computer use, Codex Cloud and Security Cloud, plugin extensions with MCP Events, @ChatGPT in Slack/Teams, and a Pro 500 tier with 25× Plus allowance.
Gemini 4 Argon
Google announced Gemini 4 Argon, its next frontier model for long-horizon software engineering, enterprise knowledge work, and cybersecurity defense—rolling out first to trusted cyber defenders via the Fairwind Program (1M output tokens; intro API pricing $2/$10 per MTok), with broader access planned next for paid API customers and Google AI Ultra after phased safety review.
Strands Decider 2B
Strands Agents open-sourced Strands Decider 2B, a 2B-parameter decision model (Qwen3.5-2B torso with a pointer head) that picks among fixed options with confidence scores in tens of milliseconds locally—weights and training data on Hugging Face/GitHub, aimed at agent routing, tool selection, guardrails, and hybrid workflows with the Strands harness.
📡 10 AI News You Need to Know About
Meta launched Meta Enterprise Platform, a new business pillar to sell its AI stack—including Muse, Meta Business Agent, Muse API, and Muse Code—to companies and developers, hiring Chirantan “CJ” Desai from MongoDB as Chief Enterprise Platform Officer reporting to Mark Zuckerberg (MongoDB named Dev Ittycheria interim CEO).
AMD agreed to acquire World Labs—Fei-Fei Li’s spatial-intelligence lab behind models that generate and simulate interactive 3D environments—in an all-stock deal valued at about $8.2 billion, expected to close by year-end, with Li joining as executive vice president and chief scientist reporting to Lisa Su.
Instinct raised $1 billion in a Series C at a $10 billion valuation from investors including Sequoia Capital, Benchmark, and Coatue—about a month after a round that valued the invite-only personal AI agent (own phone and computer for bookings, bills, and calls) at $2.5 billion.
EliseAI raised $350 million at a $4 billion valuation—roughly double its August Series E mark—in a round led by Andreessen Horowitz and Bessemer Venture Partners (Ontario Teachers’ Pension Plan, Sapphire Ventures, and Navitas Capital also in), to push its housing and healthcare ops automation deeper after saying it passed $200 million ARR and reaches about 1 in 6 U.S. apartments.
Meta expanded Muse for Small Business with new connectors—Shopify, Slack, Dropbox, QuickBooks, Stripe, Zoom, and more, plus Instagram professional analytics, Facebook Pages, and Meta ads—so the agent already knows what a shop sells and how the brand sounds; free with usage limits, with paid plans for more capacity.
Flow Engineering raised $50 million in a Series B at a $750 million valuation, co-led by Valor Equity Partners’ Antonio Gracias and Atreides Management’s Gavin Baker, with Sequoia Capital returning and Roelof Botha joining the board—backing its AI agents that keep CAD, requirements, and simulation aligned for hardware teams (customers include Anduril, Rivian, and Joby Aviation).
Restate raised $20 million in a Series A led by Singular, with Redpoint Ventures and Capital One Ventures participating (about $27 million total), to scale its durable-execution infrastructure for long-running AI agent and backend workflows—used by Replit and Fortune 500 customers, from the Apache Flink creators.
TSMC is weighing a Texas campus with multiple fabs to meet strong AI-chip demand, but talks remain early, alongside a company-reported $265 billion Arizona plan for six logic fabs, two advanced-packaging facilities and an R&D center.
DoorDash introduced a U.S. beta that lets users text requests to find food and groceries, build personalized carts, and check out in the thread. Bloomberg reports that the experience works in Apple’s Messages/iMessage.
DeepSeek open-sourced Ascend infrastructure—including compute and communication libraries—for Huawei’s AI chips; Bloomberg reports the companies are also jointly optimizing a 128-card Ascend 950 supernode as a potential Nvidia alternative.
How it works
Once you click Generate, Ollama reads this article and crafts 5 comprehension questions. Your answers are graded against the article content — general knowledge won't be enough. Score 70+ to count toward your certificate.
Questions are cached — you'll always get the same 5 for this article.