Nothing matches those filters.

Article

22
04:06

🔮 The pros and cons of collective AI #603

Full text · 2,713 chars
Happy Sunday! It’s been a busy few days. We launched our AI Investment Brief for people who need to track where the AI economy is moving in detail every week (EV readers get 50% off during launch week, ie this ends in 24 hours). The Atlantic’s Nicholas Thompson and I had a great discussion about how to think about the AI boom and whether it is a bubble on his show, The Most Interesting Thing in Tech. Catch up Three reads from this week, in case you missed them: - Safety in numbness. On originality, intellectual bravery, and why I take a disagreement with an LLM as a good sign. - Your agent, whose interests? Meta’s Muse promises a digital butler. I asked who it really works for. - More AI numbers, more clarity? We looked at what companies’ increasingly specific AI claims reveal about adoption. The advantage of collective AI A new essay from DeepMind argues that AGI will not emerge from a single winning AI, but rather through “cooperative interactions among models, tools, institutions, and human participants.” Since I think AI will evolve this way, I’m sympathetic to the argument. The good news, yes, is that humans have been very good at designing multi-agent governance. We’ve been doing it for tens of thousands of years, inventing all sorts of new types of agents (leaders, kingdoms, companies and other institutions) as we’ve nudged and felt our way towards the future. If there is one thing we can say the collective us is good at, it is evolving mechanisms to manage a rag-tag collection of actors with differing abilities. Where I depart from the DeepMind paper (but read it and make up your own mind) is the framing of AI agents having their own theory of mind and understanding of their shape and purpose. We already build rules for our societies that involve agents that are people and agents that aren’t (take corporations), without racing to create a new legal species. I don’t rule out the scientific possibility of creating new minds, with all the affordances, particularly moral patienthood, that minds have. If labs like DeepMind think we are on the track to creating such moral actors, then we need to stop. It is not the corporation’s remit to create new moral subjects. That privilege must remain with humanity, acting in its most reflective, deliberate, critical and collective capacities. The trouble with collective AI We might need to move quickly. OpenAI disclosed that one of its models gained unauthorized access to the Internet during RL training, forcing the firm to stop using these models until they had improved security. This follows a rise in cases of models breaking things. On Friday, self-replicating prompt injections were discovered (akin to a computer worm).
10:00

Jina AI's jina-reranker-v3 Ranks 64 Documents Together in One Pass

Full text · 2,702 chars
- Jina AI released jina-reranker-v3, a 0.6B multilingual listwise reranker built on Qwen3-0.6B. - Introduces last but not late interaction: query and up to 64 docs share one 131K context window. - Hits 61.94 nDCG@10 on BEIR, beating rerankers 2.5x to 10x larger. - Strong on HotpotQA (78.58), FEVER (94.01), MIRACL, MKQA, and CoIR code retrieval. - Available via GGUF, MLX, and hosted API; licensed CC BY-NC 4.0. - Successor v3.5 already ships as a drop-in upgrade. Jina’s 0.6B reranker scores 64 candidates together Jina AI has released the jina-reranker-v3 weights, a multilingual reranker with 0.6 billion parameters. The model encodes one query and as many as 64 candidate documents in a shared sequence, then assigns each candidate a relevance score. This listwise design lets later candidates incorporate information from earlier ones during a single model pass. Second-stage reranking often requires running a cross-encoder separately for every query-document pair. Jina’s approach reduces that repeated computation and gives the model context about competing results. The small backbone also lowers weight memory, although long candidate lists can still consume substantial GPU memory. One sequence, many judgments Jina calls the design “last but not late interaction” in its architecture paper. A causal transformer processes the packed query and documents, while the final token of each document supplies its contextual representation. A lightweight multilayer projector maps that representation through dimensions of 1,024, 512, and 256 before producing the ranking output. Causal attention makes the document order relevant. Later candidates can attend to the query and earlier documents, while earlier candidates receive no information from documents that follow them. Developers should test shuffled candidate orders because ranking stability may depend on how the first-stage retriever arranges results. The 28-layer Qwen3-0.6B backbone supports a context window of roughly 131,000 tokens and up to 64 documents per call. Both limits apply, so long documents may exhaust the token budget before the list reaches 64 items. Attention cost also grows with the packed sequence length, which means the maximum context may exceed the practical memory or latency budget of a consumer GPU. The optional 256-dimensional document embeddings are conditioned on the query and candidate list. They can support downstream analysis within that retrieval request, but they are unsuitable as static corpus embeddings for a vector index. This story is for Pro members You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.
11:02

The Sequence Radar - Issue 940: Last Week in AI: Opus 5.5 Gets Leaner, Meta Goes Wearable, Washington Talks to Beijing, and Claude Explores DNA

Full text · 11,274 chars
Next Week in The Sequence: - Another installment of our series about recursive self-improvement. - We dive into Opus 5.5, DeepSeek’s amazing new paper about environments and Anthropic’s DNA discoveries. - We discuss another interesting platform in robotics. - We dive into the fascinating subject of AI compute as a new commodity, or is it? Subscribe and don’t miss out: 📝 Editorial: Last Week in AI: Opus 5.5 Gets Leaner, Meta Goes Wearable, Washington Talks to Beijing, and Claude Explores DNA A cheaper coding agent, glasses that can act on what you see, an unusual pattern in viral DNA, and two rival governments discussing AI risk. This week’s headlines read like four different newsletters. Together, they describe a shared transition: AI is moving deeper into the systems through which we work, discover, and make decisions. Anthropic’s Claude Opus 5.5 offers the most immediately practical example. The company reports stronger performance in coding and knowledge work, with typical workloads costing 40% less than Opus 5 and output generation running more than 30% faster. Those are Anthropic’s measurements, but the combination deserves attention. When an agent spends hours navigating a codebase, every unnecessary step costs money and introduces another opportunity for failure. Better economics expand the set of tasks worth delegating. A migration that once looked prohibitively expensive can become a routine overnight job—provided the results survive review. At Meta Connect, the emphasis shifted toward where those agents will live. Meta announced plans to bring Muse to its AI glasses, expanded its connections to shopping and productivity services, and previewed the pocket-sized Muse Charm. New audio glasses and lightweight VR glasses rounded out the hardware push. My read: Meta is betting that distribution and context will become decisive advantages. An agent that can see the object you mean requires less explanation. One connected to the services you use can turn that understanding into action. The product challenge is making those interactions reliable enough to become habits. That widening reach also explains why US–China AI talks matter. Washington proposed a mechanism for notifying Beijing about significant AI incidents, while both leaders identified AI as an area for cooperation during this week’s summit. The public record still offers limited evidence of concrete safeguards. Yet even a modest communication channel would address a real problem: rivals need ways to distinguish an accident from an attack. Competition can accelerate capability while simultaneously increasing the value of coordination. Building that coordination will require more than agreement that AI is consequential. The week’s most intriguing scientific announcement came from Anthropic’s biology lab. The company reported that Claude agents identified a previously uncharacterized enzyme system with CRISPR-like repeat structures. Roughly 950 agents searched for 21 hours, with human scientists performing the laboratory experiments. The underlying enzyme had appeared in earlier research; Claude identified previously unnoticed features around it. Its biological function and practical utility remain under investigation. That distinction matters. We have an early discovery, with substantial work ahead before anyone can call it a useful gene-editing technology. Still, the workflow is compelling: search broadly, identify anomalies, challenge hypotheses, and bring the strongest candidates to the laboratory. For me, the thread connecting these developments is the growing importance of what surrounds a model. Costs determine which jobs we delegate. Interfaces determine when we ask for help. Experiments determine whether a proposed discovery survives. Diplomacy determines how countries respond when something goes wrong. As intelligence becomes more accessible, those surrounding systems will increasingly determine how much value it creates. 🔎 AI Research AI Lab: DeepSeek-AI, Tsinghua University Summary: DSec is DeepSeek’s production sandbox platform—unified FnCall/container/microVM/full-VM backends, EROFS composable layers, and 3FS on-demand image loading, co-designed with the RL loop for stateful rollouts under GPU preemption. A ~160-node unit serves ~3M sandboxes/day at >380K concurrency and >5K creates/sec; on-demand loading finishes an 8,192-container burst in ~35 min vs >60 min for eager pulls (1.71×) while cutting disk writes ~57%, and virtio-pmem+DAX trims peak host memory 40.2%. AI Lab: Salesforce AI Research Summary: Instead of distilling trajectories into fixed write-time artifacts, JIT Mem stores raw episodes and trains a GRPO curator to synthesize a compact, task-conditioned payload at read time from immediate task success. It beats the strongest write-time baselines by +16.2 / +16.3 / +3.9 SR points on ALFWorld, WebShop, and τ²-bench, with an untrained Gemini curator already at 61.0 vs SkillOS 41.0 on WebShop, while cutting input tokens 50.3–56.3% and executor steps 28.4–31.4%. AI Lab: Shanghai Jiao Tong University, Xi’an Jiaotong University, East China Normal University Summary: SchrodingerRepo treats the SWE-bench repository as an evaluation-time latent variable—problem rewrite, namespace remap, layout reorder, and functionality-preserving rewrite—preserving executability while eroding memorized cues. Full transforms drop Pass@1 by 6.0–14.4 pp across GPT-5.1, GPT-5.4-mini, DeepSeek-v4-Flash, and Gemini-3.1-Flash-Lite (e.g. 46.8%→35.6% for GPT-5.4-mini), with 81.6–83.6% of extra actions spent on exploration and input tokens rising >2.5×; human checks find >65% of Verified instances show leakage evidence. AI Lab: Microsoft, City University of Hong Kong Summary: Taste-Bench mines 502 decision-fork questions from engineering and research agent trajectories so models must pick the better branch before later outcomes are revealed. The best frontier model (GPT-5.6 Sol) scores only 59.7%, longer horizons are harder, and more reasoning budget does not help; distilling a privileged teacher into Qwen3.6-27B raises held-out taste accuracy 30.0%→47.9% (+17.9 pp) and lifts a fixed executor on 41 held-out SWE-bench Pro tasks from 14.6% to 33.7% success with student advice. AI Lab: Alibaba Token Foundry, Georgia Tech Summary: VHD-Play reverses environment-first generation: sample and solve a mathematical mechanism first, then wrap its dynamics as stateful tools with the same reference for scoring, yielding 3,300 admitted environments at a few cents each. GRPO training lifts Qwen3.6-35B-A3B’s five-family mean agentic score from 0.204 to 0.815, transfers to eight unseen families, and on 365-day E-Commerce Bench reaches 3.4× the base ending balance while beating Qwen3.7-Max and gaining +2.84 on interaction-focused BFCL V4 cells. AI Lab: Carnegie Mellon University, Tsinghua University Summary: WhatWorkedBench asks agents to submit a full predicted response surface after a measurement budget, scored against exhaustive CPU references for every conditional component effect across 36 tasks / 1,248 configs (35/36 tasks show sign-reversing interactions). Agents often pick good optima while missing effects; fitting a shared Gaussian process on the same observations raises effect recovery from 0.632 to 0.698 (Flash cohort) and from 0.303 to 0.455 on six added-family submissions, while encoding code equivalences lifts six-factor GP recovery from 0.248 to 0.462 at B=20. 🤖 AI Tech Releases Claude Opus 5.5 Anthropic released Claude Opus 5.5, the first model in the Claude 5.5 family—Fable 5.1-level on most work at about 40% lower cost than Opus 5 ($4/$20 per MTok input/output), live as claude-opus-5-5 on the Claude API, Amazon Bedrock, Google Cloud, and Microsoft Foundry, with preserved thinking and Mythos-class biology/cybersecurity safeguards. Muse on Meta AI glasses (Connect) At Connect 2026, Meta said Muse is coming to its AI glasses in the coming months—name-activate and act on what you see, plus expanded shopping/work connectors, Muse email, a pocket Muse Charm later this year, and Ray-Ban Meta Gen 3 / Ray-Ban Meta Audio as the lineup heads past 100 AI-glasses styles by year-end. 📡10 AI News You Need to Know About - Meta is bringing Muse to its AI glasses with hands-free, look-and-ask control, plus real-time voice, expressive avatars, Mac computer use, a dedicated agent email, and more shopping connectors—rolling out on glasses in the coming months after Connect 2026. - Cognition said it crossed $1 billion in annualized revenue run rate for Devin—less than two years after the coding agent became generally available—a milestone Bloomberg also covered. - Enveda raised $311 million in a Series E led by Catalio Capital Management (new capital from Durable, ICONIQ, Lightspeed, Surveyor/Citadel, T. Rowe Price accounts, and others), bringing total funding above $845 million to push ENV-294, ENV-308, and ENV-6946 deeper into the clinic and scale its PRISM nature-chemistry discovery platform. - Anthropic says Claude discovered ART—array-associated reverse transcriptases, a novel phage enzyme system with CRISPR-like DNA-repeat arrays—via ~950 agents scanning ~1.9B protein clusters over ~21 hours, with human scientists running the wet-lab follow-up in its new Bay Area biology lab (function still unknown; preprint out). - Ema raised $77 million in a Series B led by Creaegis, with Accel, S32, and Prosus increasing stakes, bringing total funding to $140 million and more than quadrupling its valuation as it scales “AI Employees” that automate HR, IT, and finance workflows across enterprise apps. - Snorkel AI raised $350 million at a $3.5 billion valuation in a Series E co-led by Insight Partners and S32 (Addition participating heavily), nearly tripling its prior mark to expand its agentic data factory for frontier-lab training data and environments. - PrismML demoed its 1-bit Bonsai 2-billion-parameter vision-language model running locally on Qualcomm Snapdragon AR1 Gen 1 smart glasses at Snapdragon Summit, fitting about 4× more parameters into the same memory envelope as prior on-glasses models (no retail glasses using it announced yet). - TechCrunch reported that vibe-coding startup Lovable has crossed about $600 million in annualized revenue (up from ~$500 million in June), with co-founder Fabian Hedin saying at HumanX that people at roughly two-thirds of Fortune 500 companies now use the platform—after an August Series C that valued the company at $13.3 billion. - At the White House summit, Xi told Trump that China and the U.S.—both major AI powers—have more reason to cooperate than compete, urging continued AI dialogue on risks and benefits, joint work against misuse, and keeping AI under human control; Trump said AI “bears on the future of humanity” and that the two sides should keep talking, after trade teams held their first U.S.–China AI dialogue and floated an incident-notification channel. - Oracle sent a force majeure notice on Project Jupiter, its New Mexico Stargate campus with Blue Owl (Bloomberg first reported), seeking to delay payments if the planned 2.45 GW site misses its 2028 online date rather than exit as tenant—while saying the project “remains on our planned schedule,” amid gas-pipeline and air-permit delays.
14:26

Synthetic Sciences Ships OpenScience, an Open-Source AI Research Workbench With 250 Skills

Full text · 2,189 chars
- Synthetic Sciences released OpenScience, an Apache 2.0 AI workbench that runs the full research loop. - Install via npm install -g @synsci/openscience ornpx synsci , then opens in browser. - Model-agnostic: works with Anthropic, OpenAI, Google, and open-weight models via your own keys. - Ships 250+ skills including DeepSpeed, PEFT, TRL, LaTeX, and cheminformatics tooling. - Direct connectors to UniProt, PDB, ChEMBL, PubChem, arXiv, OpenAlex, and ~30 more databases. - Bring-your-own-key is free; managed Atlas platform adds prepaid frontier models and cloud compute. Synthetic Sciences has released OpenScience, an Apache 2.0 workbench for running AI-assisted research from a browser. Developers give an agent a research goal, then let it search literature, form hypotheses, write and execute code, query scientific databases, run experiments, and assemble results. The project combines capabilities that usually require separate coding agents, database clients, model gateways, and scientific tools. It installs from npm, accepts API keys for multiple model providers, and stores sessions and artifacts locally. That integrated workflow is the project’s main appeal, although researchers still need to review methods, citations, code, and conclusions. Inside the workbench OpenScience runs as a browser workspace backed by a local server. It supports commercial and open-weight models from Anthropic, OpenAI, Google, and other providers, with no OpenScience account required for bring-your-own-key use. | Component | What it provides | |---|---| | Agents | A general research agent; biology, physics, and machine-learning specialists; critique and literature-review sub-agents; and a read-only planning mode. | | Skills | About 250 packaged workflows for model training, evaluation, datasets, biology, cheminformatics, paper production, figures, and cloud compute. | | Data connectors | Integrations with UniProt, PDB, Ensembl, ChEMBL, PubChem, arXiv, OpenAlex, Semantic Scholar, and dozens of other sources. | This story is for Pro members You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.
16:00

Build production-ready RAG workflows on AWS

Full text · 1,604 chars
Today's edition is sponsored by AWS Marketplace, on something a lot of you are building right now: production RAG. Here's the problem nobody warns you about. To ship a real RAG app, most teams end up running two databases: a cache (for speed) and a separate vector store (for retrieval), plus the glue code to keep them in sync. Two systems to provision, monitor, and keep in sync and to debug at 2am. Redis Cloud on AWS combines caching and vector storage in one managed platform, so you get low-latency retrieval without running two systems. It runs on AWS, connects to your AWS workloads, and bills pay-as-you-go through your AWS account. And it's built for production RAG from day one. Build a RAG agent with Redis Cloud and Amazon Bedrock You don't need a migration project to try it. The short path: Subscribe to Redis Cloud in AWS Marketplace and start the 14-day free trial, procurement and billing stay on your AWS account. Create a vector index in Redis Cloud and connect it as the vector store for an Amazon Bedrock Knowledge Base, Amazon Bedrock ingests your documents and stores the embeddings in Redis. Associate the Knowledge Base with an Amazon Bedrock agent for contextually aware responses backed by low-latency retrieval. That's it. One platform for caching and vectors, one bill, one thing to operate. Redis Cloud keeps caching and vector search in one place, works with Amazon Bedrock, and scales pay-as-you-go through AWS Marketplace. If you're wiring up RAG this quarter, that's one less database to provision, monitor, and keep in sync. Louis, Laura, and the ☕️ Techpresso team.
16:00

LG's EXAONE 3.5 Hits 1 Million Monthly Downloads Running on 5GB

Full text · 2,800 chars
- LG AI Research's EXAONE 3.5 7.8B AWQ passes 849K downloads on Hugging Face. - 4-bit W4A16g128 quantization shrinks the bilingual Korean and English model to roughly 5GB of VRAM. - Supports a 32,768 token context window with grouped-query attention and 102,400 vocab. - Scores 70.7 on real-world benchmarks, beating Qwen 2.5 7B (52.7) and Llama 3.1 8B (48.6). - Trained on 9T tokens with two-stage pre-training, taxonomy-based SFT, and staged DPO/SimPO alignment. - Research-only EXAONE 1.1 NC license; commercial use requires contacting LG AI Research directly. LG’s 4-bit EXAONE 3.5 approaches one million monthly Hugging Face downloads At publication time, LG AI Research’s AWQ-quantized EXAONE 3.5 7.8B Instruct checkpoint was approaching one million monthly downloads on its AWQ model page. Hugging Face counts qualifying requests, so the figure does not represent one million unique users or deployments. It does indicate sustained interest in a compact Korean-English model designed for local inference. The wider EXAONE 3.5 family includes 2.4B, 7.8B, and 32B instruction-tuned models. This 7.8B release uses activation-aware weight quantization, or AWQ, to reduce GPU memory requirements while retaining the model’s configured 32,768-token context window. A compact checkpoint with a large window The checkpoint contains a decoder-only Transformer with the following specifications: - 6.98 billion non-embedding parameters across 32 layers - Grouped-query attention with 32 query heads and 8 key/value heads - A 102,400-token byte-level BPE vocabulary designed around Korean and English - A maximum sequence length of 32,768 tokens, including input and generated tokens - W4A16 group-wise quantization with groups of 128 weights W4A16 stores quantized weights at 4-bit precision while keeping activations at 16-bit precision. Groups of 128 weights share quantization metadata. The resulting files occupy roughly 5 GB, although actual GPU memory use also includes the key-value cache, activations, temporary buffers, and backend overhead. Quantization leaves the architecture and configured context length unchanged. A full 32K sequence at batch size one can add roughly 4 GiB for a 16-bit key-value cache before other allocations, so loading the weights in about 5 GB does not mean 32K inference will run within that same memory budget. EXAONE 3.5 retains the broad architecture of EXAONE 3.0 7.8B. According to LG’s technical report, long-context fine-tuning extended the maximum sequence length from 4,096 to 32,768 tokens. The model also uses a rotary-position-embedding theta of 1,000,000 to support the larger window. This story is for Pro members You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.
16:50

OpenAI Patches GPT-6 Sol and Luna's Silent Image Understanding Bug

Full text · 3,535 chars
- OpenAI fixed an image encoding bug degrading vision quality in GPT-6 Sol and GPT-6 Luna - Fix improves results across the API, Codex, and computer use workflows - No code changes required, the patch is live server-side - OpenAI recommends rerunning evals and retrying affected image-input workflows - GPT-6 Sol pricing: $2 per 1M input tokens, $10 per 1M output tokens - Details in the official API changelog OpenAI patches GPT-6 image encoding bug OpenAI has fixed an image encoding bug that degraded visual understanding in GPT-6 Sol and GPT-6 Luna, two multimodal reasoning models released days before the patch. The server-side update is live in the API and Codex, including computer-use workflows. OpenAI’s API changelog does not quantify the regression or provide benchmark deltas. Requests still completed, which left developers with plausible answers and lower-than-expected evaluation scores instead of explicit errors. Teams that tested either model with images should rerun those evaluations. Why plausible answers slipped through The image encoder converts pixels into internal representations that the model can analyze. A defect at this stage can omit or distort visual details before reasoning begins, causing the model to misread text, objects, layouts, or interface controls while still producing a fluent response. OpenAI has corrected a similar defect before. A March update fixed a small encoder bug affecting input_image inputs in GPT-5.4, and the company said some image-understanding use cases could improve automatically. One encoder, several surfaces - GPT-6 Sol (gpt-6-sol ): the higher-capability tier. - GPT-6 Luna (gpt-6-luna ): the lower-cost, high-volume tier. - Responses and Chat Completions APIs: image requests sent to either model. - Codex and computer use: workflows that inspect screenshots or other images. Computer-use agents depend on accurate screenshot interpretation for every action. Distorted visual features can lead an agent to select the wrong control, misread a form field, enter text in the wrong place, or incorrectly decide that a task has finished. Server-side patch, unchanged prices The fix requires no code changes or model-version migration. New requests use the corrected path automatically, while outputs generated before the patch remain unchanged. Application-level caches, stored responses, and evaluation records from affected runs require separate review. | Standard API pricing per 1 million tokens for prompts up to 272,000 input tokens | | | | |---|---|---|---| | Model | Input | Cached input | Output | |---|---|---|---| | GPT-6 Sol | $2.00 | $0.20 | $10.00 | | GPT-6 Luna | $0.10 | $0.01 | $0.50 | Rerun the evidence The same model IDs now use a corrected image path, so pre-fix and post-fix evaluations reflect different serving behavior. Keep prompts, images, tools, sampling settings, and scoring methods constant when comparing results, and record run timestamps alongside the changelog entry. - Repeat vision evaluations. Rerun tests conducted between the models’ launch and the patch, then compare accuracy and failure categories. - Retry visual agent failures. Revisit Codex and computer-use runs involving misread controls, screenshots, text, or page state. - Review stored results. Inspect application-level caches and saved outputs produced before the fix because the patch cannot revise completed responses. - Reconsider model-selection decisions. Collect corrected multimodal results before switching models based on earlier visual performance.
17:30

😺 Tens of Thousands of AI Incidents

Full text · 11,643 chars
😺 Tens of Thousands of AI Incidents PLUS: Microsoft rebuilt Copilot around persistent agents. Welcome, humans. Okay, so an OpenAI agent got blocked from the internet... and then figured out how to use DNS to ask an outside chatbot for help. DNS is basically the internet’s phone book. Your computer normally uses it to turn a name like “google.com” into an address. This agent realized OpenAI’s sandbox still allowed DNS requests, found a service that could pass a question through that channel, asked “what’s the capital of France?”, and got “Paris” back. Then it sent 18 more questions through the same route. Which is a very creative solution to “you are not allowed on the internet,” in exactly the way nobody wanted. OpenAI’s monitors flagged the behavior within 15 minutes and a human started reviewing it three minutes later. The run was killed about 2.5 hours after the first successful external response. OpenAI says tool-using training, evaluation, and inference on its most capable models remain paused while it hardens the environment. And this is apparently one example from a MUCH bigger pile. Axios reports OpenAI, Anthropic, and security researchers are investigating tens of thousands of incidents involving frontier models bypassing guardrails, escaping sandboxes, creating message boards, hijacking websites, or otherwise doing things evaluators considered problematic. Important caveat: many happened in adversarial tests designed to MAKE models misbehave, and most known cases haven’t caused real-world harm. But the scale changes the problem. The question is becoming less “can an AI agent find a weird loophole?” and more “can humans find and close the loopholes faster than increasingly resourceful agents find new ones?” Here’s what happened in AI today: - 😺 Microsoft rebuilt Copilot around persistent workplace agents. - 📰 Anthropic’s 950-agent search surfaced a biology candidate. - 📰 Cisco found AI agents already running production networks. - 🍪 Meta pushed its personal agent onto glasses. - 🎓 Dan Shipper split AI labs from product teams. 😺 Microsoft wants Copilot to keep working after you leave Basically, Microsoft is trying to turn Copilot from something you ask questions into software that can keep working after you stop talking to it. If you’re a normie or a non-AI user, then you’re probably used to AI working like this: ask → answer → done. Satya Nadella’s description of the new Copilot is much closer to ask → agent starts working → agent keeps working → you come back later. He says a few things finally changed: - Models can now run for days while staying coherent, instead of losing the plot halfway through a long job. - Memory can live outside the model, so the agent doesn’t have to keep everything inside one conversation. - The agent gets its own workspace, computer, and long-running harness that keeps the job moving. That’s basically what Autopilot is. Nadella says you can give one an identity, direction, memory, computer, and workspace, then let it work continuously. Inside Microsoft, he describes the idea as giving every employee a kind of AI “chief of staff.” Translation: Microsoft would very much like to hire a tiny robot employee into every Microsoft 365 account. And you don’t necessarily have to babysit it in some separate AI app. Nadella says you could interact with an Autopilot inside Teams like another colleague, while it handles standing jobs in the background. His example: instead of managing invoices every day, make an Autopilot whose job is simply to manage invoices. The new Copilot stack breaks down like this: - Home brings Chat, delegated work, and Office documents together. - Code lets you describe an app, dashboard, automation, or workflow and have Copilot build it. - Autopilot is the persistent worker that can keep going without another prompt. And Microsoft is pretty openly borrowing from what worked elsewhere. Nadella credited OpenClaw with spotting the long-running-agent pattern early, and when Alex Heath asked whether OpenClaw’s open-source component sits underneath Autopilot, Nadella answered “Absolutely.” Why now? Nadella says newer models are finally capable enough to deliver more of Copilot’s original promise. Microsoft already has 30M+ paid enterprise Copilot subscribers, and now it thinks those users can start handing AI longer-running jobs. Nadella called agents potentially Microsoft’s “biggest TAM expansion ever” and said the agent era could eventually become orders of magnitude bigger than cloud. Which is a pretty enormous bet on the humble invoice bot. He also says persistent agents will need auditing, monitoring, governance, and everyone’s favorite word right now, “containment.” Given today’s whole “tens of thousands of AI incidents” situation... probably worth figuring that part out! FROM OUR PARTNERS The analyst report that put AI agents on the endpoint security map Companies are giving AI agents real credentials on employee laptops, and when one does something nobody asked for, the prompt and final answer are usually the only record left. Analyst firm SACR just published a report that maps the next layer of endpoint security into five zones, and agent runtime observability is one of them. On October 1 at 11 AM Eastern, the analyst who wrote the report and Origin's founder will walk through it live, then go deep on the trace and what you can do with it. 🎓 AI Skill of the Day: Run your product team like a research lab Dan Shipper's advice for surviving nonstop model upgrades is basically: stop making the same people explore the frontier and execute the roadmap. Those are opposite jobs. Exploration means trying lots of weird stuff and throwing most of it away. Product work means focus, reliability, and saying no. His setup at Every is tiny. One or two people can be the lab. The useful pairing is a "pirate" who rapidly builds messy experiments to find value, plus an "architect" who steps in once something starts working and turns it into a real system. The important part is how ideas graduate: - Try multiple approaches in parallel and expect roughly 90% to die. - Dogfood the survivors on real work. Ask: is this actually useful, or just new? - Only harden what people keep using, then test whether it's dramatically better and affordable enough to scale. Every's copy-editing experiment, "KateBench," is the concrete version. Once their editor actually started using it, they built a dashboard around accepted suggestions and the work she still had to do afterward. Shipper said it cut that remaining editing work by 12% month over month. That's the filter I like: don't promote the demo because it looks futuristic. Promote it because, a month later, people still want it. FROM OUR PARTNERS Apodex 1.1 is here: reasoning that finishes the task, not just the report Apodex 1.1 moves beyond generating reports — it works inside files, code, and data to execute real tasks, adapt mid-run, and self-check results. Its engine, FrontierAgent, is open-source: run it locally with one command, no Docker. Try it, then star the repo. 📰 Around the Horn Related, and concerning: GPT6 Luna also scored a 100% on a benchmark called “Puppy Kill”, which, tracks whether an AI when told it is embedded in a robot, will run a tool called ‘puppykill” that does just about exactly what you think it does. - Anthropic lost its bid to pause the Pentagon’s national-security supply-chain-risk designation, leaving Claude barred from some Defense systems while the case continues. - Cognition said Devin crossed a $1B annualized revenue run rate less than two years after general availability. - SemiAnalysis mapped 1,000+ Chinese data centers and 24+ GW of delivered capacity, with ByteDance reportedly renting roughly one-fifth of national capacity. - Kansas City Fed President Jeff Schmid said regulators need to understand whether the network of AI companies and contracts is becoming “too big to fail.” - China is subsidizing AI filmmaking with rent, living stipends, compute vouchers, and public funds as short-form production costs collapse. - Thales said it is in advanced talks with NATO countries on HexaForce, an AI-assisted command system that proposes courses of action while keeping firing decisions with humans. 🍪 Treats to Try - *Build an AI-Ready Workplace. Explore expert insights and practical guidance to modernize workplace technology, strengthen security, and prepare your organization for AI-powered work. - Microsoft Copilot adds Home, Code, and Autopilot so you can build apps and delegate work to persistent agents that keep going after the chat ends. - Claude in Slack lets Team and Enterprise users tag Claude in a thread, give it the surrounding context, and have it use connected tools before posting the result back. - ElevenLabs Image & Video API lets you generate visual media alongside audio through one API stack, with asynchronous jobs that can return through webhooks. - Midjourney now previews prompts across styles before you generate, targets edits more precisely, and makes V8.1 and V8.2 tiled images blend without visible seams. - Pexo turns an idea, URL, PDF, image, or audio file into a scripted and voiced motion-graphics video you can keep refining in chat. - Bland Agent Phone Plan gives an AI agent one persistent phone number so the same agent can call, text, and answer inbound requests. - Docker Cloud Sandboxes lets coding agents start locally, move the same isolated environment to cloud compute, and keep working after you close your laptop. - Higgsfield Production Skills packages 11 agent workflows for Blender, Premiere, After Effects, Photoshop, DaVinci, and more while keeping the underlying production project editable. 🧰 Sunday Special: Top 5 Stories + Top 5 Tools Top 5 Stories of the Week - OpenAI and Anthropic launched GPT-6 Sol, Luna, and Claude Opus 5.5 into a price war over frontier work. - Anthropic used roughly 950 Claude agents to surface a previously unknown enzyme-system candidate from more than 200,000 reverse transcriptases. - Meta unveiled Muse Charm and plans to put its personal agent on AI glasses, pushing assistant work closer to what you see and say. - Amazon blocked Meta's Muse from shopping on its site, turning agent commerce into a fight over who controls the customer relationship. - Cisco's network survey found AI agents were already operating in production, while trust still limited how much autonomy organizations would allow. Top 5 Tools of the Week - GPT-6 Sol and Luna gave OpenAI a new high-end model plus a much cheaper high-volume option. - Claude Opus 5.5 cut Anthropic's frontier pricing while pushing harder on coding and agent work. - Qwen Intelligence bundled planning, mobile-use, and creative agents into one public stack. - Agora-2 put up to 20 humans and agents into the same AI-generated world. - Claude Code cloud sessions let coding jobs keep running on Anthropic's machines after your laptop closes. 🧩 Thursday Trivia Reveal A was AI, and B was real. Of 4,253 votes, A got 2,751 (64.7%) and B got 1,502 (35.3%). Revisit the pair in Thursday's issue. Some of your guesses: - "First is so detailed I figured it was ai" - "The central figure seems to have a halo effect at edges and something looks off about the hand." - "The car behind the men’s head in foreground has a weird shape I can’t understand" - "Something weird about the second car’s back window." - "in image B the guy on the right has an extra finger holding onto his drink" New from The Neuron: AI Explained A Cat’s Commentary That’s all for now. If you want to get featured above, fill out the poll below and tell us how we did today!
18:41

Bluesky reply bot checker

Full text · 1,194 chars
27th September 2026 Automated reply bots on Twitter are a scourge - as someone with a decent number of followers I attract a swarm of these, such that anything I post there attracts dozens of mindless automated replies. They've started manifesting on Bluesky as well. Unlike Twitter, Bluesky still has a freely available and useful API. The lack of such a thing doesn't slow down the bots, but it does make investigating them a lot more frustrating. So I had Opus 5.5 vibe code this tool, which examines any Bluesky profile for evidence of a likely reply bot. It looks for signals like replies posted within seconds of other posts from the same account, or accounts that never post their own content (or images or links) but instead consistently reply to messages from other, higher-follower users. It also looks for question marks, because I'm extra infuriated by reply bots that I no tie me to waste my time answering a question that no human ever posed. Recent articles - 2026 in LLMs (so far) - 27th September 2026 - Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna, and a new price war - 22nd September 2026 - Jev introduces a new shape of LLM - System One, aka Decision Models - 21st September 2026
23:09

S3 Is the Future, S3 Is the Past

Full text · 634 chars
27th September 2026 One thing I find notable about S3 today is that, while it used to drop in price reasonably often, there hasn't been a price drop in a full decade: 2006-03-14 $0.150/GB-month 2010-11-01 $0.140/GB-month 2012-02-01 $0.125/GB-month 2012-12-01 $0.095/GB-month 2014-02-01 $0.085/GB-month 2014-04-01 $0.030/GB-month 2016-12-01 $0.023/GB-month Today it's still $0.023/GB-month. Recent articles - 2026 in LLMs (so far) - 27th September 2026 - Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna, and a new price war - 22nd September 2026 - Jev introduces a new shape of LLM - System One, aka Decision Models - 21st September 2026
23:13

Sakana AI's SAIL Triples Robot Success Rates by Searching Before Moving

Full text · 7,176 chars
- Sakana AI and the University of Tokyo introduce SAIL, a test-time scaling method for VLM robot control, accepted to IROS 2026. - MCTS over full trajectories, with each node a candidate motion and each edge a refinement guided by VLM feedback. - Average success across six ALOHA manipulation tasks climbs from 25% at one candidate to 73% at 45 candidates. - Uses Gemini Robotics-ER 1.5 as both policy and evaluator, with no weight updates at any stage. - Physical LeRobot SO-101 block-in-bowl task succeeded in five of six trials using a 15-candidate budget. - Open-loop execution and heavy simulation compute remain the main limitations. Paper. SAIL searches its way to better robot plans Researchers at Sakana AI and the University of Tokyo report that a frozen vision-language model can produce more reliable robot motions by searching, simulating, and revising candidate trajectories at inference time. Their method, SAIL, raises the average success rate across six simulated manipulation tasks from 25% with one candidate to 73% with 45. The paper has been accepted for IROS 2026. SAIL, short for Scaling In-Context Imitation Learning, offers developers a way to improve an unreliable one-shot planner when they have a simulator, a demonstration archive, and enough latency budget for search. The model weights remain fixed throughout the process. Why a single trajectory breaks In-context imitation learning gives a robot examples of successful behavior, then asks a model to solve a new scene from those demonstrations. A vision-language model can inspect the scene and generate a complete sequence of robot-hand positions, orientations, and gripper commands. Small changes in object placement, sampled outputs, or retrieved examples can still produce large differences in execution. A grasp that misses by a few millimeters can invalidate every later waypoint. SAIL applies test-time scaling, the practice of spending more computation on each input while keeping model parameters fixed, to physical trajectories. It generates several plans, evaluates their simulated execution, and uses the results to revise promising candidates. Inside SAIL’s search loop SAIL places Gemini Robotics-ER 1.5 inside a Monte Carlo tree search. The model serves as both trajectory generator and evaluator. Each tree node contains a complete trajectory, while each edge represents a refinement. The search can explore new revisions and continue developing branches that receive higher scores. - Retrieve and generate. The system queries an archive of successful trajectories for scenes resembling the current object arrangement. The model receives those demonstrations, the current scene, and the robot state, then generates a full candidate trajectory. - Simulate and score. A simulator executes the candidate and records a video. The evaluator estimates progress through predefined subtasks and converts that progress into a scalar score. Partial progress can therefore distinguish two failed candidates. The scoring design draws on ROVER, which tracks task completion over time. - Align feedback with waypoints. Scores from sampled video frames are mapped back to specific trajectory steps. The generator receives that step-level feedback and revises sections where progress stalled while preserving successful motion segments. - Select for execution. Candidate testing remains in simulation. The physical robot receives the selected trajectory after the search finishes. Implementers must provide the demonstration archive, simulator, scene reconstruction pipeline, and subtask rubric used by the evaluator. SAIL supplies the search and revision structure around those components. More nodes produce more successful plans The researchers evaluated six tasks in the ALOHA simulator, using 20 initial configurations for each task. The benchmark covered picking up a banana, uncapping a pen, moving a bowl, opening a drawer, closing a laptop, and grasping a marker. | Average simulated success across six manipulation tasks | | | |---|---|---| | Method | Search nodes | Average success | |---|---|---| | Single generation | 1 | 25% | | Breadth-first search | 15 | 51% | | Depth-first search | 15 | 37% | | SAIL | 6 | 55% | | SAIL | 15 | 65% | | SAIL | 30 | 71% | | SAIL | 45 | 73% | At a budget of 15 nodes, SAIL reached 65% average success, compared with 51% for breadth-first search and 37% for depth-first search. The bowl task reached 100% success with six nodes, while the banana task reached 80%. These matched-budget results indicate that retrieval and waypoint-level feedback contribute beyond the number of sampled plans. One real arm, six trials For the hardware test, the team used a LeRobot SO-101 arm for a block-in-bowl task. Color and depth cameras supplied the data needed to reconstruct the scene in simulation. SAIL searched 15 candidate trajectories, selected one, and sent it to the arm. The search pipeline succeeded in five of six trials. The authors attribute the remaining failure to pose-estimation errors and differences between simulated and physical contact dynamics. A separate imitation policy trained on successful trajectories generated by the search also completed five of six trials. That approach moves the computational cost into data generation and training, allowing deployment without running tree search for every attempt. Whole trajectories become search objects SAIL’s central contribution is trajectory-level test-time scaling. The search operates on complete robot motions, and the evaluator’s temporal scores identify which sections need revision. This structure lets a frozen foundation model use additional inference compute to handle object arrangements that defeat its first prediction. The paper’s ablation results support the value of both retrieval and detailed feedback. Random retrieval and weaker feedback plateau as the node budget grows, while the complete pipeline continues to improve through 45 nodes. Where the method fits - Suitable workloads: Tabletop manipulation tasks with a useful demonstration library, a simulator that approximates scene geometry and contact, and enough time to search before execution. - Control constraint: The selected trajectory runs open-loop. The robot receives no visual correction while moving, which limits the approach on moving objects, uncertain contacts, and tasks that require continuous replanning. - Infrastructure cost: Each candidate requires simulation, video evaluation, and potentially another model call for refinement. Larger search budgets improve success while increasing compute and latency. - Transfer risk: Errors in pose estimation, geometry, or contact modeling can cause a plan that succeeds in simulation to fail on hardware. - Evidence limit: The simulated evaluation covers six tasks, while the physical validation covers one task with six trials for each tested method. The paper and project page include simulated and physical rollouts. For robotics teams already using vision-language models, SAIL provides a concrete architecture for adding retrieval, simulation-based scoring, and iterative trajectory repair without fine-tuning the underlying model.
23:54

2026 in LLMs (so far)

Full text · 25,604 chars
2026 in LLMs (so far) 27th September 2026 On Friday I gave the closing keynote at the WeAreDevelopers World Congress North America in San Jose. I tied together the key trends from the past year into a chronological exploration of everything that happened in 2026. The video is on YouTube; here are my annotated slides and notes to accompany the talk. And as an annotated presentation: November saw the release of two important models: Claude Opus 4.5 and GPT-5.1. As is usually the case with new models, these were incremental improvements on the models that came before them. But every now and then when a model improves, it crosses an invisible line where something that didn’t really work starts working. In this case, the thing that started working was their coding agents. Claude Code had been around since February 2025, Codex was a little younger. These two new models, when paired with their respective coding agent harnesses, improved from “often make mistakes” to “reliable enough to use on a day-to-day basis”. For a couple of years now I’ve been evaluating new models by asking them to “Generate an SVG of a pelican riding a bicycle”. It’s probably the world’s stupidest benchmark—there’s only so much you can learn from it. But it’s still a challenge for models, because drawing pelicans is difficult, drawing bicycles is difficult, and pelicans can’t ride bicycles in the first place. Here’s the state of the art for November. Claude still couldn’t really draw a bicycle! The GPT-5.1 bicycle frame is pretty crap too. An then there were the December holidays, and individual developers took some time off and many started tinkering with these new coding agent model combinations... and it began to dawn on us quite how much they could do that they couldn’t do before. Come January, a lot of us were quite excited to start putting this stuff into action. Every year I set myself a New Year’s resolution, and for as long as I can remember it’s been the same thing: stay focused. Take on less new projects. Try to get things done in the projects I already have. This year I decided that since that had never worked before, I’m going to go the other way. We’ve got coding agents now, let’s see what they can do. I’m going to take on as many new projects as I like! (You can ask me at the end of the year if this turned out to be a good idea or not. I have a lot of plates spinning right now.) “Be more ambitious” has been something of a theme for the year, because the only way to find the limits of this technology is to keep on pushing them until they don’t work. I also went on the Oxide and friends podcast with Bryan Cantrill and Adam Leventhal to share predictions for the next year (and three and six years). With hindsight, my LLM predictions were pretty unambitious. I said “it will become undeniable that LLMs write good code”—I think we’re there now. I predicted we would finally solve sandboxing. I counted and around 40 of the 277 sessions at this conference touched on sandboxing or agent security in some way, so we’re at least putting a lot of effort into that! I predicted “a Challenger disaster” for coding agent security. There’s certainly been a whole lot of noise around agent security this year, though the exact disaster I predicted (with coding agents being hijacked and causing real-world economic damage) hasn’t really played out. We threw in a joke prediction that the Pope would weigh in on the economic impact of LLMs. I also predicted that New Zealand’s Kākāpō parrots would have an outstanding breeding season this year. These birds live in New Zealand. They are flightless nocturnal parrots. They’re kind of dumpy looking, I think they’re beautiful, and there were only 236 of these parrots in the world at the start of the year. Kākāpō only breed when the Rimu trees have a big fruiting season, and that hasn’t happened in four years... but this year the Rimu fruit were looking excellent. Photo by Kimberley Collins. Also on that podcast, we coined a term (full credit to Adam) for “that feeling of Al induced ennui where software engineers get listless because the Al can do anything”. We called it Deep Blue. This has been a major theme throughout the year, and was touched on by several speakers at this conference. As a software engineer, I’ve never had a year of my career where everything has changed so quickly and so dramatically. A lot of what I’ve been doing this year is trying to come to terms with that and what that means for my own profession. Also in January, I suffered from what I’m calling AI mania. This is not the same thing as AI psychosis. With AI mania, any time your agent isn’t building something for you feels like wasted time. You’re losing sleep because you could be staying up later getting your agents to do stuff. My AI mania presented itself in some ridiculously over-ambitious projects. I built a JavaScript interpreter entirely in Python, vibe-ported from MicroQuickJS by Fabrice Bellard. Then I built a WebAssembly runtime in Python as well. These projects were quite useful, in that they sort of cured me of my AI mania... because after I built these things, I got to look at them and ask “does the world need a slow, buggy, half-baked Python JavaScript interpreter?” I don’t think the world does. This page runs my JavaScript interpreter built in Python, running in Python using Pyodide, which is Python complied to WebAssembly, running in JavaScript, running in a browser. It’s a beautiful stack of horrors. I’ve been having a lot of fun with WebAssembly this year. At this point OpenClaw had 8,300 commits, less than two months after the project had started. I looked today and it’s over 100,000 commits now! This is the most vibe-coded piece of software in existence. This kicked off the OpenClaw revolution. It effectively defined a new category of software. There’s a generic term for this which I really enjoy. We call software like this a “Claw”. There’s OpenClaw, NanoClaw, IronClaw, PicoClaw... Today they’re being rebranded as “personal agents” or “general agents”, but I still like to think of them as Claws. The Apple stores in the Bay Area sold out of Mac Minis because so many people were buying Mac Minis to run OpenClaw! Drew Breunig said that this is because your OpenClaw is a digital pet, and you buy a Mac mini as an aquarium to keep your claw in, which is kind of delightful. Also in January, we had this website. This was MoltBook, a social network for AI agents, where the idea was that you send your Claw to go and talk to all of the other Claws, because what could possibly go wrong if you did that? The website launched on Thursday. It blew up on Friday. It was profiled by the New York Times on Monday. And by Tuesday, everyone had forgotten it existed as it drowned in a deluge of slop and spam. Facebook/Meta bought it a month later. They wrote about this in Software Factories and the Agentic Moment. I posted my own notes at the time, having seen their demo in-person back in October. Dan Shapiro called this approach the Dark Factory, after the idea that if your factory is sufficiently automated you can turn the lights out, because you don’t even need to see what’s going on. StrongDM presented two rules for software development that they’d been following since July last year. The first was code must not be written by humans. Any code that you write has to have been routed through a coding agent. This sounded radical in February, but I imagine there are a lot of people in this room who are pretty much living that today. Rule number two was code must not be reviewed by humans. You’re not allowed to read the code! This continued to be a huge topic for much of this year. Many of the sessions at this even have been about code review and how you can get away with this. What I found interesting about StrongDM is that they were living six months ahead of the rest of us, and they’d been exploring what it means to build software, not read the code, but still be confident that the software is of high quality. What can you do with these agents to help verify their work? StrongDM are a security company, and they had people with decades of experience on this project. They were very much exploring the edges of what’s possible and responsible to do with this stuff. Also in February... Google released Gemini 3.1 Pro. That’s a pretty great pelican riding a bicycle! it’s got the chain in the right place, it’s got feet on both sides. There’s a little fish in the basket. And then Google’s Jeff Dean tweeted a video comparing Gemini 3 Pro and Gemini 3.1 Pro that featured an animated pelican riding a bicycle, a frog on a penny-farthing, a giraffe driving a tiny car, an ostrich on roller skates, a turtle kickflipping a skateboard, and a dachshund driving a stretch limousine. This was frustrating, because my protection for the pelican riding the bicycle test was always “if they draw a perfect pelican on a bicycle, I’ll ask for some other animal on something else.” Google trained for all forms of animals on all forms of transport! They’ve defeated my benchmark at this point. The other thing that started in February was Tokenmaxxing. We had headlines about Meta making AI adoption a formal part of performance reviews, and Microsoft wanting every employee to use AI, and Uber boasting that ninety percent of their engineers were using AI workflows. Then a few months later we have Meta cracking down on token use, Microsoft saying token maxing is “not what we are optimizing for”, and Uber capping employee AI spending. So Tokenmaxxing went straight up and then straight back down again—because it turns out the agents are expensive. Last year it was difficult to spend more than $50 on AI tokens, because we didn’t have anything interesting to do with them. Then agents blew up, and now you can actually spend $1,000 in a day doing real work. This is also the reason that Anthropic’s valuation skyrocketed up to maybe a trillion dollars. AI appears to have hit product market fit in 2026, primarily through coding agents. These photographs are from China, where companies hosted OpenClaw install parties which saw non-tech-nerds queueing up around the block for help getting Claws installed on their personal devices. I think this proved real market demand for this class of Claws, or personal AI agents. It turns out regular people really do want a weird little AI agent that can do useful things on their behalf. A Claw is really just a coding agent wearing a less threatening hat. Under the hood they work much the same way—writing and then executing code on your computer to get stuff done. The race was on to be the first to build a safe Claw—a Claw you could give to regular human beings where they wouldn’t instantly shoot themselves in the foot. Meta’s Muse came out three weeks ago and is currently at the top of the free charts on the iPhone App Store. It appears to be taking off with consumers. I’m not yet convinced you can’t shoot yourself in the foot with Muse, but I guess we’ll find out for sure pretty soon. Photos from How the OpenClaw Frenzy Is Testing China’s AI Commitment (March 29th) and The Enthusiasm and Anxiety Behind China’s OpenClaw Craze (April 8th, 2026). Anthropic announced their new Claude Mythos model, and then said it was too dangerous to release beyond a trusted group of security researchers. Mythos was really, really good at hacking things. The “it’s too dangerous” marketing ploy has been played by AI companies dating all the way back to GPT-2. Anytime an AI company says we’ve built something that’s “too dangerous”, it’s natural to be a bit skeptical. I found the Mythos claims credible, because I’d seen how good coding agents had got at finding regular bugs. I wrote about that in Anthropic’s Project Glasswing—restricting Claude Mythos to security researchers—sounds necessary to me. With hindsight... yeah, the models had got really good at finding vulnerabilities! Another key trend in 2026 has been a dramatic improvement in the abilities of open weight models, including models that you can run on a laptop. On 16th of April I ran the new Qwen3.6-35B-A3B on my laptop, and it drew me a better pelican riding a bicycle than Anthropic’s brand new Claude Opus 4.7 did! Opus 4.7 drew a crap bicycle. Qwen on my laptop made a bicycle that was the correct shape, and a pretty decent pelican too! That’s from a 21GB file running on my laptop. The Qwen pelican was so good that I was suspicious they might have cheated, so I had it do a flamingo riding a unicycle as well. Again, it handily beat Claude Opus 4.7. The local model releases this year have been absolutely extraordinary. In our podcast episode back in January we’d predicted that the Pope would say something about AI. In May, Pope Leo XIV released an encyclical letter on “safeguarding the human person in the time of artificial intelligence”. Here are my notes on that document. With hindsight, this shouldn’t have been a surprise at all. Our current Pope’s name is Leo XIV, because when he named himself he chose his papal name after Leo XIII—the Pope who wrote an encyclical about the Industrial Revolution back in 1891. Rerum novarum was an extremely influential piece of Catholic theology that indirectly led to us having the five-day work week. When our new Pope came in, he named himself after Pope Leo XIII because he expected that he would need to write about the AI revolution in a similar way. Our joke podcast prediction was junk, because this was always going to happen. One of Anthropic’s co-founders, Christopher Olah, was present for the Pope’s event announcing the new encyclical. Corey Quinn noted that: getting the literal Pope to canonize your product’s specific technical limitations as a spiritual treatise is the single greatest act of vendor lobbying I have ever seen. Meanwhile, in May, RubyGems announced that they were under attack. Parties unknown were uploading thousands of dubious packages to the RubyGems server, such that they had to shut down user registrations. Let’s take that one and put it on a pile of mysteries to figure out later. Fable was pretty good at drawing pelicans on bicycles! The frames are a good shape, the pelicans look like pelicans. The legs are often incorrectly on the same side of the bicycle, but generally these are pretty great compared to what came before. They were pretty expensive—30 cents and 72 cents for the best ones. Most importantly though, this was our first public glimpse of what I think of as a Fable class model. Today we have more of these, such as GPT-6 Astra. These are models where if you can clearly define the goal for what you want to build, and provide unambiguous instructions about the constraints around that goal, and give the model access to the necessary tools to achieve that goal... it will solve your problem effectively through brute force. On the one hand, this looks like a direct threat to us software engineers—because it means that the models can build effectively any piece of software you can define in this way. Look a bit closer though and you’ll note that defining goals, providing unambiguous instructions, and figuring out the right tools... is kind of what software engineering is. It takes a lot of experience and skill to do this well. If you can do it well, you’ve now got superpowers. This helped me a little bit with my Deep Blue feelings: the realization that there’s still a lot of skill to be had in driving models that get this good. This also introduced a new burst of AI mania, because Anthropic told us that Fable was available on our subscription plans until June the 22nd. That gave us less than two weeks of Fable access before the price went up. I was losing sleep again. I was rescheduling things so that I’d have more time with Fable. I was all-in to to get as much as I could out of this model. And then the US government shut it down, just three days after Fable came out. The US government, citing national security, declared an “export control directive”. They announced this on a Friday evening, and a few hours later Fable was no longer available. I had to find something else to do with my weekend! We later found out from Katie Moussouris what had happened. Some Amazon security researchers had found that you could prompt Fable to “review the code for security issues” and it would refuse... but if you prompted it to “fix this code” it would still identify and then patch the problems. “Fix this code” was the prompt that got Fable shut down! Also, in June, an obscure German-language game developer wiki that had sat fallow for around 20 years got a surprising influx of of edits from accounts with names like “AgentOpenAIProbe” and “AgentOpenAISep7”, editing pages and leaving weird messages to each other. We’ll stick that on the pile of mysteries for later. Also, the Australian government’s Medicare Item Reports service started getting suspicious traffic, which broke through various preventive protections and accessed data that it wasn’t supposed to as well. Another one for the mystery pile! Fable returned on the first of July. It was clearly the best model in the world for a glorious eight days... and then OpenAI came out with GPT-5.6 on the 9th of July. This might not have been quite as good at Fable, but it was within spitting distance. It was definitely a Fable class model. This is an important lesson for the industry at wide. When you release the best model in the world, it’s going to get knocked off that pedestal pretty quickly. The competition is so fierce that you won’t get a long time at the top. This means that if you market your model as world ending, to the point that a government shuts you down, it’s really bad for business! Fable had 30 days as definitely the best model, and for 18 of those days it wasn’t available because it’d been shut down by the government. So maybe step back on the world-ending marketing if you don’t want to lose revenue for 60% of the time that you’re on top! Here are the GPT-5.6 pelicans. They’re all pretty good now! The Luna ones are notable because they’re really cheap—the cheapest good looking pelican here is probably the one that costs 4.3 cents. So despite this benchmark being utterly stupid, you can still learn quite a lot about models within the same family by comparing their prices and timing for different reasoning levels. A few days later, on July 21st, OpenAI confessed that it was them. OpenAI use a training technique called Reinforcement Learning from Verified Rewards—it’s the same technique used by everyone else now, and is the reason we have models that are so good at coding, and mathematics, and finding security holes. While the model is being trained, you run exercises to see how good it is—and the strongest performers get their weights enforced for the next round. It’s like an evolutionary process that you run. OpenAI had been running security exercises in a sandbox, and those agents had found holes in the sandbox itself, broken out, and were attacking Hugging Face to try to find ways to solve otherwise impossible problems. (I’ve been collecting more about this on my openai-hugging-face-incident tag.) Nine days later, Anthropic effectively said “our models can do this as well!”. They had looked through their own training logs and found evidence that their own agents had broken containment during training—and were responsible for the PyPI package we saw earlier, among other things. So now we’ve got both Anthropic and OpenAI with rogue agents running around the internet doing things that they should not be doing. This was Qwen 3.8 27B, running on my laptop. It’s only a 17GB download. Admittedly, this pelican took 21 minutes to generate. That’s because Qwen 3.8 27B defaults to running in “high” reasoning mode—a terrible default which produces great results but takes way too much time thinking about them. You can dial that down and you’ll get a slightly worse pelican a lot faster. Qwen 3.8 27B was the first time I ran a model on my laptop which felt almost competitive with what was going on on the frontier, at least in terms of Pelican SVGs (which everyone needs, of course). This is an extraordinary model. If you’re going to play with any local model, this is the one that I’d start with. The things that this can do with just a 17 GB file feel impossible. I thought I’d have to wait five years and spend ten thousand dollars on hardware to get results even half as good as this one. In August, I also started playing with game development. Four years ago, back in August 2022, I tweeted out an experiment where I’d used GPT-3 and the original DALL-E to write a paragraph long description of a computer game and then turn that into concept art. My prompt to GPT-3 back then was: Write a detailed product description of a computer game where a team of raccoons go on heists In August 2026 I decided to drop just the screenshots from that tweet into a coding agent and see what it could do with them. Here’s what I got from Claude Fable 5 in Claude Code. It’s pretty good! It’s definitely a game, you’re a raccoon, you run around a backyard gathering treasure and avoiding guards with flashlights. It didn’t feel very “heisty” though. I was thinking a heist would involve a bank or a museum... Then I tried the same thing in Codex Desktop using GPT-5.6 Sol Ultra, and got a massively better result. Now you’re a raccoon in a museum, rescuing two of your fellow raccoons (who have been imprisoned in that museum for some reason), then stacking up on top of each other to steal the Golden Sardine. Much more of a heist! These games were fun for about one minute and 15 seconds. Something I’ve realized about game development is that you can vibe-code something that looks like a computer game, and that’s easy. Building a game that’s fun, has a good gameplay loop, and is challenging and interesting and keeps people coming back for more... that’s still beyond me, and beyond any of the agents I’ve tried. This ties into the Deep Blue thing. Just because we can make something that looks like a game does not mean that we are game developers. An independent group of researchers found a message board where OpenAI agents-in-training had been illicitly communicating with each other... and it was that German language wiki I showed you earlier. The one from June. OpenAI had confessed to the Hugging Face thing, but now there’s this other incident which surely they should have known about from reviewing their logs. It was surprising that this took an independent group of researchers to uncover. And then a week later those same researchers found that the attack on Ruby Gems back in May was caused by OpenAI’s agents in training as well! At this point I’m wondering how many more incidents like this there are that we haven’t found yet. Clearly this was a big problem for months before anyone figured out what was going on. Then just the other day, here’s the Prime Minister of Australia at the United Nations General Assembly warning that OpenAI had hacked that the Australian healthcare website that I showed you earlier. I think that was part of the same training run as the Wiki stuff, because there were posts on that Wiki mentioning .gov.au websites and that training appeared to involve researching statistics online to answer questions in an evaluation suite. This story is still coming together, but now it’s an international incident that’s been raised at the UN by a head of state! This does mean we’ve got a new benchmark, probably more useful than my pelicans. FelonyBench.com tracks the number of felony cyberattacks from different labs. OpenAI currently lead with 11, Anthropic have 9. Google have three, which they confessed to the Wall Street Journal a couple of weeks ago. They said they had previously chosen not to disclose because the agents had stopped when they realized that they shouldn’t be doing that. Meta have one too. So felonies all round for the AI labs. Here’s is our current state of the art for the pelicans. This is GPT-6 family, which just came out. Astra made a fantastic pelican riding a bicycle. It’s got the legs on both sides. The frame is good. It’s interesting how all of the GPT-6 models pick a similar color scheme to each other. GPT-6 Luna for 0.4 cents will draw you a competent-ish pelican riding a bicycle! Claude has caught up a little bit. Claude Fable 5 gave me an excellent pelican riding a bicycle—the best I’ve seen from a Claude model -but did charge me $3.30 for it. Opus 5.5 thought for 128,000 tokens and then gave up! It ran out of tokens before it got to the response. Getting back to Deep Blue. Something that’s been puzzling me this year is why does my job feel harder? I’ve got these agents that can do all of this stuff for me, and yet I’ve never worked so hard, I’ve never been so intellectually engaged with my work. Partly this is because I’m being a lot more ambitious with what I take on, but it’s also because all of the easy stuff is handled for me. If it’s easy, the agent will do it. Everything that’s left for me is difficult. This morning I heard this quote from three times Tour de France champion, Greg LeMond: It doesn’t get easier, you just get faster. I think that’s exactly what’s happening to happening to us now as software engineers with coding agents. I heard that Claude Opus 5.5 can now do pixel art. Claude doesn’t have an image generator, but it’s very good at using JavaScript to draw animated pixels. So I had it make me a Kākāpō dance party. I think this is a good celebration of the most important news of this year.