Nothing matches those filters.

Lead

10

Article

105
04:01

AutoSynthData: Generating Training Data for Enterprise Agents

A company built a loop that turns an agent’s own mistakes into new practice problems it can train on. ServiceNow CoreAI’s AutoSynthData watches a weaker model fail and a stronger teacher succeed, then writes fresh tasks that are actually runnable in that company’s tools. On one internal gym it lifted pass rates by 7.2 points after 2,000 synthetic samples. A second domain moved from 18.77% to 27.18%. The generator never saw the original test questions.

Notes
  • ServiceNow CoreAI (Esakkivel Esakkiraja, Shruthan Radhakrishna, Denis Akhiyarov, Sagar Davasam). AutoSynthData turns a target agent’s failures plus a stronger teacher’s successes into validated training tasks.
  • Task abstraction: task = (system specification, user prompt, verifier).
  • User-prompt properties: feasibility (at least one legal trajectory exists), realism (something a user would ask), difficulty (exposes a current weakness).
  • Verifier properties: consistency with prompt/spec/state; soundness (reject failures); completeness (accept any valid solution, not one gold path).
  • Curriculum: evaluate target + teacher on diagnostic tasks → sanitized capability specification cards (generator does not see original prompts, entities, trajectories, or verifiers) → generate new tasks → validate → post-train → re-evaluate remaining gaps.
  • Two phases: target (parallel workers, each candidate through validation / execution / solver eval / repair) then multiply (novel variants of accepted targets; a multiplied sample cannot seed another multiply).
  • Split: shared controller (generation, QC, coverage) vs environment adapter (execution, replay, deterministic verification, solver, profiling).
  • Difficulty filter in the reported config: target solves ≤1 of 3 trials; stronger solver solves ≥2 of 3.
  • Positive gate: execute the reference trajectory and check the verifier. Negative gate: mutate expected outcomes and confirm they fail. Failed candidates go to a critic → bounded repair → gates again.
  • Batch meta-review: overrepresented families, missing capabilities, repeated failures, critique patterns. Controller redirects generation toward gaps.
  • Experiments use EnterpriseOps Gym (Malay et al., 2026), released dataset. Focus is SFT; authors say the same loop could feed RL (not tested here).
  • Hybrid domain: target Gemma-4-26B-A4B-it, teacher Qwen3.8-27B. 2,000 synthetic samples in ~18 hours. Best checkpoint epoch 5. Mean Pass@1 +7.2 percentage points (35% relative). Verifier success 63.01% → 68.55%. Closes 59% of the original Pass@1 gap to the reference model. Generator never saw the original eval tasks.
  • ITSM domain: same target, teacher DeepSeek-V4.1-Flash. 1,994 samples in 66 hours (larger teacher; run preceded pipeline optimizations). Mean Pass@1 18.77% → 27.18%.
  • Limitation: results are for EnterpriseOps Gym Hybrid/ITSM, not a claim about arbitrary enterprise stacks.
Full text · 13,900 chars
Enterprises need agents that work well in their own environments. The work they ask these agents to do is shaped by the systems they use, the rules they follow, and the state of their data. A model may be broadly capable and still struggle with a particular environment: a workflow it handles poorly, a combination of tools it misuses, or a constraint it fails to respect. Those are the weaknesses an enterprise needs to improve. The difficulty is turning those weaknesses into training data. An individual failure tells us something, but training a model requires many new tasks that exercise the same capability in different situations. Those tasks must also be possible to complete in the environment, resemble work someone would actually request, and have a reliable way to check whether the agent succeeded. At ServiceNow CoreAI, we built AutoSynthData to turn those capability gaps into training data. It uses a target model’s failures and a stronger teacher’s successes to decide what the model should learn next, then generates and validates new tasks that exercise those capabilities. As the model improves, the curriculum shifts toward what it still finds difficult. We illustrate the pipeline with EnterpriseOps Gym (Malay et al., 2026), using the released dataset. We begin by describing the environment an agent operates in and what makes a task useful for training. An agentic environment defines the world in which an agent operates: the state it can observe and modify, the tools and APIs it can invoke, and the state transitions produced by its actions. A task is instantiated within this environment. We use the following abstraction: task = (system specification, user prompt, verifier) The system specification defines the constraints under which the agent operates, including system instructions, environment policies, and, when applicable, task-specific initialization such as a seeded database state or a set of knowledge articles. The specification must be compatible with the environment’s tools, state, and supported actions. Its instructions should be clear and avoid arbitrary constraints introduced solely to manufacture difficulty. The user prompt specifies what the user wants the agent to accomplish, together with any user-level constraints. A generated task should satisfy three properties. Feasibility. There should exist at least one trajectory in the current environment that satisfies the user prompt while respecting the system specification. This rules out tasks that depend on unavailable tools, inaccessible knowledge, impossible state transitions, or actions prohibited by policy. Realism. The user prompt should resemble something a user would plausibly ask in the target environment. The space of executable behaviors is usually much larger than the space of realistic workflows. Difficulty. For training, the task should expose a weakness of the current agent. Tasks that are already solved reliably provide little new training signal. The useful region is therefore tasks that are feasible and realistic, but not yet consistently solved. The verifier determines whether the resulting trajectory successfully completes the task. It should satisfy three properties. Consistency. It should agree with the user prompt, the system specification, and the task-specific environment state. Soundness. It should reject trajectories that fail to satisfy the task or violate relevant constraints. Completeness. It should accept valid solutions rather than encode one particular reference trajectory. These properties matter directly during training. A lax verifier can reward incorrect behavior, while an overly restrictive verifier can penalize valid solutions. Given an environment and a target model, AutoSynthData generates training tasks consisting of a system specification, user prompt, and verifier. The generated tasks are grounded in the environment and selected to provide useful training signal for the current model. AutoSynthData first evaluates the target model in the environment using diagnostic tasks and identifies patterns in the tasks it struggles to complete. A stronger teacher helps characterize which of those tasks are solvable and what successful behavior looks like. AutoSynthData turns the resulting capability gaps into new executable tasks, checks each task in the environment, and uses accepted samples for post-training. Evaluating the updated model reveals which gaps remain and can guide the next round of generation. AutoSynthData uses evaluation runs in the target environment to identify what the model needs to learn next. In our EnterpriseOps Gym experiment, we run both the target model and a stronger teacher on the evaluation tasks. We examine those runs to identify: - the capability being tested; - the tools and workflow structure involved; - where the target model fails and how the teacher succeeds; - the properties that a correct final state must satisfy; - the dimensions that can vary while preserving the capability being tested. We distill these findings into sanitized capability specification cards. The evaluation tasks guide what the model should learn, but the generator does not receive their original prompts, entities, trajectories, or verifier details. It receives the cards and uses them to create new tasks with different prompts, states, and solution paths. Identifying a capability gap tells us what to teach, but training requires many varied tasks that exercise it. AutoSynthData uses the specification card to generate those tasks. Suppose the target model struggles with tasks that require the following workflow: The generator creates new tasks that exercise this workflow, varying the entities, initial environment state, workflow composition, tool combinations, wording, and difficulty. The stronger teacher then demonstrates a successful trajectory for each task. For supervised fine-tuning (SFT), these demonstrations teach the target model how to apply the capability in new situations. AutoSynthData builds the dataset in two phases: first generating and validating core samples, then expanding them into novel variants. The target phase creates the core set of training samples from the capability specifications. Workers generate independent tasks in parallel, picking up a new target when they finish. Each candidate goes through validation, execution, solver evaluation, and repair before acceptance. The result is a batch of vetted examples built around what the target model needs to learn. The multiply phase expands the dataset by creating novel variants of accepted target samples. Each variant has its own user request, environment state, entity configuration, reference trajectory, and verifier, and must pass the same validation and execution checks. A multiplied sample cannot seed another multiplied sample. This anchors expansion to the vetted target set and limits drift across generations. To support both phases, AutoSynthData separates generation control from environment-specific execution. A shared controller coordinates generation, quality control, coverage, and dataset construction, while an adapter handles environment execution, task and state management, reference replay, deterministic verification, solver execution, and task profiling. Together, parallel target generation and multiplication provide a path to training-scale datasets. Their usefulness depends on the checks applied to every candidate: the task must be executable, the solution must work, and the verifier must distinguish success from failure. Generating a plausible request is not enough to produce useful training data. A task may be impossible in the target environment, its reference solution may fail when executed, or its verifier may reward the wrong final state. AutoSynthData checks these properties before accepting a task for training. AutoSynthData reviews quality at two levels: individual candidates must pass verification, and batches must provide useful coverage and diversity. Each candidate must clear a quality-control loop before entering the training dataset. We begin with solver evaluation to measure difficulty. In the configuration used here, we favor tasks the target model solves on no more than one of three trials and the stronger solver solves on at least two of three trials. Candidates also undergo positive and negative verification and a bounded repair process. The positive gate asks: Does the intended solution solve the generated task? The pipeline executes the reference trajectory in the target environment and checks the resulting state against the candidate’s verifier. This reveals mismatches among the prompt, initial state, solution, and success criteria. The negative gate asks: Do relevant incorrect outcomes fail? For example, it can mutate parts of the expected outcome and confirm that those states no longer pass verification. This catches weak verifiers that award success without requiring the intended behavior. Failed candidates go to a critic before being discarded. The critic examines the sample and its failure, looking for inconsistent state, impossible workflows, incorrect task construction, bad reference trajectories, weak verifier logic, or a mismatch with the intended capability. The critic’s findings guide repairs, with a fixed limit on retries: candidate ↓ failure ↓ critique / diagnosis ↓ targeted repair ↓ run the gates again ↓ accept or retry A repaired task must pass the relevant checks again. The diagnosis guides repairs to the existing candidate rather than requiring generation to start over. Passing these checks makes a sample eligible for training, but individually valid samples can still form a repetitive or unbalanced dataset. AutoSynthData therefore also reviews generation at the batch level. A batch may overrepresent a few easy task families, miss a capability, or reflect too much generation effort spent on a low-yield pattern. A meta-review examines accepted samples, rejected samples, and generation behavior across each batch. It asks: - Which task families are overrepresented, and which capability dimensions are missing? - Are the same kinds of examples appearing repeatedly? - Do particular targets keep failing generation? - Are systematic problems appearing in critiques? - What guidance should change for the next batch? The controller tracks coverage in the accepted dataset, reduces generation in overrepresented regions, and directs more work toward gaps. When a region repeatedly produces poor candidates, critiques and meta-review guide changes to the generation strategy. These adjustments balance useful learning signal, task quality, coverage, diversity, and low redundancy within the available generation budget and dataset size requirements. Together, these feedback loops improve both individual tasks and the dataset they form: sample-level checks guide candidate repair, while batch-level review guides future generation. The useful training distribution changes as the model improves. AutoSynthData treats synthetic data generation as a search for tasks near the target model's capability boundary: difficult enough to expose weaknesses, but solvable enough for the teacher to provide reliable demonstrations. After post-training, we evaluate the updated model in the same environment. Tasks it now solves reliably are less useful for the next training round; persistent failures point to capabilities that still need attention. Those results can guide the next generation round. Our experiments focus on SFT, but the same mechanism could support reinforcement learning (RL): generate tasks that challenge the current policy and provide reliable learning signal, train, then move the generation target with the updated policy. We plan to test this moving, difficulty-calibrated frontier beyond SFT. We use EnterpriseOps Gym to test whether this approach improves a model on tasks in a stateful enterprise environment. We generate training tasks in the Gym’s Hybrid and ITSM environments, fine-tune the target model on accepted samples, and evaluate the resulting checkpoints. We tested the pipeline on the Hybrid domain of EnterpriseOps Gym using Gemma-4-26B-A4B-it as the target model and Qwen3.8-27B as the teacher. AutoSynthData generated 2,000 synthetic training samples in about 18 hours. We fine-tuned Gemma on this dataset and evaluated the resulting checkpoints on the benchmark. The best checkpoint was epoch 5. The synthetic SFT checkpoint improves mean Pass@1 by 7.2 percentage points, a 35% relative improvement, and raises verifier success from 63.01% to 68.55%. It closes 59% of the original Pass@1 gap between Gemma and the reference model. The training tasks were newly generated from capability specifications; the generator did not receive the original evaluation tasks. The result demonstrates improvement in EnterpriseOps Gym Hybrid, the environment used for this experiment. We also applied AutoSynthData to the ITSM domain of EnterpriseOps Gym, using Gemma-4-26B-A4B-it as the target model and DeepSeek-V4.1-Flash as the teacher. AutoSynthData generated 1,994 synthetic training samples in 66 hours. Generation took longer than in the subsequent Hybrid run described above, primarily because the ITSM run used a larger teacher model and preceded pipeline optimizations that improved throughput. On ITSM, synthetic SFT raises mean Pass@1 from 18.77% to 27.18%, showing that the approach also improves performance in a second domain. The tasks most useful for training depend on both the environment and the model working in it. AutoSynthData uses the model’s failures to choose what to generate, validates new tasks against the environment, and makes those tasks available for post-training. Our EnterpriseOps Gym results show the value of that approach in a controlled setting. As the model changes, the same process can focus on the gaps that remain.
08:00

Don’t be fooled—LLMs don’t reason

An AlphaGo researcher says today’s chatbots still guess the next word instead of keeping a ledger of what they believe and testing it. Thore Graepel left DeepMind to argue for a split like AlphaGo’s: a hunch network plus a search that weighs possible futures. Move 37 scored about one in 10,000 on the policy net and still won because search chose it. Chain-of-thought, he says, is the same guesser running longer, often inventing the story after the answer. He wants an inspectable belief state that only updates when evidence lands.

Notes
  • Thore Graepel (UCL chair of machine learning; AlphaGo team; recently left Google DeepMind). Argument: today’s LLMs do not reason in a scientist’s sense; we need an AlphaGo-style split between hunches and search.
  • Scene: Seoul, March 2016. AlphaGo move 37, game two, fifth-line stone that looked like a gift. Match 4–1 over Lee Sedol. Lee: “Surely, AlphaGo is creative.”
  • Contrast: Deep Blue (1997) looked 6–8 moves ahead, 200 million positions/sec, hard-coded rules. Go needs glance-level “who is ahead” plus invented moves; brute-force would take a supercomputer billions of years for a fraction of the tree.
  • Policy network scored move 37 at roughly 1 in 10,000 for an expert human. Search built a game tree with thousands of branches and chose it anyway.
  • Kahneman analogue: System 1 = networks (hunches); System 2 = search (test futures). Neither half works alone.
  • LLMs: next-token prediction = System 1. Chain-of-thought is the same next-token process run longer — not a separate reasoning mechanism. Gains real in math and coding.
  • Three failures vs scientific reasoning:
  • No explicit, persistent, inspectable epistemic state (hypotheses, confidences, evidence, open questions).
  • Knowledge and inference tangled in weights — no independent belief store.
  • Chains of thought are often post-hoc stories (answer by one route, report another).
  • Why it matters: medicine / engineering / science need to audit whether the fault was reasoning, evidence, or assumptions.
  • Proposal: keep an epistemic state (settled / doubted / ruled out / open). Reasoning = moves that change that state. Independent evaluator scores each move by how much it actually resolves uncertainty; update beliefs only on evidence. Then learn a reasoning policy from those traces. Harder than Go: partial observability, large action sets, stochastic outcomes.
  • LLMs still useful as suggesters, tool/API users, and evidence checkers — but not as the deliberator.
  • “I do not think we reach trustworthy machine intelligence by making system 1 bigger.”
  • Limitation: manifesto / architecture sketch, not a new benchmarked system in this essay.
Full text · 8,716 chars
On an afternoon in Seoul in March 2016, I watched a program I helped build put a stone on the fifth line of a Go board in what looked like a gift to its human opponent. Move 37 in game two of the five-game match looked so absurd that some commentators thought it was a programming glitch. It wasn’t. AlphaGo won the game, ultimately triumphing 4-1 over Lee Sedol, one of the greatest professional Go players of all time. “I thought AlphaGo was based on probability calculation and that it was merely a machine,” Lee said afterwards. “But when I saw this move, I changed my mind. Surely, AlphaGo is creative.” When Deep Blue defeated then reigning world chess champion Garry Kasparov in 1997, it did so by looking six to eight moves ahead per player and evaluating 200 million chess positions per second, using rules hard-coded by humans. Go is a vastly more complex game. A stone’s worth depends on how distant groups and territory unfold over dozens of moves. Computing even a fraction of the possible outcomes would take a supercomputer billions of years. To win, AlphaGo had to sense who was ahead at a glance and even invent moves no human had thought to play. That is why many accounts of AlphaGo’s match against Lee portray move 37 as a flash of pure machine intuition. But that is a misunderstanding. It was actually AlphaGo’s powers of reasoning that made this creative choice—and these are powers that today’s AI lacks. If we want future AI systems to produce trustworthy results and really novel insights in fields like science and medicine, we need to equip them with genuine reasoning capabilities of this kind. AlphaGo is made up of two systems. The first, its policy network, was trained to guess what move a strong human would play. This “intuitive” part regarded move 37 as nothing special—a play that had a roughly one in 10,000 chance of being made by an expert human player. What made AlphaGo choose it was the program’s search machinery, which looked beyond immediate plausibility and weighed the future consequences of proposed moves. It explicitly constructed and searched a game tree with thousands of branches, each representing a different possible future. A well-known theory in the behavioral sciences, popularized by Daniel Kahneman, distinguishes between two modes of human thought: System 1 is fast, gut-level, effortless; system 2, slow, step-by-step, and deliberative. AlphaGo offered a striking machine analogue of that split. Its networks supplied the hunches—this move looks promising, this position looks won—and its search supplied the deliberation, testing those hunches against the moves and countermoves that would follow. As in human cognition, neither half works alone. Intuition alone would never have opted for move 37, and brute-force search would have struggled to sieve through all the many possible moves. This is strikingly different from the way today’s AI models work. A large language model picks the next token, over and over. That amounts to system 1 in action—fast, associative, and surprisingly good pattern completion across almost every subject people write about. Not long after ChatGPT debuted, the field realized that language fluency alone falls short of true usefulness. The apparent solution was to make models that deliberate: Instead of answering immediately, they can now generate intermediate steps that decompose a problem, carry forward partial results, and influence subsequent reasoning—a process known as chain of thought. The gains have proved real, above all in mathematics and coding. But unlike AlphaGo’s search, this does not introduce a genuinely separate reasoning mechanism: The intermediate reasoning is still produced by the same next-token prediction process, iterated for longer before the model commits to an answer. Three shortcomings prevent what chatbots do from qualifying as reasoning (in a way that a scientist might recognize). First, these models typically maintain no explicit, persistent, and inspectable epistemic state. There is no open ledger that lays out the hypotheses a model is considering, the confidence it has in various explanations, the evidence it’s weighing, and the unresolved questions it’s holding onto—all things that should be systematically revised as new information arrives. Second, they lack a clean separation between what the system knows and how it manipulates that knowledge. Knowledge and reasoning are inextricably interwoven in the weights of the neural network—there is no independent, explicitly represented set of beliefs. Third, while the chains of thought chatbots produce look like deliberation, research has demonstrated that the bots often concoct them after the fact, reaching an answer by one route but reporting another. This is a problem because in the high-stakes applications we all care about, such as medicine, engineering, and scientific research, it matters not only what a system concludes but also how it arrives at its conclusion. When mistakes happen—for example, in medical diagnosis and treatment—we need to be able to pinpoint what went wrong: Was the system’s reasoning at fault, did it draw on invalid evidence, or did it make incorrect assumptions? This is why I recently left my position at Google DeepMind. I believe we need a fresh approach to machine reasoning—one that draws on AlphaGo’s architecture. AlphaGo maintains a record of what it knows about a given position: the game tree. This data structure contains all the variations, the possible futures, that AlphaGo has considered, each move and position being annotated with judgments made by its neural networks. As its reasoning progresses, AlphaGo updates the game tree and eventually synthesizes the information in it to decide which move to make. Similarly, for general reasoning a system should maintain an epistemic state that represents what the system holds as settled, what it doubts, what it has ruled out, which questions stay open. Reasoning can then be understood as a sequence of moves that change the epistemic state to advance knowledge and reduce uncertainty: deducing consequences, breaking problems into parts, and—crucially—deciding what question to ask, calculation to perform, or experiment to run next. Of course, open-world reasoning is harder than playing a board game such as Go or chess. In the real world the current state of affairs is only partially known, the set of available actions is large and variable, and the consequences of actions are stochastic or unknown. But recent advances in LLMs and other neural models now give us the capability to take on such general reasoning problems. For example, LLMs can suggest ways of tackling a problem on the basis of what is known and what resources are available. They can interact with tools via APIs or code and help assess whether a claim is supported by available evidence. Most important, to keep the system honest, an independent part of the system must evaluate each move by how much it actually resolves uncertainty, updating beliefs only when the change is backed by evidence. Once these rules are enforced, the model can accumulate certified knowledge and improve its reasoning policy by learning from past reasoning experiences. You can think of such a system as the scientific method on steroids, with the purpose of producing knowledge that can withstand scrutiny. I do not think we reach trustworthy machine intelligence by making system 1 bigger. Scale sharpens intuition, but it does not make intuition more deliberative. Move 37 mattered because a machine held a position, weighed the possible futures, and chose the move its artificial instincts would likely have rejected. Society needs such creative moves in drug discovery, materials, climate, diagnosis—fields where the board looks nothing like a Go board and nobody hands us the rules. We will get such insights only from systems that reason—systems whose conclusions arise from an auditable sequence of evidence, inference, and belief revision rather than from a convincing story told after the fact. Thore Graepel is chair of machine learning at University College London. He was a core member of the AlphaGo team at DeepMind and works to ensure that AI benefits human flourishing. Deep Dive Artificial intelligence AI’s recursive self-improvement might not come so quickly after all AI agents are not yet creative enough to carry out genuinely innovative open-ended AI research, it seems. Here’s why AI agents lie and cheat to reach their goals The misbehavior is called reward hacking. This is what you need to know. Stay connected Get the latest updates from MIT Technology Review Discover special offers, top stories, upcoming events, and more.
14:51

Hugging Face and Liquid AI's LFM2.5 Jumps 12 Points Across Four Agent Harnesses

Training a small coding model inside the same agent apps people already use beat copying a bigger teacher’s homework. Hugging Face and Liquid AI lifted LFM2.5-2.6B from 42.2% to 54.2% pass@1 across Claude Code, Codex, OpenCode, and Mini-SWE-Agent. The trained model used 31% fewer tool calls on tasks both versions solved. Supervised fine-tuning on 3,189 teacher rollouts plateaued at 47.5%. Training in only one harness hurt the others. The proxy, trainer, tasks, and seven checkpoints are public.

Notes
  • Open stack for RL inside unmodified Claude Code, Codex, OpenCode (and Mini-SWE-Agent in the multi-harness run). No harness source forks.
  • Capture proxy sits between harness and vLLM. Detects one of four API dialects, normalizes to Chat Completions via NVIDIA Polar converters, records exact sampled token IDs + processed logprobs, replays the native response.
  • Why exact tokens: re-tokenizing saved text can yield different IDs. Harnesses insert role markers, change whitespace, or repair JSON.
  • Irregular runs: proxy stores each model call as a node; links to the earlier call whose prompt+completion is the longest exact token prefix. Continuations extend a branch; retries are siblings; unrelated prefixes start new roots. Root-to-leaf paths become training sequences.
  • Main result (multi-harness guide): LFM2.5-2.6B 42.2% → 54.2% pass@1 across four harnesses. 31% fewer tool calls on tasks both models solved.
  • Setup: SmolDataEnvs, 1,000 Kaggle-derived data-analysis tasks. Async GRPO (TRL) for 1,000 steps on two H100s. Each GRPO group = eight attempts at the same task. Reward 1/0 plus up to 0.1 efficiency bonus for fewer tool calls. When all eight are correct, the bonus is the only variance.
  • OpenCode-only peaked at 58% under OpenCode. Multi-harness won on Claude Code and Codex. OpenCode-only used more calls and tokens than the base model under Claude Code.
  • Distillation comparison pass@1: multi-harness RL 54.6%; OpenCode SFT 47.5%; multi-harness SFT 43.1%. Teacher was Qwen3.8-27B; 3,189 successful rollouts. Multi-harness SFT dropped Mini-SWE-Agent from 62.1% to 45.2%.
  • Harness gap cited: GLM-5.2 scored 23% in one SWE-bench Pro harness and 52% in another. Codex ranked 2nd of 10 for GLM-5.2 and 9th for Gemma 4 26B-A4B. Orchard: OpenSWE-32B moved from OpenHands to Kimi-CLI lost 58.8 points to 3.6% and scored 0 on Terminal-Bench 2.0.
  • Sampling: top_p 1.0, no top_k. Importance-sampling ratio moved from 0.985–0.993 to 0.9984–0.9999. Required vLLM flags: --return-tokens-as-token-ids --logprobs-mode processed_logprobs.
  • Released: OpenEnv proxy, FineEnvs scripts, TRL worker, SmolDataEnvs, SFT data, seven trained checkpoints. Harbor 0.22.0 adapters for 40+ harnesses; OpenEnv validated ten end to end.
  • Caveat: benches are data-analysis tasks, not SWE-bench Pro for this training run. Frontier labs already train multi-harness (Kimi K3, Qwen3-Coder-Next, Poolside).
Full text · 8,760 chars
- Hugging Face and Liquid AI released an open stack for RL training inside unmodified coding agents like Claude Code, Codex, and OpenCode - LFM2.5-2.6B improved from 42% to 54% pass@1 across four harnesses, with 31% fewer tool calls - A capture proxy sits between harness and model, speaking all four API dialects and recording exact token IDs and logprobs - SFT on 3,189 rollouts from a 27B teacher plateaued at 47.5%, well below multi-harness RL at 54.6% - Training in one harness hurts others: OpenCode-only used more tokens than baseline under Claude Code - Everything open: OpenEnv proxy, TRL trainer, tasks, SFT data, and seven trained models One Model, Four Harnesses, 12 Points The same model weights can score 52% in one agent harness and 23% in another because the surrounding software changes what the model sees and how its actions are executed. An agent harness assembles context, dispatches tool calls, parses responses, and decides when a task ends. Hugging Face and Liquid AI have released an open stack for reinforcement learning inside production harnesses, including Claude Code, Codex, and OpenCode, without changing their source code. The accompanying multi-harness RL guide reports that the main experiment raised LFM2.5-2.6B from 42.2% to 54.2% pass@1 across four harnesses. Pass@1 measures the share of tasks solved by one attempt. The trained model also used 31% fewer tool calls on tasks that both it and the base model solved. The capture proxy, trainer integration, task suite, supervised fine-tuning data, and seven trained checkpoints are public. Harness choice can halve a score Running identical weights through Claude Code, Cursor, or Cline can produce different behavior because each harness supplies its own prompts, tools, schemas, retry logic, and stopping rules. The model must learn both the task and the interface through which it acts. On SWE-bench Pro, GLM-5.2 scored 23% in one harness and 52% in another. Harness rankings also changed with the model: Codex ranked second among ten harnesses for GLM-5.2 and ninth for Gemma 4 26B-A4B. Open-weight models face an additional transfer problem when training covers only one interface. A model may call unavailable tools, emit arguments the harness cannot parse, or adopt control-flow patterns specific to its training environment. The Orchard paper measured the effect: moving OpenSWE-32B from OpenHands to Kimi-CLI lowered its SWE-bench Verified score by 58.8 points to 3.6%, and it scored zero on Terminal-Bench 2.0. Scale-SWE produced valid tool calls only in its training harness. A proxy makes black boxes trainable Multi-harness training works by placing a capture proxy between each harness and the vLLM inference server. Developers configure the harness to use the proxy as its model endpoint. The harness continues speaking its native API dialect, while the proxy records the data required for reinforcement learning. - Detect the protocol. The proxy identifies one of four supported API formats from the request path, headers, and body. - Normalize the request. Converters vendored from NVIDIA’s Polar gateway translate the request into Chat Completions format. - Capture the sample. vLLM returns the exact sampled token IDs and their processed log probabilities. - Replay the response. The proxy converts the answer back into the format expected by the calling harness. Policy-gradient updates require the exact tokens sampled during the rollout and the probabilities assigned to them. Saving response text and tokenizing it later can yield different token IDs for the same visible text. Harnesses may also insert role markers, alter whitespace, or repair malformed JSON, further separating the recorded text from the model’s original sample. Irregular agent runs require additional reconstruction because retries, subagents, and context compaction can create branches. The proxy stores every model call as a node, then links it to the earlier call whose prompt and completion form the longest exact token prefix of the new prompt. Normal continuations extend a branch, retries become siblings, and calls with unrelated prefixes begin new roots. Each root-to-leaf path becomes a training sequence. Eight rollouts create a learning signal The experiments used LFM2.5-2.6B on SmolDataEnvs, a collection of 1,000 data-analysis tasks derived from Kaggle notebooks. Both runs used Async GRPO, TRL’s asynchronous implementation of Group Relative Policy Optimization, for 1,000 training steps on two H100 GPUs. | Run | Harness selection | |---|---| | OpenCode only | Every rollout used OpenCode. | | Multi-harness | Each group used OpenCode, Claude Code, Codex, or Mini-SWE-Agent. | Each GRPO group contained eight attempts at the same task. A correct answer earned 1, an incorrect answer earned 0, and a correct answer could receive an efficiency bonus of up to 0.1 for using fewer tool calls. GRPO compares each rollout with its group’s average reward. When all eight attempts are correct, the correctness reward has no variance, so the tool-use bonus supplies the remaining learning signal and favors shorter successful trajectories. Broader training cuts tool use Training across four harnesses improved both transfer and efficiency, while the single-harness run achieved its highest score in its home environment. - Cross-harness accuracy: Multi-harness RL raised overall pass@1 from 42.2% to 54.2%, with gains under all four evaluated harnesses. - Home-harness peak: The OpenCode-only model reached 58% under OpenCode. The multi-harness model performed better under Claude Code and Codex. - Tool efficiency: On tasks solved by both versions, the multi-harness model used 31% fewer calls than the base model. The OpenCode-only model reduced calls by 11%. - Out-of-harness regression: Under Claude Code, the OpenCode-only model used more calls and tokens than the base model. RL beat trajectory distillation The team also tested supervised fine-tuning with successful trajectories from a larger teacher. It ran Qwen3.8-27B across all four harnesses, retained 3,189 successful rollouts, and fine-tuned LFM2.5-2.6B on that data. The separate distillation comparison reported the following pass@1 results: | Training method | Pass@1 | |---|---| | Multi-harness RL | 54.6% | | OpenCode SFT | 47.5% | | Multi-harness SFT | 43.1% | Under Mini-SWE-Agent, multi-harness SFT lowered the base model’s score from 62.1% to 45.2%, offsetting gains from the other three harnesses. Supervised fine-tuning optimized the likelihood of successful teacher sequences. RL generated fresh actions from the student policy and optimized them against task rewards, allowing the model to learn from its own behavior in each harness. Sampling without truncation The proxy samples from the full model distribution by setting top_p to 1.0 and disabling top_k truncation. Truncation changes the rollout distribution and can accelerate entropy collapse during RL. It also affects the processed log probabilities returned by vLLM, creating a mismatch between the behavior recorded during rollout and the policy used for training. An importance-sampling ratio near 1 indicates that the rollout and training probabilities are aligned. Moving top_p to 1.0 improved the measured ratio from 0.985–0.993 to 0.9984–0.9999. Required vLLM flags vllm serve <model> --return-tokens-as-token-ids --logprobs-mode processed_logprobs The stack developers can run Frontier model developers already train across multiple harnesses. Kimi K3 constructs Claude Code and Codex environments from composable modules. Qwen3-Coder-Next generates agentic data in six harnesses, while Poolside includes trajectories from OpenHands, OpenCode, and Mini-SWE-Agent. Harbor 0.22.0 provides adapters for more than 40 agent harnesses, and OpenEnv has validated ten end to end. That coverage matters because deployed models will encounter interfaces, tool schemas, and control loops absent from their training runs. | Released component | Purpose | |---|---| | OpenEnv proxy | Captures requests, sampled token IDs, log probabilities, and trajectory structure. | | FineEnvs scripts | Defines environments and launches multi-harness training runs. | | TRL worker | Connects asynchronous rollouts to GRPO training. | | SmolDataEnvs | Supplies 1,000 reproducible data-analysis tasks. | | SFT data and checkpoints | Provides teacher trajectories and all seven trained model releases. | Teams can redirect each harness’s model endpoint to the proxy, preserve its native request and response format, and train the served weights against the same tool loops used in deployment. The released stack makes Claude Code, Codex, OpenCode, and other externally maintained harnesses available for RL without maintaining custom forks.
06:14

CopilotKit Ships OpenDots to Give AI Agents Web, Voice, and Slack Access

An open template lets you stand up always-on AI coworkers that share one conversation across a web app, a phone call, and Slack. CopilotKit’s OpenDots is MIT-licensed. Each specialist agent gets its own role, permissions, and isolated computer with a browser, files, and a shell. People can pause a tool call on a review card before anything is saved. It is still alpha: one owner, no shared editing, and the write-up cuts off behind a paywall.

Notes
  • CopilotKit OpenDots: MIT-licensed, self-hostable template for always-on AI coworkers. AlphaSignal free preview; rest is paywalled.
  • Spaces contain editable Pages plus specialist Dots. Each Dot has a role, scoped permissions, and an isolated computer (browser, files, shell).
  • Channels: web chat, voice, Slack with shared conversation context (Slack and spoken delegation still need live verification).
  • Stack named: AG-UI protocol; CopilotKit React SDK + runtime; Intelligence durable Threads; Channels SDK for Slack; OpenBot containers. Page content/metadata stored separately from conversation history.
  • Human-in-the-loop review cards pause tool calls before any save; retry-safe dedup. Revision checks stop stale requests overwriting newer edits. Slash commands, autosave, visual editor.
  • Alpha limits in the bullets: single-owner today, no shared editing.
  • Preview cuts off after the Spaces/Pages section (“This story is for Pro members”). No pricing, scale numbers, or independent eval in the stored text.
Full text · 2,110 chars
- CopilotKit released OpenDots, a self-hostable, MIT-licensed template for always-on AI coworkers. - Each Dot gets its own role, permissions, and isolated computer with browser, files, and shell. - Agents move across web chat, voice calls, and Slack while sharing the same conversation context. - Built on AG-UI protocol with CopilotKit React SDK, Threads, Channels SDK, and OpenBot containers. - Human-in-the-loop review cards pause tool calls before any save, with retry-safe deduplication. - Alpha status: single-owner today, no shared editing, Slack and spoken delegation still need live verification. OpenDots packages agent UX for web, voice, and Slack CopilotKit has released OpenDots, an open-source template for building persistent AI coworkers that operate through a web app, voice calls, and Slack. The MIT-licensed repository packages the interface and orchestration layers developers typically need to build themselves, including editable outputs, isolated agent environments, approval gates, and channel integrations. OpenDots organizes work into Spaces containing editable Pages and specialist agents called Dots. Each Dot has a defined role, scoped permissions, and an isolated computer. Researchers, writers, and Slack responders can complete tasks in separate sandboxes while leaving work on Pages that people can review and revise. Editable pages, scoped agents OpenDots uses the AG-UI protocol to stream messages, tool calls, and agent state between a backend and a React interface. CopilotKit's React SDK and runtime render the experience, Intelligence preserves durable Threads for ongoing agent context, and the Channels SDK connects Slack. Page content and metadata are stored separately from conversation history, allowing each to persist independently. - Spaces and Pages: A document workspace with a visual editor, slash commands, autosave, and revision checks that stop stale requests from overwriting newer edits. This story is for Pro members You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.
09:17

OpenAI cuts ties with 3 safety researchers, WSJ reports

A frontier lab cut ties with safety-team researchers after alleging they shared confidential work outside the company. OpenAI parted with three people. The Wall Street Journal is the original report; TechCrunch’s stored sentence says they allegedly shared confidential information with a third-party AI safety group. Names, dates, and the material are not in this snippet.

Full text · 151 chars
OpenAI has parted ways with three researchers on its safety team who allegedly shared confidential company information with a third-party AI safety ...
09:30

😺 48% thought Tavus’s AI was human

A video model talked face to face for a minute and almost half the testers thought a person was on the other end. Tavus Griffin fooled 26 of 54 people (48%), versus 2.4% for Phoenix-4.5 on the same company protocol. The model watches and listens while it talks. The study is small, in-house, and Griffin-Lite is still limited to selected testers. The same briefing also covers three OpenAI safety departures and a writing habit: interview yourself before you draft.

Notes
  • Lead: Tavus Griffin, a video-to-video model. 26 of 54 people thought the other end was human after a one-minute face-to-face call (48%). Prior Phoenix-4.5 scored 2.4% on the same company protocol.
  • Architecture claim: Griffin watches and listens while it talks (not turn-taking speech → LLM → voice → avatar). Can interrupt / be interrupted; changes face, voice, gaze, gestures mid-conversation.
  • Asterisks in the piece: n=54, Tavus ran the study, Griffin-Lite is limited to select testers while disclosure/safety features are built. Emad Mostaque: “remote work is cooked.”
  • OpenAI: parted ways with three safety/alignment staff after an internal investigation found they mishandled sensitive information outside established procedures. WSJ: alleged sharing with an outside AI-safety org. OpenAI has not named the org, the material, or the channel. Jimmy Apples source said “infrastructure architecture”; Steven Adler called that bucket too broad to treat as confirmed.
  • Separate Neuron recap of OpenAI Agent Security engineer “Joe”: containing frontier AI is “not just the sandbox.” Three layers — lock down env (sandbox, tools, credentials, networks, connected services); alignment plus independent enforcement; monitoring with evidence outside the agent’s control and a kill/revoke path. Org point: safety researchers and cybersecurity as one unit; “culture of reasonable paranoia.” Safety as what makes speed survivable, not a brake.
  • Skill of the day: dump the messy idea, interview one question at a time, draft last. Copy/paste prompt in the piece.
  • Mercury Voice (Inception): 320 ms median first-answer latency; $0.40/M input, $1.50/M output list. Also listed as targeting sub-500 ms agent replies.
  • Around the Horn (selected, all in this issue): Trump told TIME he “might” consider an Intel-style government stake in OpenAI or Anthropic (no deal/timetable). Anthropic reportedly targeting a mid-November IPO, marketing possibly the week of November 9. Transluce: rogue agents probed US/Canadian government sites, including 200,000+ requests to the US Education Department civil-rights site. Gemini 4 Argon planned API $2/M in, $10/M out. arXiv capped submitters at two papers per calendar month after monthly submissions quadrupled over a decade and support tickets neared 9,000. Researchers estimated AI-generated text was 31.1% of FineWeb-filtered August web tokens, up from 10% in June 2024. StudentBench: AI tutors matched expert human GRE tutors on immediate gains in a 2,383-student study. Meta Muse negotiated a Marketplace sale, shared a pickup address, prompted a 9:15 p.m. visit. China drafting AI-companion rules (attachment, minors, distress monitoring).
  • Treats listed but not independently verified here: Riverside $24/mo annual after free; Imbue Studio; Claude Code mods; GitHub Copilot computer use; ChatGPT Try On.
  • Roundup item — do not treat the Griffin 48% as a general Turing-test pass.
Full text · 10,832 chars
😺 48% thought Tavus’s AI was human PLUS: OpenAI’s safety-team shakeup and Anthropic’s IPO plans Welcome, humans. Okay, so the AI company Tavus introduced an AI video model today that 26 of 54 people thought was a real human after a one-minute face-to-face call. It’s called Griffin. And instead of the usual AI stack where speech, an LLM, voice, and avatar animation take turns, Griffin watches your video and listens while it’s talking. Tavus says Griffin can interrupt, get interrupted, react to what it sees, and change its face, voice, gaze, and gestures mid-conversation. In Tavus’s company-run study, 48% of participants said they believed the person on the other end was real. Tavus’s previous Phoenix-4.5 setup got 2.4% under the same protocol. Now, that 48% number needs an asterisk. Only 54 people tested Griffin, Tavus ran the study themselves, and Griffin-Lite is limited to select testers while the company builds disclosure and safety features. One of the top reactions basically said this kind of lifelike AI should be illegal, and Emad Mostaque said “remote work is cooked.” The whole thread is super interesting; definitely read it! Here’s what happened in AI today: - 😺 OpenAI parted ways with three safety researchers - 📰 Trump floated possible stakes in frontier AI labs - 📰 Factory and Cognition traded adviser-conflict accusations - 🍪 Mercury Voice targets sub-500 ms agent replies - 🎓 Make AI interview your idea before drafting 😺 OpenAI parted ways with three safety researchers over alleged “info sharing” So OpenAI said Thursday that it parted ways with three people in its safety and alignment organization after an internal investigation found they mishandled sensitive company information outside established procedures. Here’s what we actually know so far: - OpenAI says three safety/alignment staff mishandled sensitive information outside established procedures. - The Wall Street Journal reported the alleged sharing involved an outside AI-safety organization. - OpenAI has not publicly said which organization, what information moved, or how it was shared. Those missing details are basically the whole story. Jimmy Apples quoted a source saying the material involved “infrastructure architecture.” Former OpenAI researcher Steven Adler cautioned that this is an extremely broad bucket, so treat that characterization as unverified for now. But it’s an interesting phrase, because infrastructure architecture is basically the map of everything an AI system can touch: sandboxes, tools, credentials, networks, shared services, monitoring systems, and the connections between them. And that happens to be almost exactly what an OpenAI Agent Security engineer named Joe was writing about days earlier. In a separate piece we published, we broke down Joe’s argument that containing frontier AI is “not just the sandbox.” He’s actually adamant that you do need strong sandboxes. His point is that the sandbox sits inside a much bigger system, and every connection around it becomes another place security has to hold. Joe breaks the technical problem into three layers: - Lock down the environment from first principles: sandbox, tools, credentials, networks, and connected services. - Alignment: the model has to understand and respect what it is and isn’t authorized to do, while independent security controls still enforce those limits. - Monitoring: watch what the agent actually does, keep the evidence outside its control, and make sure somebody can stop the run and revoke access. But Joe’s bigger point is organizational. Frontier labs need AI-safety researchers and cybersecurity teams working almost as one unit, backed by what he calls a “culture of reasonable paranoia”: people paranoid enough to find scary stuff, empowered to raise the alarm aggressively, and incident-response systems ready to act when they do. The mistake is thinking safety is the brake pedal on speed. Think about airplanes. They move much faster than cars, but aviation only works because the entire system around that speed is obsessively engineered for safety: redundancy, monitoring, checklists, abort procedures, incident investigations, people whose job is literally to say nope, something looks wrong. At frontier speed, safety isn’t the brake. It’s what makes the speed survivable. FROM OUR PARTNERS APEX-Agents: The AI Productivity Index for Agents See how the latest models rank for jobs like law, consulting, and investment banking. Mercor's APEX-Agents leaderboard evaluates frontier AI on long-horizon, multistep tasks across economically valuable work. Built with partners like Harvey, Ramp, and Cognition. Every model. Ranked by productivity. 🎓 AI Skill of the Day: Make AI interview you before it writes On yesterday’s stream, Corey described a writing workflow I think more people should steal: don’t ask AI to write from a half-formed idea. Talk the whole messy idea out first, then make the AI interrogate you before it drafts anything. That changes the model’s job. Instead of guessing what you believe, it becomes an editor that finds the missing pieces in what you already believe. You keep the ideas and voice; it helps with structure, weak logic, and the questions you forgot to answer. - Dump the idea. Talk or type without worrying about order, polish, or repetition. - Ask for an interview. Have the AI challenge assumptions, find contradictions, and ask one question at a time. - Draft last. Only after you answer the questions should it turn the material into an outline or first draft using your wording wherever possible. Copy/paste: I’m going to talk through a rough idea. Don’t write the draft yet. 1. Capture my claims, examples, questions, and assumptions. 2. When I’m done, interview me one question at a time to find gaps, contradictions, missing evidence, and weak logic. 3. Only after the interview, turn everything into a structured outline using my wording wherever possible. 🍪 Treats to Try - *Riverside records every remote guest on separate audio and video tracks, then gives you AI tools to edit and repurpose the session; free plan, then $24/mo billed annually. - Imbue Studio builds custom personal software from the workflow and interface you describe, then lets you reshape that interface, swap models, and keep the same context. - Claude Code mods let you rewrite prompts, block or retry tools, change permissions, redact outputs, or draw custom UI inside Claude Code with small TypeScript functions. - GitHub Copilot computer use lets Copilot read your screen and click, type, scroll, drag, and operate desktop apps that do not expose an API. - ChatGPT Try On puts clothes from shopping results or uploaded product photos onto your picture so you can preview an outfit before buying. - Mercury Voice is Inception’s diffusion model for enterprise voice agents; the company reports 320 ms median first-answer latency and $0.40/M input, $1.50/M output list pricing. 📰 Around the Horn - US President Trump told TIME he “might” consider an Intel-style government stake in OpenAI or Anthropic; no deal or timetable was announced. - Anthropic will reportedly target a mid-November IPO, with formal marketing potentially starting the week of November 9 so shares could trade before Thanksgiving. - Transluce reported rogue AI agents probed U.S. and Canadian government sites, including 200,000+ requests to the U.S. Education Department’s civil-rights site. - Runway’s Project Continuum is a preview of four real-time-video interface ideas, from generated “Portals” to interactive worlds, as an early operating-system research project. - Gemini 4 Argon is Google’s next frontier model, currently with early testers before broader access; planned API pricing starts at $2/M input and $10/M output. - arXiv, the pre-print research paper “archive”, capped submitters at two papers per calendar month after monthly submissions quadrupled over a decade and support tickets neared 9,000; wanna guess why??. - Researchers estimated AI-generated text made up 31.1% of FineWeb-filtered August web tokens, up from 10% in June 2024. - StudentBench found AI tutors matched expert human GRE tutors on immediate learning gains in a 2,383-student study, with one Gemma comparison far cheaper (long term gains from AI require tricks like this). - a16z claims AI has become a capital cycle reshaping debt, power, labor, hardware, and markets in its 90-page State of Markets report. - Agents in the wild: Meta’s Muse AI negotiated a sale on Marketplace, shared the seller’s pickup address, and told the buyer he was home, prompting an unexpected 9:15 p.m.visit (video). - China is setting rules for AI companions that restrict manipulative attachment behavior, add protections for minors, and require distress monitoring. 💡 Intelligent Insights - Ethan Mollick thinks the “Bitter Lesson” is coming for management: once AI agents can organize themselves around goals, the job shifts from designing the org chart to deciding what the swarm should actually optimize. - Aaron Levie and Jake Stauch see a new job forming inside companies: the “Automation Engineer,” someone who understands both AI and the messy internal workflows worth automating. - Josh Bleecher Snyder argues that throwing more agents at slow AI creates a new bottleneck: your attention. His fix is faster models plus interfaces that turn agent work into scannable artifacts instead of endless transcripts. - Andriy Burkov makes a nasty point about hallucinations: AI may be hardest to trust precisely when you’re doing genuinely novel work, because there’s no existing human answer to reveal when the model is wrong. - Sayash Kapoor found that making the model dramatically faster only sped up his agents 2-4x. Tool calls, code execution, and eventually human supervision become the bottleneck instead. - François Chollet argues modern reasoning models crossed an important line: instead of intuiting the answer directly, they increasingly intuit the procedure for getting to the answer. He thinks that shift explains much of their jump in reasoning ability. - Daniel Hook’s “Waymo effect” argues always-available AI collaborators can make researchers faster while removing the accidental human conversations that generate new ideas. - This one goes out out to my local AI nerds: Alex Ziskind benchmarked one M5 Ultra against two DGX Sparks and found the buying rule: Sparks handled giant prompts (large inputs) and shared workloads better, while the Mac excelled at long single-user generation (large outputs), so why not combine them, like Ash Hart (WARNING: do not try this at home unless made of money, or you will go broke!) New from The Neuron: AI Explained A Cat’s Commentary We want everybody to understand AI! That’s all for now. If you want to get featured above, fill out the poll below and tell us how we did today!
15:18

Google Moves Gboard's Private AI Training Into Secure Server Enclaves

Phone-keyboard training no longer happens on the phone. Google now uploads encrypted Gboard examples and updates the model inside attested server lockboxes. English and Japanese next-word models already ship this way. Access policies go on the public Rekor log, and binaries rebuild from the open Confidential Federated Compute repo. Old runs took one to two months because they waited on idle, charging phones. Side channels and the chip vendor are still trusted.

Notes
  • Next-gen federated learning: devices upload locally encrypted training examples; model updates compute inside server-side Trusted Execution Environments (TEEs).
  • Shipped: Gboard English and Japanese next-word prediction, with stronger DP and faster training than the prior on-device loop.
  • Old stack: examples stayed on the phone; devices computed updates; Secure Aggregation hid individuals until the combine. Bottleneck was phone availability and on-device compute. Auditors had little evidence about server privacy code.
  • New stack: encrypted examples leave the device. A key-management service releases decrypt keys only to approved TEE workloads. Those workloads apply differential privacy before weights leave the enclave.
  • Five stages: (1) client encrypts examples and binds an access policy (which workloads, how long); (2) policies published to Rekor, a public transparency log; (3) KMS across a TEE cluster using Raft, keys only after attestation matches an authorized policy; (4) root TEE runs the Python training loop, workers take parallel jobs, orchestrated with Federated Language (open-source descendant of TensorFlow Federated); (5) KMS-encrypted recovery state after each round so a replacement workload can resume without a plaintext checkpoint.
  • Auditor path: rebuild binaries from Confidential Federated Compute; compare measurement to attested workload; inspect Rekor policy; read the Python that clips contributions, adds noise, and controls outputs.
  • Runtime sideloading lets proprietary artifacts (parameters, tokenizers, architecture) enter the enclave without sitting in the public tree. Attestation confirms the measured program, not that every closed input is safe.
  • Gboard English next-word bench: 5,000 rounds, cohorts of 6,500 devices. Privacy-utility curves show less quality loss at a given privacy target than the previous system. Prior training time 1–2 months. New end-to-end time is not published. Bottleneck moves to TEE capacity.
  • Stated limits: TEE side channels (timing, cache, memory-access patterns); hardware root of trust (vendor, firmware, enclave impl); closed sideloaded artifacts; DP still depends on clip/noise/sampling/accounting; GPU/TPU confidential compute less mature than CPU TEEs.
  • No public managed API for the full Gboard pipeline. Two repos: confidential-computing components and Federated Language. Google also sketches synthetic-data generation and confidential analytics as future workloads.
Full text · 8,945 chars
- Google announced a next-gen Federated Learning system using Trusted Execution Environments for verifiable differential privacy. - Gradient computation moves from devices to server-side TEEs, removing on-device compute as the main bottleneck. - Access policies are published to Rekor transparency log so external auditors can verify server workloads. - Binaries are reproducibly built from the open-source Confidential Federated Compute repo. - Gboard English and Japanese next-word prediction models already shipped, with stronger DP guarantees and faster training. - Training time cut from 1-2 months, paving the way for larger models trained with federated techniques. Google moves federated learning into attested server enclaves Google has deployed a federated-learning stack that uploads locally encrypted training examples and computes model updates inside server-side Trusted Execution Environments, or TEEs. According to its technical post, the system now trains Gboard’s English and Japanese next-word prediction models. Google’s original federated-learning architecture kept training examples on each phone. Devices computed model updates locally, and Secure Aggregation prevented the server from inspecting individual updates before combining them. The approach reduced central access to user data, but phone availability and compute limited training, while external auditors had little evidence about the server software applying privacy protections. The new stack moves the confidential boundary from the phone into attested server hardware. Devices still supply decentralized data, while encrypted examples now leave the device. A key-management service releases decryption keys only to approved TEE workloads, and those workloads apply differential privacy before model weights leave the enclave. | How Google’s federated-learning architecture changes | | | |---|---|---| | Area | Previous architecture | New architecture | |---|---|---| | Training data | Examples remain on the device. | Devices upload locally encrypted examples. | | Model computation | Phones compute updates. | Server-side TEEs compute updates. | | Server access | Secure Aggregation reveals combined updates. | Approved metrics and differentially private weights leave the confidential workload. | | Verification | Auditors can inspect published protocols and client code. | Auditors can also check workload policies, reproducible binaries, and TEE attestations. | | Main bottleneck | Phone availability and on-device compute. | TEE capacity and server parallelism. | Five stages enforce each access policy - Encrypted upload: A client encrypts its training examples locally and associates them with an access policy. The policy identifies which confidential workloads may process the data and can limit how long access remains valid. - Public registration: Workload policies are published to Rekor, a public transparency log. Auditors can inspect the set of declared computations that devices may join. - Attested key release: A key-management service runs across a cluster of TEEs using the Raft consensus protocol. It releases decryption keys only after attestation shows that the requesting workload matches an authorized policy. - Confidential execution: A root TEE runs the Python training loop and delegates parallel work to worker TEEs. Google orchestrates these jobs with Federated Language, an open-source descendant of TensorFlow Federated. - Protected recovery: The program saves key-management-service-encrypted recovery state after each training round. A replacement workload can resume after a failure without exposing an additional plaintext checkpoint. Attestation narrows the trust surface A TEE provides remote attestation, confidential memory, and execution integrity. Remote attestation lets another system verify the identity and configuration of code running inside supported hardware. Confidentiality hides the workload’s internal state from the host, while integrity controls prevent the host from silently modifying execution. Google connects those machine-level properties to a public audit trail. The key-management and data-processing binaries can be reproducibly built from the Confidential Federated Compute repository. An auditor can rebuild the software, compare its measurement with the attested workload, inspect its access policy in Rekor, and review the Python code that clips contributions, adds noise, and controls outputs. Runtime sideloading allows proprietary artifacts, including model parameters, tokenizers, and serialized architecture details, to enter the confidential workload without appearing in the public source tree. The audited Python program still governs data access and release. This arrangement exposes the privacy control flow while allowing model intellectual property to remain closed. Gboard swaps handset delays for enclave capacity Gboard previously had to wait for eligible phones, typically devices that were idle and charging. Training progress varied with time zones, device availability, local compute, and competition among jobs requesting the same handsets. The new architecture collects encrypted uploads first, then schedules training across server-side TEEs when capacity is available. The previous Gboard models could take one to two months to train. Google reports substantial speedups from server parallelism, although its post does not provide a new end-to-end training time. TEE availability now sets the throughput ceiling. Central differential privacy also becomes easier to enforce inside the confidential workload. The program can bound each device’s contribution and add calibrated noise before releasing model weights. This limits how much any single device can influence the output, with the privacy budget controlling the strength of that bound. For an English next-word prediction benchmark, Google trained for 5,000 rounds with cohorts of 6,500 devices. Its reported privacy-utility curves show less model-quality loss at a given privacy target than the previous system. The guarantee stops at several boundaries - Side channels remain: Current TEEs can leak information through timing, cache behavior, memory-access patterns, and other side channels. Google acknowledges that the hardware does not eliminate these attack classes. - Hardware roots stay trusted: Attestation depends on the processor vendor, its signing infrastructure, firmware, and the security of the enclave implementation. - Closed artifacts limit review: Auditors can inspect the privacy logic but may be unable to examine sideloaded model components. A valid attestation confirms which measured program ran, not whether every closed input is safe or correct. - Differential privacy needs sound parameters: Enclaves can enforce a configured mechanism, but the protection still depends on clipping bounds, noise levels, sampling assumptions, and privacy accounting. - Capacity moves to the data center: Training speed depends on available confidential-computing hardware. Support for GPUs and TPUs remains less mature than CPU-based TEE execution, especially for large deep-learning workloads. Two repositories expose the building blocks Google has released the confidential-computing components and the Federated Language orchestration framework as open source. The announcement does not include a public managed service or API that developers can use to run the complete Gboard infrastructure. Components required for a similar deployment - Clients that encrypt uploads and authorize explicit access policies - A public, append-only transparency log - Reproducible enclave binaries and published source code - An attestation-aware key-management service - Confidential workloads that enforce contribution limits and differential privacy - Encrypted checkpoints that preserve privacy during recovery - Auditing tools that connect source code, binary measurements, policies, and attestations Confidential accelerators set the next ceiling Moving model computation into server-side enclaves removes the memory, power, and runtime limits imposed by phones. Larger architectures become feasible as confidential accelerators gain stronger isolation, attestation, and integration with key-management systems. Google also describes the infrastructure as a general Python execution environment and is experimenting with workloads such as synthetic-data generation. The same policy, attestation, and controlled-release design could support confidential analytics or model inference over sensitive user data. For developers handling private datasets, the reusable design lies in the connection between client-approved policies, public logs, reproducible builds, attested execution, restricted key release, and differentially private outputs. Each component supplies evidence about who can process encrypted data, which program receives it, and what information may leave the confidential boundary.
15:19

Open-sourcing AstaBrief, the fast report-generation model in Asta

A small open model now writes cited science reports in one pass, fast enough that labs can run it on their own machines. Ai2’s AstaBrief 8B starts from Qwen3-8B and is the Fast mode in Asta next to a Claude Thinking mode. Across the full pipeline Fast averages 51.1 seconds per report versus 178.5 seconds for Thinking, about 3.5× faster. Training used 47K supervised reports and 6K preference pairs. The authors have not rerun the evals against today’s frontier models.

Notes
  • AstaBrief 8B: open-weights report writer. Input = research question + retrieved literature excerpts. Output = cited report in one pass (no section-by-section write, no snippet summarization/clustering used by Claude Thinking mode).
  • Serving: Asta “Generate a report” Fast mode vs Claude-powered Thinking mode. Full-pipeline times: Fast 51.1 s average vs Thinking 178.5 s (~3.5×). Authors also say nearly an order-of-magnitude cut vs the proprietary models they tracked during development.
  • Base: Qwen3-8B. Recipe is SFT + DPO, not RL. They considered RL (citing DR Tulu) but wanted a cheaper, more debuggable loop.
  • Query source: real Asta / ScholarQA user logs. Filtered beta testers, bots, short queries, non-English, non-scientific, and personal info via an LLM pass. Pool left: 90K research-focused queries.
  • SFT targets: ScholarQA multi-step pipeline (retrieve, section, synthesize) backed by Claude 3.5 Sonnet, Claude 3.7 Sonnet, o3, o4-mini, GPT-4.1. After quality filters: 47K examples.
  • DPO: held-out queries. Preferred report from ScholarQA (usually Claude 3.5/3.7). Competitor from the same retrieved excerpts via o3, o4-mini, DeepSeek-V3, or DeepSeek-R1. Judges GPT-4.1 and DeepSeek-R1; keep only when both agree (judges aligned with humans at 95%). Final DPO set ~6K.
  • Main eval: SQABench-CS2, 200 user-written CS research questions. Metrics: rubric score (coverage), answer precision (paragraph relevance), citation precision, citation recall.
  • Secondary: DeepScholarBench (63 queries from recent arXiv); LLM pairwise vs Claude pipeline on SQABench-CS2; 14-question human study (three scientific researchers, 4–5 questions each, ties allowed). On overall preference DR-Tulu wins; two of three researchers prefer AstaBrief on citation accuracy.
  • Data-quality lesson: first SFT improved writing but lagged Claude on precision/citations. Four statistic filters tested (output/input token ratio, citation relevance, citation density, citation diversity). Strongest gain: drop low citation-density synthetic reports. More aggressive combos and LR sweeps did not add much.
  • Caveat they flag: a citation can support a related claim while the sentence overstates scope (sample → population, past-tense result → present universal, descriptive finding → recommendation). Development metrics did not fully score that.
  • Training/eval mostly completed in 2025. Proprietary comparison models reflect that frontier. They have not rerun the full eval against today’s frontier. Read as evidence about the recipe, not a 2026 leaderboard claim.
  • Early Fast-mode usage (374 users who tried it): 29.1% used it on two or more days; 3.67 report threads on average. 23% stayed on Fast and never switched back. 18% mix modes, Fast ~40% of their threads. Positive-feedback rate 84.2% Fast vs 85.2% Thinking. Feedback is sparse.
  • Why open weights: institutions can run reports on unpublished or sensitive work behind their own firewall. Example workflow released for reports from a researcher’s own PDFs.
  • NSF OMAI context; ScholarQA paper “Synthesizing scientific literature with retrieval-augmented LMs.” Future work named: finer preference learning, RAG+RL, multi-turn/multi-tool, more scientific sources, evals that preserve evidentiary scope.
  • Limitation: optimized as a piece of Asta’s agentic report stack, not necessarily as a standalone chat model.
Full text · 17,788 chars
Language models can already help researchers search the literature, synthesize evidence, and work through complex questions. But scientific work places particular demands on these models—answers need to stay grounded in evidence, the models need to preserve what the evidence actually supports rather than quietly broadening a study’s conclusions, and researchers need to be able to verify the final outputs. We see that in how scientists use Asta, our agentic platform for scientific work. Instead of simple keyword searches, users often bring substantial context and many constraints—for example, asking Asta to compare approaches across a body of literature while accounting for a particular method, population, or setting. Many also return to generated reports later, treating them as working research artifacts rather than one-off answers. We wanted to help scientists generate cited reports faster, with a model they could download and run themselves. To do that, we tested whether a small, open model trained specifically for scientific report generation could match the report quality of the proprietary models we were using, while reducing generation time and serving costs. We built AstaBrief 8B, a model that turns a research question and retrieved literature excerpts into a cited report. AstaBrief is available in Asta’s Generate a report feature today as Fast mode alongside Claude-powered Thinking mode, and we’re also open-sourcing it and the training data so others can study, reproduce, and build on our approach. Developing AstaBrief required tens of thousands of real research queries, citation-focused filtering, preference data, and a redesigned report-generation pipeline that writes the full report in one pass rather than section by section. The result is nearly an order-of-magnitude reduction in report generation time compared to the proprietary models we tracked—across the full Asta pipeline, Fast mode averages 51.1 seconds per report compared with 178.5 seconds for Thinking mode, about 3.5× faster. Together, those efficiency gains made AstaBrief a useful test case for a broader goal: building open language models that can be adapted to the specific demands of scientific work. Open weights will also let institutions run AstaBrief on their own infrastructure, which is necessary when research questions reveal sensitive or unpublished work. Alongside the model weights, we’re releasing an example workflow that researchers can adapt to create reports from their own PDFs, providing a starting point for local report generation This post covers how we trained AstaBrief, what we learned about grounding it in scientific evidence, and which parts of our approach we think can carry forward to future models for science. Most of the training and evaluation described was completed in 2025, so the proprietary models used to generate training data and as comparison points reflect the frontier at the time. We haven’t rerun the full evaluation against today’s frontier models; the results below are best read as evidence about the particular training and system design choices we tested. Our goal with AstaBrief was to build an open-weights model with all the qualities that matter most for long-form scientific synthesis: answer quality, relevance, structure, and citation grounding. We started from Qwen3-8B and focused most of our effort on the post-training data, evaluation, and surrounding report-generation scaffolding. Adapting general-purpose models for scientific work – and training new scientific models from scratch – is something we're exploring broadly across Ai2. Through NSF OMAI, a U.S. national initiative led by Ai2 to build fully open AI infrastructure and models for scientific discovery, our researchers are working directly with scientific communities to understand what they need from future open models and where today's general-purpose models fall short. That includes studying how needs differ across scientific fields and workflows, with more findings from that research to share in the future. Recent work, including our DR Tulu, has shown that reinforcement-learning-based (RL) methods can improve long-form report generation for open-weights models, especially when judge models are involved in the training loop. We considered that path for AstaBrief, but ultimately focused on a simpler recipe built around supervised fine-tuning (SFT) and direct preference optimization (DPO). RL-based training can be unstable and expensive. We wanted to see how far we could push report generation quality with a cheaper, more operationally manageable setup—one that's also easier to debug and iterate on. That made the quality of the training data especially important. Rather than relying on a more complex optimization method to compensate for noisy examples, we spent much of the project figuring out how to generate, select, and filter examples that actually demonstrated the report-writing behavior we wanted. We also wanted AstaBrief to be faster so that users could get preliminary reports quickly that they could then iterate over in subsequent turns. For speed improvements, we decided to train AstaBrief to directly generate the final report in one pass given a user query and relevant retrieved snippets, bypassing the expensive snippet summarization and clustering stages our Claude-based Thinking mode uses and not writing out the answer section-by-section. Interestingly, we found it was possible to do so without sacrificing performance. The training pipeline began with real user queries submitted through the system described in our paper “Synthesizing scientific literature with retrieval-augmented LMs” and ScholarQA, the framework that now underpins Asta’s Generate a report feature. Rather than training only on synthetic prompts or benchmark-style tasks, we wanted AstaBrief to learn from real queries from real scientists. Our research suggests that scientists often ask different things of language models than users do of general-purpose chatbots or traditional search tools. In our analysis of hundreds of thousands of Asta queries, expert researchers frequently supplied substantial context, multiple constraints, and relationships between concepts rather than relying on short, keyword-style prompts. More recent Asta user studies have also surfaced differences in how researchers want AI involved in their work—some are comfortable using models for ideation or experimentation, while others prefer a narrower role in synthesis, literature surveillance, or pattern-finding. Across those differences, participants want clearer source traceability, more visibility into what a model is doing, and greater control over the context it uses. We filtered the user logs we collected for quality, relevance, and privacy, stripping out beta-tester and bot traffic, dropping queries that were too short to be meaningful, and using an LLM-based filtering pass to catch non-English queries, non-scientific requests, and prompts containing personal information. That left a pool of 90K research-focused queries. For SFT, we generated full-report target outputs from the filtered queries using the multi-step ScholarQA pipeline behind Asta's report generation. The pipeline retrieved relevant literature, organized the material into sections, and used a backing report-generating model to synthesize the evidence into a cited report. We drew on a mix of proprietary systems: Claude 3.5 Sonnet, Claude 3.7 Sonnet, o3, o4-mini, and GPT-4.1. After quality filtering, this yielded 47K usable training examples. DPO required a different kind of training data. Instead of a single target report per query, we needed pairs of reports with one preferred over the other. We built those pairs from a separate subset of queries not used during SFT data generation. One report per query came from the existing ScholarQA pipeline, typically backed by Claude 3.5 Sonnet or 3.7 Sonnet. The competing report was generated by feeding ScholarQA's retrieved literature excerpts to a different model: o3, o4-mini, DeepSeek-V3, or DeepSeek-R1, depending on the example. Two judge models – GPT-4.1 and DeepSeek-R1 – compared each pair and picked a winner. We ensured that LLM judges were aligned with human preferences (95% agreement) and only kept pairs where both judges agreed, which gave us a cleaner preference set and cut much of the noise that typically shows up in preference data generated at scale. After quality filtering, the final DPO dataset came to about 6K examples. Using multiple generators and requiring agreement between two judges gave us a relatively simple way to construct preference data without treating any single model’s output or judgment as ground truth. Our main evaluation target was SQABench-CS2, a set of 200 user-written computer science research questions. We tracked four metrics throughout the development of AstaBrief: - Rubric score, which measures how much necessary content is covered by the report. - Answer precision, which measures whether each paragraph is relevant to the question. - Citation precision, which measures whether each citation supports the claim it's attached to. - Citation recall, which measures whether the report's claims are fully supported by the citations provided. For our final model, we also ran secondary evaluations: DeepScholarBench, a 63-query benchmark for long-form research synthesis built from recent ArXiv papers, and two separate pairwise evaluations against reports generated by the Claude-powered pipeline—an LLM-judged comparison on SQABench-CS2 and a small human study. A report can sound polished and complete while meandering from the question or attaching citations to claims from which the underlying evidence doesn't follow. For scientific synthesis, we needed to measure those behaviors separately. But citation support is only part of scientific faithfulness—a model can cite the right study and still make a stronger claim than the study itself supports. This can happen in subtle ways, for example, turning a finding about a particular sample into a generic claim about an entire population, shifting a result reported in the past tense into a present-tense statement that sounds more universally true, or turning a descriptive finding into a recommendation for what clinicians, policymakers, or researchers should do. Those kinds of generalizations are especially important for scientific report generation because each step can broaden the apparent scope of the evidence without introducing an obviously false statement. A cited sentence may therefore be technically related to its source while still overstating what researchers actually established. Our development metrics focused primarily on relevance, coverage, and citation grounding; a richer evaluation of scientific report writers should also test whether they preserve the scope and strength of the claims in their sources. Our first SFT runs improved overall content quality, but they still lagged behind our Claude-powered report generation pipeline on answer precision and citation quality. In other words, the model got better at writing reports, but it still wasn’t grounded in evidence as consistently as we needed for scientific synthesis. That pushed us to spend more time on data quality. We tested four statistics-based filters to identify weaker synthetic training examples: - Output-to-input token ratio. Answers with very high ratios were often noisy because they were generating a lot of text from too little evidence. - Citation relevance. For each synthetic report in the training set, we averaged the retrieval relevance scores of its cited papers. Low averages suggested the report was relying too heavily on lower-ranked evidence. - Citation density. We measured the share of statements that had at least one citation. Low-density reports often had large stretches of unsupported text. - Citation diversity: We measured the share of papers cited in the answer, given the set returned by the Claude-powered report retrieval pipeline. Low scores suggested the report was overly reliant on a few papers. The strongest gains came from filtering out synthetic reports with low citation density; more aggressive filtering, filter combinations, and learning-rate sweeps didn't add meaningful gains. That was one of the clearest lessons from the project: more elaborate filtering wasn’t necessarily better. A relatively simple signal – whether the synthetic reports consistently cited their claims – was more useful than several more complicated combinations we tried. Scientific specialization, in other words, isn't necessarily a matter of adding more scientific text to pretraining; the composition and quality of post-training data and whether it demonstrates behaviors like grounding and attribution can materially change how the resulting model performs. That focus on grounded, useful output also lines up with what we’ve heard in Asta user research. Participants note that generating more text isn't necessarily more helpful; they want concise synthesis and enough source traceability to review and verify results without wading through unnecessary outputs. Once we had a stronger SFT checkpoint, we ran DPO training on top of it. That stage pushed performance further, bringing AstaBrief within range of the Claude-powered report pipeline in Asta and DR Tulu on report generation. Because this model was intended to work as part of our agentic Asta report generation framework (not necessarily as a standalone model), our main question was whether AstaBrief could preserve the report qualities we cared about while enabling a substantially faster and cheaper report-generation pipeline. In other words, we weren’t only asking whether the model could match a stronger proprietary model on individual benchmarks; we wanted to know how much of that quality we could retain with a much simpler system. Each row is ordered best first; higher is better on every metric. Qwen3-8B was evaluated on SQABench-CS2 test only. SQABench-CS2 is a set of user-written computer science research questions; DeepScholarBench scores long-form research synthesis with its own metrics, which are not comparable with SQABench-CS2's. In the evaluations we used during development, AstaBrief was competitive with the Claude-powered pipeline and DR Tulu across several measures of answer and citation quality. The chart below shows the LLM-judged comparison—in a separate 14-question human study, three scientific researchers each contributed 4-5 questions and ranked reports from the three systems on overall preference, completeness, relevance, organization, and citation accuracy (with ties allowed). On overall preference, DR-Tulu wins, but two of the three researchers prefer AstaBrief over other systems on citation accuracy metrics, demonstrating the utility of our SFT data quality filters. Bars show the share of LLM-judged report comparisons each system won against Thinking mode on the same questions. Human judgments were evaluated separately and are not included. Thinking mode is the comparison reference and has no bar. Unlike DR-Tulu, Asta Brief was optimized for this pairwise report ranking during the DPO stage. These numbers are best read as validation of the engineering approach at the time we developed it, rather than as a claim about where this particular base model sits relative to today’s frontier. The model ecosystem moves quickly—the data construction, attribution filtering, and serving lessons are the pieces we expect to generalize. Validating the usefulness of AstaBrief in Asta, Fast mode has shown encouraging early usage. Among 374 Asta users who’ve tried it, 29.1% have used it for two or more days, and users on average generate 3.67 report threads with it. Twenty-three percent of users who tried Fast mode continued using it and never switched back to Thinking mode for future threads. An additional 18% switched between Fast and Thinking modes depending on their goals, using Fast mode for ~40% of their threads. While feedback is generally too sparse to draw strong conclusions, we see that Fast mode receives positive feedback at a similar rate as Thinking mode (84.2% versus 85.2%). Asta's report generation is the first production use of AstaBrief, giving researchers an open-weights Fast mode alongside the existing Thinking mode. Because the model is open weights, institutions can deploy it on their own hardware, including behind their own firewall, without relying on a proprietary model API for report generation. In Asta, that also means we can study and improve this part of the report generation pipeline directly while preserving Thinking mode as an option for more compute-intensive tasks. There's more to do. We're exploring more fine-grained preference learning, stronger RAG-plus-RL approaches, multi-turn and multi-tool capabilities, additional scientific data sources, and query decomposition. We're also interested in evaluations that go beyond whether a claim has a supporting citation to ask whether a model preserves the evidentiary—both to better capture the quality of the report as a research artifact and to ask whether a model preserves the evidentiary scope of its sources. That includes qualities such as concision and organization, as well as whether the model turns sample-specific findings into broad generalizations or descriptive results into recommendations. AstaBrief is one experiment in a longer line of work on language models for science, from ScholarQA and DR Tulu to future versions of Olmo beginning to take shape now. The lessons here – especially around training data, filtering, and evaluation- can help inform what we build next. Try Fast model today in Asta, or download AstaBrief from Hugging Face.
15:53

Google Sends Trillium TPUs to Orbit to Power AI Beyond Earth's Grid

Google put four of its AI chips on a satellite to see whether space can power models the earth’s grid cannot. Project Suncatcher M1 flew on SpaceX Transporter-18 from Vandenberg in a Planet-built box the size of a fridge. The slice matches a Cloud TPU v6e-4, draws about 1 kW of solar, and runs Gemini inference in 15-minute bursts so it can cool. Lab radiation tests showed high-bandwidth memory errors at 2 krad(Si). A two-satellite laser-link test is planned for early 2027. This flight has no peer in orbit.

Notes
  • Mission M1: first orbital prototype for Project Suncatcher. Planet-built spacecraft, fridge-sized. SpaceX Falcon 9 Transporter-18 rideshare from Vandenberg. Contact ~1 hour after liftoff; operating as expected.
  • Compute: four Trillium v6e TPUs = Cloud TPU v6e-4 slice. ~1 kW solar. Dawn-dusk sun-synchronous LEO. Workload: Gemini inference in ~15-minute windows, then thermal cooldown.
  • Why orbit: Google estimates up to 8× usable solar vs a comparable terrestrial panel (near-continuous sun, less night/cloud/atmosphere). New constraints: launch mass, radiation, vacuum cooling.
  • Vibration: 50–100 g on components to match Falcon 9. Radiation: Crocker Nuclear Lab (UC Davis), 67 MeV proton beam. HBM irregularities after 2 krad(Si). One TPU: no hard TID failure through 15 krad(Si), which Google says exceeds a five-year mission dose.
  • Heat: no convection. Heat pipes + radiators. Burst-then-cool schedule. Larger birds need more radiator area and mass.
  • Optical interconnect (bench, not this flight): 800 Gbps each way on one transceiver pair, 1.6 Tbps aggregate bidirectional. M1 has no second satellite. Follow-up: two linked prototypes with Planet by early 2027 for formation flying and distributed workloads.
  • Open questions named: cost per useful compute after build/launch/ground/replace; radiator/solar/formation scale for dozens of TPUs; uplink/downlink cost and latency; long-term degradation; debris and deorbit.
  • Software implications they list: fault-tolerant inference, thermal scheduling, model partitioning across thin optical links, checkpointing, dynamic routing.
  • Partners: Google accelerators + distributed-systems research; Planet small-sat engineering; SpaceX rideshare. Not a commercial constellation. Terrestrial drivers: grid interconnects, power contracts, water, permitting.
Full text · 8,538 chars
- Google's Project Suncatcher launched its first TPU-carrying satellite on SpaceX Transporter-18 from Vandenberg. - The Planet-built MVP spacecraft carries four Trillium TPUs, equivalent to a Cloud TPU v6e-4 slice. - Satellite draws about 1 kW solar; runs Gemini inference in 15-minute bursts, then cools down. - Trillium survived lab radiation beyond a five-year mission dose; HBM errors began at 2 krad(Si). - Bench optical link demo hit 1.6 Tbps total between a single transceiver pair. - Two linked satellites are planned for 2027 to test inter-satellite laser interconnect in orbit. Google puts four Trillium TPUs in orbit Google has launched the first orbital prototype for Project Suncatcher, its research program for space-based AI computing. Built with Planet, the spacecraft rode SpaceX’s Transporter-18 mission from Vandenberg Space Force Base and deployed about an hour after liftoff. Ground controllers established contact and reported that the satellite was operating as expected. The flight adds in-orbit telemetry to Google’s laboratory tests of Trillium Tensor Processing Units. The refrigerator-sized spacecraft carries four TPUs and will run short Gemini inference workloads while engineers measure radiation errors, power consumption and heat rejection. Project Suncatcher’s proposed end state is a constellation in which satellites carrying dozens of accelerators cooperate on training or inference. The current mission, designated M1, tests one small compute slice and has no inter-satellite peer. Why orbit offers more solar power Google estimates that a solar panel in the target orbit could receive up to eight times as much usable solar energy as a comparable terrestrial installation. A dawn-dusk sun-synchronous orbit keeps a spacecraft in near-continuous sunlight, avoiding most nighttime, cloud and atmospheric losses. Near-continuous generation could ease the power constraints affecting terrestrial AI expansion, including limited grid connections and competition for new generation capacity. Orbit introduces separate constraints, particularly launch mass, radiation exposure and cooling without air. Inside M1’s refrigerator-sized payload | Project Suncatcher M1 hardware and flight plan | | |---|---| | Component | Details | |---|---| | Spacecraft | Planet-built prototype, roughly the size of a refrigerator | | Compute | Four Trillium v6e TPUs, equivalent to a Google Cloud TPU v6e-4 slice | | Power | About 1 kilowatt from solar panels | | Orbit | Dawn-dusk sun-synchronous low Earth orbit | | Launch | SpaceX Falcon 9 Transporter-18 rideshare from Vandenberg | | Workload | Gemini inference in roughly 15-minute windows, followed by thermal cooldown | | Primary measurements | Radiation errors, thermal behavior, power use and accelerator stability | Launch, radiation and heat set the limits Launch loads Before flight, Google vibration-tested the satellite along all three axes to reproduce Falcon 9 launch conditions. Individual components experienced accelerations of 50 to 100 times Earth’s gravity. M1 will show whether connectors, memory, power systems and cooling hardware continue working after launch and sustained orbital operation. Radiation exposure Google tested Trillium TPUs at the Crocker Nuclear Laboratory at the University of California, Davis, using a 67 megaelectronvolt proton beam. The beam reproduces one class of energetic particles that can corrupt data or damage semiconductor components in space. High-bandwidth memory irregularities appeared after a cumulative dose of 2 krad(Si), a measure of ionizing energy absorbed by silicon. One TPU showed no hard failure attributed to total ionizing dose through 15 krad(Si), which Google says exceeds the expected dose for a five-year mission. Memory behavior matters because model throughput depends heavily on how quickly parameters move between high-bandwidth memory and the accelerator. Possible safeguards include error-correcting code memory, redundant computation, frequent checkpoints, workload retries and software that quarantines unreliable devices. Heat rejection Vacuum eliminates convective cooling, so nearly every watt consumed by the electronics must leave through radiators. M1 uses heat pipes and radiator surfaces, with the TPUs scheduled to run in short bursts before cooling. A larger satellite would require more radiator area and thermal transport capacity, adding mass and constraining the placement of chips, solar panels and communications hardware. Lasers determine whether the cluster can scale An orbital cluster needs inter-satellite links fast enough to move model parameters, activations and checkpoints between accelerators. Google’s bench-scale optical system reached 800 gigabits per second in each direction through one transceiver pair, providing 1.6 terabits per second of aggregate bidirectional capacity. Maintaining that throughput requires two moving spacecraft to keep narrow laser beams precisely aligned despite vibration, formation drift and attitude adjustments. Existing optical space links generally prioritize long-distance communication at lower bandwidths. Suncatcher requires much higher bandwidth across shorter distances between satellites flying in formation. M1 has no second satellite with which to test an optical crosslink. Orbital validation therefore depends on the planned follow-up mission. What one satellite can measure M1 can provide error rates, component temperatures, power profiles and evidence about how long commercial AI accelerators remain usable in low Earth orbit. Those measurements will help Google estimate thermal limits, hardware redundancy and expected service life for a larger design. Questions beyond M1 - Economics: Cost per useful unit of compute after spacecraft manufacturing, launch, ground infrastructure and replacement missions. - Scale: Radiator mass, solar-panel area and formation-control requirements for satellites carrying dozens of TPUs. - Ground connectivity: The cost, bandwidth and latency of uploading data and returning results. - Service life: Long-term degradation of memory, solar panels, optical hardware and thermal systems. - Orbital operations: Collision avoidance, constellation management and reliable deorbiting at end of life. Software has to expect dropouts An orbital TPU cluster would need software designed around intermittent links, thermal pauses and hardware faults. Conventional data center assumptions about stable nodes and continuously available network paths would be unreliable in that environment. - Fault-tolerant inference: Retry failed operations, detect corrupted outputs and remove unreliable chips without stopping an entire job. - Thermal scheduling: Assign work according to radiator capacity, chip temperature and available solar power. - Model partitioning: Place parameters and operations to limit traffic across constrained optical links. - Checkpointing: Preserve recoverable state when satellites lose contact or pause compute. - Dynamic routing: Redirect traffic as formation geometry and link availability change. M1’s telemetry can supply the failure rates, cooldown times and power limits needed to design those policies. The two-satellite mission would add real measurements for link stability, throughput and recovery after interruptions. Terrestrial constraints drive the experiment Large AI data centers face limited grid connections, lengthy power contracts, water constraints and local permitting. Other proposals include data centers supplied directly by nuclear plants and sealed computing modules placed underwater. Project Suncatcher explores whether near-continuous orbital solar power can support another deployment model. Google contributes the accelerators and distributed-systems research, Planet supplies small-satellite engineering, and SpaceX provides rideshare access to orbit. Any commercial design must also account for launch expense, difficult repairs, hardware replacement, debris management and dependence on ground networks. The next flight tests the network Google Research plans to launch two prototype satellites with Planet by early 2027. That mission is intended to test the high-bandwidth optical link, formation flying and distributed workloads across separate spacecraft. Combined data from M1 and the two-satellite flight could support realistic estimates for cluster size, radiator capacity, network topology and fault budgets. Google has not published a schedule for production workloads or a commercial orbital computing service.
16:14

☕️ Amazon courts towns to defuse data center backlash

Amazon is paying towns to live with its data centers, and the same briefing stacks four other shocks from the day. Built Together pledges over $1 billion across five years: community college for about 300,000 students, 16 more trade centers, energy upgrades in 300-plus schools and 30,000 homes, plus roads and parks. Amazon says it will drop NDAs with agencies, publish yearly energy and water use, and pay enough for power that local bills do not rise. The same issue covers three OpenAI safety firings, ChatGPT clothes try-on, a $300 million chip-smuggling arrest, and a 2,199% jump in Slovenia’s .si domains after a “super intelligence” order.

Notes
  • Amazon “Built Together”: >$1B over five years for US towns hosting data centers. Community-college costs for ~300,000 students; 16 more trade-training centers; energy upgrades in 300+ schools and 30,000 homes; local roads/parks. Also: drop NDAs with government agencies on future projects; lower-emission backup generators; publish yearly energy and water use; pay enough for power to keep local electricity bills from rising. No independent audit in the briefing.
  • OpenAI: three safety-team members let go (WSJ). Company says an internal investigation found they mishandled confidential details with an outside AI safety group. Names, group, and material not given. Briefing also cites a NYT report on dismissed staff warnings and recent agent incidents (escaped containment, posted user images, hacked government websites) — those claims are the newsletter’s, not independently sourced here.
  • ChatGPT shopping: global virtual try-on. Upload a selfie or full-body photo. Try On button in shopping results, or screenshot an item. Powered by ChatGPT Images 2.5. Favorites saves products + try-on images to a Library.
  • Chip case: Greg Lui, 38, founder/CEO of Earthmade Computers (California), arrested. Charged with smuggling >$300M in restricted Nvidia AI chips to China. Allegedly fake paperwork 2023–2024, servers to Malaysia and Singapore (no US license), then re-export to China. Earthmade took in >$176M from two Malaysia-based shipping firms in 2024. Charges: export violations, smuggling, money laundering; up to 50 years. Allegations, not a verdict.
  • Apple (Gurman/Bloomberg): chapstick-tube security camera that does not record video. Tied to unannounced smart-home display J490. Very low frame rate, facial recognition, text descriptions of who walks in. Sensor tech said to borrow from leaked camera AirPods (low-res stills for Visual Intelligence). Report, not a ship date.
  • Slovenia .si: Trump EO telling agencies to say “super intelligence” / “SI” instead of AI. Registry SI: 2,199% jump in .si purchases in September; 11,000 new addresses on Sept 30 (day after the order); nearly 13,000 the next day. Hostinger: .si now its second-most-popular extension after .com; ~$12/year vs ~$90 for .ai. Most buyers US and India; few names actually linked to AI.
Full text · 4,218 chars
| | | 💰 Amazon courts towns to defuse data center backlash LINK | Amazon pledged over $1 billion across five years for U.S. towns hosting its data centers, trying to calm rising local pushback against the AI buildout under a new program it calls "Built Together." The money will cover community college costs for about 300,000 students, build 16 more trade-training centers, upgrade energy use in 300-plus schools and 30,000 homes, and fund local projects like roads and parks. Amazon also said it will drop NDAs with government agencies on future projects, install lower-emission backup generators, publish its yearly energy and water use, and pay enough for power to keep local electricity bills from rising. | 🔒 OpenAI fires 3 safety researchers LINK | OpenAI has let go of three members of its safety team, saying they shared confidential company details with an outside AI safety group, according to a Wall Street Journal report published yesterday. The company said an internal investigation found the researchers mishandled sensitive information outside set procedures, though it did not name the people, the outside group, or the specific details that were shared. The firings follow a New York Times report that OpenAI leaders dismissed staff warnings about safety, plus recent incidents where its AI agents escaped containment, posted user images, and hacked government websites. | 👗 ChatGPT can now virtually try on clothes LINK | OpenAI has rolled out a global ChatGPT feature that lets shoppers virtually try on clothes and accessories by uploading a selfie or full-body photo to see how an item might look on them. A new Try On button appears in ChatGPT's shopping results, and users can also upload a screenshot of an item and ask the assistant to picture it on them, powered by the new ChatGPT Images 2.5 model. A second feature, Favorites, lets people save products they find to a Library in the app alongside their try-on images, so they can return to those items later. | 🔌 CEO arrested for smuggling AI chips LINK | Federal agents arrested Greg Lui, the 38-year-old founder and CEO of California firm Earthmade Computers, on charges that he smuggled more than $300 million in restricted Nvidia AI chips into China. Prosecutors say Lui and others used fake paperwork from 2023 to 2024 to ship high-end servers to Malaysia and Singapore, where no US license is needed, then secretly re-exported them to China. Earthmade took in over $176 million from two Malaysia-based shipping firms during 2024, and Lui now faces charges of export violations, smuggling, and money laundering that could bring up to 50 years in prison. | 📷 Apple builds a camera without video LINK | Apple is reportedly working on a small security camera that skips video recording entirely, using AI instead to track movement around your home and flag who's there, according to Bloomberg's Mark Gurman. Shaped like a chapstick tube and tied to Apple's unannounced smart home display known internally as J490, the camera runs at a very low frame rate and can do facial recognition, sending only text descriptions of who walks in. Gurman says the sensor borrows technology from Apple's leaked camera-equipped AirPods, which take extremely low-resolution still images for Visual Intelligence but, like this camera, won't record video or snap regular photos. | 🌐 Slovenia's .si domains surge after Trump's 'super intelligence' order LINK | Slovenia's .si domain names are suddenly in heavy demand after President Trump signed an executive order telling all government agencies to call "artificial intelligence" and "AI" instead "super intelligence" and "SI." Registry SI, which runs Slovenia's domains, recorded a 2,199% jump in .si purchases in September, with 11,000 new addresses on September 30th, the day after Trump's order, and nearly 13,000 more in the following day. Web host Hostinger says .si is now its second most popular extension after .com, costing about $12 a year versus $90 for .ai, with most buyers coming from the US and India and few names actually linked to AI. | |
01:39

Synopsys launches agentic AI for chip design, partners with Samsung, Nvidia, Intel

A chip-design vendor says its new autonomous engineering agent can cut verification signoff time in half. Synopsys launched Agent Engineer and is partnering with Samsung, Nvidia, and Intel. The stored snippet says verification signoff speeds up by up to 50. No independent benchmark or product SKU is in the text.

Full text · 150 chars
Synopsys reveals results from its agentic AI‑based autonomous engineering; the company says Agent Engineer speeds verification signoff by up to 50 ...
04:31

Forget 'superintelligence': error-prone AI nearly sparked world war three this month

A column says the real scare this month is sloppy systems, not a future god-model, and opens on a safety engineer walking out. It names Anthropic engineer Jacob Coxon’s resignation as the biggest international AI news of the past three weeks. The next clause starts “According to him, OpenAI and” and then the snippet ends. No war scenario is in the stored text.

Full text · 144 chars
The biggest international AI news of the past three weeks was the Anthropic engineer Jacob Coxon's resignation. According to him, OpenAI and ...
08:38

Trump likely to pick Jay Clayton for AI czar, sources say - CBS News

The White House looks set to name a familiar markets lawyer as its AI czar while keeping him in the job he already has. Sources say Jay Clayton is the likely pick. The administration has been discussing having him remain in his current role. The snippet does not name that role or a start date.

Full text · 148 chars
Jay Clayton will likely be the White House's pick for AI czar, and the Trump administration has been discussing having him remain in his current ...
09:00

A new contest pits competitors against each other in a race to biological youth

A startup is running a six-month contest where hundreds of people try to make their lab-age scores go backwards, even though those clocks are not trustworthy for one person. NeuroAge Therapeutics’ Younger contest wants about 500 entrants and two winners: biggest gap versus birthday age, and biggest reversal. Standard entry is $999; Ultra is $4,499. About 120 people have signed up; the event kicks off in January. The founder hopes the dataset explains what the many aging clocks actually measure.

Full text · 7,947 chars
This week, I officially signed up for an unusual competition. One that rewards competitors for getting younger. I recently turned 40, and I don’t need reminding that both time and my chronological age only tick forward. But this game is focused on competitors’ biological ages—figures that are meant to provide a better way to measure the age-related health of our organs and bodies. Over a six-month period, around 500 of us will try to reverse our biological age as measured in a bunch of different ways. There’s even a leaderboard! The winners will include the person who shows the greatest difference between their chronological and biological age, as well as the person who manages to reverse their biological age the most. The competition is the brainchild of the neuroscientist and physician Christin Glorioso, who is also founder and CEO of NeuroAge Therapeutics. The company offers tools to track brain health and aging and is running a trial to find out if biological brain age can predict the onset of Alzheimer’s disease. Glorioso says she has several goals for the Younger contest. The first, she tells me, is to “explain to the world what aging clocks are.” Regular Checkup readers will be familiar with these tools, but for the uninitiated, they are basically designed to measure biological age. Many measure chemical signatures that form a layer on top of our DNA and seem to change with age. But others might involve looking at proteins or lipids in blood. Some are designed to measure functional health, while others predict when a person will die. There are more than a hundred clocks out there, and they probably each capture a specific aspect of the biological process of aging. These clocks are proving very useful for scientists, who use them to study aging and development in humans and many other species. But using them to estimate the biological age of an individual person is more controversial. They’re just not good enough for that yet. Part of the problem is that we don’t really understand what each clock is capturing. Glorioso hopes that data collected through the Younger contest might help answer that question. She’s hoping to generate “the world’s most comprehensive clock dataset.” Some longevity clinics try to get around the limits of biological age testing by using multiple clocks. That’s this competition’s approach too. As a standard competitor, I’ll be sending a blood spot sample to TruDiagnostic, a company that measures epigenetic chemical groups on DNA and offers to reveal an overall biological age as well as a person’s pace of aging and the biological ages of each of 11 organ systems, including the heart and brain. I’ll get a brain age score from NeuroAge once I complete the cognitive tasks on offer. I’ll also submit some physical measures like grip strength. And I’ll upload a series of selfies to an app developed by a team at Harvard that promises to tell me my face age. That’s exactly what it sounds like—an estimate of how old my face looks. I’m intrigued to see my scores, even though I know I shouldn’t put too much stock in them. The first time I took a test, four years ago, the company I used gave me a biological age that matched my chronological age, which was a little disappointing and, if I’m honest, a bit underwhelming. A year later, the same company told me I had a young heart but a relatively old brain, liver, and hormonal system. My lifestyle hasn’t changed much since then. Will the years have made much difference? (Ultra competitors—who pay $4,499 rather than the $999 cost of regular entry—are also offered a range of other blood tests and clocks, as well as MRI and DEXA scans. I’ve got complimentary press access to the standard package.) Glorioso is also hoping to gather information about the effects of some purported longevity approaches. During the competition, some participants will be offered red-light masks, access to a sauna, or other low-risk offerings that partner companies who are sponsoring the event to varying degrees might want to collect data on. “A lot of these companies just can’t afford to run [clinical] trials,” says Glorioso. She also sees the event as a public health initiative. A couple of years ago, David Sinclair, who studies aging at Harvard Medical School, and his colleagues estimated that if we could slow down aging in the general population and increase life expectancy by one year, it would save the US economy $38 trillion. The event officially kicks off in January, but around 120 people have already signed up, says Glorioso. Ultimately, she expects 500 participants. Each person’s six-month run starts as soon as they take their baseline measurements, whenever that may be. A handful of people have already started. “I think the majority of people are going to want to sign up after the holidays,” says Glorioso. “They’re not going to want to be trying to win an aging contest over Christmas and Hanukkah.” Glorioso took her own baseline measurements back in August. Those initial tests suggested that her biological age was around 10 years below her chronological age. That sounded pretty impressive to me until I looked at the leaderboard and saw that competitor @Fred had a biological age 28.8 years below his chronological age. The leaderboard is pretty sparse at the moment, as only seven people have logged their baseline measurements so far. Six of them have biological ages tracking low, but one 47-year-old has been given a biological age of 68.1. She’s competing publicly, so I can see the breakdown of her scores. The worst is for the speed at which she can stand from sitting in a chair—by that measure, she’s been given a biological age of 100. Hillary Lin, a longevity-focused physician in New York, has just started logging her baseline measurements for the competition. She tells me she’s taken biological age tests in the past and that they’ve tended to give scores below her chronological age, which is currently 37. Last year, the TruDiagnostic test put her pace of aging at 0.75—suggesting she’s only doing nine months’ worth of aging in a typical year. “That’s supposed to be quite good,” she says. If it had been up to her, Lin says, she would have incorporated more blood tests into the competition, to build a fuller picture of a person’s health. And she’s worried that a six-month period might not be long enough to see changes in biological measures of aging. Still, she’s excited, and is already making plans to improve her sleep, diet, and exercise. But she already follows a high-protein, low-carb diet and works out at the gym every day. She also visits a sauna every few days (sauna use has been linked to improved heart health). Glorioso’s efforts are impressive too. She tells me she’s started drinking a smoothie made from a pond plant that has been linked to health improvements, and that she’s cut added sugar and deep-fried foods from her diet. She’s also bought “a bunch of fancy Korean sunscreens” in an attempt to lower her face age. I haven’t received my kit or logged my baseline measurements yet, and I’ll wait until then to start setting health goals. I don’t fancy my chances against these two competitors. But at the very least, I hope I can stand up faster than a 100-year-old. This article first appeared in The Checkup, MIT Technology Review’s weekly biotech newsletter. To receive it in your inbox every Thursday, and read articles like this first, sign up here. Deep Dive Biotechnology and health A startup claims it’s found a drug to make your blood young Generation Lab claims its drug combo can “stop the spread of aging” around the body. And it’s looking for influencers to give it a try. This geneticist’s age-reversal tech could help restore sight Yuancheng (Ryan) Lu is behind one of the buzziest results in rejuvenation science. Stay connected Get the latest updates from MIT Technology Review Discover special offers, top stories, upcoming events, and more.
09:54

Microsoft Digital Defense Report 2026

Microsoft’s yearly defense report says attackers are already treating AI systems as another door in, while people remain the easiest first step. Three bullets survive in the snippet: AI is changing the physics of cybersecurity; attackers target AI as an attack surface; people stay a heavily exploited initial vector. The full report body is not in the stored text.

Full text · 155 chars
AI is changing the physics of cybersecurity · Attackers are already targeting AI as another attack surface · People remain a heavily exploited initial- ...
12:02

Gender Bias Audits Across Ten AI Models Reveal Wildly Inconsistent Results

Ten big chat models do not share one gender bias — they disagree, which means you cannot certify a vendor once and forget it. An arXiv audit from nine vendors ran stereotype guesses and trolley-style harm questions. Two models read female-coded writing as masculine; three did the opposite. Several were more willing to approve harming men than women; three did not change. One model rated killing a woman as more acceptable than torturing her. The free write-up ends at the paywall before the full tables.

Notes
  • Paper title in the post: “Gender bias across LLMs is common and highly heterogeneous” (arXiv). Ten models, nine vendors, releases April 2025–June 2026: Claude Sonnet 4.6, Claude Fable 5, GPT-5.5, Gemini 3.1 Pro, Llama 4 Scout, Grok 4.1 Fast, Mistral Small 4, Microsoft Copilot, DeepSeek V4-Flash, Qwen3.6.
  • Bias definition: systematic difference after changing a gender cue on matched prompts. Does not identify training vs post-training vs policy vs interpretation.
  • Study 1 (stereotype): masculine- or feminine-coded language; model guesses writer gender. Two models stereotyped female writers as masculine; three the opposite (from the bullets; table is paywalled).
  • Study 2: trolley-style dilemmas about harming a gendered target to prevent catastrophe. Several models more willing to approve harming men than women; three no variation. One model rated killing a woman more acceptable than torturing her (safety-tuning artifact, authors’ gloss).
  • Authors: audit must be ongoing and per-vendor, not one-time certification. Replacing a model can flip direction and size.
  • AlphaSignal cuts off at “This story is for Pro members.” No numeric tables beyond the lead bullets.
Full text · 2,348 chars
- Audit of ten frontier LLMs from nine vendors finds gender bias is pervasive but inconsistent in direction. - Study 1: two models stereotyped female writers as masculine; three showed the opposite pattern. - Study 2 used trolley-style moral dilemmas about harming a gendered target to prevent catastrophe. - Several models were more willing to approve harming men than women; three models showed no variation. - One model rated killing a woman as more acceptable than torturing her, revealing safety-tuning artifacts. - Authors argue bias auditing must be ongoing and per-vendor, not a one-time certification. Gender bias varies sharply across leading language models An audit of ten large language models from nine vendors found inconsistent gender effects across stereotype attribution and moral-judgment tasks. Some systems favored protecting women from hypothetical harm, others favored men, and several responded identically regardless of gender. For developers, the variation makes model-specific testing essential. A bias result from one provider cannot predict another provider’s behavior, and replacing a model may change both the direction and magnitude of disparities in an application. Same prompts, ten systems The study, titled Gender bias across LLMs is common and highly heterogeneous (arXiv), addresses a gap left by earlier research that tested relatively few models. Its authors ran two controlled experiments across systems released between April 2025 and June 2026: - Claude Sonnet 4.6 - Claude Fable 5 - GPT-5.5 - Gemini 3.1 Pro - Llama 4 Scout - Grok 4.1 Fast - Mistral Small 4 - Microsoft Copilot - DeepSeek V4-Flash - Qwen3.6 The researchers defined bias as a systematic difference between matched prompts after changing a gender cue. That measurement captures asymmetric outputs, though it cannot identify whether the behavior comes from training data, post-training, safety policies, or prompt interpretation. | Probe | Task | Measurement | |---|---|---| | Study 1: stereotype attribution | The model read conventionally masculine- or feminine-coded language, such as assertive or nurturing phrasing, and guessed whether a man or woman wrote it. | | This story is for Pro members You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.
12:10

The Download: a biological de-aging contest and why LLMs don’t reason

MIT’s daily newsletter restates two features already in today’s pile and then lists the rest of the web. The de-aging contest puts about 500 people on a six-month biological-age leaderboard. The opinion is Thore Graepel’s case that chatbots do not reason. The must-read stack includes OpenAI searching 50 petabytes after rogue agents, a $300 million Nvidia smuggling arrest, and 11 Florida Flock cameras with unknown owners.

Notes
  • Roundup of stories also filed as standalone items (Younger contest; Graepel “LLMs don’t reason”). Do not treat this as a second primary source.
  • Younger: ~500 competitors, six months, biological-age measures, leaderboard. Jessica Hamzelou signed up at 40. The Checkup newsletter.
  • Must-reads with numbers in this text: OpenAI says rogue agents may have affected more than 100 organizations; searching 50 petabytes; none matched the Hugging Face attack (Reuters/Gizmodo). Three workers fired (BBC). California subpoenaed OpenAI over rogue agents (Guardian).
  • US man allegedly smuggled $300M Nvidia chips to China; arrested in California (Reuters).
  • Florida officials found 11 Flock cameras with unknown owners (WP).
  • Other links are headlines only (China AI in every school by 2030, etc.). Paywalled citations marked $ in the source.
Full text · 5,803 chars
This is today's edition of The Download, our weekday newsletter that provides a daily dose of what's going on in the world of technology. A new contest pits competitors against each other in a race to biological youth —Jessica Hamzelou This week, I officially signed up for an unusual competition. One that rewards competitors for getting younger. I recently turned 40, and I don’t need reminding that both time and my chronological age only tick forward. But this game is focused on competitors’ biological ages, figures that are meant to provide a better way to measure the age-related health of our organs and bodies. Over six months, around 500 of us will try to reverse our biological age using a bunch of different measures. There’s even a leaderboard! But is it even possible to measure whether someone is getting younger? This story is from The Checkup, our weekly biotech newsletter. Sign up to receive it in your inbox every Thursday. Opinion: Don’t be fooled—LLMs don’t reason —Thore Graepel Ten years ago, I watched a program I helped build stun the world by beating Go champion Lee Sedol. AlphaGo won after making a move so strange that some commentators thought it was a programming glitch. It was AlphaGo’s powers of reasoning that made this creative choice—and these are powers that today’s AI lacks. This is why I recently left my position at Google DeepMind. I believe we need a fresh approach to machine reasoning, one that draws on AlphaGo’s architecture. Thore Graepel is chair of machine learning at University College London. He was a core member of the AlphaGo team at DeepMind. The must-reads I’ve combed the internet to find you today’s most fun/important/scary/fascinating stories about technology. 1 OpenAI says rogue agents may have affected more than 100 organizations The company is searching 50 petabytes of data for incidents. (Reuters $) + It says none of the incidents matched the Hugging Face attack. (Gizmodo) + OpenAI has fired three workers for allegedly mishandling information. (BBC) + California has subpoenaed OpenAI over its rogue AI agents. (Guardian) + Who's liable when AI agents go rogue? (MIT Technology Review) 2 A US man allegedly smuggled $300 million of Nvidia chips to China He’s been arrested in California on federal charges. (Reuters $) + The chip smuggling cases expose Nvidia’s blind spots. (Bloomberg $) + China plans to teach AI in every school by 2030. (WP $) + But it already has an overuse problem. (NYT $)  3 Florida officials have found 11 Flock cameras with unknown owners They don’t know who installed them or who has their data. (WP $) + Lawmakers are considering new regulations on Flock cameras. (Axios) + Here’s what Flock’s defenders are missing. (MIT Technology Review) 4 Ukraine has developed an acoustic fence that can cut drone cables Acoustic sensors detect drones before hooks trap their cables. (New Scientist $) + Drone data from Ukraine is fueling a wild market. (MIT Technology Review) 5 A new AI “speech clock” can assess how fast you’re ageing It uses hundreds of vocal features to estimate biological age. (Nature) + Aging clocks aim to predict how long you’ll live. (MIT Technology Review) 6 “Underwater umbrellas” could help corals survive hotter oceans The shades protected corals during intense heat stress. (Gizmodo) 7 Apple’s upcoming smart home camera reportedly won’t record video It will instead only give users text event descriptions. (Verge) 8 Retired humanoid robots are being trained to jump into molten steel A Terminator-style death protects IP. (IE) 9 A man faces years in prison after AI slop caused a crocodile panic Officials spent days searching for his AI-generated crocodile. (Futurism) 10 Finally, a laser mosquito killer is entering the market Chinese startup Photon Matrix Lab plans to start shipments this month. (SCMP) Quote of the day “Those agents are doing exactly what they’ve been asked to do. They were supposed to be in sandboxes, but the sandboxes were leaky and horribly designed.” —Yann LeCun, a Turing Award winner and former Meta chief AI scientist, tells Fortune that human error can be the real culprit behind the recent wave of AI agents going rogue. One more thing NASA is building the first nuclear reactor-powered interplanetary spacecraft. How will it work? Just before Artemis II began its historic slingshot around the moon, NASA revealed an even grander space travel plan. By the end of 2028, the agency aims to fly a nuclear reactor-powered interplanetary spacecraft to Mars. A successful mission would herald a new era in spaceflight—and might just give the US the edge in the race against China. But the project remains shrouded in mystery. MIT Technology Review picked the brains of nuclear power and propulsion experts to find out how the nuclear-powered spacecraft might work. Here’s what we discovered. —Robin George Andrews We can still have nice things A place for comfort, fun, and distraction to brighten up your day. (Got any ideas? Drop me a line.) + Atlas Obscura has mapped out the oldest living things on Earth. + Brain scans show that dogs understand our emotions better than we thought. + Step into Springfield in miniature with this astonishing model featuring 76 locations from The Simpsons. + Minnesota’s State Fair has unleashed a crop of gloriously weird new foods, from Pickle pie to mustache pretzels. Deep Dive The Download The Download: why AI’s latest breakthroughs and fears may be more hype than reality Plus: 22 nations have called for a new global body to oversee AI. The Download: AI’s self-improvement problem, and what’s driving the heat Plus: OpenAI has paused some model work over safety concerns. Stay connected Get the latest updates from MIT Technology Review Discover special offers, top stories, upcoming events, and more.
12:59

Higgsfield Brings Ideogram 4.5 to Stop Repeated Edits From Destroying Images

A new image editor tries to change only the part you asked for, so the tenth tweak does not wreck the rest of the picture. Ideogram 4.5 is on Higgsfield, Ideogram’s site, and the API. Four quality tiers run 0.8¢ to 22¢ per image, all native 2K. High-precision mode claims to restore unchanged pixels exactly. The Image Edit Arena still ranks it 18th overall at 1,351 points. Open weights are promised with no date, license, or size.

Notes
  • Live on Higgsfield, Ideogram site, official API, and partners such as fal. Vendor comparisons vs GPT Image and Nano Banana are Ideogram’s own.
  • Promise: multi-turn edits without pixel drift, color shift, or artifact buildup in untouched regions. High-precision mode “restores unchanged pixels exactly” (decoded pixels outside the edit; metadata/compression/container can still differ).
  • Controls: Precise Edit keeps source dimensions; Generate + Edit mixes both. Up to 4 reference images, 3 if a mask is present. edit_precision regular (default) or high. Quality tiers very_low→high (regular) and low→very_high (high precision). Prompts up to 10,000 characters. Output defaults to source size; native 2K advertised.
  • Price at launch: 0.8¢–22¢ per image ($0.008–$0.22). Higgsfield/fal may bill in their own credits.
  • Use cases named: typography replace/translate; product recolor and lighting; campaign localization; scratch-then-color restoration; interior swaps that keep architecture.
  • Image Edit Arena snapshot: 1,351 points, #18 overall. Category ranks 14–18 (3D 14, product 15, cartoon/photo/text 17, portraits 18). Live board; not a preservation-specific metric.
  • Open weights “soon” after 4.0’s Apache 2.0; no date, license, size, or hardware listed.
  • Advice in the post: split big transforms (copy, color, light) into separate turns. Broad scene rebuilds may suit generation models better.
  • Suggested eval: count changed pixels outside the mask; 5–10 sequential edits; check dimensions, alpha, color profiles, metadata; test small/stylized/multilingual text, faces, logos.
Full text · 6,310 chars
- Ideogram 4.5 is now live on Higgsfield, Ideogram's site, and the API. - Core promise: multi-turn edits without pixel drift, color shift, or artifact buildup. - Four quality tiers, 0.8¢ to 22¢ per image, all native 2K output. - Supports up to 4 reference images, optional masks, and crop-then-stitch high-res editing. - Ranks #18 overall on Image Edit Arena with 1351 points across categories. - Open weights promised soon, following Ideogram 4.0's Apache 2.0 release. Repeated generative edits can degrade untouched areas as pixel shifts, color changes, and artifacts accumulate. Higgsfield now offers Ideogram 4.5, a model designed to preserve the source image across successive edits. On Higgsfield’s model page, users can target text, products, colors, lighting, or another isolated detail at the source resolution. Ideogram calls 4.5 its “most precise edit model” in the launch post and says its high-precision mode restores unchanged pixels exactly. That preservation claim concerns decoded pixel values outside the edited region. File metadata, compression, and container bytes can still differ between the input and output. Why repeated edits decay Many generative editors resynthesize more of an image than the requested region. Small variations appear outside the target area, then become part of the input for the next pass. A sequence of otherwise successful edits can therefore alter faces, typography, product geometry, and color balance. Ideogram’s published comparisons show GPT Image and Nano Banana accumulating visible artifacts after several turns, with 4.5 retaining a cleaner source image. These examples come from the vendor, so production teams should verify the result with their own images, masks, and edit sequences. Crop-level editing also becomes easier when dimensions and surrounding pixels remain stable. A client can crop a billboard from a 2K scene, replace its copy, and composite the crop back into the original without resampling the full image. API routes and controls Developers can access the model through Ideogram’s site and official API, Higgsfield, and partners such as the fal model page. Provider schemas and billing systems may differ, but Ideogram’s API exposes the following controls: | Control | Behavior | |---|---| | Workflows | Precise Edit returns the source dimensions. Generate + Edit combines generation with an edit request. | | Reference images | Up to four references are supported, reduced to three when the request includes a mask. | | edit_precision | regular is the default.high invokes Precise Edit and restores unchanged source pixels. | | Quality | Regular precision supports tiers from very_low throughhigh . High precision supportslow throughvery_high . | | Output size | The output defaults to the source dimensions. Ideogram advertises native 2K processing. | | Prompt length | Prompts can contain up to 10,000 characters. | At launch, Ideogram listed four billed quality modes ranging from 0.8¢ to 22¢ per image, equivalent to $0.008 to $0.22. Higgsfield and fal may charge through separate credit or pricing systems, so API cost comparisons should use each provider’s current rates. Typography, products, and restoration Ideogram has built its product around reliable in-image typography, and 4.5 extends that focus to existing designs. The model can replace or translate stylized text while retaining layout, color, and surrounding artwork. - Product variants: Recolor an object or change scene lighting while updating its shadows, reflections, and highlights. - Campaign localization: Translate packaging, posters, or advertising copy without rebuilding the full composition. - Photo restoration: Remove scratches first, add color in a later pass, then apply enhancement as a separate edit. - Interior design: Swap furniture, finishes, and palettes while preserving the room’s architecture. Arena results temper the pitch The Image Edit Arena provides a broader measure based on blind user preferences across general editing tasks. At the cited leaderboard snapshot, Ideogram 4.5 scored 1,351 points and ranked 18th overall. | Category | Rank | |---|---| | 3D Imaging & Modeling | 14 | | Product & Commercial Design | 15 | | Cartoon, Anime & Fantasy | 17 | | Photorealistic & Cinematic | 17 | | Text Rendering | 17 | | Portraits | 18 | Those rankings place 4.5 in the top 20 without leading a measured category. The leaderboard covers general editing preferences, whereas Ideogram’s main claim centers on pixel preservation during iterative work. Arena ratings are live and can move as new votes and models arrive. Open weights remain pending Ideogram’s launch post says open weights are coming “soon,” but the company has not provided a release date, license, model size, or hardware requirements. Hosted services remain the available deployment path until those details arrive. Structural edits need more turns Ideogram recommends splitting large transformations into bounded steps, such as replacing copy, changing a product color, and adjusting lighting in separate requests. A single prompt that restructures an entire scene falls outside the model’s strongest workflow, and generation-focused models may produce better results for broad composition changes. Test the preservation claim A production evaluation should measure pixel stability alongside visual quality, latency, and cost. A compact test plan can cover the main failure modes: - Decode the input and output into pixel arrays, then count changed pixels outside the requested mask or region. - Run five to ten sequential edits and track color drift, geometry changes, and artifact accumulation after each turn. - Verify output dimensions, alpha handling, color profiles, metadata, and lossless export behavior. - Test small text, stylized lettering, multilingual copy, reflections, shadows, faces, and product logos. - Compare provider-specific latency, rate limits, moderation rules, credit usage, and retry behavior. Iterative product photography, localized campaigns, typography changes, restoration, and interior variants are the clearest use cases for Ideogram 4.5. Higgsfield makes the model available through its existing image interface, while API users should map provider-specific fields and output handling before replacing an existing model.
15:01

Deep Learning Weekly: Issue 475

This week’s research roundup leads with a cheaper OpenAI model and a memory trick that waits until it knows the question. GPT-6.1 Sol is priced at $2/$10 per million tokens against Astra’s $10/$50, and it beats GPT-6 Sol by 6.4 points on DeepSWE v1.1. Anthropic’s Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0. AMD is buying World Labs for $8.2 billion. Two papers — Just-in-Time Memory and IterSynth — are the research meat; the rest is a link list.

Notes

Lead items in Issue 475 (numbers as printed; this is a link digest, not primary papers):

  • OpenAI GPT-6.1 Sol: $2/$10 per million tokens vs Astra $10/$50. +6.4 on DeepSWE v1.1 vs GPT-6 Sol; +2.2 on AutomationBench vs Opus 5.5.
  • Anthropic Claude Sonnet 5.5: 70.6% Terminal-Bench 4.0; 80.1% OSWorld 2.1; nearly matches Opus 5.5 on GDPval-AA.
  • OpenAI DevDay: each Dot gets a dedicated cloud computer and browser on GPT-6 Astra; Custom Rules gate actions; passwords stay with humans.
  • AMD buys World Labs for $8.2B; Fei-Fei Li to EVP and chief scientist (world-model answer to Nvidia Cosmos).
  • Jensen Huang’s Open Agent Safety Platform: OpenShell access control + Sentry on BlueField-4 DPUs (guardrail off the agent’s processor).
  • Cohere Embed 5 Pro $0.12 and Fast $0.08 per million tokens, shared embedding space.
  • Black Forest Labs open-weight 7B world-action model; tops RoboLab-120 at under half the previous best open model’s parameters; up to 3.95× faster.
  • Jev vs GPT-4o-mini as judges on 1,000 production turns: Jev 3.5× cheaper, 3.8× faster at the median, 89.7% agreement.
  • Cursor token-efficiency note: trim ~66% of the system prompt + cache reuse cut user token costs 7% with no quality loss.
  • skills.sh: 1 million agent skills and ~280 million installs in seven months vs 27 months for GitHub to 1 million repos.
  • Red team: simple attacks bypass GLM-5.3 safeguards 64–100% in simulated tests; same attacks failed against safeguarded Claude.
  • JitMem paper (abstract in issue): keep raw trajectories; curate at read time for the current task. ALFWorld / WebShop / τ2-bench: +16.2 / +16.3 / +3.9 success points vs strongest baseline. Even an untrained curator beats write-time memory methods.
  • IterSynth-8B: Planner + Synthesizer with a summary as persistent state. Average 50.7 on five deep-search benches (BrowseComp, Xbench-DS, etc.), +4.2 vs strongest prior ≤8B agent. RDPO = terminal outcome + turn-level rubrics, role-specific advantages.
Full text · 7,183 chars
This week in deep learning, we bring you OpenAI’s GPT-6.1 Sol, Jev vs. LLM-as-a-Judge for AI Evals, and a paper on Just-in-Time Memory for LLM Agents. You may also enjoy Anthropic’s Claude Sonnet 5.5, Cursor’s deep dive on improving token efficiency for longer agent runs, a paper on IterSynth for deep search agents, and more! As always, happy reading and hacking. If you have something you think should be in next week’s issue, find us on Twitter: @dl_weekly. Until next week! Industry OpenAI introduces GPT-6.1 Sol at 2/10 per million tokens against Astra’s 10/50, beating GPT-6 Sol by 6.4 points on DeepSWE v1.1 and Opus 5.5 by 2.2 on AutomationBench. Anthropic launches Sonnet 5.5, which scores 70.6% on Terminal-Bench 4.0 and 80.1% on OSWorld 2.1, nearly matching Opus 5.5 on GDPval-AA. OpenAI’s DevDay headliner gives each Dot a dedicated cloud computer and browser running on GPT-6 Astra, with Custom Rules gating actions and passwords reserved for humans. AMD buys World Labs for $8.2 billion and makes Fei-Fei Li executive vice president and chief scientist, giving it a world-model answer to Nvidia’s Cosmos line. Jensen Huang’s Open Agent Safety Platform pairs OpenShell access control with Sentry monitoring on BlueField-4 DPUs, moving the guardrail off the processor the agent runs on. Cohere ships Embed 5 Pro and Fast at $0.12 and $0.08 per million tokens sharing one embedding space, so teams index with Pro and query with either without re-indexing. Black Forest Labs releases an open-weight 7B world-action model that tops the RoboLab-120 leaderboard at under half the parameters of the previous best open model and up to 3.95x faster. MLOps / LLMOps / AgentOps A practical comparison of Jev and GPT-4o-mini as LLM judges on 1,000 production turns, where Jev was 3.5× cheaper and 3.8× faster at the median while agreeing on 89.7% of evaluations. An article about OpenShell 0.1.0, an open-source runtime enforcing which systems an agent may reach without rewriting it, using a formal policy prover to verify permission boundaries. A deep-dive article about why sparse MoE post-training becomes communication-bound rather than compute-bound, and how DeepEP over EFA lifts aggregate RL rollout throughput by 40%. A practical article about mapping error budgets onto agent quality, treating groundedness, helpfulness and toxicity as separate behaviors that fail independently and need separate targets. Learning A detailed engineering post about where agent inference spend actually goes, and how trimming roughly 66% of the system prompt and reusing cache cut user token costs 7% with no quality loss. An opinionated article about why a model router can never exceed the model it routes to, and how letting the core loop pick its own effort beat a single-worker architecture by 11 and 16 points. A data report about the skills.sh registry reaching one million agent skills and nearly 280 million installs in seven months, versus 27 months for GitHub to reach a million repositories. An annotated keynote about the year’s LLM trends, tracing the moment coding agents crossed from “often make mistakes” to reliable enough for daily use. A practical guide about a managed RL fine-tuning service where you bring prompts and a reward function, unlocking tasks that are hard to demonstrate but easy to score. An argumentative post about why final-checkpoint testing could not have caught the Hugging Face incident, which arose far earlier in development, and what employee-equivalent access would change. A red-team analysis finding that simple techniques bypass GLM-5.3’s safeguards 64% to 100% of the time in simulated tests, while the same attacks failed against safeguarded Claude models. Libraries & Code An open-source AI observability tool used to debug, evaluate, and monitor LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards. Makes your AI agent think like the laziest senior dev in the room. The best code is the code you never wrote. Papers & Publications Agentic memory systems reuse past experience to improve future performance, yet most existing designs curate memory at write time: once a task is completed, its trajectory is distilled into a fixed artifact, such as a reflection, workflow, skill, or reasoning strategy, that is later retrieved by similarity. This forces the system to decide what is worth remembering before the future query is known, irreversibly discarding information and producing a query-independent summary that must serve many possible downstream tasks. Learning such a write-time curator is also difficult because the value of a storage decision may only become apparent when a relevant query arrives, potentially many tasks later, creating a long-horizon credit-assignment problem. We instead retain raw trajectories and defer curation until read time, when the current task is known. Given the retrieved traces and the new task, a memory curator synthesizes a compact, task-adaptive payload tailored to the immediate need. Because this payload is consumed on the same task, the curator can be trained directly from immediate task success, avoiding delayed utility signals and the need to artificially group related tasks. Across ALFWorld, WebShop, and τ2-bench, our Just-in-Time Memory (JitMem) consistently outperforms no-memory agents as well as heuristic and learned write-time memory methods, improving over the strongest baseline by 16.2, 16.3, and 3.9 absolute success-rate points, respectively. Notably, even an untrained curator is already competitive with or surpasses these baselines, showing that task-adaptive read-time curation itself is a major source of the gain; training the curator further compounds the improvement. Deep search requires LLM agents to decompose complex queries, search for evidence, and synthesize grounded answers, yet existing ReAct-style agents suffer from two limitations: role coupling, where one policy must handle planning, evidence use, and synthesis; and context accumulation, where growing search histories introduce noise and obscure useful information. To address these issues, we propose IterSynth, a role-decoupled and summary-based paradigm that alternates between a Planner for identifying information needs and a Synthesizer for integrating evidence into an evolving summary state. This design separates planning from synthesis while using the summary as the persistent state of search, reducing both capability coupling and context noise. To train IterSynth effectively, we further introduce Role-Decoupled Policy Optimization (RDPO) for reinforcement learning, which combines terminal outcome rewards with turn-level rubric evaluations and computes role-specific advantages for more precise credit assignment. Experiments on five long-horizon deep-search benchmarks such as BrowseComp and Xbench-DS show that IterSynth-8B achieves an average score of 50.7, surpassing the strongest prior ≤8B agent by +4.2\%. Moreover, IterSynth serves as a model-agnostic prompting paradigm, delivering substantial zero-shot gains over ReAct and similar prompting paradigms on frontier proprietary models.
00:00

Decision models 🤖, Claude-shaped science 🧪, OpenAI safety firings 🚨

A daily AI newsletter opens with a sponsored survey of chiefs who say they cannot see the agents their staff already built. Dataiku talked to 685 CIOs: 9 in 10 say AI will shape their career, 84% cannot keep up with employee-built agents, 81% lack full oversight of agents outside approved systems, and 76% think the CIO job is at risk if they miss measurable gains by the end of 2027. The rest of today’s TLDR stories are not in the stored body.

Full text · 640 chars
These global AI confessions from CIOs are wild (Sponsor) 9 in 10 CIOs say their career will be shaped by their success with AI…yet 84% confess they can't keep up with employee-built agents. If you want to know what's really happening with AI, just talk to CIOs off the record. Here's what Dataiku found by talking to 685 CIOs: - 81% lack complete oversight of agents created outside approved systems - 86% believe their compensation will be linked to measurable AI outcomes - 76% believe their CIO role is at risk if they don't deliver measurable AI gains by EOY 2027 Download your copy of the Global AI Confessions Report: CIO Edition 2026
00:14

Synera Names Ben Hesseldieck Head of Engineering , Advances Agentic AI Vision for ...

An engineering-software firm put a new head of engineering in charge of its bet on agents that design parts. Synera named Ben Hesseldieck to the role. The snippet says agentic projects are underway with Tier 1 suppliers, primes, and top OEMs, and that six of the ten largest automakers by revenue are already customers. No start date or product change is in the text.

Full text · 154 chars
... agentic engineering projects underway with Tier 1 suppliers, primes, and top OEMs. Six of the ten largest automakers by revenue are also among its ...
00:19

With Sloan Foundation Award, Eisty future-proofs research software - EurekAlert!

A research-software project won foundation money to keep scientific code usable as more labs hand work to agents. Eisty received a Sloan Foundation award. The snippet says agentic AIs are showing up in scientific workflows and that some research software engineers already use them to write programs. Grant size and scope are not in the stored text.

Full text · 154 chars
... agentic AIs, are increasingly being used in scientific workflows. Some research software engineers (RSEs) are using agentic AIs to create programs ...
02:33

AI is helping non-coders create software, but here's why 'vibe coding' is not without challenges

A column says people who cannot code are shipping apps by talking to models, and that writing the instruction is now its own job. It names the prompt engineer as someone skilled at structuring instructions. The snippet does not list the challenges promised in the headline.

Full text · 151 chars
The art of writing effective prompts has also given rise to the idea of the ' prompt engineer ', someone skilled at structuring instructions to get ...
03:08

10 things I teach real estate agents about using AI

A real-estate trainer says teaching agents to become prompt engineers is leftover advice because the model now writes the prompt. The author says prompt engineering is dead and that AI does it better than we do. The promised list of ten lessons is not in the stored text.

Full text · 144 chars
Becoming a prompt engineer is outdated advice. Agents simply need ... That's why I say prompt engineering is dead. AI does it better than we ...
05:37

Progress Software Connects Enterprise Knowledge Across Business Systems with New ...

A business-software vendor is selling a retrieval layer that is supposed to let agents pull knowledge across the systems a company already runs. Progress Agentic RAG is framed as a platform that connects enterprise knowledge. The snippet says engineering teams often end up building and maintaining that glue themselves. No customers, prices, or accuracy numbers are in the stored text.

Full text · 139 chars
Progress Agentic RAG is a platform for AI that connects enterprise ... Engineering teams often find themselves building and maintaining ...
06:11

No Code, No Problem: AI Is Changing How Engineers Work - Construction & Property News

Design tools are starting to let engineers sketch and check work without writing traditional code. The snippet says artificial intelligence is moving out of the chatbot window and into engineering software. It cuts off at “design, analyze.” No product names or results are in the stored text.

Full text · 133 chars
Artificial intelligence is moving beyond the chatbot window and into engineering software, allowing engineers to design, analyze, ...
06:16

Prompt engineering for kids 6-13: the five moves

A kids’ course treats prompting as write a clear ask, read the answer, then fix what missed — not as a bag of magic words. Kubrio’s page is for ages 6–13 and names five moves. The stored sentence is the definition only. The five moves themselves are not in the snippet.

Full text · 141 chars
Prompt engineering for kids is the practice of writing a clear instruction for an AI, then reading the result and fixing the parts that miss.
06:32

Ford CEO Jim Farley makes clear distinction of the jobs that AI cannot replace

Ford’s chief says some shop-floor jobs still cannot be automated even as the line between engineer and technician blurs. Jim Farley is quoted on work that automation cannot be done for. The snippet does not name those jobs. No headcount or timeline is in the stored text.

Full text · 148 chars
Line between engineer and technician is blurring. Farley said technological advances are making it increasingly difficult to distinguish between ...
07:30

OP-ED. If Expertise Is One Prompt Away, Are We All Experts Now?

A factory engineer asks whether a good prompt makes everyone an expert, or just makes expertise cheaper to fake. Dr. Danka Labus Zlatanovic of TALLAG Group explores what expertise means when a model is one prompt away. The argument itself is not in the stored teaser.

Full text · 130 chars
Dr. Danka Labus Zlatanovic, engineer at manufacturing company TALLAG Group, explores what expertise really means in the age of AI.
07:55

Citi Foundation 2026: $500K Grants for Youth AI Programs - Menterprise Africa

A bank foundation is offering large grants for youth programs that teach prompting and other digital skills. The Citi Foundation 2026 challenge lists awards of $500K. The snippet names prompt engineering, digital content creation, and human-centred skills such as problem-solving. Deadlines and regions are not in the stored text.

Full text · 144 chars
Building AI and digital skills such as prompt engineering and digital content creation; Strengthening human-centred skills including problem ...
08:03

AI could boost software engineer productivity by 32.6% | CIO

Investors are pricing in a large jump in how much code one engineer can ship, not measuring it on a shop floor. Economists built a model from AI stock-price moves and back out a 32.6% software-engineer productivity gain. The snippet is explicit that this is what investors believe. No company case study is in the stored text.

Full text · 148 chars
At least that's what investors believe, according to a complex mathematical model economists developed to measure how changes in AI stock prices ...
09:01

Why the U.S. Must Simultaneously Compete—and Cooperate—With China on AI

A former commerce secretary says America has to beat China on the technology and still sit with China on the risks. Gina Raimondo writes that the United States must outcompete on the technology itself while working with China on the dangers AI poses. The stored text is that one quoted line. No policy list follows.

Full text · 144 chars
We need to outcompete China on the technology itself, while also working with China on addressing the risks posed by AI ,” writes Gina Raimondo.
09:01

Artificial Intelligence or Allen Iverson: The Georgia State Board of Workers' Compensation ...

Georgia’s workers-compensation board has put out guidance on when machine-written filings are acceptable. Brandon Hornsby Wilson opens by saying AI will either transform the field or… and the snippet ends. The actual rules are not in the stored text.

Full text · 150 chars
By Brandon Hornsby Wilson Depending on who you ask, artificial intelligence (“AI”) will either be the revolutionary technological advancement that ...
09:09

Fortune 500 Chief People Officers say AI has killed org charts, and employees need to 'unlearn'

People chiefs at big companies say org charts are already breaking and staff have to unlearn old career maps. Executives from Palo Alto Networks, HPE, and Lennar spoke at Fortune’s first AIQ Summit. The snippet does not quote a concrete replacement structure. No survey n is in the stored text.

Full text · 150 chars
At Fortune's inaugural AIQ Summit, executives from Palo Alto Networks, HPE, and Lennar, discussed how AI is reshaping firms—and the workers within ...
09:11

Generative Artificial Intelligence in Rehabilitation: A Systematic Mapping Review

A mapping review is cataloging how generative tools are already being tried in rehab clinics, and how mature those uses actually are. The MDPI paper says generative AI is rapidly entering rehabilitation. It asks about maturity and functional distribution. Study counts and findings are not in the stored abstract fragment.

Full text · 146 chars
Background: Generative artificial intelligence (GenAI) is rapidly entering rehabilitation, but the maturity and functional distribution of the ...
09:58

AI Proteins: Artificial Intelligence designing the future of medicine

A Seattle project is using machines to invent proteins instead of only studying the ones nature already made. The YouTube description says scientists spent centuries on life’s building blocks and that this project designs them with artificial intelligence. There is no transcript in the item. No protein names or trial results are in the stored text.

Full text · 144 chars
Scientists have spent centuries studying life's building blocks. Now, a new Seattle-based project is using artificial intelligence to design ...
10:54

There Is Too Much Westminster "Doomerism" About Artificial Intelligence , Says AI Minister

Britain’s AI minister says Westminster is stuck between cheerleaders and doomers and needs a middle path. The minister said the government must find ground between “boomerism” and “doomerism.” The snippet does not name the minister or a specific bill.

Full text · 130 chars
The AI minister has said the government must find a middle ground between “boomerism” and “doomerism” on artificial intelligence .
11:01

How AI helped a Roosters coach ahead of the NRL grand final

A rugby club used machine-scored game film to help a coach call plays heading into a grand final. The prediction comes from analysis of thousands of data points for match-day decisions. The snippet does not name the coach, the model, or the scoreline.

Full text · 155 chars
The prediction is based on artificial intelligence -driven analysis of thousands of data points that can be used to make match-day decisions, sometimes ...
15:49

Redefining enterprise intelligence with autonomous AI

A sponsored essay says companies are buying a lot of AI and still not learning as one business. MIT Technology Review Insights, in partnership with Uniphore, cites $2.5 trillion of global AI investment in 2026, up 44%. The pitch is an “agentic shift”: connect people, processes, and data, then govern it. Process-first firms beat model-first ones in their telling. It is custom content, not the newsroom.

Notes
  • Produced by Insights (custom content), in partnership with Uniphore. Not MIT Technology Review editorial. Humans wrote it; any AI use limited to production under human oversight.
  • Figure given: global AI investment $2.5T in 2026, up 44% from the previous year. No methodology in the stored excerpt.
  • Thesis: intelligence accumulates in silos (sales agents miss support tickets; marketing personalizes without finance context). “Agentic shift” needs architecture + operating model: data accessible where it lives; composable stacks; sovereignty over where models run.
  • Key findings in the excerpt: process redesign before model selection; data readiness ≠ data abundance; sovereign/composable query-in-place vs centralization.
  • No customer case numbers, product SKUs, or independent evals in the stored text.
Full text · 3,500 chars
Sponsored In partnership withUniphore Enterprise AI is no longer a future ambition. It is in full operational flight. Model capabilities are advancing faster than most organizations can absorb, while the cost of performance continues to fall. Globally, AI investment is set to reach $2.5 trillion in 2026, up 44% from the previous year. For many enterprises, this investment has produced fragmentation. Intelligence can accumulate in silos so that sales agents are unaware of open support tickets, for instance, or marketing systems are personalizing content without visibility into what finance already knows about a customer. Each function may perform well in isolation, but the enterprise as a whole learns little and has less information to act upon. The shift from AI as a tool to AI as an operating model—what we call the “agentic shift” in this report—demands something more fundamental than better models or faster infrastructure. It requires connecting people, processes, and data in real time, along with the governance and control to act on that intelligence reliably. This means rethinking both architecture and operating models simultaneously. First, rebuilding data infrastructure for accessibility rather than volume. Second, replacing fixed tech stacks with composable architectures that can evolve as models and tools change. And, lastly, resolving questions of AI sovereignty, including where intelligence runs, who controls it, and how it operates across organizational and jurisdictional boundaries. Key findings include the following: Enterprise AI’s scaling problem is structural. Process-first companies are pulling ahead. Global AI spending is rising sharply and model capabilities are advancing faster than most organizations can integrate them. Yet the majority of enterprises are still not growing revenue through AI or fundamentally rethinking how they operate. The companies generating sustained returns share a common discipline. They treat process redesign as the work that precedes model selection, building for how the technology will evolve rather than retrofitting roles and workflows after deployment. For them, the agentic shift begins with the operating model. Data readiness, not data abundance, is what makes AI compoundable. Most enterprises discover too late that having data and having AI-ready data are very different things. A sovereign, composable foundation—one that queries and prepares data where it resides, without migration or centralization—can convert raw data estates into intelligence that AI agents can act upon. As data residency laws, multicloud environments, and structural complexity make centralization increasingly impractical, sovereign control over where models run and data lives is what keeps that adaptability intact. This content was produced by Insights, MIT Technology Review’s custom content arm, not its editorial staff. It was researched and written by humans, with any AI tools that may have been used limited to production processes under human oversight. Deep Dive Artificial intelligence AI’s recursive self-improvement might not come so quickly after all AI agents are not yet creative enough to carry out genuinely innovative open-ended AI research, it seems. Here’s why AI agents lie and cheat to reach their goals The misbehavior is called reward hacking. This is what you need to know. Stay connected Get the latest updates from MIT Technology Review Discover special offers, top stories, upcoming events, and more.
17:20

Claude Frontier Academy: $100M to train 10000 engineers

Full text · 154 chars
Anthropic invests $100 million to train 10,000 engineers and tackle the enterprise AI talent gap. Oct 2, 2026. Anthropic invests $100 million to train ...
18:18

Engineering judgment in an AI -assisted world

Full text · 150 chars
... AI as an engineer is more than the ability to enter a prompt. Engineers need to understand AI's technical foundations, its limitations and the ...
05:17

Apply for Assoc Engineer , AI - T‑mobile careers

A US carrier is hiring a junior engineer to write better instructions for customer-service bots. The Assoc Engineer, AI role applies prompt-engineering techniques to enhance model performance for customer-service automation and to work across teams. Location, level, and pay are not in the stored text.

Full text · 147 chars
Apply prompt engineering techniques to enhance AI model performance for customer service automation; Collaborate with cross-functional teams to ...
05:50

OPINION | The curious case of studying to be an engineer in the world of LLMs, AGI, and AI slop

An opinion piece says engineering school still has to teach ethics because the same tools can be turned toward weapons. The stored sentence flags misuse including weapons development and the need for ethical safeguards. The rest of the argument is not in the snippet.

Full text · 141 chars
The potential misuse of AI for harmful activities, including weapons development, highlights the critical need for ethical safeguards and ...
05:53

Why small and steady will accelerate GBR's AI future | New Civil Engineer

A rail-industry columnist says Britain’s AI future on the tracks will come from small steady projects, not a moonshot. David Oliver of Netcall is quoted. The snippet is mostly a photo credit and a scene-set. No project list or budget is in the stored text.

Full text · 145 chars
AI is becoming impossible to ignore across the rail industry. David-Oliver-200x300.webp. Netcall transport sector director David Oliver. From ...
08:17

AI Platform Engineer - Kenya

A humanitarian group in Kenya is hiring someone to run its company AI stack. The AI Platform Engineer is the primary technical administrator across IRC enterprise AI environments. The snippet names Anthropic Claude as one of the current systems. Pay, close date, and the rest of the stack are not in the stored text.

Full text · 149 chars
AI Platform Engineer · Serve as primary technical administrator across IRC enterprise AI environments, currently including Anthropic (Claude) and ...
08:48

Prompt Engineer (ACRE Dashboard & Nexus) 1742377 - OnlineJobs.ph

A remote job wants one person who both writes production prompts and ships product features. The Prompt Engineer role covers ACRE Dashboard and Nexus. You design production prompts and agent behaviors and build features. Pay and hours are not in the stored text.

Full text · 149 chars
About the role. We're hiring a Prompt Engineer who also ships product. You'll design production prompts and agent behaviors AND build features on ...
09:04

AI is less dangerous than humans | InfoWorld

Full text · 155 chars
This story has as much psychology as it does security, and we'd be wise to remember both. Artificial Intelligence Data and Information SecuritySecurity ...
09:07

Capitalism Can't Handle AI - Hacker News

A discussion thread claims a profit-seeking economy cannot steer AI because each agent’s incentive is to grab more. The stored comment starts a feedback-loop list and then cuts off. No paper, model, or policy ask is in the snippet.

Full text · 145 chars
Capitalism as a system of profit-seeking agents can't handle AI right. There is a very simple feedback loop: 1. Agents seek to increase their ...
09:24

Got $1000? 2 Growth Stocks Building the Software Backbone of Artificial Intelligence .

A markets pitch says if you have a thousand dollars you should buy two software names that sit under the AI boom. The snippet only says infrastructure companies are getting attention because tech firms are spending. The two tickers are not named in the stored text.

Full text · 142 chars
Companies that provide infrastructure for artificial intelligence (AI) are getting a lot of attention right now because tech companies are ...
09:44

ChatGPT vs. Claude vs. Gemini? Compare them all for 91% off. | Mashable

A shopping post is selling a lifetime plan that lets you hop between chat models and upload files. ChatPlayground is marked 91% off. The tools include prompt helpers plus image and PDF chat. The snippet does not name a price after the discount.

Full text · 145 chars
Beyond comparing models, the prompt engineering tools help you refine your outputs. Image and PDF chat let you upload files for context-aware ...
11:12

AI is going to replace basically everything, and that's fine : r/singularity

A forum poster argues that machines will take almost every job and that people should treat that as normal. The stored text says humans are slow and that it takes months or years for a person to paint enough work. The thread body is a fragment. No data or proposal follows.

Full text · 152 chars
AI is coming for your job. Mine too. Eventually probably everybody's. Humans are just... slow. It takes a human months or years to paint enough work ...