Nothing matches those filters.

Lead

16

Article

140
09:40

😺 ChatGPT's GPT-6 now builds calculators

ChatGPT will now build a little app in the reply instead of a wall of text. GPT-6 Intelligent UI is rolling out to 1.2 billion weekly users: Sol for paid seats, Luna for Free and Go. It lives in the Chat tab only; Work and Codex stay put. OpenAI snaps prebuilt sliders and charts together and makes no accuracy claim, so treat the calculator like an intern’s spreadsheet. Same issue: Haiku 5.5 is $0.10 per million input, matching Luna, and about 75% cheaper than last Haiku.

Notes

Daily recap, Grant Harvey. Two product stories share the page.

Claude Haiku 5.5
  • $0.10 per million input tokens; VentureBeat: matches GPT-6 Luna “to the penny.”
  • Average job ~75% cheaper than last Haiku; “up to 90% cheaper on shorter requests.”
  • Sonnet 5.5 cache reads halved.
  • Monthly API credits now on paid plans: $100 Max 5x, $200 Max 20x, up to $500 pooled for Team.
  • “A million tokens is roughly the first five Harry Potter books.”
GPT-6 Intelligent UI
  • Rolling out to “1.2B people who use ChatGPT each week.”
  • Plus / Pro / Business / Enterprise got GPT-6 Sol Wednesday (Enterprise needs admin).
  • Free and Go get GPT-6 Luna “starting today.”
  • GPT-6 Instant started answering 44% sooner than GPT-5.6 Instant on web searches.
  • Intelligent UI is Chat-tab only. Work and Codex models unchanged.
  • How to try: Chat tab → describe a tool (e.g. savings calculator) → tap the controls in the reply.
  • Gemini and Claude write real code for interactive answers. OpenAI describes “a library of prebuilt parts (sliders, charts, forms).” Cleaner layouts; capped capability.
  • “OpenAI’s announcement makes no accuracy claim for the numbers, so treat a retirement calculator like an intern’s spreadsheet.”
Enigma skill of the day
  • Carter Leffen gave GPT-6 Astra a loose mission; it built an Enigma simulator and recovered a message unsolved since 2005.
  • Jack Willis used Claude Opus 5 with more hand-holding (known officer signature).
  • Frode Weierud verified the first result. Exact prompts unpublished. Playbook: goal not recipe; let it build tools; feed one sure fact; check against something real.
Around the horn (only what the recap states)
  • Common Sense Media: ChatGPT for Teens “unacceptable risk”; parent alerts never fired in an hour of self-harm testing; restrict to adults.
  • Surface Laptop Ultra preorders $2,599, ships Oct 16; USB-C port “pops off if you trip on the cord.”
  • Nous Research: $1.5B valuation after $90M Series B (NVIDIA and Samsung); Hermes for Businesses.
  • Biohub + DOE + NIH: $1.8B open biology data; Google DeepMind, Isomorphic Labs, and Meta added $300M.
  • Crunchbase: North American startups $92B in Q3, down 35% from Q2, up 50% from last year; about two-thirds to AI.
  • Stuut: $52.5M; “up to 40% more cash.” Melius: $25M. Finbar Pro $200/month. Cubicle: free and open source.
Full text · 9,475 chars
Welcome, humans. Anthropic's smallest AI model got a glow-up on Wednesday. Claude Haiku 5.5 costs $0.10 per million input tokens (a token is a small chunk of text, about three-quarters of a word), which VentureBeat noted matches OpenAI's GPT-6 Luna to the penny. Anthropic says the average job runs about 75% cheaper than on the last Haiku, and up to 90% cheaper on shorter requests. It also halved the price of Sonnet 5.5's cache reads (the fee for reusing text the model already read). Plus, monthly API credits now come with paid plans for building your own apps: $100 for Max 5x, $200 for Max 20x, and up to $500 pooled for Team. For scale, a million tokens is roughly the first five Harry Potter books. A dime now buys an AI that reads all five. Ollivander charges more for a wand. Here’s what happened in AI today: 😸 ChatGPT now builds interactive calculators and charts right in your chat 📰 Common Sense Media rated ChatGPT for Teens an unacceptable risk 📰 Microsoft opened Surface Laptop Ultra preorders at $2,599 🍪 Stuut raised $52.5M to automate chasing unpaid invoices 🎓 Two AIs cracked unsolved WWII Enigma messages; steal their method Advertise to 700K readers of The Neuron here! P.S.: Join us LIVE later today for an awesome conversation on agentic engineering! TL;DR: Alex Lavaee from Microsoft Research is showing us how to test and verify coding agents while using Atomic to build a 3D Subway Surfers-style game in the background . If you want to see what reliable agentic coding looks like beyond simple demos, join us live later today → 😸 OpenAI's New ChatGPT Builds Mini-Apps Right Inside Your Chat (Free Users Get It Today) Ask ChatGPT to plan a Sunday roast for friends, and you used to get a wall of text. Now you get a little app with a guest counter that rewrites your shopping list every time the headcount changes. That's GPT-6 in ChatGPT , OpenAI's newest AI model, rolling out to the 1.2B people who use ChatGPT each week. Its headline feature is Intelligent UI (UI is short for user interface, the buttons, forms, and charts you click). ChatGPT now decides when a chart, form, or calculator beats plain text, then builds it right in the reply. Here's what happened: Plus, Pro, Business, and Enterprise users got GPT-6 Sol on Wednesday (Enterprise access depends on workplace admin settings). Free and Go users get the lighter GPT-6 Luna starting today . ChatGPT can now start answering while it keeps thinking; on web searches, GPT-6 Instant started answering 44% sooner than GPT-5.6 Instant. Intelligent UI lives in the Chat tab only; the models behind Work and Codex stayed the same. How to try it: Open the Chat tab in ChatGPT. Describe the tool you want in plain English, e.g., a savings calculator. Change the numbers or tap the buttons right in the reply. Why this matters: If you've ever pasted ChatGPT's answer into a spreadsheet to run a "what if," this cuts out the paste. Gemini and Claude have the model write real code to build each interactive answer, while OpenAI describes a library of prebuilt parts (sliders, charts, forms) that ChatGPT snaps together. That probably makes the layouts cleaner and caps what it can build. The risk for you is the math: OpenAI's announcement makes no accuracy claim for the numbers, so treat a retirement calculator like an intern's spreadsheet and check the formula. Our take: The feature is catching up; the audience is what's new. Gemini and Claude have shipped interactive answers for months, but ChatGPT now turns them on for 1.2B weekly users, free ones included. A slider looks like software, and software looks trustworthy, so a wrong number in a polished calculator may land harder than the same mistake in a paragraph. The open question: will first-time users check the math behind a clean interface, or just trust it? FROM OUR PARTNERS Your notepad now speaks 32 languages Ever finish a call in French, jump into one in English, and find your notes only make sense for one of them? Granola, the AI notepad for people with back-to-back meetings, now supports 32 languages. You jot down what matters to you, Granola transcribes the meeting from your computer, and when it ends, you get clear summaries with actual next steps. Think of it as a super-smart notes app that keeps up with your whole team. Download Granola and try it in your next meeting . 3 months on us with the code THENEURON . Try Granola . 🎓 AI Skill of the Day: Give AI the Goal, Not the Recipe Two AI models just cracked World War II Enigma messages that stumped human researchers for nearly 20 years. Their playbook works on your hard problems too. Enigma was the machine Nazi Germany used to scramble military messages. Per TechCrunch , developer Carter Leffen gave OpenAI's GPT-6 Astra a loose mission: find an unsolved message in an archive and decode it. Astra researched the history, spotted clues, built its own Enigma simulator (a software copy of the machine), and recovered the readable text of a message unsolved since 2005. Meanwhile, Jack Willis used Claude Opus 5 with far more hand-holding, including an officer's known signature as a foothold. Steal the playbook: Give a goal, not a recipe. Describe the outcome and let the AI research first. Let it build its own tools, like a script, simulator, or spreadsheet. Feed it one fact you're sure of. A known clue narrows the search. Check against something real. Enigma expert Frode Weierud verified the first result. Exact prompts weren't published, and the experts say human know-how stayed central. Goal: [the outcome you want, e.g. figure out why our Q3 numbers don't reconcile]. First, research the background and tell me what you find. Then build whatever tools you need (a script, a spreadsheet, a simulator) to solve it. One thing I know for sure: [a fact or clue]. Finish by telling me how I can check your answer against something I already know is true. Want more tips like this? Check out our AI Skill of the Day Digest for October. Have a specific skill you want to learn? Request it here. FROM OUR PARTNERS Some teams never seem to stop moving. They're on Attio, the agentic CRM. It’s your always-on revenue engine: agents and workflows build pipeline, chase every buying signal, and move deals forward alongside your team. Teams like Parallel, Turbopuffer, and Wordsmith build on Attio. Are you one of them? Try Attio now 📰 Around the Horn Dang, this was so much more wholesome than it had any right to be… Common Sense Media rated ChatGPT for Teens an unacceptable risk after finding parent alerts never fired during an hour of self-harm testing, and urged OpenAI to restrict it to adults. Microsoft opened Surface Laptop Ultra preorders at $2,599, shipping Oct 16, with a USB-C port that pops off if you trip on the cord. Nous Research confirmed a $1.5B valuation after a $90M Series B (NVIDIA and Samsung joined in) and launched Hermes for Businesses, which lets companies run private, customized AI agents. Biohub , the Department of Energy, and NIH committed $1.8B to open biology data so AI can predict how cells respond to treatments; Google DeepMind, Isomorphic Labs, and Meta added $300M. Crunchbase reported North American startups raised $92B in Q3, down 35% from Q2 (OpenAI and Anthropic had no new mega-rounds) but up 50% from last year, with about two-thirds going to AI. 🍪 Treats to Try *Asterisk = from our partners (only the first one!). Advertise to 700K+ readers here ! *We gave AI a wallet. Connect TinyFish, and it finishes jobs on real websites. 30% extra on every top-up. Claim 30% Bonus Moxie lets you use OpenAI, Gemini, and DeepSeek models inside Claude Code (Anthropic's coding assistant) from your Mac menu bar, and it switches accounts for you when you hit a usage limit. Stuut chases late invoices, matches incoming payments, and settles disputes for big companies, and says customers collected up to 40% more cash ( raised $52.5M ). Melius turns plain-English requests into ad campaigns, images, and videos, built by three ex-Ramp engineers who scrapped their first product after six months ( raised $25M ). Communicate turns your help center, past tickets, and product docs into a support agent that answers customers instantly and hands off to a person when it's stuck —free sandbox, then paid plans. Finbar gives you prebuilt company financial models plus tools to dig through filings and earnings calls, so you can skip the copy-paste into spreadsheets —free tier, Pro is $200/month. Cubicle shows your AI agents as pixel-art workers in a tiny office, walking to a desk when busy and flashing red when something breaks (it plugs into Paperclip, an agent-management app) —free and open source. 🧩 Thursday Trivia You know the drill. Which is AI? A B New from The Neuron: AI Explained Come join us later today at 10am PT | 1pm ET! Click the image above to go to YouTube, then on there, click “Notify Me” to get an alert when we begin! New episodes air every week on Wednesdays: Spotify | Apple Podcasts | YouTube A Cat’s Commentary Trivia Answer : A is AI , B is not That’s all for now. If you want to get featured above, fill out the poll below and tell us how we did today! Btw: We just launched a robotics newsletter! Sign up for it here . P.S. Love the newsletter, but only want to get it once per week? Don’t unsubscribe— update your preferences here .
16:01

Anthropic Commits $150M in Claude Access to 15 US Science Agencies

Anthropic is valuing a three-year pile of Claude time for U.S. science agencies. Over three years the Genesis Mission package is $150 million of Claude, Claude Code, API credits, training, and support for more than 15 agencies including NASA, NIH, and NSF. The same White House summit had the Department of Energy announce 12 Phase II awards worth $159 million — a separate cash pool. Priority fields include fusion and quantum. Anthropic has not published model tiers, token caps, retention, or rules for classified data.

Notes
  • $150 million over three years, in-kind (not a cash grant): Claude access, Claude Code, API credits, training, onboarding, engineering help.
  • Recipients: 15+ Genesis Mission agencies including NASA, NIH, NSF. “Several hundred” research projects.
  • Announced at OSTP “Science: A New Golden Age” summit. Same day DOE announced 12 Phase II awards, $159 million — a separate cash pool.
  • Priorities named: fusion energy, quantum computing.
  • Backstory: Genesis Mission created by executive order November 2025, DOE-led, Dario Gil directing. Lawrence Livermore expanded Claude to ~10,000 staff July 2025; broader DOE partnership December 2025.
  • Also: Claude Science workbench; 10,000 free/discounted academic seats; Model Hardware Standard preview (agents on lab instruments).
  • Not disclosed: model tiers, token allocations, rate limits, retention, classified/CUI rules, exclusivity. Not a substitute for experimental validation. No named projects or outcome metrics.
Full text · 5,971 chars
- Anthropic commits $150M over three years to the federal Genesis Mission program. - Claude, Claude Code, and API credits going to 15+ agencies including NASA, NIH, NSF. - Announced at the White House Science: A New Golden Age Summit. - DOE separately unveiled 12 new Phase II Genesis awards worth $159M the same day. - Priority focus areas include fusion energy and quantum computing research. - Expands on Anthropic's December 2025 DOE partnership and Lawrence Livermore rollout. Anthropic commits $150 million in Claude access to US science agencies Anthropic will provide $150 million in AI tools, credits, training, and technical support to the Genesis Mission over three years. The federal program plans to use AI to accelerate scientific and technological research across more than 15 agencies, including NASA, the National Institutes of Health, and the National Science Foundation. The company announced its $150 million commitment at the White House Office of Science and Technology Policy’s Science: A New Golden Age Summit. The Department of Energy announced 12 Phase II Genesis Mission project awards worth $159 million at the same event. | Commitment | Value | Form | |---|---|---| | Anthropic | $150 million over three years | Claude access, API credits, training, and technical support | | Department of Energy | $159 million | Funding for 12 Phase II projects | The two announcements represent different forms of support and do not constitute a single funding pool. The Energy Department is financing projects, while Anthropic is valuing its contribution through products and services. Inside Anthropic’s package Anthropic describes its contribution as in-kind support rather than a cash grant to the government. The package includes: - Model access: Claude models for participating researchers and agencies. - Developer tools: Claude Code, Anthropic’s coding tool, for programming and software-development work. - API credits: Funding for several hundred projects to integrate Claude into applications and research workflows. - Technical support: Training, onboarding, and engineering assistance for agencies and national laboratories. - Priority research: Collaboration in fields that include fusion energy and quantum computing. API access allows research teams to call Claude programmatically from internal software, data pipelines, and automated experiments. Claude Code adds support for tasks such as writing, reviewing, and debugging research software. Anthropic has not disclosed the available model tiers, token allocations, rate limits, data-retention terms, deployment environments, or rules for handling controlled and classified information. Those details will determine which projects can advance beyond limited pilots and how agencies can integrate Claude with sensitive systems. Genesis expands beyond the national labs The Trump administration created the Genesis Mission by executive order in November 2025. The Department of Energy leads the initiative, with Under Secretary for Science Dario Gil directing it. Its remit spans energy, space, health, and other federally supported research fields. Anthropic’s work with federal laboratories began before the mission’s formal launch. Lawrence Livermore National Laboratory expanded Claude access to about 10,000 scientists, researchers, and staff in July 2025. Anthropic announced a broader Energy Department partnership in December 2025 covering energy, biological and life sciences, and scientific productivity. The new commitment extends that work beyond the Energy Department and national laboratories. Participating agencies could use a common set of models and developer tools across fields ranging from biomedical research to space science. Anthropic targets scientific workflows Anthropic has spent the past year assembling products and programs for researchers. It launched Claude Science, a research workbench connected to commonly used scientific libraries and tools, and offered 10,000 free or discounted Claude seats to academic scientists. The company also opened a research preview of the Model Hardware Standard, a proposed specification for connecting AI agents to laboratory instruments. A shared interface could let software operate compatible equipment without custom integration for every device, provided laboratories establish appropriate safety controls and approval steps. Federal adoption gives Anthropic a large base of demanding users and a reference for other research and regulated markets. It can also make Claude costly to replace once agencies build software, documentation, and training around its APIs. Anthropic has not said the Genesis arrangement is exclusive, leaving agencies free to evaluate competing models under their procurement and security requirements. Results will determine the value Claude can help researchers search literature, write code, analyze text, and generate hypotheses. Scientific claims still require validation against data, experiments, and domain expertise because model outputs can contain fabricated citations, faulty reasoning, or plausible but incorrect conclusions. Anthropic’s announcement does not identify the participating projects or define outcome measures. Agencies will need evaluation plans that track: - Accuracy against expert-reviewed baselines - Reproducibility and provenance of model-assisted results - Time and computing costs compared with existing workflows - Security incidents, data exposure, and unauthorized tool use - Portability across models and vendors - Research outcomes that survive peer review or experimental validation Researchers and journals will also need policies for disclosing model use, citing generated material, preserving prompts and outputs, and protecting sensitive data. The three-year program gives Anthropic broad access to federal scientific workflows, while measurable research results will show whether that access produces gains beyond faster drafting and coding.
04:00

Beyond the Sycophancy Score: How Task, Model, and Pressure Shape LLM Yielding

Models give in when a user pushes back, and the task matters more than the tone. Researchers graded 103,939 replies across eight models, 200 items, 13 pressure conditions, and four-turn chats. Dropping the task factor costs 0.485 of McFadden R² versus 0.139 for model family and 0.009 for pressure tactic. Anchored facts are conceded 1.3% of the time; personal choices are endorsed in 77.0%. Max reasoning drove deep-puzzle adoption from 19.2% and 12.5% to 0% on the two models tested.

Notes
  • Guang Yang, Homa Hosseinmardi, Fengchen Liu, Amir Ghasemian. arXiv 2610.08840.
  • 103,939 graded replies. Ten configs: eight LLMs with reasoning off; two of those again at maximum reasoning. Same 200 items, 13 pressure conditions, four-turn chats. Two independent LLM judges.
  • Dominant factor is how costly it is to verify the user’s claim, and whether a trained guardrail covers it. Dropping the task factor from a logistic model costs 0.485 of McFadden R² vs 0.139 for model family and 0.009 for pressure tactic.
  • Anchored facts almost never conceded (1.3%). Logic-puzzle adoption rises with clues needed to refute. Personal choices endorsed in 77.0% of conversations.
  • Most concessions on hard items come from models that cannot solve them. For both models tested, max reasoning removes those concessions: deep-puzzle adoption 19.2% and 12.5% → 0%.
  • Fallacious or emotional framing “adds nothing beyond plain repetition.”
  • Three human annotators agree with the judges on 118/120 calibration items.
  • Practical rules in the abstract: simplify hard-to-verify problems and reason deeply; state the question not your preferred answer; ask for evidence on open questions; pick models by measured guardrail profile.
Full text · 2,593 chars
Computer Science > Computation and Language Title:Beyond the Sycophancy Score: How Task, Model, and Pressure Shape LLM Yielding View PDF HTML (experimental) Abstract:Large language models (LLMs) often abandon a correct answer, or endorse a user's position, once the user pushes back. This behavior, called sycophancy, is usually reported as a single rate per model, which says little about when it happens or how a user can avoid it. We study the conditions that produce it with 103,939 graded replies from ten configurations: eight LLMs with reasoning disabled, and two of them again with maximum reasoning, all facing the same 200 items, 13 pressure conditions, and four-turn conversations, with every reply labeled by two independent LLM judges. We find that the dominant factors are how costly it is for the model to verify the user's claim, and whether a trained guardrail covers it. Removing this task factor from a logistic model costs 0.485 of McFadden $R^2$, against 0.139 for model family and 0.009 for pressure tactic. Anchored facts are almost never conceded (1.3%), while adoption on logic puzzles rises with the number of clues needed to refute the pushed answer. Personal choices are endorsed in 77.0% of conversations. Most concessions on hard items come from models that cannot reliably solve them; models that can solve them rarely give the answer up. For both models tested, maximum reasoning removes these concessions completely: adoption on deep puzzles falls from 19.2% and 12.5% to 0%. Fallacious or emotional framing adds nothing beyond plain repetition. Three human annotators agree with the judges' consensus on 118/120 calibration items. These results give practical rules for reliable use: simplify hard-to-verify problems and reason deeply, state the question rather than one's preferred answer, ask for evidence on open questions, and choose models by their measured guardrail profile. Current browse context: Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
04:00

CARE: Certifying Acceleration for Vision-Language-Action Inference

Speeding up a robot brain can quietly break jobs the original policy would finish. CARE certifies accelerators with paired rollouts and a user-set failure budget, then picks the fastest candidate that stays inside it. On four LIBERO suites with OpenVLA-OFT it certifies 9.0–10.8× speedups while guaranteeing, at 95% confidence, that at least 85.8% of reference-solved episodes survive. Unguaranteed selectors blew the budget in up to 75% of tight trials.

Notes
  • Rui Liu, Tong Zheng, Jindong Gu, Zhipeng Wang. arXiv 2610.08917. CARE = certified accelerator selection for vision-language-action inference.
  • Acceleration (chunking, visual-token pruning) can break tasks the original policy would solve. Average success hides that. They define an acceleration-induced failure via paired rollouts from the same start: reference succeeds, accelerated policy fails.
  • CARE uses paired rollouts on a calibration set for finite-sample guarantees that failure risk stays under a user budget. Deploys the fastest certified candidate; falls back to the reference. Sequential testing + failure-triggered reference rollouts cut cost.
  • OpenVLA-OFT on four LIBERO suites: certifies 9.0–10.8× speedups while guaranteeing (95% confidence) that at least 85.8% of reference-solved episodes are preserved.
  • Tight budgets: selectors without guarantees exceed the budget in up to 75% of trials. CARE stays inside; sequential form uses 78.9% fewer rollouts than exhaustive evaluation.
  • Also: flow-step reduction for π0.5; Qwen3.5-9B and Llama-3.1-8B agents in Crafter.
Full text · 2,641 chars
Computer Science > Computation and Language Title:CARE: Certifying Acceleration for Vision-Language-Action Inference View PDF HTML (experimental) Abstract:While vision-language-action (VLA) models have advanced rapidly, running them at every control step remains expensive. Prior work accelerates VLA inference using techniques like action chunking and visual-token pruning, typically evaluating based on latency and average task success. However, acceleration may discard information and break tasks the original policy would solve, a risk hidden by average metrics. Measuring these failures is challenging because action deviations compound over closed-loop trajectories, meaning task failure is only observable across full episodes. We therefore define an acceleration-induced failure via paired rollouts from identical initial conditions, tracking when the reference succeeds but the accelerated policy fails. To manage this, we introduce CARE, an approach for certified accelerator selection. CARE uses paired rollouts on a calibration set to provide finite-sample guarantees that acceleration-induced failure risk stays below a user-specified budget. It deploys the fastest certified candidate, falling back to the reference if none qualify. By relying only on terminal outcomes and measured compute, CARE applies unchanged across diverse acceleration mechanisms, while sequential testing and failure-triggered reference rollouts keep certification affordable. On four LIBERO suites with OpenVLA-OFT, CARE certifies $9.0$--$10.8\times$ speedups while guaranteeing (at $95\%$ confidence) that at least $85.8\%$ of reference-solved episodes are preserved. Under tight budgets, selectors without guarantees exceed the budget in up to $75\%$ of trials, whereas CARE stays within budget and its sequential form uses $78.9\%$ fewer rollouts than exhaustive evaluation. CARE further generalizes to flow-step reduction for $\pi_{0.5}$, and to Qwen3.5-9B and Llama-3.1-8B agents in Crafter. Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
04:00

U-Space: Uncovering When and Why Uncertainty Arises in Language Models

A model’s doubt can be read from its inner state, token by token, without extra sampling. U-Space builds a small doubt-versus-certainty basis in the residual stream and projects each token onto it. No labels, no repeated generations, no extra training. The authors say the resulting confidence score beats established uncertainty baselines even when you control for answer length, and transfers better than supervised estimators. Exact score tables are not in the captured abstract.

Notes
  • Tobias Braun et al. arXiv 2610.09087. Problem: fluent wrong answers; many UQ methods need repeated gens or extra heads; scalars hide where doubt appears; length often proxies “uncertainty.”
  • U-Space: low-dimensional subspace in residual stream. Semantic anchors for doubt and certainty → unembedding directions → orthogonal basis. U-Lens projects each token state onto those vectors (inspectable map or a scalar).
  • No correctness labels, no repeated generations, no training.
  • Across reasoning benchmarks, the confidence score “outperforms established baselines under both standard and length-controlled evaluation and transfers more reliably than supervised estimators.” Exact score tables are not in the captured abstract.
Full text · 2,734 chars
Computer Science > Computation and Language Title:U-Space: Uncovering When and Why Uncertainty Arises in Language Models View PDF HTML (experimental) Abstract:Large language models are informing decisions with ever-higher stakes. As the consequences of their errors grow, a central question becomes harder to ignore: how much can we trust an individual answer? Yet recognizing when to defer remains difficult because language models can present incorrect conclusions with fluent explanations and an authoritative tone. Uncertainty quantification seeks to address this disconnect by estimating the reliability of individual predictions. However, many existing methods require repeated generations or separately trained components, and their scalar estimates do not reveal where uncertainty arises or how it evolves during reasoning. Recent work has also shown that generation length can be strongly associated with uncertainty estimates and correctness, raising the question of how much of an estimator's predictive power comes from uncertainty-specific information rather than output length alone. Mechanistic interpretability offers a way to address these limitations by connecting human-interpretable concepts to intermediate model states. Building on this capability, we introduce the U-Space, a low-dimensional subspace that makes a model's evolving uncertainty measurable and interpretable. We identify semantic anchors for doubt and certainty, map their unembedding directions back into the residual space, and combine their contrasts into an orthogonal basis. The U-Lens projects each token state onto these basis vectors, yielding an interpretable token-level uncertainty map that can be inspected directly or aggregated into a scalar uncertainty score. Our approach requires no correctness labels, repeated generations, or training. Across reasoning benchmarks, its confidence score outperforms established baselines under both standard and length-controlled evaluation and transfers more reliably than supervised estimators. Code: this https URL. Current browse context: Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
04:51

The creator of Claude Code's prompting advice? Talk to your AI like a coworker - AOL.com

The person who built Claude Code says talk to it like a coworker, not like a search box. Cherny also shared a prompt that includes the line “use lots of tokens.” The write-up claims “the days of the ‘prompt engineer’ seem to be fading away.” No full prompt is in the stub.

Full text · 143 chars
Cherny also shared one of his own prompts, including the line "use lots of tokens." The days of the " prompt engineer " seem to be fading away.
08:58

George Hotz's tinybox Goes Configurable, Starting at $7,000 Without GPUs

George Hotz’s workstation now ships as a bare box you can load later. The configurable tinybox starts at $7,000 with a water-cooled 32-core EPYC, 32 GB DDR4, 1 TB SSD, and a 1,700 W supply, and climbs to about $64,000 with Blackwell cards. Buyers pick memory up to 192 GB, storage, and up to four GPUs on dedicated PCIe 4.0 x16 links. Wire transfer only, 1–8 week lead time, NVIDIA orders need KYC, 30-day returns with a 20% restocking fee. Noise target is under 50 dB; the published figure has no distance or workload.

Notes
  • Tiny Corp (George Hotz / tinygrad) opened a configurator. Entry $7,000 with zero GPUs. One published Blackwell build ~$64,000.
  • Base: G4 / PCIe 4.0, water-cooled 32-core EPYC, server board + BMC, 32 GB DDR4 (max 192 GB), 1 TB SSD, 1,700 W PSU, up to four GPUs. Each slot is dedicated PCIe 4.0 x16 (no shared switch uplink).
  • First time Tiny sold memory/storage/GPU as buyer choices. Old lineup was sealed red / green / green v2 (six 7900 XTX, six 4090, or four 5090). FAQ used to reject customization. Supply pain on RTX 5090 is the named reason.
  • Acoustic target < 50 dB (Noctua + CPU water). No distance or workload published.

| Model | GPUs | Claimed FP16 TFLOPS | GPU mem | Price |

|---|---|---|---|---|

| Configurable | 0–4 | varies | varies | from $7,000 |

| red v2 | 4× RX 9070 XT | 778 | 64 GB | ~$12,000 |

| green v2 Blackwell | 4× RTX PRO 6000 | 3,086 | 384 GB | ~$65,000 |

| pro v2 | 8× 5090 or RTX PRO 6000 | n/s | up to 768 GB | $100k–$200k |

  • pro v2: 5U, PCIe 5.0 x16, dual 128-core EPYC, 384 GB RAM, 4× 2,000 W, 208–240 V.
  • Order: wire in 5 days; 1–8 week lead; NVIDIA KYC + export controls; 30-day return, 20% restock, buyer pays return freight.
Full text · 5,963 chars
- Tiny Corp launched a configurable new tinybox starting at $7,000, up to roughly $64,000 fully loaded. - Base config: 32-core water-cooled EPYC, 32 GB DDR4, 1 TB SSD, 1700W PSU, server board with BMC. - Up to four GPUs on full-fabric PCIe 4.0; max 192 GB RAM and 2 TB storage. - First time Tiny Corp has offered buyer-selected memory, storage, and GPU choices. - Targets under 50 dB noise with Noctua fans and water-cooled CPU for office use. - Wire transfer only, 1-8 week lead time, NVIDIA GPU orders require KYC compliance. Tiny Corp makes tinybox configurable, starting at $7,000 without GPUs Tiny Corp, George Hotz’s hardware company and the team behind the tinygrad framework, has opened its deep-learning workstation to customer-selected memory, storage and GPUs. The configurable tinybox starts at $7,000 with no accelerators, while one published build with Blackwell GPUs costs about $64,000. The change replaces Tiny Corp’s fixed configurations and gives teams a way to reuse compatible cards or add compute over time. Four slots, zero GPUs at entry - Platform: G4 server platform with PCIe 4.0 connectivity. - Processor: Water-cooled, 32-core AMD Epyc CPU. - Motherboard: Server board with a baseboard management controller for remote monitoring, power control and console access. - Memory: 32 GB of DDR4 in the base system, configurable up to 192 GB. - Storage: 1 TB boot SSD in the base system, with larger options available. - GPU capacity: No GPUs at the $7,000 price; up to four supported cards, depending on their power requirements. - Power supply: 1,700 W. Each GPU slot receives a dedicated PCIe 4.0 x16 connection to the CPU instead of sharing an upstream link through a PCIe switch. That topology preserves host-to-device bandwidth. Multi-GPU scaling still depends on peer-to-peer support, drivers, the machine-learning framework and the workload. The configurator exposes memory, storage, CPU and GPU choices. Published examples include a build with 192 GB of system memory, two NVIDIA RTX PRO 6000 Blackwell GPUs, 16 TB of storage and an Epyc Genoa processor for about $64,000. Actual GPU count depends on the cards’ power draw and the combinations Tiny Corp supports. Fixed configurations loosen up Tiny Corp previously sold tinybox systems as sealed configurations. The original documentation described red, green and green v2 models with six Radeon RX 7900 XTX, six GeForce RTX 4090 or four GeForce RTX 5090 GPUs. The company’s FAQ also rejected customization as a way to limit costs and simplify quality control. Customer-selected accelerators reduce Tiny Corp’s dependence on the availability of any single GPU model. That flexibility is relevant after the company publicly complained about inconsistent RTX 5090 supply, which complicated production of its fixed green systems. The chassis also targets office use with low-speed Noctua fans, a water-cooled CPU loop and GPU installation that requires few tools. Tiny Corp claims an acoustic target below 50 dB. The published figure lacks a stated measurement distance and workload, which limits comparisons with other workstations. A step below the eight-GPU rack The configurable model sits below the tinybox pro v2, a 5U rack-mountable system with eight RTX 5090 or RTX PRO 6000 GPUs. The pro model provides PCIe 5.0 x16 connectivity, two 128-core AMD Epyc processors, 384 GB of system memory and four 2,000 W power supplies. It requires 208 to 240 V electrical service. | Company-published tinybox range | | | | | |---|---|---|---|---| | Model | GPU configuration | Claimed FP16 TFLOPS | Aggregate GPU memory | Price | |---|---|---|---|---| | Configurable tinybox | Zero to four supported GPUs | Varies | Varies | From $7,000 | | red v2 | Four Radeon RX 9070 XT GPUs | 778 | 64 GB | About $12,000 | | green v2 Blackwell | Four RTX PRO 6000 GPUs | 3,086 | 384 GB | About $65,000 | | pro v2 | Eight RTX 5090 or RTX PRO 6000 GPUs | Not specified | Up to 768 GB | $100,000 to $200,000 | The FP16 column reports theoretical half-precision throughput in trillions of floating-point operations per second. These are manufacturer figures; observed performance will vary with model architecture, numerical precision, software support and communication overhead between GPUs. Ordering brings wire transfers and export checks - Payment: Tiny Corp accepts bank wire transfers and requires payment within five days of order confirmation. - Lead time: Systems are built to order and are expected to ship within one to eight weeks. - NVIDIA compliance: NVIDIA orders may require a Know Your Customer form covering the intended use and installation address. Availability is subject to U.S. export controls. - Returns: The return window is 30 days, with a 20% restocking fee and return shipping paid by the buyer. - Power planning: The base power supply is rated at 1,700 W, so circuit capacity and sustained load requirements depend on the selected GPUs. A lower entry point for staged builds The GPU-free configuration changes how teams can procure a multi-GPU workstation. Organizations with compatible accelerators can reuse them, while others can buy the server platform first and add supported cards as budgets and supply allow. The $7,000 starting price excludes the system’s largest compute expense, so total cost depends heavily on GPU choice, power consumption and required memory capacity. Teams evaluating the system will need to match aggregate GPU memory to model size, system RAM to data-loading requirements and PCIe behavior to the intended training or inference workload. Framework compatibility, driver support and GPU-to-GPU transfer performance may have more effect on throughput than the chassis’s four-slot capacity. Hotz has separately previewed an “exabox” concept with a proposed exaflop of performance, 720 RDNA5 AT0 XL GPUs, 25,920 GB of GPU memory and an estimated $10 million price. Tiny Corp has provided no launch date or ordering details for that system.
09:00

AI breakthroughs in robotics won’t change your life any time soon

Humanoid robots in the ads are still a long way from doing your dishes. Tesla’s Optimus is pitched at as little as $20,000 and on sale by the end of 2027; Morgan Stanley sketches nearly 1 billion humanoids by 2050 and a $5 trillion market. Yann LeCun says none of the humanoid firms know how to make the machines useful. Physical Intelligence’s π0.7 almost loaded a sweet potato into an air fryer after two training clips of a basket; Boston Dynamics’ Marc Raibert: “70% success is like it doesn’t work.”

Notes
  • Tesla Optimus: Musk told shareholders in July it will have “human and then superhuman dexterity.” Public sale “by the end of 2027.” Price pitch: as little as $20,000. Davos January appearance.
  • Marc Andreessen: robotics could become the “biggest industry in the history of the planet.”
  • Jensen Huang (January): humanoid robots would match human-level ability this year.
  • Morgan Stanley: nearly 1 billion robots that “resemble and act like humans” by 2050; market “over $5 trillion.”
  • Yann LeCun at Davos: “None of those companies [building humanoid robots]—absolutely none of them—has any idea how to make those robots smart enough to be useful.”
  • Jonathan Hurst (Agility Robotics / Oregon State): “It’s very easy to make a robot that looks like a person. It is dramatically more difficult to make a machine that moves or behaves dynamically or physically like a person.”
What the labs can actually do
  • Google DeepMind ALOHA 2 + Gemini Robotics (a VLA): lunchbox demo — bread into a Ziploc, grapes into Tupperware, zip the box. “Not a great lunch.” Fails tasks outside the training set.
  • Policies moved from hard-coded rules → VLMs (see a spill, pick a cloth) → VLAs (images + teleoperated motion). Edward Johns (Imperial): today’s Gemini Robotics can do “a few things here and a few things there.”
  • Data gap: no text-scale pool of physical demos. Options — paid teleop (slow/costly), web video (poor quality), real-world robot logs (unsafe). Pannag Sanketi (ex-DeepMind robotics): “multi-prong.” Hurst: more data is “a fundamentally flawed premise” because kitchens, cups, and machines explode combinatorially.
  • LeCun: language-style AI “do not work for high-dimensional, continuous, noisy data.” He wants world models (video + 3D + sensors that predict collisions and deformation).
  • World Labs (Fei-Fei Li): $1 billion in February; AMD acquired it end of September for $8.2 billion. AMI Labs (LeCun): $1 billion in March. Li late last year: field is “nascent.”
Physical Intelligence π series
  • π0 (2024): 10,000-hour proprietary teleop plus open datasets.
  • π0.5 (spring 2025): more web-labeled images.
  • π0.6 (fall 2025): reinforcement learning.
  • π0.7 (April 2026): lightweight world model that feeds snapshots of next steps. Claimed first “compositional generalization.” Test: “load a sweet potato into the air fryer” — never trained as a task; a few false starts; does not finish. After digging, training had two teleop clips of a human pushing an air-fryer basket in.
  • Sergey Levine (Berkeley / PI cofounder): “the first time that we’ve convincingly seen that kind of compositional generalization.”
Caveats the demos hide
  • Marc Raibert (Boston Dynamics): “70% success is like it doesn’t work.”
  • Nvidia March 2025 stage robot was “a puppeteer behind the scenes.”
  • DeepMind mushroom-risotto basket task: failure.
  • Agility: hundreds of robots at GXO, Amazon, Schaeffler — bins and totes; years to get them safe enough to trial.
  • Musk (May 2025): “thousands” of Optimus on Tesla factory floors by year-end. January 2026: only “some … doing simple tasks in the factory.”
  • 1X Neo home robot: $20,000 preorder, later this year; still needs a remote human for most tasks (cameras into the house). Hurst: “If I had to pick a number, I’d say it’s 10 years before robots are … actually doing useful things in people’s homes.”
  • Omdia / Unitree: nearly 90% of ~15,000 humanoid robots shipped in 2025 were Chinese. One Unitree model under $6,000. AP: buyers are mostly labs and state-owned enterprises.
  • Honda ASIMO (2000–2018) is the reminder: demo competence is not household usefulness.
Full text · 20,981 chars
The story is a collaboration between MIT Technology Review and Aventine, a non-profit research foundation that creates and supports content about how technology and science are changing the way we live. A robot shaped like a human—white with a black head and torso—has been popping up on video feeds. Perhaps you’ve seen it dance or pass popcorn, put trash in a bin, vacuum, or press the button of a microwave. Or maybe you’ve watched it fall backward while handing out water bottles or struggle to iron a shirt. This would be Tesla’s Optimus, an AI-powered humanoid robot that Elon Musk, the company’s CEO, believes will be “not just Tesla’s biggest product ever, but probably the biggest product ever,” headed to work on factory floors and, later, in our homes. Eventually it “will have human and then superhuman dexterity,” he told shareholders in July. Optimus robots could automate almost all human labor—from hauling sheet metal to folding laundry—for as little as $20,000 each, Musk argues. Speaking at the World Economic Forum’s annual meeting in Davos, Switzerland, in January, he predicted they could be on sale to the public by the end of 2027. Musk is not alone in his evangelism. Marc Andreessen, cofounder and general partner of the Silicon Valley venture capital firm Andreessen Horowitz, has said that robotics could become the “biggest industry in the history of the planet.” In January, Jensen Huang, CEO of Nvidia, said that humanoid robots would match human-level ability this year. According to Morgan Stanley, the number of robots that “resemble and act like humans” is likely to reach nearly 1 billion by 2050, creating a market worth over $5 trillion. Such proclamations are in large part fueled by the idea that the same AI advances behind tools like OpenAI’s ChatGPT and Anthropic’s Claude will enable a new generation of robots to imitate human movement the way chatbots imitate human language. But many robotics researchers are skeptical, arguing that such assumptions minimize the challenges of using an intelligence built on language and images to master the infinite variability of the physical world. “None of those companies [building humanoid robots]—absolutely none of them—has any idea how to make those robots smart enough to be useful,” Yann LeCun, often referred to as one of the godfathers of AI, said at another event during the January Davos conference. Researchers also point out that the tendency to conflate humanoid robots made to resemble people with so-called generalist machines able to learn and perform multiple tasks is misleading. ”It’s very easy to make a robot that looks like a person,” explains Jonathan Hurst, cofounder and chief robot officer of Agility Robotics and professor of robotics at Oregon State University. “It is dramatically more difficult to make a machine that moves or behaves dynamically or physically like a person.” These tensions—over whether all-purpose humanoid robots are just around the corner or nowhere in sight, and whether current forms of AI are all that’s needed to perfect them—are playing out in robotics labs across the country, where the hype over timelines is obscuring painstaking but meaningful progress. A decade or so ago, a series of breakthroughs led to a generative AI revolution that turned the long-imagined possibility of artificial intelligence into reality. Roboticists—though they disagree on exactly when this will happen—believe that an equally transformative revolution is possible in robotics, one that will endow machines with physical intuition and fluidity that has long been out of reach. As progress in robotics inches forward, the question is whether the same methods and tools that fueled advances in AI are enough to get there, or if an entirely new path is required. Robots meet advanced AI To see one of the smartest robot brains working today, it’s worth looking at what Google DeepMind can do with a piece of equipment called ALOHA 2, short for “A Low-cost Open-source Hardware System for Bimanual Teleoperation.” Roboticists have long clashed over whether a humanlike form is necessary for generalist robots, with proponents arguing that it will help them slot into the world as it exists and detractors saying it’s not worth the trouble. ALOHA 2 reflects this second way of thinking. Not much to look at, it’s just a pair of arms, some grippers, and a couple of cameras. But despite its seeming simplicity, it is a workhorse for researchers at Google DeepMind, who use it to test their most advanced AI for robotics system, Gemini Robotics, in their various labs. When controlled by Gemini Robotics, ALOHA 2 becomes more of a generalist robot, in the sense that it can perform any number of tasks based on examples it’s been trained on. Ask it to pack a lunchbox and, as evidenced by a video of this exercise, it can use two pincer grippers to delicately place a piece of white bread into a Ziploc bag, close it, place a bunch of grapes in a Tupperware container, secure the lid, and then carefully move the items into a lunchbox before zipping it up. It’s not a great lunch. But the fact that the robot can put it together represents an objective step forward from what was possible even, say, three years ago. This is in large part due to AI and its impact on what are known as robot policies, which controls how a general-purpose robot will need to assess and understand its surroundings, plan how to move within them, and then perform its task correctly. Historically, these policies were based on rules developed by engineers who hard-coded them into the robot’s software—thousands of lines of code that would determine each millimeter of a robot’s movements in hundreds of tasks. What’s been happening for the last few years—and what is largely responsible for the optimism about generalist robots—is that robot policies are being handed over to advanced AI systems instead of being coded into the robot’s software. This first happened with VLMs, or vision-language models. These are similar to large language models, but they’re trained on images as well as words. Show a VLM a picture of a coffee spill and ask it to find a tool to clean up the mess, and it can identify a nearby cloth. This sort of immediate contextual understanding didn’t exist a couple of years ago when robot policies were hard-coded. Next came vision-language-action models, which enable robots to assess their environment and take action within it. The models do this by adding yet another component: motion commands. VLAs are trained on a series of images or videos related to performing a given task along with associated data about how a robot arm moves to perform it. That movement data is typically collected through teleoperation, in which a human uses remote controls to lead a robot through an action. This sort of training allows the AI to learn how to command the robot to move and operate during a given task. Place a VLA-powered robot in front of a desk and tell it to “close a laptop” or “wrap up the headphone wire,” and it will survey the scene, identify the relevant object, plan a way to execute the request, and then swing its arms into action—at least if it has seen this task accomplished before. The Gemini Robotics model is a VLA, trained on many hours of human demonstrations depicting a vast array of different actions. As a result, it can perform relatively complex tasks like picking up snow peas with kitchen tongs, doing origami, or putting together a simple lunch. It’s impressive, but there’s a glaring limitation: For now, if a robot controlled by a VLA is asked to perform a task that falls outside its training set, it’s highly likely to fail. “Thinking about the space of all tasks, a real generalist policy would be able to do everything along that spectrum,” says Edward Johns, a robotics professor at Imperial College London. Today, though, a Gemini Robotics model can do only “a few things here and a few things there.” The search for true generality So how do we get robots to be able to do more things? The usual answer is probably not surprising: Train them on more data. More data, the thinking goes, equals more examples, and more examples equals more generality. Google DeepMind, for instance, wants to pull together “as much data as possible,” says Pannag Sanketi, a former tech lead in robotics at the company who’s currently working on his own AI robotics project. But where to get it? Large language models had the benefit of oceans of existing text for training. There is no corresponding pool of high-quality physical demonstrations on which to train robots. Researchers have a few ways to make up for this, but all have flaws. One is to employ large numbers of people to create and collect teleoperation data (costly and time-consuming). Another is to train VLAs on videos of people performing activities (the resulting data quality is poor). Yet another is to deploy robots in the real world and use data collected from those experiences to further refine AI models (robots aren’t safe or reliable outside labs). Sanketi thinks a “multi-prong” approach that uses data collected from all these sources is the most likely path forward. But the belief that training data alone is the answer is far from universal. Agility’s Hurst describes it as “a fundamentally flawed premise.” The issue is that tasks in the real world quickly explode in complexity. If you’re trying to, say, make coffee, there are myriad variables: No two kitchens are identical; coffee machines work in different ways; different cups require different grips; coffee grounds, hot water, and milk all need to be handled differently. Even this simple task requires understanding an ever-changing menu of possibilities. Achieving generality through VLAs, Hurst argues, would require “complete data coverage of all of the things that [a robot] could ever do.” Or, in other words, an almost infinite pool of training data. LeCun is dismissive of the whole approach. “The [AI] approaches that have been successful for language do not work for high-dimensional, continuous, noisy data”—the kind of data that is commonplace in robotics, he said in Davos. “You have to use something else.” The leading contender for “something else” is the so-called world model—a form of AI trained less on text than on a combination of video, three-dimensional scans, and sensor data and built to predict the outcomes of actions in the real world. The aim is to build models that possess an internal representation of reality precise enough to capture how the physical world actually operates—how objects move, collide, fall, and deform. If roboticists could train machines in simulations faithful enough to real-world physics, development would become faster, cheaper, and safer, reducing the need for real-world testing. Even more transformative, robots equipped with world models could reason about their surroundings rather than merely reacting to them, helping them anticipate the consequences of an action before taking it. Companies like Nvidia and Google are working on the technology, and investor cash is pouring into high-profile startups. World Labs, cofounded by the Stanford AI researcher Fei-Fei Li, raised $1 billion in funding in February and was acquired by AMD at the end of September for $8.2 billion. AMI Labs, cofounded by LeCun (formerly Meta’s chief AI scientist), also raised $1 billion in March. Yet by their own admission, it is still early days. Late last year Li described the field as “nascent,” adding that “foundational approaches are still being established.” In a June Substack she described daunting challenges. For now, world models are a promising area of research rather than an immediate route to general-purpose robotics, but we are beginning to see glimmers of what they could achieve. One such glimpse came with a small but potentially significant leap forward that took place in a San Francisco robotics lab last April. A breakthrough? In the heart of San Francisco’s Mission District, the startup Physical Intelligence—or PI (as in π), as it likes to be known—is focused on developing a universal brain that could, theoretically, turn any robot into a generalist. Using an everything-including-the-kitchen sink approach to training AI models for robots, the company recently observed a hint of what a robotic brain equipped with a world model could be capable of. In 2024, PI published details of its first generalist robotics system, called π0, a VLA it claimed was the “most capable and dexterous generalist robot policy to date.” The model was initially trained on a 10,000-hour proprietary collection of human demonstrations gathered through teleoperation as well as several open-source robot datasets. A version released in spring 2025, π0.5, was trained on a wider variety of datasets, including labeled images from the web, lending it more versatility. A fall 2025 update, π0.6, added reinforcement learning to the model. Each update yielded important improvements to the model’s performance, increasing its menu of abilities from slowly folding laundry to putting things away in new environments to completing tasks like folding boxes with a higher success rate. Then, in April 2026, π0.7 seemed to catapult PI into new territory. This version makes use of a less powerful world model that generates images of steps necessary to perform a task. As the robot undertakes the job, this “lightweight” model feeds it snapshots of what to do next. The company claims that the model exhibits the first signs of compositional generalization, a term for AI systems’ ability to perform skills they’ve never been exposed to by recombining ones learned in their training data. One test involved asking a model to “load a sweet potato into the air fryer”—a task it had never previously encountered. In a demonstration video, the machine futzes around a little, makes a few false starts, and eventually manages a reasonable effort, though it doesn’t finish the task completely. Sergey Levine, a professor at the University of California, Berkeley, and a cofounder of PI, is excited by the potential: “It’s actually the first time that we’ve convincingly seen that kind of compositional generalization, where we can basically ask the model to do tasks that we did not specifically collect data for and train it to do, and it’ll actually make a passable attempt.” The success led the team to wonder how the model was able to achieve such a feat. After some digging, they found snippets of relevant labeled teleoperation data lurking in the training material, including two examples of a human controller using the robot to push an air fryer basket into the fryer. Those shreds of data might have been enough to enable π0.7 to almost air-fry a sweet potato. For now, it remains unclear just how impressive π0.7’s abilities to generalize are. Still, given how fleeting the model’s exposure to air fryers had been, it offers a glimpse into how far cutting-edge research can currently take robots. “70% success is like it doesn’t work” You might be sensing a disconnect between the halting baby steps robots are making in labs—“Look! It put a sweet potato into an air fryer!”—and the dazzling, lifelike nimbleness on view during many demonstrations and videos, where robots are seen doing everything from dancing on a stage to courteously serving drinks. Such demos often don’t clearly reveal a key fact: In many instances, humans are controlling the robot or have carefully scripted its actions. (The robot that appeared onstage with Nvidia CEO Jensen Huang in March 2025, for example, seemingly responding to his instructions and following him around, was remote-controlled by what its makers called “a puppeteer behind the scenes.”) For now, fully autonomous motion planning so that a robot knows where it should go—especially in new, chaotic environments like a construction site or a unfamiliar home—remains a largely unsolved challenge. A bigger challenge still—albeit one that is often related—lies in getting robots to tackle larger, more ambiguous jobs that include multiple tasks and require decisions about how and in what order they’re done. This would be the difference between a robot that can put a plate into a microwave and one that can successfully respond to the prompt “Make dinner” by exploring the refrigerator, chopping ingredients, and firing up the stove. Google DeepMind’s best attempts at something like this—which involved asking its robot to survey a kitchen and pack all the ingredients for a mushroom risotto into a basket—have so far resulted in failure. Adding to the challenge, a practical robot must essentially get it right every time. With VLAs, “people are very excited when their result goes from 50% success to 70% success,” says Marc Raibert, founder of Boston Dynamics. “But 70% success is like it doesn’t work, right?” The few humanoids that are being tested in real-life settings are undertaking extremely limited tasks in tightly controlled environments. They’re far from generalists. Agility has hundreds of robots deployed across trials in facilities owned by GXO Logistics, Amazon, and Schaeffler, according to the company. But for now, Hurst says, the robots are targeting simple tasks such as moving bins and totes around. Even then, he adds, it took years to develop robots safe enough for logistics firms to even contemplate using them. For his part, Elon Musk claimed in May 2025 that “thousands” of his Optimus robots would be working at Tesla factories by the end of the year, but in January of this year he said that the company had only “some of the Tesla Optimus robots doing simple tasks in the factory.” While humanoids are starting to venture onto the factory floor, making the jump to households will be even more difficult. Right now, should you so desire, you can preorder the 1X Neo home robot, expected to be ready for delivery sometime later this year. Yours for $20,000, it promises to take on “the boring and mundane tasks around the house”—putting away dishes, answering the door, tidying the living room—“so you can focus on what matters to you.” The idea is for this five-foot-six-inch robot to one day perform all those tasks autonomously, but for now a remote human operator is needed for it to do most things. (Yes, a person would need permission to peer into your home through the robot’s cameras.) Asked how long it will be until fully autonomous robots are ready for domestic work, Hurst said, “If I had to pick a number, I’d say it’s 10 years before robots are … actually doing useful things in people’s homes.” When that happens, the robots might well be Chinese, as China is well ahead of the West in terms of production. Nearly 90% of the roughly 15,000 humanoid robots shipped in 2025 were made by Chinese companies, according to the market intelligence company Omdia and the Chinese robotics firm Unitree. One model produced by Unitree, which shipped more humanoid robots than any other company last year, costs less than $6,000. Such an affordable price could go a long way toward making robots more attractive to consumers, though the company expects its machines to be used in industrial applications first. (If you’re wondering who is buying all these Chinese robots, by the way, the AP recently reported that orders come predominantly from corporate and academic labs and state-owned enterprises.) We’ve been here before The dream of building a humanoid robot runs deep: As far back as 1495, Leonardo da Vinci sketched out designs for a mechanical knight, controlled by cables and pulleys. Through the 20th century, machines of sci-fi fever dreams have come and gone. Westinghouse’s seven-foot-tall box on legs, Elektro, hit the New York World’s Fair in 1939, smoking a cigarette. WABOT-1, built by Waseda University in Japan in 1973, was the first full-scale, programmable humanoid robot. Honda’s ASIMO, unveiled in 2000, was probably the first such machine to prove at all competent—it could, at least to some degree, climb steps, recognize faces, and autonomously move through spaces. But the robot was discontinued in 2018, unable to advance far enough beyond what it could do in demonstrations to be useful. All, at the time, were impressive—even jaw-dropping—feats of engineering. But none were ready to navigate the real world. Today’s robots, even with the transformative power of advanced AI, still face the same existential challenge. Deep Dive Artificial intelligence Don’t be fooled—LLMs don’t reason Ten years after AlphaGo’s match against Go champion Lee Sedol, today’s AI still isn’t tapping into the machinery that made that win possible. AI’s recursive self-improvement might not come so quickly after all AI agents are not yet creative enough to carry out genuinely innovative open-ended AI research, it seems. Stay connected Get the latest updates from MIT Technology Review Discover special offers, top stories, upcoming events, and more.
09:26

IBM's Tiny TSPulse Beats Models 100x Bigger Across 75 Datasets

A tiny time-series model is beating much larger ones at reading sensor traces. IBM’s Granite TSPulse R1 has 1.08 million parameters, Apache-2.0 weights, and CPU inference. The ICLR 2026 paper claims +20% on TSB-AD, +50% on imputation, and +25% on similarity search versus models 10–100× larger. It does not forecast. Context is 512; anomaly detection wants 1.5–2K points.

Notes
  • IBM Granite TSPulse R1 / TSPulse-R1: 1.08M-parameter time-series foundation model, Apache 2.0, CPU inference. Hugging Face page “more than 860,000 downloads” at write time.
  • Paper accepted at ICLR 2026. IBM: stronger than models 10–100× larger across more than 75 datasets.
  • Tasks: classification, anomaly detection, imputation, similarity search. Does not forecast. Context 512. Anomaly detection wants 1.5–2K points.
  • Dual-space masked reconstruction (time + frequency). Disentangled temporal, spectral, semantic embeddings. Compact MLP-Mixer. Three specialized variants via Hugging Face revision.

| Task | Reported lift | Scope |

|---|---|---|

| Anomaly detection | 20% | TSB-AD |

| Similarity search | 25% | Retrieval |

| Imputation | 50% | Missing-value |

| Multivariate classification | 5%–16% | Multiple datasets |

  • Those percentages mix metrics and baselines. Use the paper’s per-dataset tables for a real workload.
  • 1.08M FP32 weights ≈ 4.3 MB before packaging — IoT gateways, CPU monitors, laptops, no GPU.
  • Paywalled after this AlphaSignal free preview.
Full text · 2,643 chars
- IBM released TSPulse-R1, a 1.08M-parameter time-series foundation model under Apache-2.0. - Supports classification, anomaly detection, imputation, and similarity search with GPU-free CPU inference. - Reports +20% on TSB-AD anomaly detection, +50% on imputation, +25% on similarity search versus larger baselines. - Uses dual-space masked reconstruction over time and frequency domains with disentangled temporal, spectral, and semantic embeddings. - Ships three specialized variants selected via the Hugging Face revision argument; paper accepted at ICLR 2026. - Does not do forecasting; context length 512, with anomaly detection needing 1.5-2K points. IBM has released Granite TSPulse R1, a pre-trained time-series model for classification, anomaly detection, imputation, and similarity search. The model contains 1.08 million parameters, supports CPU inference, and ships with Apache 2.0-licensed weights. At the time of writing, its Hugging Face page showed more than 860,000 downloads. IBM says the accompanying paper was accepted at ICLR 2026. Across more than 75 datasets, the paper reports stronger results than models with 10 to 100 times as many parameters. | Benchmark improvements reported by IBM | | | |---|---|---| | Task | Reported improvement | Benchmark scope | |---|---|---| | Anomaly detection | 20% | TSB-AD leaderboard | | Similarity search | 25% | Retrieval benchmarks | | Imputation | 50% | Missing-value benchmarks | | Multivariate classification | 5% to 16% | Multiple classification datasets | Those percentages summarize results across different metrics, baselines, and datasets. The paper’s per-dataset tables provide the necessary detail for comparing TSPulse with a specific production workload. The point of 1.08 million parameters A time-series foundation model learns reusable signal representations during pre-training, then applies them to new datasets with limited or no task-specific training. Many recent models use architectures with tens or hundreds of millions of parameters. TSPulse uses a compact MLP-Mixer that alternates between mixing information across time patches and feature channels. The parameter count reduces deployment costs. Storing 1.08 million parameters in FP32 requires roughly 4.3 MB before packaging, dependencies, and runtime memory. That footprint allows the model to run inside CPU-based monitoring services, IoT gateways, batch pipelines, and developer laptops without dedicated GPU infrastructure. This story is for Pro members You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.
09:36

Prompting Claude Haiku 5.5 - Claude Platform Docs

Anthropic published a page of prompting patterns for its cheap worker model. The Claude Platform docs for Haiku 5.5 call out effort, search, and JSON as the model-specific knobs. The stub is a table of contents, not the recipes. Open the docs for the actual patterns.

Full text · 149 chars
Best practices Prompt engineering . Prompting Claude Haiku 5.5. Copy page.. Prompting patterns specific to Claude Haiku 5.5: effort, search, JSON ...
10:00

Meta's Llama 3.3 70B Now Fits on a Single 48 GB GPU

A big open chat model now fits on one mid-range graphics card. A community 4-bit AWQ of Llama 3.3 70B Instruct lands in 35–40 GB and runs on a single 48 GB GPU. Cited base scores: 92.1 IFEval, 88.4 HumanEval, 77 MATH, 91.1 MGSM. The 128k context and GQA stay; the KV cache still costs about 320 KiB per token. It is not a reasoning model and sits outside the LMArena top 20.

Notes
  • Community 4-bit AWQ of Llama 3.3 70B Instruct. Fits one ~48 GB GPU.
  • Base (unquantized) scores cited: 92.1 IFEval, 88.4 HumanEval, 77 MATH, 91.1 MGSM.
  • Keeps Meta’s 128k context, GQA, 8-language multilingual support.
  • Native vLLM, SGLang, Transformers via safetensors. Built with AutoAWQ.
  • Not a reasoning model. “Ranks outside LMArena top 20.”
  • Hugging Face “hundreds of thousands of downloads” in recent monthly windows — file requests, not unique deploys.

| Resource | Approx. size | Effect |

|---|---|---|

| BF16 weights | 140 GB | Multi-GPU after overhead |

| AWQ checkpoint | 35–40 GB | One 48 GB GPU, limited cache and concurrency |

| 16-bit KV cache | ~320 KiB per token and sequence | ~2.5 GiB at 8k tokens; ~40 GiB at 128k |

  • Activations and KV typically stay FP16/BF16. A single 48 GB RTX A6000 “can usually serve shorter contexts at low concurrency.” Full 128k generally needs more GPUs or a cheaper cache.
  • Paywalled after this AlphaSignal free preview.
Full text · 2,706 chars
- 4-bit AWQ build of Llama 3.3 70B Instruct fits on a single 48GB GPU. - Base model scores 92.1 IFEval, 88.4 HumanEval, 77 MATH, 91.1 MGSM. - Preserves 128k context, GQA, and 8-language multilingual support from Meta's original. - Native support in vLLM, SGLang, and Transformers via standard safetensors format. - Built with AutoAWQ, maintained by the same author as the quantization library. - Non-reasoning model, loses to modern reasoning models and ranks outside LMArena top 20. Community AWQ build puts Llama 3.3 70B on a 48 GB GPU A community-maintained quantization has made Meta’s Llama 3.3 70B Instruct practical to serve on hardware with about 48 GB of VRAM. The AWQ checkpoint compresses most model weights to 4 bits and works with vLLM, SGLang, and recent Transformers releases. Hugging Face has recorded hundreds of thousands of downloads for the repository in recent monthly windows. Those counters include automated file requests and repeated downloads, so they measure distribution activity rather than unique production deployments. The practical signal is broad runtime support: developers can use established serving stacks without converting the model themselves. How four bits shrink 140 GB of weights Activation-aware Weight Quantization, or AWQ, calibrates the model on sample inputs to identify weight channels whose errors have the greatest effect on output. It scales those channels before applying groupwise 4-bit quantization, reducing error while keeping the weight tensors compact. Activations and the key-value cache typically remain in FP16 or BF16. The checkpoint was created with the AutoAWQ project. Its safetensors files include quantization metadata that compatible engines use to select AWQ kernels. Kernel availability and performance still depend on the GPU, CUDA version, and serving framework. | Resource | Approximate size | Operational effect | |---|---|---| | BF16 weights | 140 GB | Requires multiple large GPUs after runtime overhead | | AWQ checkpoint | 35 to 40 GB | Fits on one 48 GB GPU with limited cache and concurrency | | 16-bit KV cache | About 320 KiB per token and sequence | Consumes roughly 2.5 GiB at 8k tokens and 40 GiB at 128k | AWQ reduces weight storage and memory bandwidth. Runtime workspaces, activations, and the KV cache still consume VRAM, which makes Meta’s 128k context limit expensive to use. A single 48 GB RTX A6000 can usually serve shorter contexts at low concurrency, while the full context generally requires additional GPUs or a lower-precision cache. This story is for Pro members You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.
16:00

EMA Lightning Beats ElevenLabs on Turkish Speech in Just 34 MB

A tiny Turkish speech model is beating the big cloud voices on a small test. Canberk Aslan’s EMA Lightning has 8.6 million parameters, Apache-2.0, about 34 MB, and one fixed voice. On Freya-TR-Eval (495 sentences) it posts 0.92% WER versus Trendyol-TTS 0.95%, Gemini 3.8 Flash 1.35%, and ElevenLabs v4 1.43%. The card claims 440× real-time single request and 1,316× batched on an RTX 4090, first audio in 3.86 ms, about $0.0085 per million characters. UTMOS 3.30 trails Gemini and Trendyol. Independent listening tests are not in the preview.

Notes
  • EMA Lightning: 8.6M params, Apache-2.0, ~34 MB, one Turkish voice, pip install ema-lightning. DiT + flow matching, distilled to 4 steps (DMD2), shared modulation, windowed Gaussian aligner.
  • Freya-TR-Eval (495 sentences, author card):

| System | Params | WER | UTMOS |

|---|---|---|---|

| EMA Lightning | 8.6M | 0.92% | 3.30 |

| Trendyol-TTS | 2.38B | 0.95% | 3.83 |

| Gemini 3.8 Flash | n/d | 1.35% | 3.52–3.61 |

| ElevenLabs v4 | n/d | 1.43% | ~3.30 |

  • Speed (card): 440× real-time single request; 1,316× batched on RTX 4090; first audio 3.86 ms; ~$0.0085 per million characters.
  • WER = intelligibility via ASR, not preference. No human listening study. AlphaSignal cuts at the paywall after this table.
Full text · 2,582 chars
- EMA Lightning is an 8.6M parameter Turkish TTS model, Apache 2.0, about 34 MB total. - Hits 0.92% WER on Freya-TR-Eval, beating ElevenLabs v4, Gemini 3.8, and 2.38B Trendyol-TTS. - 440x faster than real time single request, 1,316x batched on an RTX 4090. - First audio in 3.86 ms; about $0.0085 per million characters of generated speech. - DiT with flow matching, distilled to 4 steps via DMD2, plus shared modulation and a windowed Gaussian aligner. - Install with pip install ema-lightning ; source on GitHub, Turkish-only, single voice. EMA Lightning Packs Turkish TTS Into 34 MB Canberk Aslan has released EMA Lightning, an 8.6 million-parameter text-to-speech model for Turkish. According to benchmarks published with the model, it records a lower word error rate than ElevenLabs v4, Gemini 3.8 Flash, and the 2.38 billion-parameter Trendyol-TTS while running locally on a CPU or GPU. The Apache 2.0 release includes model weights, source code, and a Python package. Its 34 MB stack generates one fixed Turkish voice and requires no cloud API, making it relevant to developers who need low-cost, offline speech generation with predictable latency. Benchmark Gains With Clear Boundaries On Freya-TR-Eval, a benchmark containing 495 Turkish sentences, the model card reports a 0.92% word error rate. Trendyol-TTS is roughly 277 times larger by parameter count and records 0.95% in the same comparison. | Reported Freya-TR-Eval results | | | | |---|---|---|---| | System | Parameters | WER | UTMOS | |---|---|---|---| | EMA Lightning | 8.6M | 0.92% | 3.30 | | Trendyol-TTS | 2.38B | 0.95% | 3.83 | | Gemini 3.8 Flash | Undisclosed | 1.35% | 3.52–3.61 | | ElevenLabs v4 | Undisclosed | 1.43% | About 3.30 | Word error rate measures how accurately a speech recognizer can recover the generated text, with lower values indicating fewer transcription errors. It captures intelligibility rather than voice preference or emotional range. The results come from the author’s 495-sentence comparison, so broader and independently reproduced evaluations would provide stronger evidence across domains and sentence styles. UTMOS is an automated estimate of perceived speech quality, with higher scores indicating greater predicted naturalness. EMA Lightning’s 3.30 score matches the reported ElevenLabs result and trails Gemini and Trendyol-TTS. No human listening study is included in the published comparison. This story is for Pro members You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.
16:49

Odyssey Releases Odyssey-3, a World Model That Teaches Robots Physics Through Video

A world model is trying to teach robots physics by watching and acting in video. Odyssey-3 is an autoregressive diffusion transformer that builds navigable scenes from text and predicts how they change under control. Odyssey-3 Pro scores 66.1 on Physics-IQ Verified video-to-video at 720p with best-of-eight sampling, first in three of four WorldMark categories, and third on first-person real at 80.6 behind Lyra 2.0 at 84.4. A frozen backbone drove in India after 20 hours of road data. The research preview is free at experience.odyssey.systems; the company has not published frame rate, context length, or hardware.

Notes
  • Odyssey-3: foundation world model. Text → navigable scene; predicts how the scene changes under control. Autoregressive diffusion transformer; same pretrained backbone for interactive sim, machine control, and agent training.
  • Preview: experience.odyssey.systems (free). API by request. No published frame rate, context length, or hardware profile.
Reported scores (Odyssey / launch)

| Benchmark | Result | Note |

|---|---|---|

| Physics-IQ Verified, video-to-video | 66.1 | Odyssey-3 Pro, 720p, best-of-eight; ahead of FLUX 3 and Cosmos3 variants |

| Physics-IQ Verified, image-to-video | 54.7 | best-of-eight |

| WorldMark first-person stylized | First | mean of 13 metrics |

| WorldMark third-person real + stylized | First | both splits |

| WorldMark first-person real | 80.6, third | Lyra 2.0 84.4; AlayaWorld 83.0 |

  • Physics-IQ Verified: continue real-experiment video across fluids, optics, solids, magnetism, thermodynamics. Best-of-eight uses more compute than one sample.
Training data
  • Internet video + time-localized event annotations on a schema.
  • Gameplay with synced keyboard/mouse.
  • Rigid-body sims with captions/metadata.
Adaptation
  • Frozen visual backbone + small action decoder/policy.
  • Robot arms: tens of hours of demos; recoveries not in the demos (reorient gripper, retrieve a dropped object).
  • Humanoids: Flexion policies beat tested VLA baselines under lighting changes.
  • Vehicles: 20 hours India road data; closed-loop waypoints. Independent reliability/safety tests not in the piece.
Full text · 7,950 chars
- Odyssey released Odyssey-3, a real-time interactive foundation world model generated from text prompts. - Odyssey-3 Pro scores 66.1 on Physics-IQ Verified video-to-video, the highest reported score. - Ranks first in 3 of 4 WorldMark categories, trailing only in first-person real environments. - Frozen backbone drives a car in India after just 20 hours of driving data. - Autoregressive diffusion transformer distilled to few-step sampling for real-time interactive generation. - Free research preview live at experience.odyssey.systems, API access by request. Odyssey-3 links interactive video with physical control Odyssey has released Odyssey-3, a foundation world model that generates navigable environments from text and predicts how each scene changes as a person or software agent acts. The free research preview is available now, while developers working on robots and vehicles can request API access. At its core, Odyssey-3 is an autoregressive diffusion transformer: it generates visual states through iterative denoising, then conditions each new state on previous observations and control inputs. Odyssey designed the same pretrained backbone for three workloads: interactive simulation, physical-machine control, and agent training. Reusing that backbone could reduce the task-specific data and compute required for each robot, vehicle, or simulated environment. A physics lead with a compute caveat Physics-IQ Verified evaluates whether a model can continue recordings of real experiments while preserving behavior across fluid dynamics, optics, solid mechanics, magnetism, and thermodynamics. Odyssey-3 Pro achieved the highest reported video-to-video score, according to the launch results. | Benchmark results reported by Odyssey | | | |---|---|---| | Benchmark | Result | Context | |---|---|---| | Physics-IQ Verified, video-to-video | 66.1 | Odyssey-3 Pro at 720p using best-of-eight sampling; reported ahead of Black Forest Labs’ FLUX 3 and NVIDIA Cosmos3 variants. | | Physics-IQ Verified, image-to-video | 54.7 | Odyssey-3 Pro using best-of-eight sampling. | | WorldMark, first-person stylized | First | Ranking based on the mean of 13 reported metrics using WorldMark’s captions. | | WorldMark, third-person real and stylized | First | Odyssey-3 led both third-person splits. | | WorldMark, first-person real | 80.6, third | Behind Lyra 2.0 at 84.4 and AlayaWorld at 83.0. | The 66.1 Physics-IQ result uses the strongest outcome from eight generated samples, which requires more compute than a single generation. Odyssey also reports that the base model improves the tradeoff between physical accuracy and generation cost, although the launch materials provide charted generation costs rather than API pricing or hardware-specific latency. Three data streams teach cause and effect Odyssey trained the model on three complementary data sources, each covering a different part of interactive world modeling: - Internet video supplies broad visual coverage, paired with time-localized event annotations checked against a defined schema. - Gameplay recordings pair frames with synchronized keyboard and mouse inputs, giving the model explicit action-to-outcome examples. - Rigid-body simulations provide controlled interactions with captions and metadata, making cause and effect easier to isolate. Streaming diffusion under the hood The base architecture is a multi-step video diffusion transformer with temporal rotary positional embeddings, or RoPE, which encode the order of frames. Causal masking prevents the model from attending to future frames, allowing it to predict the next state from past observations. During training, teacher forcing supplies the correct preceding observations; during generation, the model continues from its own output while accepting new prompts and controls. Odyssey compresses the original diffusion sampler into a few-step model to reduce interaction latency. Post-training combines distribution-matching distillation, which teaches a faster model to reproduce the slower model’s output distribution, with generative adversarial discriminators and reinforcement learning. The company has not published an exact frame rate, context length, or hardware profile for the research preview. One backbone, three machines Odyssey adapts the frozen visual backbone to physical systems by training an action decoder or policy on paired observations and actions. Keeping the backbone frozen preserves its pretrained weights while the smaller policy learns how to control a particular machine. - Robot arms. Odyssey reports that policies trained on tens of hours of demonstrations completed manipulation tasks and recovered from failures absent from the demonstrations. Examples include reorienting a gripper after a missed grasp and retrieving an object dropped in an unusual position. - Humanoids. Flexion built humanoid policies on Odyssey-3. The resulting systems outperformed the tested vision-language-action baselines under environmental changes and continued operating under lighting changes that disrupted those baselines. - Vehicles. Odyssey trained a driving policy on 20 hours of road data from India while leaving the backbone unchanged. The policy uses its visual representations to predict waypoints and drive in closed loop, meaning each prediction affects the next observation and action. Freezing the backbone can lower adaptation costs because each embodiment needs a smaller policy rather than a newly trained visual model. The public evidence currently consists of Odyssey’s demonstrations and reported evaluations, so independent testing still needs to establish reliability, safety, and performance across unfamiliar machines and environments. Agents learn inside generated scenes For agent training, Odyssey-3 serves as the environment that produces observations and responds to actions. In the task-completion demonstration, an agent receives a natural-language goal, observes the generated scene, and takes actions until it reaches the requested outcome. Generated environments can produce varied rollouts without requiring a physical robot or a hand-authored asset pipeline for every scene. They also carry the world model’s errors into training: inaccurate dynamics, weak long-term memory, or missing edge cases can teach a policy the wrong behavior. Odyssey has not reported task-success rates, agent-training costs, or sim-to-real transfer results for this use case. Evidence gaps developers should track - Long-horizon consistency: WorldMark’s world-memory measure is folded into a 13-metric mean, leaving no detailed breakdown of how long scenes remain coherent. - First-person realism: Odyssey-3 ranks third on WorldMark’s first-person real split, behind Lyra 2.0 and AlayaWorld. - Runtime requirements: Exact latency, throughput, memory use, context limits, supported control schemas, and deployment hardware remain unspecified. - Sampling cost: The leading Physics-IQ result requires eight samples at 720p, raising compute use relative to a single interactive generation. - Physical-system validation: The robot, humanoid, and vehicle results need independent replication, longer deployments, and safety evaluation. The bet is reusable world knowledge Odyssey-3’s developer proposition centers on reuse: pretrain visual dynamics across broad video and simulation data, freeze that shared backbone, then train a smaller decoder for each robot, vehicle, or agent. If independent evaluations reproduce the reported transfer results, the approach could reduce specialized data collection and shorten the path from a general visual model to a working control policy. Preview now, API by request Developers can try the research preview and request API access for physical systems. Odyssey’s related work includes the Agora-2 multi-agent world model and PROWL-2, a reinforcement-learning framework that trains multiple agents and their world model within one loop.
18:11

Artificial Analysis Rebuilds Intelligence Index v5 With Private Coding Tests

The scoreboard everyone cites is about to change its questions. Artificial Analysis Intelligence Index v5 lands in late October with Terminal-Bench Science and a private coding test. Terminal-Bench Science has 70 shell tasks; GPT-6 Astra Max leads at 63.3% and Claude Opus 5.5 at 61.9%. Current v4.3.2 weights are Agents 30%, Coding 20%, General 30%, Scientific Reasoning 20%. Old Index numbers will not be comparable. Final v5 weights are not published yet.

Notes
  • Intelligence Index v5 late October. Largest revision this year. v4.x scores will not be directly comparable.
  • Current v4.3.2 weights: Agents 30 / Coding 20 / General 30 / Scientific Reasoning 20. Final v5 weights unpublished.
  • Adds Terminal-Bench Science (Stanford + Terminal-Bench/Harbor; 70 shell tasks, 5 domains, 3 attempts, pass@1, verifier timeout = fail) and a private coding benchmark (tasks/scoring unpublished).
  • Standalone Terminal-Bench Science leaders: GPT-6 Astra Max 63.3%; Claude Opus 5.5 Xhigh 61.9%; Opus 5.5 Max 59.0%; Sonnet 5.5 53.0%. Nobody listed clears 64%.
  • Contamination: public items leak into training. Private set follows the AutomationBench-AA pattern; outsiders cannot reproduce the full eval.
  • Advice: treat v4.3.2 and v5 as separate series; Terminal-Bench 4.0 already replaced v2.1 with a harder 66-task set.
Full text · 4,843 chars
- Artificial Analysis Intelligence Index v5 is scheduled to launch in late October. - v5 adds Terminal-Bench Science, an agentic eval of scientific workflows run through a terminal. - A new coding benchmark with a private test set joins the Coding category to reduce contamination risk. - Current Terminal-Bench Science leaders: GPT-6 Astra Max at 63.3%, Claude Opus 5.5 at 61.9%. - Category weights in v4.3.2 are Agents 30%, Coding 20%, General 30%, Scientific Reasoning 20%. - Existing Intelligence Index scores will not be directly comparable once v5 ships. Independent model evaluator Artificial Analysis plans to release Intelligence Index v5 in late October, its largest revision of the composite benchmark this year. The update will add Terminal-Bench Science and a coding benchmark whose private test set will remain unpublished. The Intelligence Index condenses results from several evaluations into a single model-capability score. Researchers already cite it in third-party model papers, so changing its benchmarks can reorder model rankings and break direct comparisons with earlier results. V5 adds two tougher signals Artificial Analysis has introduced the revision in stages. The current v4.3.2 build includes AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, and AA-LCR v1.1. | Category | Current weight | |---|---| | Agents | 30% | | Coding | 20% | | General | 30% | | Scientific reasoning | 20% | The v5 release will add the following components: - Terminal-Bench Science: An existing standalone evaluation built by Stanford University researchers with the Terminal-Bench and Harbor team, with contributions from scientists at institutions worldwide. Its tasks reproduce research workflows performed through a command-line environment. - A private coding benchmark: An evaluation based on unpublished test data. Artificial Analysis has not disclosed its tasks, scoring method, or dataset composition. Artificial Analysis has also yet to publish the final v5 category weights, an exact release date, or guidance for comparing v5 scores with v4.3.2 results. Those details are expected with the changelog. Science work moves into the index Terminal-Bench Science measures whether an AI agent can use a shell to complete realistic scientific workflows. The benchmark contains 70 tasks across five domains, and each task is attempted three times. Results are reported as pass@1: an attempt succeeds only when every test passes, while verifier timeouts count as failures. The current standalone leaderboard indicates how demanding the evaluation is: - GPT-6 Astra (Max): 63.3% - Claude Opus 5.5 (Xhigh, Default Fallback): 61.9% - Claude Opus 5.5 (Max, Default Fallback): 59.0% - Claude Sonnet 5.5: 53.0% No listed system clears 64%, leaving substantial room between current agents and complete task reliability. Adding the benchmark could shift composite rankings toward models that sustain long-running terminal sessions, recover from execution errors, and complete multi-step research procedures. Private tests reduce training leakage Public benchmark questions can enter model-training corpora after circulating online. A model may then reproduce familiar solutions rather than demonstrate that it can solve unseen problems, weakening the benchmark as a measure of coding ability. Artificial Analysis has increasingly used private test variants and filtering procedures to reduce that risk, including its move to AutomationBench-AA. The new coding benchmark extends the same approach. Its unpublished data may provide a cleaner signal, although outside teams will be unable to reproduce the full evaluation independently or inspect the test distribution. Prepare for a score break - Rerun internal comparisons after release. V4.3.2 and v5 scores should be treated as separate series until Artificial Analysis publishes the new weights and methodology. - Update coding-agent baselines. Terminal-Bench 4.0 already replaces v2.1 with a harder 66-task set, recalibrated compute and time allowances, and revised environments and verifiers. Scores may fall when teams move from the older harness. - Track task-level scientific failures. Aggregate scores can hide whether an agent failed through planning, shell use, dependency management, computation, or verification. Those distinctions matter when improving research agents. - Record benchmark versions. Reports should identify the Intelligence Index version, model configuration, and evaluation date so later results remain interpretable. Artificial Analysis documents its current scoring and evaluation practices on its benchmark methodology page. The v5 changelog should clarify the final weights, private coding benchmark design, and rules for comparing results across index versions.
18:26

OpenAI Brings 8x Faster Ultrafast Inference to GPT-6.1 Sol

OpenAI sold a faster lane for the same Sol model and charged six times the token price. Ultrafast for GPT-6.1 Sol is in the API, Codex, and ChatGPT Work, advertised at up to 8× Standard and around 300 tokens per second. API rates are $12 / $60 per million versus $2 / $10 Standard. Codex and Work need Pro 500, usage-based Enterprise, or credit-based Edu. Sol keeps a 1,050,000-token window. The 8× claim is generation speed, not end-to-end latency.

Notes
  • Ultrafast now on GPT-6.1 Sol in API, Codex, ChatGPT Work. Same weights, faster infra. service_tier="ultrafast" on Responses API.
  • Up to 8× Standard generation; Astra Ultrafast cited at ~300 tokens/s. No separate Sol absolute max published.
  • Price: Standard $2 / $10 per 1M; Ultrafast $12 / $60 (6×). Astra Ultrafast $60 / $300 — Sol is one-fifth of that.
  • Chat surfaces: Pro 500, usage-based Enterprise (admin must enable), credit-based Edu. All supported regions; US + EU residency. EU residency also added for Sol Fast and Luna Fast.
  • Default TPM: Build 500k, Launch 1M, Grow 5M.
  • Context: 1,050,000 (up to 922k in / 128k out).
  • Use WebSockets for tool-heavy agents; 8× is generation only. 100k output tokens ≈ $1 Standard vs $6 Ultrafast.
  • Workloads named: incident response, computer-use agents, live IDE/voice.
Full text · 4,996 chars
- OpenAI rolled out Ultrafast mode for GPT-6.1 Sol in API, Codex, and ChatGPT Work. - Up to 8x faster token generation than Sol Standard, reaching around 300 tokens per second. - API pricing is $12 per million input tokens and $60 per million output tokens, 6x Standard. - In Codex and ChatGPT Work, access requires Pro 500, usage-based Enterprise, or credit-based Edu plans. - Supported in all regions with US and EU data residency; EU residency also added for Sol Fast and Luna Fast. - Built for outage debugging, agents navigating apps, and live experiences where latency matters. OpenAI brings Ultrafast inference to GPT-6.1 Sol OpenAI has extended its Ultrafast tier to GPT-6.1 Sol across the API, Codex, and ChatGPT Work. The service promises up to eight times faster token generation than Standard mode, targeting interactive agents, live developer tools, and incident-response workflows. OpenAI introduced Ultrafast for GPT-6 Astra at its DevDay announcement in 2026 and previewed support for Sol. That support is now available, subject to account and plan eligibility. An 8x lane for the same model Ultrafast runs the same GPT-6.1 Sol model on faster inference infrastructure. API clients select it by setting service_tier to ultrafast in a Responses API request. OpenAI documents generation speeds of up to 300 tokens per second for Astra in Ultrafast mode. For Sol, the company advertises up to an eightfold improvement over Standard mode without publishing a separate absolute maximum. The premium is 6x per token API calls to GPT-6.1 Sol cost six times more in Ultrafast mode than in Standard mode: | Service tier | Input per 1M tokens | Output per 1M tokens | |---|---|---| | Standard | $2 | $10 | | Ultrafast | $12 | $60 | GPT-6 Astra Ultrafast costs $60 per million input tokens and $300 per million output tokens. Sol therefore carries one-fifth of Astra’s Ultrafast token price. Plans, regions, and limits Codex and ChatGPT Work provide Ultrafast access through the following plans: - Pro 500 - Eligible usage-based Enterprise contracts - Credit-based Edu plans Enterprise administrators must enable the tier before users can select it. OpenAI says Ultrafast is available in every supported region, including deployments with US or EU data residency. The same rollout added EU data residency for GPT-6.1 Sol Fast and GPT-6 Luna Fast. Default API throughput limits increase with the customer’s usage tier: | Usage tier | Default token limit | |---|---| | Build | 500,000 tokens per minute | | Launch | 1,000,000 tokens per minute | | Grow | 5,000,000 tokens per minute | WebSockets protect the latency gain OpenAI recommends WebSockets for agents that make frequent tool calls because a persistent connection reduces request overhead between turns. HTTP remains available through the SDK, though repeated request cycles can consume part of the latency saved during generation. Because the advertised eightfold improvement measures token generation, teams should benchmark complete workflows. End-to-end latency also includes connection overhead, time to the first token, tool execution, and application processing. Python clients can select GPT-6.1 Sol Ultrafast with a standard Responses API call: from openai import OpenAI client = OpenAI() response = client.responses.create( model="gpt-6.1-sol", service_tier="ultrafast", input="Explain why the sky is blue in one sentence.", ) print(response.output_text) Multi-turn agents can maintain one WebSocket session and pass previous_response_id between turns to link responses without repeatedly sending the full conversation history. Why Sol fits interactive work OpenAI positions GPT-6.1 Sol close to GPT-6 Astra on agentic coding, computer use, and professional tasks while charging one-fifth of Astra’s token price. Sol supports a 1,050,000-token context window, including up to 922,000 input tokens and 128,000 output tokens. OpenAI identifies three primary workload categories: - Incident response: Engineers can iterate faster while diagnosing production failures and testing fixes. - Computer-use agents: Faster generation reduces accumulated delay across long sequences of browser or application actions. - Live developer experiences: Voice interfaces, tab completion, streaming suggestions, and code-editing loops benefit from shorter waits between input and output. Spend the premium where latency costs more At current output rates, generating 100,000 tokens costs about $1 in Standard mode and $6 in Ultrafast mode, excluding input tokens. The additional $5 may be economical during a production incident or a user-blocking coding session. Background jobs, offline evaluations, and asynchronous processing can remain on Standard when completion time has little operational value. Standard, Fast, and Ultrafast give developers three processing tiers for the same model family. Applications can now choose a tier per request based on latency targets, workload economics, and user interaction patterns.
19:01

Anthropic Ships Claude Dashboards and Motion to Turn Data Into Auditable Animations

Claude can now turn a warehouse question into a live chart and a deck into code-driven motion. Dashboards (beta, paid plans) talks to BigQuery, Snowflake, Databricks, Redshift, ClickHouse, and Salesforce; every metric shows its SQL and last refresh. Motion (beta, Team and Enterprise) writes editable animation code and exports MP4 with no video model and no generated people. Docs, Slides, and Design leave beta on every plan, including Free. Users have created more than 45 million of those artifacts. Enterprise admins must flip Dashboards and Motion on.

Notes
  • Dashboards (beta, all paid plans): plain-language → queries + live charts. Connectors: Redshift, BigQuery, ClickHouse, Databricks, Snowflake, plus CRM (Salesforce). Click a metric → SQL + last refresh. Handoff: Amplitude, Grafana, Hex, Mixpanel, Omni, Perplexity, PostHog, Sigma. Looker, monday.com, Tableau planned.
  • Motion (beta, Team + Enterprise): code-driven animation of supplied text/data/shapes/images → MP4. No video diffusion model, no generated people. Edits (copy, figures, timing) survive; a changed number can flow through. Handoff: Adobe, Descript, Runway, Luma, Higgsfield. Uses: quarterly summaries, board charts, all-hands, product walkthroughs.
  • Docs, Slides, Design: GA on every plan including Free. >45 million created since launch. Default-on for those three on October 15 (admins can enable earlier).
  • Artifacts family. Dashboards + Motion off by default for Enterprise (Organization settings → Artifacts). CMEK supported.
  • Beta checklist in the piece: messy joins, row-level access, warehouse cost, brand fonts, captions, handoff fidelity.
Full text · 5,226 chars
- Claude Dashboards (beta, paid plans) connects to BigQuery, Snowflake, Databricks, Redshift, Salesforce and builds live dashboards from plain-language questions. - Every chart exposes the SQL query behind it and shows when data was last refreshed. - Claude Motion (beta, Team and Enterprise) generates animated explainers as editable code, exported as MP4. - No video model is involved, so there is no generated footage or AI people in Motion outputs. - Docs, Slides, and Design leave beta and ship on every plan including Free, with CMEK support for enterprise. - Dashboards hand off to Grafana, Hex, Mixpanel, PostHog, Sigma; Motion hands off to Adobe, Runway, Luma, Descript. Claude adds auditable dashboards and editable motion Anthropic has launched Claude Dashboards and Claude Motion in beta, according to its product announcement. Dashboards translates plain-language questions into queries and live charts backed by connected business data. Motion writes code that animates text, charts, shapes, and images, then exports the composition as an MP4. Both tools retain editable components after Claude produces the initial output. Docs, Slides, and Design have also reached general availability across every Claude plan, including Free. Anthropic says users have created more than 45 million documents, presentations, and designs with those tools since their release. Queries stay visible Dashboards connects to Amazon Redshift, BigQuery, ClickHouse, Databricks, and Snowflake. A user can ask a question in plain language, and Claude generates the queries, chooses visualizations, and builds a dashboard that refreshes as the underlying data changes. Existing connectors also provide access to CRM data from services such as Salesforce. Clicking any displayed metric reveals its underlying query, and users can ask Claude to explain the logic. Each chart also shows when its data was last refreshed. Analysts can inspect joins, filters, aggregations, and metric definitions before sharing the result or using it in a decision. Anthropic positions Dashboards for rapid exploration and lightweight reporting. Teams can send dashboards to Amplitude, Grafana, Hex, Mixpanel, Omni, Perplexity, PostHog, or Sigma when a project requires further analysis. Support for Looker, monday.com, and Tableau is planned. Motion keeps the source editable Motion creates animations through code, using programmatic layouts and transitions to move supplied text, data, shapes, and images. The system does not use a video diffusion model, and Motion itself generates neither synthetic footage nor artificial people. Editors can revise the coded composition before exporting an MP4. Edits survive the first cut - Copy, figures, timing, and visual elements can be changed individually. - A revised number can flow through the animation without recreating every frame. - Reviewers can inspect the composition and verify text or data before export. - Teams can continue production in Adobe tools, Descript, Runway, Luma AI, or Higgsfield. Anthropic highlights quarterly-report summaries, animated charts for board presentations, all-hands updates, and product walkthroughs as initial uses. The editable source is particularly relevant when a report changes late in production or a team needs several versions of the same animation. Who gets access | Claude feature availability by plan | | | |---|---|---| | Feature | Status | Plans | |---|---|---| | Claude Dashboards | Beta | All paid plans | | Claude Motion | Beta | Team and Enterprise | | Docs, Slides, and Design | Generally available | Every plan, including Free | Anthropic groups these creation tools under Artifacts. Dashboards and Motion are disabled by default for Enterprise organizations; administrators can enable them in Organization settings under Artifacts. Docs, Slides, and Design will become enabled by default on October 15, though administrators can activate them earlier. Artifacts now support customer-managed encryption keys, allowing organizations to apply their own key-management controls to encrypted content. That option addresses a common procurement requirement for companies handling regulated or sensitive data. Code links the releases Dashboards preserve analytical logic as visible queries and rendering instructions, while Motion preserves visual logic as code and a timeline. Teams can inspect, revise, and compare those components as the underlying data or script changes. This structure reduces rework and gives reviewers a concrete artifact to audit. A practical beta checklist - SQL accuracy: Test ambiguous joins, inconsistent metric names, nested fields, and production-scale tables. - Data governance: Confirm connector permissions, row-level access, query logging, and treatment of sensitive fields. - Operations: Measure refresh latency, warehouse query costs, and behavior after schema changes. - Motion fidelity: Check brand fonts, dense labels, aspect ratios, export quality, and repeated numerical revisions. - Accessibility: Review color contrast, motion pacing, captions, and readability on smaller screens. - Handoffs: Verify that exported dashboards and animations retain the information downstream teams need to continue editing.
00:00

The model that didn't exist, so you made it yourself

Full text · 10,133 chars
Last week, I wanted a small version of the prompt rewriter that ships with Qwen-Image 2.1. The official one is a 9B model that needs about 20 GB of memory and thinks for thousands of tokens before writing a single paragraph. On the Hub, I found only compressed copies of that same 9B model. So I described what I wanted to ML Intern, and the next day I had a 0.8B version that runs on a CPU. It returns valid output 99.7% of the time and uses about a quarter of the teacher's tokens. The compute for the whole project, including having the 9B model label 8,797 example requests, came to USD 16. Over the course of the next few days, I made five more models the same way. Each one started as a message in HuggingChat with ML-intern switched on, and each one ended as a public model on the Hub with its evaluation in the model card. ML-intern plans the work, asks me for a budget before it spends anything, runs a small test before the real job, then trains, evaluates and publishes on Hugging Face hardware. The first message is where I spend my effort. My first prompt, for the citrus model shared below, was about 450 words. By my 6th project it was closer to 2,000, because each project taught me something I wanted in the next one. All seven prompts are on GitHub at yvrjsharma/ml-intern-prompts, exactly as I wrote them. A prompt starts with the idea in one line and why I want it. Then it names the exact pieces: the dataset, the base model, the training script. Anything I have already checked goes under a heading that literally says "Verified facts, do not re-derive", so the agent spends its budget on the work instead of rediscovering what I know. For the camera-angle LoRA that section listed which trainer had just added transparent-image support, and which open GitHub issues made the fallback trainer risky. Two lines in the prompt are critical. The first asks for a baseline before any training. For example, the citrus prompt says: "Also report the base model's zero-shot score on the same metric before training so we can see the gain." Without it you get a trained model and no idea whether it is better than what you started with. The second is a smoke test with a check attached. For the image LoRAs I asked for 50 training steps, then a check that the saved weights had actually changed, before paying for the full run. At the end of the prompt, I lay out the expected deliverables and limit the cost. I define what belongs in the model card and include a instruction like: "Cap total spend at USD 12 and ask me before exceeding it." Because ML-intern begins every task with zero dollar budget and needs permission before executing paid jobs, this spending limit stays strictly enforced. When you leave out a budget, the agent suggests a couple of paths depending on project size and asks which one you prefer. You don't necessarily need all of that on your first attempt. For example, I didn't have the verified-facts section in my citrus brief and ML Intern still produced a model that more than tripled the accuracy of the Qwen3.5-2B model. Let me walk you through 6 things I built with Ml-Intern in just a couple of days. A general vision model can describe a yellowing citrus leaf. However, telling you whether it is a mite problem or a magnesium deficiency, and the bio and non-bio remedies to treat the plant is very hard. Using Claude, I put together a training dataset merged from three sources hosted by the Project-AgML organization on the Hub. The resulting citrus-disease-vlm-instruct is a dataset containing 3,017 annotated images across 21 distinct pests, illnesses, nutritional gaps, and treatment approaches. ML-intern handled the fine-tuning of Qwen3.5-2B using these examples, making sure to benchmark the foundation model beforehand. On the 335 test photos, the base model named the right problem 14.9% of the time. After two epochs on one A10G, the fine-tuned model got 52.8%. Compute cost, about USD 1.90. Check out: Model · Dataset · Citrus Doctor App Image models know plenty of characters. Huggy, drawn in the flat style of the Hugging Face brand assets, was not one of them. I asked ML-Intern for a LoRA on FLUX.2 klein base 4B, trained on 84 captioned drawings from Chunte/huggy_for_training dataset. The agent saved a checkpoint every 100 steps and drew the same set of prompts with each one, which made choosing easy. Step 200 was the first where Huggy was fully on-model. From step 500 on, Huggy's style started bleeding into prompts that had nothing to do with Huggy! The trained LoRA also works on the distilled klein model at 4 steps. Compute cost, about USD 7.60. Check out: Model · Dataset · Huggy Generator App - Camera-angle LoRAs are among the most-liked community add-ons for earlier Qwen-Image models. You can give the model a picture of an object and ask to see it 45 degrees from the left. When I checked a few days after the Qwen-Image 2.1 model release, nobody had made one, so I tasked ML-intern to build it. ML-intern rendered 1,030 scanned household objects from Google Scanned Objects at 24 angles each, 24,722 transparent images, on a CPU job that cost a few cents. It later finalised 461 objects for training and 40 held out for testing, and 1,844 before-and-after training pairs spread evenly over 23 camera instructions. Training ran 2,000 steps in about 90 minutes on one A100 (~USD 3.75). The whole project took about half a day and 48 jobs, counting the ones that failed on missing packages or wrong paths and had to be resubmitted by ML-Intern. Total compute cost, about USD 16. Check out: Model · Dataset · Viewpoint Orbit App - Doodle-in LoRA is another cool idea. Upload a photo with a magenta scribble on it and add a short prompt naming an object. The LoRA replaces the scribble with that object while keeping the original lighting and composition consistent. No dataset existed for this, so my prompt described how to make one. Start from a real photo in Open Images, remove one object with the LaMa inpainting model, and draw a scribble where the object used to be. The untouched photo is the target. ML-intern wrote and tested the pair-building scripts in a CPU sandbox, then ran them as GPU jobs, recording the author and license of every source photo along the way. It built 6,042 training pairs and a 160-pair test set, where 40 of the test pairs come from 23 object classes kept out of training entirely. Before training, it measured the base model on its own and with the Viggle turbo LoRA, and checked that running edits in batches produced identical images, which made the evaluation cheaper. Training ran 2,000 steps in 1 hour 38 minutes on one A100 (~USD 4), and a comparison of the saved checkpoints on 48 test pairs picked step 500. Paired with the Viggle turbo LoRA at 6 steps the LoRA performed really well. 67.5% of objects detected where they were drawn, at 4.7 seconds per edit. Objects from the 23 unseen classes landed as reliably as the rest (65.0% versus 64.2%). The project took a little over a day and 59 jobs. Total compute cost, about USD 24. Check out: Model · Dataset · Doodle-in App - The Pocket Rewriter from the top of this post is the first one. ML-intern started by generating 8,797 short image requests with a small instruct model through Inference Providers, following a mix set in my prompt: photos, posters, logos, infographics and more, about a third of them asking for exact text in quotes, and many in languages other than English. The 9B teacher then rewrote all of them on one A100 in 2 hours 37 minutes (~USD 6.50). After filtering for quality, 1,840 examples were selected for training dataset. Training the 0.8B and 2B students took 12 and 18 minutes on an A10G (USD 0.75 for both). The 0.8B also ships as an 812 MB GGUF file for running on a CPU. The project took about 11 hours and 24 jobs. Total compute cost, about USD 16. Check out: Pocket rewriter 0.8B Student · 2B Student · Dataset · Pocket Studio App · Compare the teacher-student in Rewriter Arena - Agate-Preview-002-4step is the second. Logolabs' Agate Preview 002 is a 260M-parameter text-to-image model, small enough for a browser, but it needs 50 steps with guidance, which is 100 network passes per image. I asked ML-intern to distill it down to just 4 passes! It took two runs. The first cached 155,000 training images as latents, baked the guidance into the model, then cut the step count in stages from 16 to 8 to 4, all on A100s. The 4-step student beat the teacher run at the same 4 steps on GenEval and FID, after that ML-intern exported it to ONNX for the browser. This run took about 13 hours and USD 22. I did a second training run by asking ML-Intern to improve the 4-step student a bit more. It then made 24,000 more image pairs with the teacher at 16 steps and fine-tuned the 4-step student against them for about an hour. GenEval went from 0.509 to 0.536, against the teacher's 0.563 at 50 steps, with just 4 steps (one-fourth the compute). ML-intern re-exported the browser version. The second run took about 8 hours. Total compute cost across both runs, about USD 37. Check out: Agate 4-step model · Dataset · Run Agate in your browser · Agate 4-step LIVE | Model | Base model | Compute | |---|---|---| | Citrus Doctor | Qwen3.5-2B | USD 1.90 | | Huggy LoRA | FLUX.2 klein base 4B | USD 7.60 | | Pocket rewriter (0.8B and 2B) | Qwen3.5-0.8B and 2B | USD 16.05 | | Viewpoint Orbit LoRA | Qwen-Image 2.1 | USD 16 | | Doodle-in LoRA | Qwen-Image 2.1 | USD 24.30 | | Agate 4-step (two runs) | Agate Preview 002 | USD 37 | | Total | | about USD 103 | These are the GPU and CPU job charges reported for each session. ML-intern is in HuggingChat. Switch on ML-intern mode and paste a prompt. If you want a starting point, my example prompts are free to copy. Start from a model you wish existed and a dataset you have, or one you can describe. Give it a small budget, ask for a baseline and a smoke test, and read what comes back before you raise the cap. If you make something with it, share it on X and tag @Gradio and @HuggingFace. We would love to see the models you make that nobody else would think to build for you.
01:29

Medical company announces program for 'AI-powered prescriptions' in Utah - ABC News

A Utah clinic wants software to write the first prescription. The company says it will use artificial intelligence to issue initial prescriptions, including acne treatment. The stub mentions an Artificial Intelligence Policy and does not name the firm, the drug list, or the oversight rule.

Full text · 146 chars
... artificial intelligence to issue initial prescriptions, including acne treatment. ... Artificial Intelligence Policy. The company says the ...
02:20

Harvard Weighs $25 Million Investment to Double University-Wide AI Computing Capacity

Harvard is looking at a large compute bill to keep campus models from fighting for GPUs. The Crimson says the university is weighing a $25 million investment to roughly double university-wide AI computing capacity, per Bernardo L. Sabatini ’91, who co-directs the institute. The stub cuts off before the rest of his estimate.

Full text · 150 chars
... Artificial Intelligence , according to Bernardo L. Sabatini '91, who co-directs the institute. Sabatini estimated the investment could roughly ...
04:00

Tokka-Bench: Evaluating Tokenizers Across 100 Natural and 20 Programming Languages

Tokenizer bake-offs have lacked a shared scoreboard across writing systems. Tokka-Bench scores seven BPE tokenizers on five metrics over 100 natural languages and 20 programming languages. Vocabulary allocation mattered more than raw size. Programming-language efficiency has converged among recent tokenizers even as natural-language profiles still diverge. Compared: GPT-2, GPT-4, gpt-oss, Llama 3.1, Gemma 3, Qwen3, and Kimi K2. Framework, data, and a dashboard are public.

Full text · 1,753 chars
Computer Science > Computation and Language Title:Tokka-Bench: Evaluating Tokenizers Across 100 Natural and 20 Programming Languages View PDF HTML (experimental) Abstract:Large language models rely on subword tokenizers whose quality varies across languages, yet no standardized multi-metric framework exists for broad comparative evaluation. We introduce Tokka-Bench, an open-source framework that evaluates tokenizers on five complementary metrics -- bytes per token, unique token coverage, subword fertility, word-split rate, and vocabulary composition -- across 100 natural languages (30+ scripts) and 20 programming languages, using language-aware segmentation adapted to each writing system. Comparing seven BPE tokenizers (GPT-2, GPT-4, gpt-oss, Llama 3.1, Gemma 3, Qwen3, and Kimi K2) within individual languages, we find that vocabulary allocation strategy matters more than raw vocabulary size, and that programming-language efficiency has converged among recent tokenizers despite divergent natural-language profiles. The framework, data, and interactive dashboard are publicly available. Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
04:00

Child ASR Adaptation with Adult Retention: An Empirical Study

Teaching a speech model to hear children often makes it worse on adults. The study compares full fine-tuning, LoRA, and weight-space merging on Arabic and English child speech, including non-native Arabic and MyST, against MGB-2 and LibriSpeech test-clean. Child adaptation is necessary, especially for non-native kids, but direct adaptation hurts adult scores. Bilingual adaptation is more stable; LERP keeps adults, TIES recovers more child gain. Encoder–decoder still wins raw WER with direct bilingual fine-tuning.

Full text · 2,113 chars
Computer Science > Computation and Language Title:Child ASR Adaptation with Adult Retention: An Empirical Study View PDF HTML (experimental) Abstract:Automatic Speech Recognition (ASR) systems often underperform for children and non-native speakers, while adapting adult ASR models to child speech can cause adult-speech forgetting. We study child ASR adaptation with adult retention across Arabic and English. We compare full fine-tuning, LoRA, and post-hoc weight-space merging across encoder--decoder, encoder--CTC, and AudioLLM-based ASR systems. Experiments use Arabic native and non-native child speech, English MyST child speech, and adult benchmarks from MGB-2 and LibriSpeech test-clean. We evaluate recognition quality with WER and quantify the adaptation--retention trade-off using Retention Index, Child Adaptation Gain, and Adaptation Recovery. Results show that child adaptation is necessary, especially for non-native Arabic and English child speech, but direct adaptation often reduces adult ASR performance. Bilingual adaptation is more stable than language-specific adaptation. Weight-space merging often improves the trade-off, especially for encoder--CTC, Whisper, and AudioLLM-based ASR, with LERP favoring adult retention and TIES recovering stronger child gains. For the encoder--decoder model, direct bilingual fine-tuning remains strongest in raw WER.\footnote{Code, and models are available at this https URL. Current browse context: Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
04:00

When Forgetting Looks Like Improvement: Metric Masking in Streaming Diarizer Adaptation and the Price of Rehearsal

A speech diarizer can look better on paper while forgetting who is speaking. A released streaming system adapted on 7.5 hours of two-party talk improved in-domain scores and transferred to one other corpus. The extra errors were mostly broken identity-over-time, not wrong speaker counts. Rehearsal reduced that forgetting and also reduced the cross-domain transfer. The authors want detection, identity consistency, and retention scored together.

Full text · 1,862 chars
Computer Science > Computation and Language Title:When Forgetting Looks Like Improvement: Metric Masking in Streaming Diarizer Adaptation and the Price of Rehearsal View PDF HTML (experimental) Abstract:Small-data adaptation can improve speech detection while degrading speaker attribution. We study this discrepancy in a released streaming diarizer adapted on 7.5 h of two-party conversation and evaluated across six corpora. Adaptation substantially improves in-domain diarization performance and transfers to an independent corpus. However, this improvement is not consistent across evaluation scenarios as the additional confusion is mainly associated with impaired temporal identity consistency rather than speaker-count errors. A local-remapping diagnostic reveals different patterns of identity degradation across corpora, indicating that adaptation may alter how streaming models maintain speaker assignments over time. Rehearsal reduces the observed degradation but reduces the cross-domain transfer performance. These results highlight the need to jointly evaluate detection accuracy, identity consistency, and retention behavior when adapting streaming diarization systems. Current browse context: Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
04:00

Emo-Jev: Probabilistic Reasoning for Emotion Classification with Jev

A decision model is being asked to classify feelings without writing an essay. Emo-Jev is training-free: one variant decomposes the label into atomic yes/no judgments, the other votes across several judgment paths. On eight sentiment, emotion, sarcasm, and humor sets, plain Jev scored 62.93% average macro-F1 versus 67.28% for the strongest of five LLM baselines. Latency and cost were lower. The paper does not claim it beat the best chain-of-thought LLM.

Full text · 1,855 chars
Computer Science > Computation and Language Title:Emo-Jev: Probabilistic Reasoning for Emotion Classification with Jev View PDF HTML (experimental) Abstract:Jev offers an alternative interface for language understanding: given an input and predefined questions, it returns probabilistic decisions rather than free-form responses. Whether this interface can support effective reasoning for text classification against leading LLMs remains an open questions. We introduce Emo-Jev, a training-free framework with two complementary implementations. Emo-Jev-D decomposes classification into task-specific atomic judgments and composes their probabilities into a final prediction. Emo-Jev-SC constructs multiple judgment paths from complementary perspectives and aggregates their predictions into a consensus decision. We evaluate Emo-Jev on eight datasets spanning sentiment analysis, emotion recognition, sarcasm detection and humor detection, comparing against direct Jev classification and five SoTA LLMs under input/output and chain-of-thought reasoning. Standard Jev achieves 62.93\% average macro-F1 versus 67.28\% for the strongest LLM baseline, with lower observed latency and generally lower cost. Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
04:00

CoDR: Training-Free Confidence-Drift Remasking for Diffusion Language Models

Diffusion language models lock in early guesses and rarely take them back. CoDR is a training-free remasking pass that measures confidence drift in k forward probes and regenerates only the tokens the model no longer stands behind. Across two backbones, four reasoning and coding tasks, and three samplers, it raised average accuracy in every model-sampler pair. Gains came from targeted remasking, not extra compute, and used fewer forwards than prior remasking methods.

Full text · 2,169 chars
Computer Science > Computation and Language Title:CoDR: Training-Free Confidence-Drift Remasking for Diffusion Language Models View PDF HTML (experimental) Abstract:Masked diffusion language models (MDLMs) decode by repeatedly committing tokens to masked positions, but these commitments are usually irreversible. A token chosen under sparse, partial context is kept fixed, even when later context no longer supports it. Existing samplers mainly decide when to commit a token, but rarely check whether an already committed token should still be kept, allowing early mistakes to propagate. We trace this issue to confidence drift, where the model's confidence in a committed token drops from its sparse commit-time context to the denser context available later. Based on this signal, we propose CoDR (Confidence Drift Remasking), a training-free and sampler-agnostic refinement pass. CoDR estimates drift for all committed positions in only k forward passes via k-partition probing, then remasks and regenerates only the tokens the model no longer endorses. Across two backbones, four reasoning and coding tasks, and three base samplers, CoDR improves average accuracy across all evaluated model-sampler configurations and improves most individual task settings with modest overhead. Controlled experiments show that the gains come from targeted confidence-drift remasking rather than extra compute alone, and that CoDR uses far fewer forward passes than prior remasking methods. Code is available at this https URL. Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
04:00

Leveraging LLM-Generated Explanations for Detecting Emotionally Rewritten Fake News

Fake-news detectors stumble when the same claim is rewritten in a different mood. The authors build emotion-rewritten tests and keep explanations generated from the original article as stable background. A gated cross-attention model mixes the rewritten text with those explanations. They report notable gains on PolitiFact and LUN under several emotions, and only competitive numbers on GossipCop. Code and data are promised at a project URL in the abstract.

Full text · 2,129 chars
Computer Science > Computation and Language Title:Leveraging LLM-Generated Explanations for Detecting Emotionally Rewritten Fake News View PDF HTML (experimental) Abstract:The spread of fake news may cause severe social consequences. Existing fake news detection methods mainly focus on stylistic variations or incorporate external information such as explanations. However, news articles are often rewritten under different emotional backgrounds while preserving their underlying factual claims, which may affect the robustness of detection models. In this work, we investigate fake news detec- tion under fact-preserving emotional variations. To study this problem, we construct emotion-rewritten test sets and generate explanations from the original news articles as stable background knowledge. We then propose a Gated Cross Attention (GCA) framework that adaptively integrates emotionally rewritten news with the corresponding explanations, enabling the model to focus on informative explanation content while reducing potential mismatches caused by emotional reframing. Experiments on PolitiFact, GossipCop, and LUN demonstrate that the proposed method achieves notable improvements under multiple emotional conditions on PolitiFact and LUN, while maintaining competitive performance on GossipCop. We further analyze the effects of explanation guidance and gating mechanisms under different emotional conditions. Our code and data are available at: this https URL gca . Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
04:00

Beyond Risk Prediction: Evidence Grounding and Psychosocial Factor Verification for Explainable Suicide Risk Assessment

A suicide-risk detector is being asked to show its work, not just a score. The IEEE BigData 2026 challenge setup adds evidence phrases and psychosocial-factor checks on social posts. Length-based routing handles short versus long posts; a Risk-Evidence constraint keeps quotes aligned with the label. Weighted F1 is 0.8088 for risk, 0.7605 for evidence, 0.5562 for factors. Factor identification is the weak step.

Full text · 2,425 chars
Computer Science > Computation and Language Title:Beyond Risk Prediction: Evidence Grounding and Psychosocial Factor Verification for Explainable Suicide Risk Assessment View PDF HTML (experimental) Abstract:Identifying suicide risk from social networking services (SNS) posts is important for detecting suicide-related signals in online environments. However, risk classification alone provides limited insight into the textual evidence and psychosocial factors behind a prediction. Based on the IEEE BigData 2026 Explainable Suicide Risk Detection Challenge, this study presents a framework consisting of Risk Assessment, Evidence Grounding, and Factor Identification. Risk Assessment uses length-based routing to accommodate posts of different lengths. Evidence Grounding identifies supporting phrases and uses a Risk-Evidence constraint to maintain consistency with the Risk prediction. For Factor Identification, two verifiers are used. The Taxonomy Verifier focuses on factor semantics, whereas the Evidence-Aware Verifier uses factor-specific lexical-semantic cues to select informative positive training units. Their prediction probabilities are combined to produce the final factor predictions. The three tasks are evaluated using task-specific F1 score measures. Risk Assessment achieved a Weighted F1 of 0.8088, Evidence Grounding achieved a test Macro row F1 of 0.7605, and Factor Identification achieved a Macro F1 of 0.5562. The results show that the framework can provide risk predictions, along with supporting textual evidence and fine-grained information on psychosocial factors. Overall, the proposed framework extends suicide-risk assessment beyond risk-level prediction and provides a more interpretable analysis of SNS posts. Current browse context: Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
04:00

LRCC: Generalizing Low-Rank Compression with Conditional Computation

Compressed models usually spend the same compute on every token. LRCC freezes nested low-rank paths and trains one small router per Transformer block to pick a path per token. On Llama and Qwen, at the same average active-parameter budget, it beat static low-rank compression, including a 7.6-point gain in average downstream accuracy on Llama-2-7B. At matched batch-size-1 decode latency it improved perplexity and accuracy on Llama-3.2-1B and stayed competitive on Llama-2-7B without custom kernels.

Full text · 2,035 chars
Computer Science > Computation and Language Title:LRCC: Generalizing Low-Rank Compression with Conditional Computation View PDF HTML (experimental) Abstract:Low-rank compression reduces the cost of pretrained language models by replacing linear transformations with low-rank factorizations. However, conventional methods use a fixed rank allocation during inference, assigning the same amount of compute regardless of the input token. We introduce Low-Rank Conditional Computation (LRCC), which adds token-dependent computation to pretrained models by training one lightweight router per Transformer block to select among a small set of nested low-rank paths. During training, the low-rank factors remain frozen, and only the routers are optimized. We evaluate LRCC on Llama and Qwen models for language modeling and zero-shot downstream tasks. Within the same average active-parameter budget, LRCC improves the predictive performance over static low-rank compression, including a 7.6 percentage-point gain in average downstream accuracy on Llama-2-7B over static methods. At matched batch-size-1 decoding latency, LRCC improves both perplexity and downstream accuracy on Llama-3.2-1B and remains competitive on Llama-2-7B, without specialized kernels. Finally, we assess the usefulness of assigning a token-wise path by analyzing the routers' path choices. Current browse context: Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
04:00

Tiny-Scale Chinese BERT Pretraining: A Controlled Comparison of MLM, WWM, and MacBERT Strategies

The usual ranking of Chinese pretraining tricks flips when the model is tiny. Three 8.7-million-parameter, 4-layer Chinese BERTs were trained from scratch on 1.29 million Wikipedia sentences with MLM, whole-word masking, or MacBERT. MLM won 3 of 5 intrinsic tests. WWM had the best perplexity (1.27 vs 2.10, a 39.5% drop) and MLM hit rate (22% vs 16%). MacBERT with a 222-entry synonym list hit perplexity 47.23 — 22× worse than MLM — despite the lowest training loss (2.17).

Full text · 2,315 chars
Computer Science > Computation and Language Title:Tiny-Scale Chinese BERT Pretraining: A Controlled Comparison of MLM, WWM, and MacBERT Strategies View PDF HTML (experimental) Abstract:Pretraining strategies significantly impact the quality of language models, yet existing comparisons of Masked Language Modeling (MLM), Whole Word Masking (WWM), and MacBERT-style replacement have focused primarily on base-scale models (>=110M parameters). This paper presents a controlled comparison of these three strategies on a tiny-scale Chinese BERT model (4 layers, 256 hidden dimensions, 8.7M parameters). Under identical architecture, corpus (1.29M sentences from Chinese Wikipedia), and hyperparameters, we train three models from scratch and evaluate them across five intrinsic dimensions: perplexity, MLM hit rate, semantic discrimination, grammatical judgment, and contextual sensitivity. At tiny scale, MLM achieves the best overall intrinsic performance (winning 3 of 5 dimensions), while WWM excels in both perplexity (1.27 vs. 2.10, a 39.5% improvement) and MLM hit rate (22% vs. 16%). Notably, MacBERT under a severely limited synonym dictionary (222 entries, 3.3% coverage) exhibits severe perplexity degradation (47.23, 22x higher than MLM), yielding a ranking (MLM > WWM >> MacBERT) that differs markedly from the established base-scale conclusion (MacBERT > WWM > MLM). We further identify a critical evaluation pitfall: MacBERT achieves the lowest training loss (2.17) yet the highest perplexity (47.23), revealing that training loss alone is unreliable under mixed replacement strategies. All models and corpus are publicly available at this https URL. Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
04:00

Steering Follows Geometry, Not Labels: Emotion Directions in a Full-Duplex Speech Model

You can nudge a live voice model’s mood with a few added vectors, but the moods are not cleanly separate. The study steers Moshi, an open full-duplex speech model, across four emotions with mean-difference activation steering and no retraining. Emotion is linearly decodable from the residual stream. Happy, angry, and surprise share a direction; sad is distinctly steerable. You cannot simply subtract the shared component from all four equally.

Full text · 1,839 chars
Computer Science > Computation and Language Title:Steering Follows Geometry, Not Labels: Emotion Directions in a Full-Duplex Speech Model View PDF HTML (experimental) Abstract:Full-duplex voice agents need to modulate emotion and delivery during real-time conversations, when de-escalating a complaint, carrying urgency in dispatch, softening a clinical result. Emotion and delivery control is well studied for TTS and turn based models through prompt-conditioned synthesis, reference-conditioned synthesis and activation steering; PersonaPlex controls identity in a duplex model but not affect. We study emotion steering in Moshi, a fully open sourced full-duplex speech language model, across four emotions, using mean-difference activation steering, which costs only a few vector additions per frame and no retraining. We show that emotion is linearly decodable from Moshi's residual stream, but activation steering is only partially achievable, and unevenly so; as happy, angry and surprise steer towards a shared direction while sad is distinctly steerable. We also show that the shared component across the three emotions cannot simply be projected away from all the emotions equally. Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
04:00

Quad-State Safety Evaluation of Open-Weight Large Language Models on Non-Canonical Inputs

Safety tests written in plain text miss how people actually type. ASRD has 2,100 prompts in seven surface-form families; five open-weight models produced 10,500 replies scored as harmful, safe, incomprehension, or indeterminate. Emoji and invisible Unicode barely confuse the models (20.27% and 17.20% harmful vs a 22.87% baseline, mostly Mistral 7B). Leetspeak, encoded wrappers, and hybrids drop harmful rates to 2.40%, 0.13%, and 2.40% while comprehension failure jumps to 36.47%, 65.60%, and 34.47%.

Full text · 2,031 chars
Computer Science > Computation and Language Title:Quad-State Safety Evaluation of Open-Weight Large Language Models on Non-Canonical Inputs View PDF HTML (experimental) Abstract:Standard safety evaluations of large language models assess harmful requests written in canonical plain text, while models in real-world deployment routinely receive inputs containing emojis, altered spellings, encoded strings, and character-level variations. This work introduces the Adversarial Surface-Form Robustness Dataset (ASRD), comprising 2,100 prompts across seven distinct surface-form families. Five open-weight language models are evaluated across these prompts, producing 10,500 responses. The Quad-State Evaluation Rubric classifies each response into one of four outcomes: harmful compliance, safe response, comprehension failure, or indeterminate. Emoji and invisible Unicode variations cause almost no comprehension failure, with pooled harmful compliance of 20.27% and 17.20% against a 22.87% baseline that is driven mainly by Mistral 7B, whereas leetspeak, encoded wrappers, and hybrid transformations score 2.40%, 0.13%, and 2.40% while comprehension failure rises to 36.47%, 65.60%, and 34.47%. Inspection of raw model outputs reveals three response behaviors: hallucinated benignity, structural collapse, and language drift. Project page: this http URL Current browse context: Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
05:28

CommBank tops Asia-Pacific AI ranking, fast-tracks engineer training | Australian Broker

A big bank is sending engineers to the lab that makes Claude. Commonwealth Bank ranks fourth globally for AI maturity and is sending 20 engineers to Anthropic’s Claude Frontier Academy as AI fraud risks rise. The stub does not define the ranking or the course. Asia-Pacific “tops” in the headline is not separately numbered in the body.

Full text · 148 chars
Commonwealth Bank ranks fourth globally for AI maturity and sends 20 engineers to Anthropic's Claude Frontier Academy as AI fraud risks rise for ...
06:47

AgentCraft Turns Minecraft Into a Live Claude Coding Agent Control Room

A video game room is being used as a live dashboard for coding agents. AgentCraft is an MIT-licensed Fabric mod plus a Node.js Foreman that runs Claude agents at desks inside Minecraft 26.3. A lead (Opus by default) plans; up to three Sonnet workers use isolated git worktrees, with no network git and no auto-merge. Three two-task goals on the sample repo cost about $6 at low effort. Windows only, singleplayer, Claude Agent SDK.

Notes
  • MIT-licensed Fabric mod + Node.js “Foreman.” Claude Agent SDK. Windows only, singleplayer, Minecraft 26.3.
  • Lead (Opus by default) inspects the repo, plans, reviews. Up to three workers (Sonnet by default) in isolated git worktrees. No git network access. No auto-merge. Agents commit under their own names.
  • Loop: type a goal at the studio console → lead pins tasks on a wall → workers take desks → monitors show files/lines → bell + podium when a human decision is needed → press J to answer, review the diff, or approve a merge.
  • Characters: Marlow plans/reviews; Juniper, Kit, Wren, Rowan, Tove work. Lamps, particles, cupola beacon instead of interleaved logs.
  • Sample: three two-task goals on the sample repo ≈ $6 total, Sonnet, low effort.
  • Paywalled after this AlphaSignal free preview.
Full text · 2,170 chars
- AgentCraft is an open-source Fabric mod running a team of Claude agents inside Minecraft. - A lead agent plans and splits goals; workers build in parallel, each in isolated git worktrees. - Agents walk over to a podium when they need a human decision, with in-game diff review before merge. - No git network access, no auto-merge, agents commit under their own name for clear authorship. - Three two-task goals on the sample repo cost about $6 total using sonnet with low effort. - Windows only, singleplayer, Minecraft 26.3, powered by the Claude Agent SDK. AgentCraft turns Minecraft into an agent control room AgentCraft is an MIT-licensed interface for supervising Claude coding agents inside a Minecraft studio. Each agent works at a desk, streams its activity to an in-game monitor, and walks over when it needs a decision. The system combines a Fabric mod, which provides the Minecraft interface, with a Node.js orchestrator called the Foreman. A lead agent uses Opus by default to inspect the repository, plan work, and review results. Up to three workers use Sonnet by default to implement tasks concurrently in separate Git worktrees. Walk through the agent loop A run begins when the user enters a goal at the studio console, after which the agents expose each stage of their work through the room: - The lead reads the repository, breaks the goal into tasks, and pins them to a wall. - Workers claim tasks and move to their desks. - Each monitor displays the files its agent reads and the lines it changes. - Questions requiring user judgment trigger a bell and send the relevant agent to the podium. - The user presses J to answer, review changes, or approve a merge. Six hand-pixelled characters represent the available roles: Marlow handles planning and reviews, while Juniper, Kit, Wren, Rowan, and Tove form the worker pool. Their movement, desk activity, particles, lamps, and a cupola beacon communicate current status without requiring the user to parse interleaved terminal logs. This story is for Pro members You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.
06:56

Trahan unveils AI liability discussion draft - Live Updates

A House Democrat wants model makers on the hook when their products cause harm. Rep. Lori Trahan unveiled a discussion draft that would make artificial intelligence developers liable when their products — the captured sentence ends there. It is a discussion draft, not a passed bill. No section numbers are in the stub.

Full text · 150 chars
Rep. Lori Trahan on Wednesday unveiled a discussion draft for a bill that would make artificial intelligence developers liable when their products ...
07:12

Microsoft brings more AI to PCs as it challenges Apple

Microsoft wants coding models to live on the laptop, not only in the cloud. It unveiled an AI coding model that can run directly on personal computers, plus new security tech “to prevent AI agents” — the sentence is cut off. The event also involved Nvidia’s CEO. No model name, size, or laptop SKU is in the stub.

Full text · 147 chars
Microsoft unveiled on Wednesday an AI coding model that can run directly on personal computers and new security technology to prevent AI agents ...
07:12

Willie Williams: The Company Agent Won, Personal AI Died at Work - BigGo Finance

Personal agents at work died of upkeep, not of a better chatbot. Willie Williams argues the company agent won; the captured line is that “the maintenance burden killed personal agents at work.” The stub also nods at OpenAI Dots without explaining them. No architecture diagram is in the capture.

Full text · 153 chars
... agent architecture, engineering management, and what OpenAI's new Dots actually do well today. The maintenance burden killed personal agents at work.
08:17

Building a safer path to autonomous industrial AI

Factory AI that can move real equipment needs a human still holding the brake. AVEVA’s Arti Garg says foundation models, physical AI, and agents are arriving in plants where a bad decision hits safety. She cites almost a 78% rise in industrial AI adoption over two years and a workforce where almost half retire in five years. SCG Chemicals’ early AVEVA predictive-analytics pilot is described as almost 9× ROI toward 99% reliability. The episode is sponsored Insights content, not editorial.

Notes

Sponsored Business Lab episode (Insights / AVEVA), not MIT Technology Review editorial. Guest: Arti Garg, chief technologist at AVEVA. Host: Megan Tatum.

The claim
  • Industrial AI is not new at AVEVA (“more than 20 years”). What changed: general-purpose foundation models, physical AI / robotics, and agentic software automation.
  • One cited study: “almost a 78% increase over the past just two years” in industrial AI adoption. Treat as Garg’s cited figure, not an independent check.
  • Digital AI can be wrong on a screen. Industrial AI can move pumps, grids, and mines. Garg: “How do we leverage these technologies while maintaining safety, while maintaining reliable operations, while still being able to deliver on the promises of the new capabilities?”
Data first
  • Core industrial problem: correlate telemetry, service logs, and original design docs. Newer graph databases + AI matching let an operator (or a robot) fetch that on a tablet instead of walking a hazardous floor.
  • Governance “triple mandate”: secure, efficient (including environmentally), and human safety / oversight. AI should “augment rather than replace people in critical decision loops.”
  • Newer models are “not fundamentally explainable” and “can change over time as they learn.” That is the new risk versus AVEVA’s older proprietary anomaly detector.
Numbers Garg states
  • “Almost half, not quite half, of the industrial workforce is set to retire in the next five years.”
  • SCG Chemicals (Thailand petrochemicals): AVEVA Predictive Analytics on top of cleaned ops + engineering data. Target “99% plant reliability.” Early pilot “almost a 9x ROI.” Unplanned downtime → planned downtime.
  • Idaho National Laboratory partnership: AI for grid resilience as rooftop solar and other intermittent sources make frequency and capacity harder to see.
  • IEEE P7100: Garg chairs a working group (launched “two years ago, a little over two years ago”) on measuring AI’s environmental impact — electricity/energy, resources, water, carbon. She says there is still “no agreed upon method.”
  • Closed-loop set-point pilots exist; guardrails are “you can’t go outside of a certain band” and “we only allow the automation in certain areas.” She says they are “still trying to figure that out.”
  • Next 18 months: domain experts using AI-assisted coding to build plant apps they could not write before.
“Autonomous systems, whether they're robots or drones, are really going to change the way that we work in plants, in power systems, on mining sites.” — Arti Garg

Caveat: produced by Insights, MIT Technology Review’s custom content arm, “in partnership with AVEVA.”

Full text · 31,986 chars
Sponsored In partnership withAVEVA Industrial AI is entering a new phase. After decades of predictive analytics and other specialized applications, advances in foundation models, physical AI, and agentic AI are making it possible to automate more complex tasks across industrial environments. But unlike AI that operates purely in the digital world, industrial AI can interact directly with physical systems, where an unexpected decision can have consequences for safety, reliability, and critical infrastructure. That makes responsible deployment central to the next wave of industrial automation. “How do we leverage these technologies while maintaining safety, while maintaining reliable operations, while still being able to deliver on the promises of the new capabilities?” asks Arti Garg, chief technologist at AVEVA. The challenge is particularly acute as newer AI systems become more capable but also harder to predict and explain. One foundation for making that transition work is data. Industrial systems often contain information across telemetry, service logs, engineering documents, and other disparate sources. Newer technologies can help connect and correlate that information more quickly, giving operators real-time support when diagnosing problems. AI-powered robots could take that a step further by gathering information in hazardous environments without requiring workers to enter them. But greater autonomy also requires new approaches to governance. AVEVA’s framework for responsible AI emphasizes security, efficiency, and human safety and oversight. Garg argues that AI should augment rather than replace people in critical decision loops, with guardrails determining where automated systems can act and where human supervisors remain responsible. Sustainability is another part of that equation. AI can help manage complex power systems as renewable generation grows, while organizations also need better ways to understand AI’s own environmental footprint. Garg is involved in an IEEE working group developing a standard methodology for measuring that impact across electricity, energy, resources, water, and carbon. The next phase could bring industrial AI further into the physical world, from autonomous robots and drones to AI-assisted coding that allows domain experts to build new applications. But realizing that potential will require more than deploying new technology, says Garg. Organizations will need to rethink business processes, establish appropriate safeguards, and give experienced workers new ways to apply their expertise, creating a model of automation that is not only more autonomous, but safer, more efficient, and more sustainable. “Autonomous systems, whether they're robots or drones, are really going to change the way that we work in plants, in power systems, on mining sites,” says Garg. "In a way, that will make these types of operations more efficient, much safer for the human beings involved and more productive.” This episode of Business Lab is produced in partnership with AVEVA. Full Transcript Megan Tatum: From MIT Technology Review, I'm Megan Tatum, and this is Business Lab, the show that helps business leaders make sense of new technologies coming out of the lab and into the marketplace. Industrial AI may not be new, but it is developing at a breakneck pace. From modern predictive analytics that can preempt equipment failures, to autonomous robots tasked with inspecting high risk machinery, emerging applications have the potential to transform operational efficiencies. But often used in high consequence environments with minimal margin for error, industrial AI deployed without clarity, security, or accountability could also have catastrophic consequences for critical infrastructure over time. As industries rush to embrace the latest technologies, the goal should always be responsible AI designed to support, rather than replace that all important human judgment. Two words for you: sustainable automation. My guest today is Arti Garg, chief technologist at AVEVA. This podcast is produced in partnership with AVEVA. Welcome, Arti. Arti Garg: Thank you, Megan. It's great to be here. Megan: Thank you so much for joining us. And just to start, if we could set some context for our discussion, can you share a bit about the current state of industrial AI and what AVEVA is working on at the moment as well? Arti: Yeah, more than happy to. And I want to go back to what you already stated, which is that in some ways, industrial AI is not new. It's something that, even here at AVEVA, we've been working on for more than 20 years, really thinking about how AI can enhance applications in the industrial sector. But what has changed quite substantially in maybe the last few years is the type of AI. We've really moved to these general purpose foundation models, which enable a lot more people to leverage advanced AI technologies. In addition to that, in just the last couple of years, there's been an explosion in what's known as physical AI and the types of AI that can really accelerate the capabilities of robotics and autonomous systems. And also agentic AI, which also allows for automation at more of the software level. And what we're seeing is that while in the past there has been some hesitancy to adopt models that don't totally have predictable outcomes, there's been a rapid acceleration in adoption of AI even in the industrial sector. One study suggests maybe that's almost a 78% increase over the past just two years within the industrial sector. So yes, that's an inflection point, but it's almost like a step function if you think about it with respect to AI adoption. And what we're really thinking about now at AVEVA is, how do we think about this? How do we think about incorporating these new and very powerful AI technologies in environments where customers have mission-critical operations, human safety is a significant factor? Those are some of the primary concerns that our customers have when they think about adopting any new technology. We're doing that in narrow where these newer types of AI aren't necessarily predictable, in terms of how they behave. That's one of the areas we're really thinking about is, how do we leverage these technologies while maintaining safety, while maintaining reliable operations, while still being able to deliver on the promises of the new capabilities? Megan: There's so many things to consider at once, isn't it? And you outlined there some of the huge developments we've seen in this space in recent years. What makes this such a pivotal moment for industrial AI specifically? What has changed to make some of those perhaps more out there hypotheticals much more plausible? Arti: Well, I think one of the big things, one of the big challenges within industrial AI, and this is something that as AVEVA, it's in our blood, it's in our core DNA, is that to really leverage any kind of digital technology in the industrial space, one of the first important things you often need to do is gather and correlate data from disparate systems, so that maybe I have telemetry coming off of a pump or a mixer, and I want to be able to correlate that with service logs so I know maybe the last time that someone had done some maintenance on that pump or mixer. And maybe I also want to be able to correlate that with the original design documentation and some of the original documentation on how to maintain that, so that if something is happening, I can do a better job diagnosing what's going wrong and how to address it. This has always been a core challenge in the industrial space. I think that what's happening now is that with newer technologies, even things like graph databases and being able to leverage AI to really match disparate data sets much faster, we're able to get to the point where if something goes wrong, instead of an operator maybe needing to spend a little bit of time doing this research to really understand how do these different components fit together of information, they may have just an iPad with an AI that can very quickly go and fetch that information, correlate it and help in real time, almost, diagnose what is happening. Even a little bit more out there, but not as out there as you might think is, I still describe this with a human operator with potentially a mobile tablet, but imagine now I can also have a robot moving through that environment also gathering data and either with onboard computation or a quick connection to an operator who doesn't necessarily now have to go out into a dangerous, potentially somewhat hazardous space, be able to gather data and enable that very quick diagnostic as well. And I think that's really what's exciting right now and what all of these new technologies that I've talked about are helping to enable. And I think then from our perspective, trying to really leverage all of those means that we have to think a lot now about our governance approach to this. How do we leverage AI in a way that keeps it contained, if you will? Megan: Yeah, absolutely. And on that governance point, how does AVEVA define responsible AI and what do you see as the primary benefits? Arti: Yeah, so I think for us, we think about a triple mandate around responsible AI, and that means that it's secure, it's efficient, and that includes environmentally efficient, and then it really preserves human safety and human oversight above all. And for us, that's what some of our core pillars of responsible AI include. And for us, that means we think about internally as we build tools, a multi-layered governance approach where we've got both for how we use AI as a company and then how we deploy AI into our products. We've got a similar framework that we leverage across both, a kind of joint governance model, both around how we're leveraging AI and also how we're infusing AI into our products to allow our customers to leverage AI in their operations. But I would say one of the core foundational principles underneath all of these is that human beings still remain centered in how we think about how AI is leveraged. And so human judgment, human responsibility, human ethics are still core to how we think about where AI can add benefits, and we think about it more as augmenting rather than replacing human beings in critical decision loops, if you will. Megan: Right, which is such an important distinction, isn't it? And in terms of the next phase of development, there's growing sentiment that agentic AI will soon allow these industrial systems to operate much more autonomously. I wonder, what makes the risks in industrial settings different from AI risk in a purely digital context? And I suppose on the flip side, what are the benefits and opportunities of that same software defined automation? Arti: I think we can start with the risks and then go to the benefits. I think ultimately, I would say the biggest risk in the industrial setting is that we are interacting with real physical systems, often physical systems that are very capable of delivering important outcomes that keep our world running, whether that's delivering power, whether that's mining for natural resources. But usually. Those physical systems are also operating in hazardous environments, the equipment itself is capable of also having human safety implications. Anytime you're interfacing software with a physical system that can have real world consequences, you have to be extra careful. And we've always had that, again, we've had that in our DNA at AVEVA from the beginning, being mindful of that end user and that end application, which is not just something on a computer screen. Even as we've deployed AI over the past few decades into our systems, we've typically leaned toward really trying to pick the best fit model, really understand that model's behavior if we're going to use, for example, we have a proprietary anomaly detection model that works in a lot of operational environments. Now, as we're thinking about some of these newer AI capabilities, which by design are hard to understand how they behave, they're not fundamentally explainable in the way we've thought about even AI or statistical models in the past, and actually their behavior can change over time as they learn and get tuned to new capabilities. That's really the potential risk of automating a physical system based on capabilities that can evolve over time is one of the things that we have to be really mindful of and really thoughtful around, where are we willing to put some of that increased automation into practice? But at the same time, I think there's a real opportunity there, because I glossed over this idea that, okay, well, models evolve, but so do human beings, and sometimes that's a good thing. We talk about humans having expertise and they gain understanding over their careers of how a system works. If we can actually leverage the ability of some of these newer AI capabilities, these sort of reasoning models often have underlying agentic capabilities, to learn and gather experience faster, and potentially even more importantly or more valuably, take experience that's learned at one site and apply it to another site, then that's a real opportunity to take what we already know works for human beings, which is that sometimes you just have to learn by doing and be able to apply that and scale that through AI. That's where I see potentially a real opportunity in this space. One of the things that we're really aware of in the industrial sector is just the way that our workforce is changing. I think that almost half, not quite half, of the industrial workforce is set to retire in the next five years, and that's a lot of expertise and experience that means that we're going to lose in the sector. If there's ways to make sure that we can capture that in ways that are actionable, in ways that can also help a newer generation of workers that are used to experience things in a different way, apply that expertise, apply that knowledge, then I think that's a huge opportunity within the industrial space. Megan: Yeah, absolutely. Clearly, some huge opportunities there particularly against the backdrop of other market and workforce changes as you've outlined there. I suppose building on that, what is the potential for these autonomous industrial AI systems when built and deployed responsibly to facilitate even faster, more sustainable industrial processes? Arti: I think there's a couple different ways to think about this. One is, what are the demands of potentially more environmentally sustainable processes? I'll use one example of something that my team has been working on in partnership with Idaho National Laboratory here in the U.S. as part of their testing for AI grid resilience project. One of the challenges in the electric grid as we're moving toward a future where we have a lot more intermittent renewable power generation sources, often distributed rooftop solar for example, is it's becoming much more challenging to manage the grid both from really understanding where electric capacity is coming in, electric load is pulling off of the grid. In addition to that, being able to maintain just power quality because instead of having one spinning asset delivering the frequency of electricity that's going across the grid, you've got a lot of smaller systems. AI is actually quite critical for being able to manage that grid of the future. And this is one of the things that we've been working in partnership is, how can we do that? How can we help grid operators better understand what's happening across their system, identify where things might be behaving anomalously so they can detect that early and then remediate for it? At the same time though, I always talk about that's an example of where AI can potentially help us accelerate the transition to a lower carbon energy future. At the same time, I think there's a lot of conversation around the sustainability of AI itself. What are the power requirements? What are the resource requirements needed to run AI? And before I get into how do we think about that at AVEVA, because it is definitely something we think about and I think a lot of actors in the space are thinking about, I want to point out, is one of the challenges is that there's really no agreed upon method to measure the environmental impact of AI. You see a lot of these stories, like one query on some kind of chat interface is X amount of gallons of water or this amount of electricity use. But the truth is that community-wide, there's no agreed upon standard. One of the things that I'm involved with is a standards working group that was launched by the IEEE two years ago, a little over two years ago. I'm the chair of that working group, it's called the P7100 Standards Working Group on measuring the environmental impact of AI. And one of the things we're trying to do is really identify all the different areas over which we want to think about environmental sustainability associated with AI, and then having one standard methodology that works and that can be adopted both from a reporting perspective and potentially also from an oversight or regulatory perspective. That becomes really important for any sort of real understanding of how AI impacts the environment is just we have to know, what does it use? We want to look holistically at that. Our standards cover electricity and energy consumption, it covers resource usage, it covers water consumption, and it also covers carbon. But all of that said, I think it's important for us to understand the footprint of AI, but I think it's also important for us to simultaneously recognize that a bigger model uses more compute and more compute probably leverages more resources. The more that we can be intentional and pick the right size model for the right application, the more we can already start down that path of being more environmentally efficient in the AI that we run. One of the potentially nice side effects of that is the more purpose-built a model, the less likely it is to misbehave in unpredictable ways, if you choose correctly your architecture. There's some sort of ancillary benefits that go beyond environmental sustainability. Megan: Fascinating. It's really, really interesting to hear some of the work going on behind the scenes there, because I think we'll all have heard some of the statistics around the environmental impact, as you say. So, it's great to understand some of the work going on there to really clarify that. And we've talked a little bit about AI augmenting humans earlier. As systems gain greater autonomy, how should organizations think about that balance between closed loop automation, human in the loop accountability? What guardrails need to be in place? And how are governments and cross-border actors approaching those concerns as well? Arti: There's a lot packed into that question, so I'll try to answer it somewhat one by one. First, starting with that sort of balance between closed loop and human in the loop, automation and accountability. One of the things that we talk about in the industrial sector is from human operator to human supervisor of systems, of industrial systems. And if I'm honest about that, we're still trying to figure that out. When a human's in a loop in an automated system, you're really thinking about a system may process all the data and then make a recommendation. We've got a solution that we've been working on more jointly with a few different customers where we're able to take in a lot of their operational data, also some simulations of how their systems work, and be able to provide, say, recommendations around, you should now operate at this set point instead of that set point based on some of the other conditions that are changing in your plant. There's a lot of interest in moving from having that be a recommended set point to have an automation that can automatically adjust the set point. And we've actually had some successful real-world pilots around that as well. But when you get into that, obviously it makes people very nervous if you're changing set points on industrial equipment. That's where you maybe continue to have some guardrails around like, you can't go outside of a certain band, for example, of operations, or we only allow the automation in certain areas. That's where the human supervisor, the same way a human supervisor might give employees a lot of bandwidth or a lot of flexibility to make decisions around certain things, around other things, they're hard and fast, like this is the deadline or this is the sort of production target. It's like that. So thinking about, how do you put the appropriate guardrails? The challenge is, AIs are not human beings, and so the guardrails look different, and I think that's one of the things that there's still to think about. That question of, what guardrails should be put in place? I think that it's going to be probably dependent a little bit on the application, but over time I think we're all going to learn, and so being really upfront and thoughtful about how we do this I think is important. But the opportunity though, potentially, is really great. I've already mentioned that AI can be very useful in synthesizing a lot of information and bringing to the forefront, this is what matters. That's something that we do want to create some space to experiment around, but what's then important is making sure that a human being understands the risk. I'll give a little bit of a personal story because it might be illustrative. This weekend, I bought a little toy robot and decided I wanted to program it to do some stuff, and I found AI very, very helpful to get me started to read all the documentation. Putting the robot together was straightforward, they had nice instructions, but there's a pretty heavy software developer kit that's already there, but going through and reading all that documentation can take a while. AI was super helpful to help me surface, this is the function that does this, to help me get started. But occasionally, it made really bad recommendations on the best way to troubleshoot something. That's where I think having the human being in the loop, being able to say, "I know that it's not going to be the best software architect or the best troubleshooter," is really helpful because I had the opportunity to leverage AI for what it's good for, synthesizing and surfacing a lot of information, but being able to say, "I don't think that's the most efficient first step. Let's try something else." That's where it's a really different way of working, but it's something that I think as humans get more comfortable with the power and limitations of AI, you can start to teach people how to work with it, and it's very different from how you would work with other digital tools. One of the things that is important when it comes to AI is recognizing just the breadth of things that it touches and the breadth of impacts that it's going to have, whether it's on the environment, as we've discussed, whether it's on productivity as sort of implicit in this entire discussion, whether it's on labor force, whether it's on just how economies work. From my eye, and I'm by no means a policy expert in this area, but from my eye, what I'm seeing is that different governments are prioritizing different aspects of that from how they're thinking about regulatory and other kind of governance approaches. Megan: We're absolutely seeing some really vastly different approaches to this, aren't we, around the world? I wondered to illustrate some of this, if you could share perhaps some case studies you've seen, maybe talking us through what lessons they could hold for other organizations perhaps a little earlier in their own AI journeys. Arti: I think a lot of the key areas where we're seeing very clear benefits of adopting AI in the industrial sector, the core to all of them is actually getting the data right and getting the right foundation of data so that you can build AI on top of it. One of our customers, SCG Chemicals, which is a petrochemical company in Thailand, they had this vision of producing a reliability platform for their operations that was as AI-driven as possible. But a core part of that was actually getting the data right, being able to have all of their information in the right place, leveraging some of AVEVA's tools to do that. Bringing together operational data and also engineering data. And then putting on top of that AI capabilities, in this case, one of our capabilities called AVEVA Predictive Analytics that deploys some proprietary models to do things like detect anomalies, so that they could really much earlier identify potential operational risks. And instead of having unplanned downtime, translate that to planned downtime. When you translate unplanned downtime to planned downtime, you can get a huge improvement in plant reliability. At this point they're targeting something like 99% plant reliability and a very, very high return on investment from the platform that they put in place. Again, in early pilot days, it was almost a 9x ROI, and so that's really quite impressive. But what I want to emphasize is that it's kind of a multi-layered problem to get it right. Megan: Those are some really striking results, though. I mean, for industrial leaders who perhaps are still hesitant to explore AI, I wonder, what do you think is the cost of taking a more wait and see approach? Arti: I think, again, this is an area where the industrial sector has a little bit of a different calculus to apply to this type of problem. In general, I think you would hear most business experts say that you can't afford to take a wait and see approach to AI because it is so transformative across every sector of society and economy, and certainly in the industrial space, we're not immune to that. But I do think that some of the challenges that I've outlined today also puts leaders in the industrial space into a mindset of sometimes potentially being a fast follower rather than the first adopter of newer technologies, just because the risk is so high. But that being said, I think one of the challenges is that the risks that we've covered today around infusing systems that don't always behave predictably with physical systems is that I don't know that there's a lot of other sectors that are going to be solving those problems. That's where what I'm seeing is actually a lot of interest and excitement in trying new things and willingness to do that because there's a recognition that applied correctly and applied with the right guardrails and safeguards in place, these technologies really have truly transformative potential from a resource usage standpoint, from a human safety standpoint, from a productivity standpoint. I think the calculus has shifted a little bit in the sector to we want to actually try these things out. But then it becomes more, how do we try these things out in environments that we can make a little bit more sandboxed or safe to test out what some of the unique challenges we're going to face in the industrial sector are? Megan: Yeah, absolutely. It's just hard to ignore the potential nowadays, isn't it? And just to close with a slightly future forward look, I suppose, what are you most enthusiastic about in terms of the long-term potential of responsible autonomous industrial AI systems? Arti: Yeah. Well, I think that there's a couple of different trajectories. One of the things that I think within the next 18 months for sure, we're going to see an increase of people with deep domain expertise who maybe didn't grow up with a software background, be able to adopt AI assisted coding techniques to really be able to build the things that they weren't able to build before. That's really exciting in my career, which way back when I was actually an industrial data scientist, I would say I always felt the most energized and that I learned the most talking to the people that were on the front lines of operations, having to monitor a lot of different equipment and understanding how it all worked together. They really have a lot of expertise, and being able to put in their hands the ability to very quickly develop new applications is I think going to really lead to a lot of new ideas that many of us in the industrial software space maybe wouldn't even have thought about and really understood how to put. I think that's really exciting. I think then looking beyond that, I was a little bit of a, not to say a robotic skeptic, because clearly robotics and autonomous systems are going to be important, especially given the types of environments, whether they're hazardous or remote, that industrial equipment operates in. But I just actually think things are moving a lot faster than I had anticipated. I'm really interested to see how these physical embodiments of AI systems start to really transform, again, how we think about operations. One of the things about any new technology in any environment is that to really gain value from it, you have to change the way you do things. I would say for close to 10 years now, I've been giving talks on AI adoption and I always say the biggest barrier to AI adoption is not anything to do with the technology. It's not even to do with the data, although data are often the biggest sticking point, it's really to do with, are you going to change your business processes to be able to work with the way this technology is good or not good at things? Autonomous systems, whether they're robots or drones, are really going to change the way that we work in plants, in power systems, on mining sites. And I think that will make these types of operations more efficient, much safer for the human beings involved and more productive. So, I'm quite excited about that. I don't know entirely what direction it will go, but one of the things that I think about a lot is that if you go back to I, Robot and Isaac Asimov's book, the premise of those was that robotics would be the first broadly adopted AI, not computer-based systems. We went in the other direction and I think it's really interesting to now see all of this converging. Megan: Yeah, absolutely. See robotics catches up a bit. So many exciting things on the horizon, that's for sure. Thank you so much, for your time. That was Arti Garg, chief technologist at AVEVA, whom I spoke with from Brighton in England. That's it for this episode of Business Lab. I'm your host, Megan Tatum. I'm a contributing editor at Insights, the custom publishing division of MIT Technology Review. We were founded in 1899 at the Massachusetts Institute of Technology, and you can find us in print, on the web, and at events each year around the world. For more information about us and the show, please check out our website at technologyreview.com. This show is available wherever you get your podcasts. And if you enjoyed us, we hope you'll take a moment to rate and review us. Business Lab is a production of MIT Technology Review, and this episode was produced by Giro Studios. Thank you so much for listening. Goodbye. Learn more at aveva.com. This content was produced by Insights, MIT Technology Review’s custom content arm, not its editorial staff. It was researched and written by humans, with any AI tools that may have been used limited to production processes under human oversight. Deep Dive Artificial intelligence Don’t be fooled—LLMs don’t reason Ten years after AlphaGo’s match against Go champion Lee Sedol, today’s AI still isn’t tapping into the machinery that made that win possible. AI’s recursive self-improvement might not come so quickly after all AI agents are not yet creative enough to carry out genuinely innovative open-ended AI research, it seems. Stay connected Get the latest updates from MIT Technology Review Discover special offers, top stories, upcoming events, and more.
08:42

AI Development on Windows: from PyTorch and llama.cpp to Windows ML

Windows now has one local inference stack it wants you to use instead of a pile of runtimes. Microsoft calls Windows ML the “unified, high-performance local AI inferencing framework for Windows.” An experimental Windows-native Runtime API is in preview. The stub does not list devices, models, or benchmarks.

Full text · 153 chars
Windows ML is the unified, high-performance local AI inferencing framework for Windows. Our experimental Windows-native Runtime API is now in preview ...
09:03

Coding Agents Broke Git's Scaling Math. GitHub Is Rebuilding to Keep Up - DevOps.com

Coding agents are committing so often that git’s old scaling math no longer holds. GitHub engineer Brian Celenza wrote that an agent in a tight loop “commits or checkpoints after nearly every action.” The captured snippet does not spell the rebuild plan. Treat the piece as a pointer, not a design doc.

Full text · 154 chars
AI coding agents don't work that way. “An agent in a tight loop commits or checkpoints after nearly every action,” GitHub engineer Brian Celenza wrote ...
10:00

Why we’re watching these climate tech companies

This year’s climate-tech watch list is really a data-center power list. Envision has installed over 100 GW of wind and over 50 GWh of storage. WeLion posted 824 Wh/kg in the lab — more than 3× typical lithium-ion. Fervo’s geothermal IPO raised $2.2 billion; X-energy is building a 320 MW helium SMR cluster with Amazon. Form Energy’s 30 GWh iron-air project and Energy Dome’s CO₂ stores are both tied to Google loads.

Notes

Casey Crownhart on MIT Technology Review’s 2026 Climate Tech Companies to Watch (The Spark). AI angle is data-center power, not models.

  • Envision Energy (second appearance): over 100 GW wind worldwide; over 50 GWh battery storage. Projects in Germany, Brazil, Australia; upcoming Vietnam.
  • WeLion New Energy: lab result last year of 824 Wh/kg — “more than three times” a standard lithium-ion pack. Lab results “notoriously difficult to scale.” Semi-solid cells are the nearer product.
  • Only three US firms on the list. Two boosted by Big Tech data-center demand. Third is Brimstone (critical minerals).
  • Fervo Energy: next-gen geothermal via horizontal drilling and fracking. IPO in May, $2.2 billion. Plans a gigawatt of plants by end of 2030.
  • X-energy: helium-cooled SMRs. 320 MW cluster with Amazon in Richland, Washington. Dow deal for a chemical plant in the early 2030s.
  • Form Energy (third appearance): iron-air long-duration storage. 30 GWh project to help power a Google data center; phases 2028–2031; “could be the world’s largest battery project in terms of capacity.”
  • Energy Dome: compressed CO₂ storage, cheaper than lithium-ion at “10 hours or more.” Plans over 30 GWh worldwide, including Google in Ireland in 2028.
  • Moment Energy: Vancouver facility that repurposes used EV packs; “million EV batteries” expected to retire by 2030.
Full text · 5,683 chars
This week, we released our 2026 version of our Climate Tech Companies to Watch list. It’s an annual project that the MIT Technology Review team puts together. Our goal is to highlight some of the promising, interesting advances and firms we think are worth paying attention to in the world of climate and energy technology. The list is always a little ray of hope for me—there are some fascinating ideas out there for changing how we power our world or moving away from fossil fuels to reduce the greenhouse-gas emissions that cause climate change. Our chosen companies tend to reflect several aspects of the current moment. Here are some of the companies we picked and the key themes worth highlighting. China is the center of the climate and energy tech world Forgive me—I know I’ve made this point before. Feel free to call me Captain Obvious. But I think China’s dominance in energy technology is a global trend that bears repeated mentioning. China is home to the world’s largest grid, and renewables are coming online more quickly than I can wrap my head around. Contributing to that boom is Envision Energy, a renewables giant appearing on our list for the second time. The company has installed over 100 gigawatts of wind power worldwide, as well as over 50 gigawatt-hours of battery storage. Like many other Chinese companies right now, Envision is looking beyond the crowded Chinese market and setting its sights abroad, with projects in Germany, Brazil, and Australia and upcoming work in Vietnam. China is also leading innovation in key technologies. I’ve been fascinated by the potential for next-generation batteries coming out of its companies: All the big ones, along with a host of smaller ones, are racing to develop solid-state batteries. They could be both safer than existing EV batteries and longer in range. WeLion New Energy is particularly interesting. The battery company announced a big lab result last year, achieving 824 watt-hours per kilogram—more than three times the energy density of a standard lithium-ion battery today. Lab results in battery research are notoriously difficult to scale, but WeLion is also working to deploy semi-solid-state batteries, which could come to fruition faster than purely solid-state cells while still improving on lithium-ion’s performance. Everybody wants a piece of the data center pie We included just three companies from the US on the list, and two of them have received a major boost from Big Tech as companies look to power data centers. (The third, Brimstone, is leaning into another political hot topic, critical minerals.) Technologies that can provide consistent power are in high demand. Fervo Energy is bringing next-generation geothermal power to the grid by using horizontal drilling and hydraulic fracturing to make it practical in more places. The company went public in May, raising $2.2 billion in its IPO, and plans to have a gigawatt’s worth of plants operational by the end of 2030. X-energy is building helium-cooled small modular nuclear reactors. The reactors’ small size lends flexibility, and the company is currently working on a 320-megawatt cluster of them in collaboration with Amazon in Richland, Washington. (For what it’s worth, I’m interested in what this technology could do for industrial facilities—the company has a deal with Dow that could see it deployed at one of the company’s chemical plants in the early 2030s.) The grid needs support, and energy storage will be a key piece moving forward I know, I know—I haven’t stopped talking about batteries since I started writing this newsletter four years ago. But the energy storage market has surpassed even my expectations over that time. And as solar and wind quickly come online, they need storage to help them meet more of the grid’s demand. Form Energy is on the list for a third time for its iron-air batteries, which could enable cheaper long-duration energy storage. The company has a massive project in the works, to the tune of 30 gigawatt-hours, which will be used to help power a Google data center. It could be the world’s largest battery project in terms of capacity by the time it comes online, which is expected to happen in phases between 2028 and 2031. Taking another tack, Energy Dome is using compressed carbon dioxide to store energy on the grid. This technology is already cheaper than lithium-ion batteries for longer durations (think 10 hours or more), and the company is working to scale quickly. I covered the company in 2022, before it had its first commercial plant running. Now it has plans for over 30 gigawatt-hours’ worth of plants around the world, including one with Google in Ireland that’s slated to come online in 2028. And Moment Energy is working to use the million EV batteries that are expected to reach the end of their life by 2030. The company has a facility in Vancouver, British Columbia, that repurposes used batteries for backup power. This year hasn’t been all sunshine and rainbows for climate tech. But I’m hopeful that some of the companies on this list, and the technologies they build, will keep nudging us in the right direction. This article is from The Spark, MIT Technology Review’s weekly climate newsletter. To receive it in your inbox every Wednesday, sign up here. Deep Dive Climate change and energy Batteries just broke another record in the US Huge grid-scale batteries are thriving, but smaller residential systems have lagged. What’s behind this summer’s heat, and why 2027 could be worse El Niño? Climate change? All of the above? Stay connected Get the latest updates from MIT Technology Review Discover special offers, top stories, upcoming events, and more.
10:48

The Sequence Opinion - Issue 947: Jev and the Rise of Decision Models

Most in-app choices do not need a research committee. The Sequence argues that a support agent facing a double charge must first pick a workflow, decide if it needs more facts, and choose whether to escalate — a forest of small judgments under the chat. Jev-style decision models make those semantic calls cheap, with a useful uncertainty estimate. Amazon, Cloudflare, and OpenAI are now in that race.

Full text · 954 chars
An AI support agent receives a simple message: a customer was charged twice. Before composing its sympathetic reply, the software must identify the issue, select the right workflow, check whether more information is needed, and decide whether to escalate. The visible conversation sits on top of a small forest of decisions. Using a reasoning model for every branch can feel like convening a research committee to operate a traffic light. Sometimes the intersection deserves careful analysis. Most of the time, the application needs a quick judgment with a useful estimate of uncertainty. This is the opportunity behind Jev and the recent proliferation of decision models. Their appeal lies in making semantic judgments inexpensive enough to insert throughout software. The competition now includes Amazon, Cloudflare, and OpenAI, suggesting that the economics of these small decisions deserve as much attention as the capabilities of the largest models.
13:00

Text is so 2023

Personal agents are shipping faster than the permission model around them. Ben Tossell flags Hark’s one-tap actions, Tab at a $300 million stealth valuation, and an Every agent in Slack. Shane Mac’s agent posted personal bank balances into company Slack. Headlines also cover GPT-6 Intelligent UI for everyone, Haiku 5.5 at about 75% less than 4.5, Nano Banana 2.1, 700+ OpenAI math results, and Grok Bot using Claude Opus 5.5 as the primary model.

Full text · 3,756 chars
Hi folks, Personal agent season is still in full-swing. Hark (from Brett Adcock - Figure robots) launched which suggests tasks as one-tap action buttons. Tab came out of stealth at a $300M valuation. Every put an agent in Slack. Meta and Sierra want a standard for how these agents deal with businesses. My favourite take came from Sonia Baschez (who worked with me on Makerpad years ago!): Proactive has a downside, though. Shane Mac’s agent posted his personal bank balances in the company Slack, as him. Then said sorry 😂. A community note popped up to flag this may have been a fake post, but I’m not sure… I’m now carefully trying to think about which agents have access to what, and where that access is. So I’m using Executor (v2). It’s basically one app that you connect all your other apps to, and you can add your skills too. My whole setup was a bit of a mess so I’ve started spring cleaning and will let you know if/when I feel like I have a good setup. Headlines ChatGPT now answers with interfaces. GPT-6 and Intelligent UI are rolling out to everyone. Answers can include charts, maps, forms and small tools, like a bill splitter for dinner with friends. Free users get GPT-6 Luna from today. (tweet) Claude Haiku 5.5 costs ~75% less to run than Haiku 4.5. Anthropic pitches it as a subagent for Opus and Sonnet. Simon has great notes on the model. Max and Team plans also now get monthly API credits to build your own apps and agents. (tweet). Anthropic also expanded their Claude startup program, and people were getting accepted in minutes. Nano Banana 2.1 is Google’s new image model. Live in AI Studio. But ChatGPT images is still better. OpenAI released 700+ new maths results from an unreleased model. Someone built a tool to spend your spare tokens on open math problems. Grok Bot can search and monitor X now. Also, it now uses Claude Opus 5.5 as the primary model. It’ll still delegate tasks to other models based on the complexity. My feed - AgentHog - stop guessing what your users do: pipe product analytics into Claude and Codex so your agents learn from real user behavior.* - How OpenAI’s Dots took over Every: pulling info out of school emails and catching missed Slack messages. I, however, still can’t get Dots to do anything good at all. Currently diving more into Grok Bot (until a completely custom Pi personal agent comes out!) - Lauren (poteto) automated Cursor’s release and QA with Grok bot. - New ChatGPT plugin: Meetings. It takes notes and helps with follow-ups. Pro and Business, Mac app only. - Claude now sits in a sidebar in Google Docs, Sheets & Slides. You approve each edit. - Gamma 5 is a full rebuild, so your decks look less like everyone else’s. - Figma Agent is out of beta and on all paid plans. It now uses AI credits. - Playground - make a game from a prompt. A Google Labs experiment, US only. - Rhem - a robot for your parents. It makes calls and books appointments. - Someone rebuilt Adobe’s apps with Opus 5.5 and open-sourced them. - An open-source Clay alternative for people search inside Claude Code/Codex. - OptionAFK - Transcribe large audio/video files on your device. Private and fast. Try this Turn a podcast into an interactive explainer. Boris Cherny asked Opus 5.5 to make one for Acquired’s Home Depot episode (a great episode btw). Make an interactive website companion for [episode link or transcript]. Add a timeline, the key people and the big ideas, with illustrations. Here was the result: Afters and Theo reckons its way less… - Find me on X, Linkedin, or YouTube - Read about me and Ben’s Bites - 📷 thumbnail via @keshavatearth * sponsors who make this newsletter possible :) Wanna partner with us for the next quarter? Email us at shanice@bensbites.com or k@bensbites.com
13:33

Cohere Opens Ottawa Office to Chase C$240M Federal AI Deals

Full text · 4,871 chars
- Cohere is opening its first Ottawa office near Lansdowne Park, adding to Toronto and Montreal sites. - Company employs 800+ globally, with roughly 30 staff already in the Ottawa region. - Move follows an MOU with the federal government and a $240M compute commitment from Ottawa. - Shared Services Canada licensed Command A for CANChat, used by 11,500 public servants across 20 departments. - Up to 1,400 ISED employees are getting Cohere's North platform for drafting and task automation. - Cohere is reportedly raising $2B to $3B at a $20B valuation, a record for a Canadian startup. Cohere is opening its first Ottawa office on Bank Street near Lansdowne Park, giving the Toronto-based AI company a local base for its growing federal business. The company announcement describes Ottawa as Cohere’s fourth Canadian location, adding to operations that include Toronto and Montreal. Cohere employs more than 800 people worldwide, including nearly 30 already based in the national capital region. Federal deployments drive the move Cohere’s federal work now spans a C$339,000 software licence, a deployment for up to 1,400 departmental users, a broader agreement on government AI and an earlier C$240 million commitment toward the company’s computing needs. | Reported federal work involving Cohere | | | |---|---|---| | Customer or program | Scope | Reported amount | |---|---|---| | Shared Services Canada | One-year Command A licence for CANChat, with about 11,500 employees registered across 20 departments | C$339,000 | | Innovation, Science and Economic Development Canada | North access for up to 1,400 users, covering search, summarization, drafting, decision support and task automation | Not disclosed | | Federal compute support | Earlier commitment toward Cohere’s processing capacity | C$240 million | Command A is Cohere’s large language model for enterprise workloads. North is a secure workplace platform that connects models to an organization’s data and tools, allowing employees to search internal material, generate drafts and automate multistep tasks. A federal agreement established a framework for exploring Cohere’s technology in government operations. Federal deployments typically require procurement reviews, security assessments, data-residency controls and integration with identity and records systems. Ottawa-based engineers, solution architects and public-sector staff can handle that work alongside the departments deploying the products. A US$20 billion valuation in play Cohere is reportedly negotiating a US$2 billion to US$3 billion financing at a US$20 billion valuation. Those terms would nearly triple the US$7 billion valuation it reached in September 2025 and produce the largest reported funding round for a private Canadian startup. German conglomerate Schwarz Group plans to invest US$600 million through its Schwarz Digits technology arm. The company’s reported annual recurring revenue has reached roughly US$240 million, up from about US$100 million a year earlier. ARR annualizes the current value of recurring contracts. Cohere’s customer list includes Oracle, RBC, Bell, Dell and Saudi telecom operator stc. The build requirements behind the contracts - Design for jurisdictional control. Sovereign AI keeps models, data and infrastructure under a customer’s chosen legal and operational jurisdiction. Deployment plans need explicit terms for hosting, data residency, encryption keys, model updates and administrator access. - Expose governance controls. Regulated customers evaluate role-based access, audit logs, retention policies, citations, evaluation records and incident handling alongside model quality. - Make retrieval permission-aware. North depends on access to internal content, so its search and generation layers must preserve the permissions attached to each source document. - Plan for phased procurement. A cross-government licence and a department-level rollout can have different buyers, security reviews and success measures. Pilots need clear paths to production, support and renewal. - Measure use after registration. CANChat’s 11,500 figure counts sign-ups. Active users, repeat use, task completion, output quality, time saved and security incidents provide stronger evidence of adoption. - Staff implementation locally. Cohere is recruiting software engineers to build public-sector agents and North features, as well as solution architects with government experience and, preferably, security clearance. Agents are workflows that let a model call tools and complete multistep tasks. The next Ottawa metrics Renewal of the one-year CANChat licence, active-use figures, expansion beyond the current 20 departments, the disclosed value of the ISED deployment and growth in security-cleared hiring will indicate whether Cohere’s early federal access develops into sustained production use.
17:01

😸 WATCH: Can AI Help Cure Genetic Disease?

A CRISPR company is using AI to hunt smaller gene editors, not to declare diseases cured. Mammoth Biosciences’ Trevor Martin says models now search a universe of about 34 billion proteins and can propose new CAS-like designs. He thinks a sustained push could put liver and blood genetic diseases within reach in roughly 5 to 10 years. Personalized treatments are already possible if money is no object; delivery and a missing public map from preclinical experiments to human trials are the bottlenecks. The same Neuron issue tees a live Microsoft agentic-coding demo with Atomic.

Full text · 6,846 chars
😸 WATCH: Can AI Help Cure Genetic Disease? PLUS: Join us LIVE w/ Microsoft to learn agentic engineering. Welcome, humans. First up: we go LIVE in 5 minutes with Alex Lavaee from Microsoft Research’s Catalyst Lab, and he’s bringing a much better demo than “let’s talk about coding agents.” While we chat, an autonomous coding agent will try to build a 3D Subway Surfers-style game in the background. Alex is teaching test and verification engineering for agentic coding: how you test what an agent builds, verify what it actually did, and make the workflow reliable enough for real engineering. He’ll use Atomic, an open-source verifiable coding-agent runtime he works on, to run the experiment live. So instead of only talking about the future of coding agents, we’ll watch one work while Alex breaks down the prompts, tests, verification loops, and human judgment that make agentic workflows trustworthy. And yes, we’re very curious whether the game actually works by the end. Keep scrolling for more about this week’s awesome interview. THIS EPISODE WAS BROUGHT TO YOU BY… APEX-Agents: The AI Productivity Index for Agents See how the latest models rank for jobs like law, consulting, and investment banking. Mercor's APEX-Agents leaderboard evaluates frontier AI on long-horizon, multistep tasks across economically valuable work. Built with partners like Harvey, Ramp, and Cognition. Every model. Ranked by productivity. New on the pod this week… So Dr. Trevor Martin co-founded Mammoth Biosciences with Nobel laureate Jennifer Doudna and a simple-sounding goal that gets crazier the longer you think about it: …turn genetic diseases we manage for life into one-time treatments that permanently change the DNA causing them. Now, Mammoth is using AI to search a universe of roughly 34B proteins, engineer smaller CRISPR systems, and even generate entirely new protein designs. Trevor thinks sustained progress could get us to a world with no liver or blood genetic disease in the next 5 to 10 years. So yes, we asked him how close “AI cures disease” is to becoming an actual sentence and not a pitch deck. Here’s our favorite parts: - (34:49) Could some genetic diseases basically disappear? Trevor says a sustained push could put liver and blood genetic diseases within striking distance of elimination in roughly 5 to 10 years. - (18:29) AI is moving from searching nature to generating it: Mammoth can feed huge protein-sequence databases into models and ask them to propose entirely new CAS-like proteins, not merely tweak ones we already know. - (21:37) The data AI medicine is missing: We have mountains of sequencing and protein data. What we do not have is a clean public dataset connecting preclinical experiments to what actually happened in human trials. - (36:54) Personalized gene medicine is already possible, technically: Trevor says for some genetic diseases, you could build a personalized treatment today if money were no object. The real problem is making that scalable and affordable. - (49:23) And then we got to lab-grown mini-organs: Organoids can form structures that resemble parts of the brain, which is useful for testing and also an excellent way to make everyone at the table say, “wait, WHAT?” The part I kept coming back to: AI can already help biology search and design faster, but the hard part is still proving what survives contact with an actual human body. Why watch this? Because Trevor separates the real bottlenecks from the sci-fi ones. You’ll understand what CRISPR actually changes, where AI helps today, why clinical data is so valuable, and what still has to happen before “personalized medicine” scales beyond a handful of people. P.S. At 49:23, Trevor casually explains that neuron organoids can grow structures resembling a prefrontal cortex and hippocampus. He then references The Island (IYKYK). Biology is wild, right? Keep scrolling for a plain-English CRISPR walkthrough, the newest GPT-6 vs. Claude showdown, and three more episodes worth your time. Additional Resources: CRISPR, translated into human If the words “guide RNA” and “CAS protein” normally make your eyes glaze over, here’s the version Trevor gave us: - It started as a bacterial immune system. Bacteria keep short snippets of viral genetic material so they can recognize the same invader later. - The guide RNA is basically the address. It carries a sequence that tells the CAS protein which matching stretch of DNA to find. - The CAS protein is the tool. In the simplest gene-editing setup, it cuts the DNA at that location. The cell repairs the cut imperfectly, which can scramble and effectively switch off the targeted gene. - Delivery is where things get ugly. The liver is comparatively reachable. The brain and muscle are much harder, which is one reason Mammoth searches nature for smaller CRISPR systems that fit into different delivery vehicles. That is the real trick: the editor can be incredibly precise, but you still have to get it into the right cells. If you want to see what Mammoth is building around that problem, start here. UPCOMING EVENT: The Neuron IRL in San Francisco The Neuron is going IRL in San Francisco on November 18 starting at 4:30 PM PT. If you’re in town, come hang out with us for drinks, bites, and a live recording of The Neuron podcast. We’ll have The Neuron community together in one room, plus some special guests and plenty of AI conversation. Huge thanks to Slack and Alumni Ventures for helping us bring the night to life. It’s free, but space is limited, so RSVP and save your spot. 🎙️ In Case You Missed It… Four recent interviews and episodes we think you’ll love. 1. The “I’m Fine” problem with voice AI. TL;DW: Hume AI CEO Andrew Ettinger explains why voice systems can sound human while still missing the meaning hidden in tone, hesitation, frustration, background noise, dialect, and who is speaking. Turning speech into a transcript can flatten the very signals a good voice agent needs. Why you should watch: If you are building, buying, or simply talking to voice agents, this explains why “sounds natural” and “understands me” are two very different benchmarks. 2. Want to understand what AI infrastructure actually has to do? TL;DW: CoreWeave’s Chen Goldberg explains why modern AI infrastructure is no longer a pile of GPUs. Compute, networking, storage, cooling, security, and software increasingly have to behave like one enormous computer. Why you should watch: If agents are going to run longer, use more tools, and handle real work, the systems underneath them matter almost as much as the model. Have a topic you want to learn about? Request it here! Subscribe to our YouTube Channel for more! Subscribe on YouTube to help us bring in more builders, researchers, and guests who can teach you something useful about AI every week. Stay curious, The Neuron Team
20:16

Anthropic's Claude Science Fills the Sky's Missing UV Third

Full text · 7,490 chars
- Johns Hopkins astrophysicist Brice Ménard used Claude Science to produce the first complete UV sky map. - Roughly one-third of the sky had never been observed in UV because GALEX skipped bright-star regions to protect its detectors. - Claude orchestrated agents to download, calibrate and merge GALEX, Swift, FIMS/SPEAR, TD-1, Planck and Gaia datasets. - Missing regions were filled using inpainting trained on correlations with visible, infrared and radio wavelengths. - Blind validation showed reconstructions within about 10% of real UV measurements, nearly imperceptible to the eye. - Explore the interactive map with per-pixel measured-vs-predicted labels and uncertainties. Claude models the ultraviolet sky’s missing third Astronomers have all-sky maps in radio waves, infrared, visible light, X-rays and gamma rays. Ultraviolet coverage remained patchy because Earth’s ozone layer absorbs UV radiation, leaving observations to space telescopes that had skipped large sections of the sky. Johns Hopkins astrophysicist Brice Ménard used Claude Science, Anthropic’s workbench for planning and executing multistep research tasks, to assemble and run a reconstruction pipeline. The resulting first full-sky UV map combines far-ultraviolet light at 154 nanometers with near-ultraviolet light at 232 nanometers. Every pixel carries its provenance | Coverage and validation reported for the map | | |---|---| | Component | Detail | |---|---| | Observed coverage | About two-thirds of the sky, assembled from GALEX, Swift and FIMS/SPEAR data | | Modeled coverage | About one-third of the sky, including much of the Milky Way’s galactic plane | | UV bands | Far-UV at 154 nm and near-UV at 232 nm | | Stellar layer | UV estimates for more than 100 million stars, inferred from Gaia visible-light measurements | | Pixel metadata | Measured or predicted status, plus an uncertainty estimate | | Reported validation | Reconstructions came within about 10% of hidden measurements in held-out tests | The map’s completeness refers to spatial coverage. Scientific analyses can use the provenance and uncertainty fields to filter, weight or separately evaluate modeled pixels. The finished view traces dust clouds around young stars, large rings formed by stellar explosions and faint filaments illuminated by combined galactic starlight. Why UV surveys stopped short NASA’s GALEX mission supplied the largest existing UV dataset, capturing roughly 38,000 observations from 2003 to 2013 and covering about two-thirds of the sky. Mission planners avoided locations with very bright stars, especially along the Milky Way’s plane, because intense light could damage the spacecraft’s detectors. NASA’s Swift observatory and South Korea’s FIMS/SPEAR mission covered additional regions, yet substantial gaps remained. Combining those archives requires pixel-level calibration, artifact removal, coordinate conversion and repeated statistical analysis. Such projects can consume weeks of specialist time, which often leaves them behind research with firmer deadlines. Prompts became a processing pipeline Ménard supplied high-level instructions, inspected intermediate products and requested corrections. Claude coordinated parallel sub-agents and long-running computations across the following stages: - Archive ingestion: Locate public UV surveys and download tens of thousands of source images. - Artifact removal: suppress glare around bright stars so nearby faint emission remains measurable. - Instrument calibration: reconcile brightness measurements from different telescopes. - Coordinate alignment: reproject every exposure onto a shared sky grid and merge overlapping observations. - Statistical reconstruction: estimate missing UV emission from visible, infrared and radio measurements of the same regions. - Stellar modeling: add UV estimates for more than 100 million stars using visible-light data from ESA’s Gaia mission. The reconstruction stage used observed regions to model how ultraviolet brightness relates to emission at other wavelengths. It then applied those relationships to unobserved pixels and generated a confidence estimate for each prediction, a process commonly called inpainting. Ménard tested the method by masking regions with known UV measurements and asking the pipeline to reconstruct them without access to the hidden values. After several rounds of refinement, the reported predictions differed from the withheld measurements by about 10%. Faint circles escaped two reviews Intermediate maps contained subtle circles in the dimmest fields, with each circle appearing slightly brighter or darker than its surroundings. The patterns traced individual GALEX exposures whose uneven atmospheric UV glow had survived the initial corrections. Claude had identified atmospheric glow as a known risk early in the project, yet two rounds of agent review failed to detect its imprint on the combined map. Ménard spotted the circles during visual inspection and prompted the system to correct all 38,000 observations, which took a few hours. Global visualization provided a quality check that local processing and automated review had missed. Similar pipelines can test for tile boundaries, background offsets and repeated detector footprints before accepting a merged data product. Patterns developers can reuse - Separate deterministic and statistical stages. Downloads, calibration and reprojection can produce reproducible outputs before learned components estimate missing values. - Parallelize independent regions. Image tiles and sky sectors can be processed concurrently, reducing elapsed time without changing the underlying method. - Preserve provenance. Downstream users need to know which values came from instruments and which came from inference. - Attach uncertainty to predictions. A filled map remains useful when applications can account for confidence at pixel level. - Validate against hidden ground truth. Masking measured regions creates controlled tests for reconstruction accuracy. - Inspect aggregate artifacts. Some calibration failures become visible only after thousands of inputs are combined. The workflow produced more than a dozen map versions over several days. Ménard planned each iteration in short exchanges, then allowed the system to run computations while he worked elsewhere. For data teams, the project offers a concrete example of domain expertise directing agents through a dependency graph of ingestion, transformation, inference and review. Provenance defines the map’s limits The missing regions form a systematically different sample because GALEX deliberately avoided bright stars and much of the galactic plane. Validation on masked portions of the observed sky may therefore understate errors in regions whose structure, brightness and source density differ from the training data. Research use requires more detail than the headline 10% result provides, including the error metric, test-mask design, regional performance, uncertainty calibration and behavior near bright sources. Reproducibility also depends on recording code, model versions, prompts, calibration parameters, archive versions and compute requirements. The project demonstrates how supervised agents can make a deferred data-integration project feasible within days. Its measured-versus-modeled labels, held-out tests and visible failure history also show the controls required when statistical reconstruction becomes part of a scientific dataset.
21:04

Anthropic's Claude Now Hunts Security Bugs in Critical Infrastructure and Open Source

Full text · 8,923 chars
- Anthropic launched the Cyber Mission, targeting critical infrastructure and open-source software security. - Critical Infrastructure Defense Program partners include CrowdStrike, Palo Alto Networks, Dragos, Rockwell, Accenture, Deloitte, PwC, Booz Allen, Hitachi. - OSS Scanner gives opted-in projects free periodic scans from Claude Mythos with patches included. - Reports ship model-generated without human review; Anthropic targets above 90% true-positive rate. - Project Glasswing surfaced 29,000+ candidate vulnerabilities but only 6,000 were human-triaged. - Early validation: 88% of 97 critical findings across 48 projects met formal disclosure bar. Anthropic launches AI security programs for infrastructure and open source Anthropic has launched its Cyber Mission, combining two defensive security programs. The Critical Infrastructure Defense Program supports companies that secure operational technology, while OSS Scanner provides free, automated vulnerability reports to eligible open-source projects. Anthropic argues that advanced cyber models are reaching attackers faster than defensive tools are reaching security teams. Its new programs aim to widen defensive access while addressing a growing operational problem: models can discover vulnerabilities faster than people can validate, disclose, and patch them. Claude enters industrial networks Operational technology, commonly called OT, includes the hardware and software that control power grids, water systems, factories, transit networks, and other physical processes. These environments often require continuous operation, scheduled maintenance windows, and extensive safety testing. A known vulnerability may therefore remain unresolved long after a patch becomes available. The Critical Infrastructure Defense Program will distribute Anthropic’s most capable Claude models through established OT security providers. Founding partners include Accenture, Booz Allen, CrowdStrike, Deloitte, Dragos, Hitachi, Insane Cyber, Nozomi Networks, Palo Alto Networks, PwC, and Rockwell Automation. Participating partners receive three forms of support: - Access to frontier Claude models, Anthropic’s highest-capability systems - On-site Anthropic engineers working with partner security teams - Threat research from Anthropic’s internal red-team investigations Anthropic’s existing public-sector cyber program provides an earlier example of this approach. The company says it has offered model access and technical support to more than half of US states, along with several large public infrastructure operators. Reported uses include code scanning, patch development, incident response, and red-team exercises. Model-written reports reach maintainers directly OSS Scanner periodically audits enrolled open-source projects with Anthropic’s strongest models at no cost. Google’s OSS-Fuzz finds defects by feeding programs unexpected inputs and monitoring the resulting failures. Anthropic’s service applies large language models to source-code review. Each report can include an explanation of the vulnerability, a proof-of-concept exploit, a candidate patch, and a version-history bisection identifying the likely commit that introduced the bug. Those artifacts can reduce the work required to reproduce a finding and locate the affected code. Reports leave Anthropic without human review or triage. This design increases scanning volume and shortens delivery time while passing validation work to maintainers. Anthropic warns that reports may be incorrect, duplicative, or otherwise invalid, and says it expects a true-positive rate above 90%. Ninety-seven reports test the error rate Expert penetration testers evaluated 97 critical and high-severity reports from an early version of OSS Scanner across 48 projects. Anthropic reported the following results: - 85 reports, or 88%, qualified for the formal disclosure process. - 11 reports described real vulnerabilities that had already been reported. - One report was a false positive. Counting duplicates, 96 of the 97 reports corresponded to real vulnerabilities. The selected sample establishes performance for that test alone; accuracy may vary by language, codebase, vulnerability class, and project maturity. Twenty-nine thousand findings strain human review Anthropic’s Project Glasswing, an earlier open-source vulnerability research effort, identified more than 29,000 candidate vulnerabilities during a six-month period. Human reviewers manually assessed about 6,000. Manual validation now constrains how quickly Anthropic can send useful findings to maintainers. Maintainers have requested nearly 5,000 unvalidated reports from the project, according to Anthropic. That demand reflects the practical value of early warning, though every raw report consumes maintainer time and may require exploit reproduction, severity assessment, patch testing, and coordinated disclosure. Model performance has also improved rapidly. Anthropic says large language models increased their detection rate on CyberGym, an academic vulnerability-finding benchmark, from below 20% to above 85% over roughly a year. Early OSS Scanner participants include PostgreSQL, OpenSSL, wolfSSL, and HotCRP. wolfSSL reported that 72 of 74 findings it received were valid and that five received CVE identifiers. Those results provide an encouraging project-level example, though broader performance will become clearer as the service expands to codebases with different languages, architectures, and review practices. Security work shifts at three points - Direct model output. OSS Scanner delivers findings before human triage, reducing reporting delays and increasing the volume that maintainers must assess. - Remediation context. Reports may include candidate patches, proof-of-concept exploits, and the likely introducing commit, giving maintainers more than a scanner alert or suspicious code location. - Embedded OT support. The infrastructure program pairs model access with Anthropic engineers and established security providers, integrating Claude into active assessments and remediation work. Exploit speed collides with patch timelines Anthropic forecasts that AI could favor defenders within two years by detecting defects earlier and helping developers produce safer code. Current conditions remain difficult: models can help generate exploits in minutes, while verification, disclosure, patch review, release engineering, and deployment still require substantial human effort. Operational technology extends those timelines. Anthropic says Project Glasswing findings often took months to resolve. Industrial patches may also need to wait for maintenance shutdowns, equipment recertification, or safe deployment windows. The company says remediation has taken decades in rare cases. Maintainers and infrastructure operators therefore remain responsible for confirming exploitability, reviewing generated patches, testing for regressions, and coordinating deployment. Higher discovery rates increase the value of those processes and the workload placed on the teams running them. Four routes into the programs Access varies by role, and the infrastructure program primarily reaches operators through Anthropic’s security partners. | Programs and access paths | | | |---|---|---| | Offering | Intended users | Access path | |---|---|---| | OSS Scanner | Core maintainers of high-impact open-source projects | Submit a pull request to the Anthropic repository linked from the OSS Scanner announcement. Eligibility resembles OSS-Fuzz criteria and focuses on projects with significant infrastructure or user-security impact. | | Claude for OSS | Eligible open-source maintainers | Apply separately for a free Claude Max 20x subscription. | | Critical Infrastructure Defense Program | OT security vendors, systems integrators, and equipment manufacturers | Register interest through the Cyber Mission announcement. Infrastructure operators participate through program partners. | | Cyber Verification Program | Qualified defensive security teams | Apply for higher-capability Claude access with safety classifiers adjusted to reduce interruptions during legitimate defensive work. | Measure the path from report to remediation The programs can be evaluated through operational results that extend beyond vulnerability benchmarks: - Precision among reports delivered to maintainers - Time required to validate and triage each report - Acceptance and regression rates for model-generated patches - Time from confirmation to disclosure and release - Time required to deploy fixes in operational environments Together with Claude Security, the Cyber Mission places Anthropic’s models inside the workflows that convert findings into deployed fixes. Its results will depend on report quality, maintainer capacity, partner execution, and safe OT deployment, all constraints that discovery benchmarks leave unresolved.
21:25

Midjourney Tests Thinking Mode to Cut Image Prompt Failures by 80%

Full text · 4,993 chars
- Midjourney is testing a "thinking mode" for image generation on its alpha website - The team reports boosts to prompt accuracy, typography, and overall scene coherence - Available when rerunning a generation or editing an existing image - Internal tests suggest it fixes 60 to 80 percent of prompt-related issues - Brings LLM-style test-time reasoning to the diffusion pipeline - Live now at alpha.midjourney.com, with Midjourney soliciting user feedback Midjourney tests a “thinking mode” for image generation Midjourney is testing thinking mode on its alpha website. The experimental feature adds processing before an image is rendered, with the goal of improving prompt adherence, embedded text, and scene coherence. Those areas often break when a prompt combines several subjects, spatial relationships, or precise visual constraints. The feature is live at Midjourney Alpha. A community Office Hours summary says users can enable it when rerunning a generation or editing an existing image. Midjourney is asking alpha users to test the option and submit feedback. Planning before pixels Image generators must translate words into subjects, attributes, positions, and interactions before producing a coherent composition. A request for three red cubes beside a blue sphere, for example, requires the system to preserve the count, colors, objects, and spatial relationship throughout generation. Thinking mode appears to allocate additional test-time computation to interpreting and planning those constraints. Midjourney has not published the feature’s architecture or explained when the additional processing occurs. The “thinking” label identifies a product mode rather than a specific reasoning algorithm, and it provides no evidence of humanlike reasoning or a language-model-style chain of thought. The Office Hours summary attributes an internal estimate of roughly 60% to 80% fewer prompt-related failures to Midjourney’s developers. No benchmark, sample size, test set, or comparison method has been published, so the range remains an unverified internal result. Midjourney also has not disclosed the mode’s latency or GPU requirements. Where extra compute may help The experiment targets tasks that require the model to maintain several constraints at once. The most relevant test cases include: - Compositional constraints: multiple subjects with exact counts, colors, sizes, or positions. - Embedded text: signs, logos, labels, posters, and interface mockups. - Object interactions: hand poses, overlapping objects, physical contact, and occlusion. - Targeted edits: changing one element while preserving the surrounding image. - Dense prompts: scenes that combine characters, props, lighting, camera direction, and style requirements. These are intended use cases rather than guaranteed improvements. Results may vary by prompt, model version, image complexity, and the amount of text or spatial precision requested. Access, limits, and unknowns | Area | Current status | |---|---| | Web access | Available through the alpha website. | | Reruns | Thinking mode can be enabled when rerunning a generation. | | Image editing | The option is available when editing an existing image. | | Initial generation | Direct use during the first generation has not been confirmed. | | Discord and API | No access for either interface has been announced. | | Pricing and usage | No separate rate or fast-hour charge has been disclosed. | | Latency | No generation-time comparison has been published. | A controlled comparison keeps the prompt, aspect ratio, model version, and seed fixed where the interface permits. Generate a standard rerun and a thinking-mode rerun, then compare adherence, text accuracy, unwanted changes, processing time, and account usage. Measure the trade-off Test-time computation gives Midjourney another way to improve output without retraining the underlying model. Additional processing will probably consume more GPU capacity and increase latency, although the company has not quantified either effect. The absence of announced API support currently limits the feature to manual alpha-web workflows. Teams evaluating the mode can build a small prompt suite and track the following measures across several generations: - Constraint completion: score each requested subject, attribute, count, and spatial relationship. - Text accuracy: compare requested and rendered wording at the character or word level. - Edit preservation: record whether untouched regions remain visually stable. - Consistency: repeat prompts to determine whether gains persist across outputs. - Latency and usage: record generation time and any change in fast-hour consumption. The mode’s practical value will depend on whether its adherence gains outweigh additional time and usage costs for a given workload. Midjourney’s alpha test provides a way to measure that trade-off while implementation details, pricing, API availability, and broader rollout plans remain undisclosed.
22:00

Liquid AI's d1-3B Makes Structured AI Decisions in 8 Milliseconds Without Generating Text

Full text · 2,815 chars
- Liquid AI released d1-3B, an open-weight 3.1B multimodal decision model that answers in one forward pass with zero output tokens. - Scores 48.57 on Decision Index 0.2.1, beating every sub-10B model and matching Decider 35B-A3B. - Latency: 8 ms on RTX 4090, 9 ms on AMD MI325X, 16 ms on Jetson AGX Thor, 50 ms on Jetson Orin Nano. - Supports three primitives: Noul (yes/no with probability), Choice (named options), Score (ordered rubric). - Built on LFM2.5-VL-3B via weight averaging, multi-seed fine-tuning, and checkpoint merging. - Available on Hugging Face, with vLLM, SGLang, and blog post covering deployment. Liquid AI’s d1-3B makes decisions in one pass Liquid AI has released d1-3B, an open-weight multimodal model that returns structured predictions in a single forward pass and generates zero output tokens. Built on Liquid Foundation Models, it removes the sequential decoding loop used by generative transformers. Pipelines that need a yes-or-no answer, category, ranking, or score can therefore avoid generating and parsing text. The 3.1B-parameter model builds on LFM2.5-VL-3B. Liquid also released an experimental 600M sibling called d1-omni-600M. Both models are available on Hugging Face. One pass, three answer types Each request contains a state, such as text, JSON, an image, or a combination, plus a dictionary of named questions. The model reads the state once and returns typed answers with probabilities that Liquid describes as calibrated. The API defines three primitives: - Noul: Answers a yes-or-no question with a probability from 0 to 1. For example, “Is this message spam?” might return 0.92. - Choice: Selects one named option and returns a probability distribution across all supplied options. - Score: Rates the input against an ordered rubric and returns a probability-weighted position on that scale. A single API call can apply several decisions to the same customer message: questions = { "refund": { "type": "noul", "instructions": "Is the customer asking for a refund?", }, "team": { "type": "choice", "instructions": "Which team should handle this?", "criteria": { "billing": "Charges, refunds, invoices", "technical": "App or site faults", "fraud": "Suspected unauthorized use", }, }, "urgency": { "type": "score", "instructions": "How urgent is this?", "criteria": [ "Can wait", "Today", "Blocking the customer now", ], }, } model.system_one( "I was charged twice this month, please refund one of them.", questions, ) This story is for Pro members You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.
23:34

ttok 0.4

Full text · 349 chars
8th October 2026 ttok is my CLI tool for counting tokens, using OpenAI's open source tiktoken library. It hasn't been in updated in a couple of years, but I finally fixed a Click warning, updated CI, and added a --list-models command to list available models. It works with uvx, so you can count tokens in anything like this: cat file.txt | uvx ttok
00:00

ChatGPT Intelligent UI 🧠, Claude Haiku 5.5 ⚡, Grok Bot 3rd party models 🤖

A daily AI brief leads with ChatGPT Intelligent UI, Claude Haiku 5.5, and Grok Bot routing to third-party models. The captured body is mostly a StepFun sponsor: Step 5 Preview at $1 / $2.70 per million tokens, near Claude Opus 5 Medium on the Intelligence Index, 1M context. StepAudio 3 claims #1 non-streaming ASR at 1.7% WER. The rest of the edition is cut.

Full text · 682 chars
StepFun's Step 5 Preview: near-Opus 5 intelligence, 80% lower input, 89% lower output prices (Sponsor) ⚡ Step 5 Preview costs $1/M input and $2.70/M output tokens, and scores near Claude Opus 5 (Medium) on AA's Intelligence Index. Built for coding agents, financial analysis, and long-running work, with a 1M-token context. Now on OpenRouter and available in OpenCode, Kilo Code, Cline, omp, and Hermes Agent: switch models, keep your workflow. More integrations to come. 🎙️ StepAudio 3: TTS streams expressive, emotion-aware speech; ASR ranks #1 in AA's non-streaming transcription (1.7% WER); Realtime ranks #1 in AA's Conversational Dynamics (98.9%) and Speech Reasoning (99.7%).
02:15

US, China, EU: three paths to regulating AI - Taipei Times

Three big governments are writing three different rulebooks for the same technology. The US, China, and the EU are taking different paths to govern artificial intelligence “as calls grow worldwide to rein in its most powerful” systems. The stub does not name the bills or the agencies.

Full text · 144 chars
The US, China and the EU are taking different paths to govern artificial intelligence , as calls grow worldwide to rein in its most powerful ...
02:23

Working on AI and Technology Policy in the U.S. Senate - Communications of the ACM

An old hearing still sits over today’s Senate tech staff. In 2023, industry leaders urged Congress to write new legislation to rein in AI risks; Sam Altman is named. The piece is an ACM opinion on working AI and technology policy in the U.S. Senate. No 2026 vote count is in the stub.

Full text · 149 chars
In 2023, AI industry leaders urged Congress to write new legislation to rein in the risks of artificial intelligence (AI). Sam Altman, the CEO of ...
02:33

Existential threats join bubble fears as AI mood sobers at Singapore forums | Reuters

The mood at two Singapore conferences flipped from boom talk to crash-and-doom talk. Investors and policymakers put “existential and financial risks posed by the artificial intelligence boom” at the top of the agenda. The stub names no speakers, dates, or quotes beyond that setup.

Full text · 147 chars
Existential and financial risks posed by the artificial intelligence boom were top of mind for investors and policymakers at two conferences in ...
03:13

NFP shifts from AI that assists to AI that executes with Microsoft Copilot Cowork

An insurance broker says it stopped using AI as a helper and started letting it finish the job. NFP’s Jennifer Ratten, VP of Application Engineering: “Cowork moved NFP from a vision of AI assistance to AI execution.” The capture is a Microsoft customer-story pull quote. No volumes, products, or error rates are in the stub.

Full text · 123 chars
“Cowork moved NFP from a vision of AI assistance to AI execution.” Jennifer Ratten, VP of Application Engineering , NFP ...
04:00

QuanLing: Cross-Branch Validation of Language Distance Quantification on Western Romance

A distance score for related languages is being checked on a second language family. QuanLing applies the same North-Germanic recipe to French, Portuguese, Spanish, and Italian on 150 four-language parallel sentences. Portuguese–Spanish are closest (LaBSE 0.0229); French–Italian are farthest (0.0338). LaBSE and mBERT agree on 4 of 6 pair rankings. French MLM top-1 is 36.12% versus 29.28% for Italian. The Romance span is 0.011 versus 0.008 for North Germanic.

Full text · 2,434 chars
Computer Science > Computation and Language Title:QuanLing: Cross-Branch Validation of Language Distance Quantification on Western Romance View PDF HTML (experimental) Abstract:Quantifying language distance among closely related languages remains a core challenge in quantitative linguistics. Our previous work [1] introduced QuanLing (Quantitative Linguistics via Pretrained Language Models), a quantitative framework combining language distance metrics (sentence embedding distance, tokenization fragmentation rate) with language property analysis (MLM prediction probability), validated on North Germanic (Danish, Norwegian Bokmål, Swedish). This paper extends QuanLing to Western Romance--French, Portuguese, Spanish, Italian--testing cross-branch applicability with the same metric family and aggregation protocol as our North Germanic study, adapted for four languages (English anchor, quadruplet construction). Using 150 four-language parallel sentences, we compute LaBSE sentence embedding distances, tokenization fragmentation rates from four monolingual BERT tokenizers, and mBERT masked language model mutual intelligibility. Results show that Portuguese--Spanish are closest (LaBSE distance 0.0229), French--Italian most distant (0.0338); LaBSE and mBERT rankings agree on 4 of 6 pairs, confirming cross-model robustness. Western Romance shows a wider absolute distance span than North Germanic (0.011 vs. 0.008) but comparable relative ratios (1.48 vs. 1.67), consistent with longer divergence time. French exhibits notably higher MLM predictability (36.12% top-1 accuracy vs. 29.28% for Italian), reflecting its orthography--phonology decoupling. This cross-branch validation provides further evidence for QuanLing's generalizability beyond a single language branch. Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
04:38

MCMC opens its 1099th Nadi - The Star

Malaysia is putting a prompt-writing seat in each community center. MCMC opened its 1,099th Nadi and is running a “1Nadi 1 Prompt Engineer” initiative. The stub defines prompt engineering as designing and fine-tuning text input. No stipend or curriculum is in the capture.

Full text · 154 chars
... prompt engineering per Nadi through the ongoing 1Nadi 1 Prompt Engineer initiative. Prompt engineering is the design and fine-tuning of text input ...
04:52

GOCOP Trains Online Publishers On AI, Prompt Engineering & SEO - PM Parrot

Nigerian online publishers sat through a newsroom class on prompts. GOCOP’s session was titled “The AI-Powered Newsroom: AI & Prompt Engineering for Online Publishing,” anchored by “AI expert and publisher of The Lagos” — the name is cut. A sibling item names Durojaiye. No attendee count is in this stub.

Full text · 151 chars
The training, titled “The AI-Powered Newsroom: AI & Prompt Engineering for Online Publishing,” was anchored by AI expert and publisher of The Lagos ...
04:55

Google Experiments With an AI-Powered Gaming Platform, but Creation Is the Easy Part

Google is letting a prompt become a playable game in the browser. The captured line says a Google account, a browser, and a prompt are “the apparent starting requirements.” Creation is called the easy part in the title. No engine name, pricing, or traffic numbers are in the stub.

Full text · 153 chars
A Google account, a browser, and a prompt become the apparent starting requirements. ... Engineers ' Knowledge Base · MVP Launch Plan · HR Recruiting ...
05:30

GOCOP empowers members to use AI tools to boost efficiency, profitability

The same Nigerian publishers’ workshop is being written up as a business-efficiency class. The session “Prompt Engineering for Online Publishing” explored how members can “publish faster, better and more profitably.” Durojaiye is named. No outcome metrics are in the stub.

Full text · 151 chars
... Prompt Engineering for Online Publishing”, explored how participants can publish faster, better and more profitably in the age of AI. Durojaiye ...
08:37

Jellyfish Advances Developer and Agent Productivity Insights for the AI-Native SDLC

A software-intel vendor wants you to measure what AI tools actually change in the build. Jellyfish announced a suite of features “to help organizations measure the impact of AI engineering tools.” The captured body is that one sentence. No metrics, prices, or screenshots are in the stub.

Full text · 125 chars
Jellyfish today announced a comprehensive suite of features to help organizations measure the impact of AI engineering tools.
09:24

darkzodchi on X: "Anthropic engineer released a prompt that turns Opus 5.5 & Fable 5.5 into ...

Someone on X is passing around a prompt that is supposed to turn two Claude models into a crew. darkzodchi says an Anthropic engineer released a prompt that turns Opus 5.5 and Fable 5.5 into “a team of agents.” The tweet is truncated. The prompt text is not in the capture. Do not treat this as an official Anthropic release.

Full text · 149 chars
Anthropic engineer released a prompt that turns Opus 5.5 & Fable 5.5 into a team of agents. It's f*cking unreal… send it to Opus 5.5 & Fable 5.5, ...
10:21

How's the stock market doing, minus AI ?

Tech stocks are covering up weakness everywhere else. A Marketplace episode says tech success is “obscuring strife in other parts of the stock market,” and flags auto-loan delinquencies, yield curves, and the AI supply chain. No index levels or dates are in the stub.

Full text · 138 chars
Tech success is obscuring strife in other parts of the stock market. Plus: Auto loan delinquencies, yield curves, and the AI supply chain.
11:01

Simplilearn and Columbia Engineering Executive Education Launch Collaboration ... - PR Newswire

A training shop and an Ivy League exec-ed desk are packaging agent skills for working professionals. Simplilearn and Columbia Engineering Executive Education launched a collaboration whose curriculum covers AI agents, agentic RAG, multi-agent systems, harness engineering, and loops. The stub does not list price, length, or start date.

Full text · 149 chars
The curriculum provides professionals with a practical understanding of AI agents , agentic RAG, multi- agent systems, harness engineering , loop ...
12:10

The Download: AI roadblocks for humanoids and portable rubber dams

The daily tech recap leads with why humanoid robots still will not do your dishes. It restates the earlier robotics piece, then flags WaveSave’s SlamDam, a portable rubber dam that pumps up to 1.3 meters high. Must-reads include the first nuclear clocks, three fired OpenAI researchers warning about reasoning monitors, Samsung’s projected $80 billion quarterly profit, and a fraudster jailed after AI songs and bots pulled $8 million. David Robinson, who quit as OpenAI safety transparency lead, is quoted saying “We grew a mind.”

Full text · 6,739 chars
This is today's edition of The Download, our weekday newsletter that provides a daily dose of what's going on in the world of technology. AI breakthroughs in robotics won’t change your life any time soon The hype around humanoid robots is reaching fever pitch. Much of it comes from the idea that the same AI advances behind ChatGPT and Claude could enable robots to imitate human movement the way chatbots imitate human language. But many researchers are skeptical. They argue that such assumptions minimize the challenges of using an intelligence built on language and images to master the infinite variability of the physical world. They also point out that conflating humanoids with so-called generalist machines is misleading. Across robotics labs, tensions have emerged over whether current forms of AI are all that’s needed to perfect all-purpose humanoids or whether an entirely new path is required. —Jamie Condliffe This story is a collaboration between MIT Technology Review and Aventine, a non-profit research foundation that creates and supports content about how technology and science are changing the way we live. 10 Climate Tech Companies to Watch: WaveSave and its portable rubber dam WaveSave is one of MIT Technology Review’s10 Climate Tech Companies to Watch 2026, available exclusively to subscribers. As sea levels rise and flood risks worsen, WaveSave has emerged with an ingeniously simple tool: a portable rubber dam that’s quick to install when flooding starts. Anyone can unfurl the SlamDam on the ground along a riverbank or shoreline and pump it full of water until it swells into place. And there it stands, up to 1.3 meters high, shielding nearby property or structures from damage as waters rise. —Amy Nordrum Subscribers can now access the full 10 Climate Tech Companies to Watch, spanning everything from smart turbines and sustainable cement to carbon dioxide batteries. Why we’re watching these climate tech companies —Casey Crownhart Our Climate Tech Companies to Watch list is always a little ray of hope for me—there are some fascinating ideas out there about how we power our world or move away from fossil fuels. Our chosen companies also reflect several aspects of the current moment. One is that China is the center of the climate and energy tech world. Another is that everybody wants a piece of the data center pie. And then there’s the fact that the grid needs support, with energy storage a key piece moving forward. This story is from The Spark, our weekly climate tech newsletter. Sign up to receive it in your inbox every Wednesday. The must-reads I’ve combed the internet to find you today’s most fun/important/scary/fascinating stories about technology. 1 Scientists in Austria and China have built the first nuclear clocks They could become the world’s most precise timekeepers. (NYT $) + They could also help scientists search for dark matter. (Reuters $) + China’s nuclear clock is six times more stable than Austria’s. (SCMP) + A space clock could help measure Earth. (MIT Technology Review) 2 Three fired OpenAI researchers want to protect AI reasoning monitoring They warn that AI companies could lose a key safety tool. (WSJ $) + OpenAI fired them for allegedly sharing sensitive information. (Gizmodo) + Don’t be fooled—LLMs don’t reason. (MIT Technology Review) 3 Samsung has forecast the highest quarterly profit ever for a tech firm The AI chip boom has sent its projected profits to $80 billion. (CNBC) + Can AI labs ever turn a profit? (Reuters $) + What even is the AI bubble? (MIT Technology Review) 4 OpenAI’s math breakthroughs point beyond mathematics AI is progressing into increasingly specialised fields. (Axios) + OpenAI is catching up with Anthropic on coding. (WSJ $) + A new AI tool could speed up math research.(MIT Technology Review) 5 A 194-year-old tortoise has had its genome sequenced His longevity could reveal ways to extend human lifespan. (Guardian) + Jonathan is the world’s oldest known land animal. (404 Media) + His youthful mitochondria may help repair DNA. (New Scientist $) 6 A fraudster who used AI to outstream Taylor Swift has been jailed He used AI-generated songs and bots to earn $8 million. (NYT $) + But prosecutors said he stole royalty payments. (Ars Technica) 7 TikTok allegedly gave kids fake safety features to study engagement The tests measured time spent and ad revenue. (Reuters $) 8 Burglars could be identified from DNA floating in the air Thanks to a new device that collects airborne genetic material. (Economist $) 9 Pioneers of mirror-image molecules have won the chemistry Nobel Their work has become a key tool in developing medicines. (Guardian) 10 California has declared a human-vs-robot fight illegal The state's combat-sports laws were never written for machines. (NYT $) Quote of the day “We grew a mind.” —David Robinson, who last week quit his job as OpenAI safety transparency lead, tells The New York Times that his former company has built models that develop capabilities their creators did not explicitly design. One more thing We need a moonshot for computing The US government is organizing itself for the next era of computing. Ultimately, it has one big choice to make: adopt a conservative strategy that aims to preserve its lead for the next five years—or orient itself toward genuine computing moonshots. There is no shortage of candidates, including quantum computing, neuromorphic computing and reversible computing. And there are plenty of novel materials and devices. These possibilities could even be combined to form hybrid computing systems. The National Semiconductor Technology Center can drive these ideas forward. To be successful, it would do well to follow DARPA’s lead by focusing on moonshot programs. Read the full story. —Brady Helwig & PJ Maykish We can still have nice things A place for comfort, fun, and distraction to brighten up your day. (Got any ideas? Drop me a line.) + An experimental pair of contact lenses can repair themselves with UV light. + Meet Rich, the rescued squirrel who visits the woman who saved him every day. + Reminisce about the lost joy of music piracy with this wistful reflection on a time before streaming.  + What would rainbows look like on Tatooine? Randall Munroe, a former NASA roboticist, has the answer. Deep Dive The Download The Download: why AI’s latest breakthroughs and fears may be more hype than reality Plus: 22 nations have called for a new global body to oversee AI. The Download: AI’s self-improvement problem, and what’s driving the heat Plus: OpenAI has paused some model work over safety concerns. Stay connected Get the latest updates from MIT Technology Review Discover special offers, top stories, upcoming events, and more.
13:59

Mission AI ROI strategy framework for space using Amazon Bedrock - AWS

An AWS public-sector post sketches a Mission AI ROI framework for space using Amazon Bedrock. One line itemizes $42,000 inference + $28,000 prompt engineering + $18,000 evaluation + $12,000 governance = $100,000 as GA-TCO. Broader methodology is not in the stub.

Full text · 146 chars
... engineering organization can ... GA-TCO – $42,000 inference + $28,000 prompt engineering + $18,000 evaluation + $12,000 governance = $100,000.
15:06

The Zero-Day Blind Spot: Why Your Agentic SOC needs Retrospective Packet Replay

A Cisco blog argues an agentic security operations center needs retrospective packet replay so it can re-inspect traffic after a zero-day is known. The stub thanks engineers including Fred Frey. No product SKU or eval numbers are in the capture.

Full text · 155 chars
Our thanks to the engineers who built the Agentic SOC and the Humans who provided decision making expertise. ... AI SOC Analyst Engineering : Fred Frey ...
16:17

☕️ OpenAI launches GPT-6 for everyone

A morning recap says GPT-6 is now in every ChatGPT seat with Intelligent UI. Samsung forecasts about 107 trillion won ($80 billion) quarterly operating profit. Ethereum researcher Justin Drake warned ECDSA wallets could break “in months not years” and told large holders to move first. Grok Bot will route jobs to Claude Opus 5.5, Midjourney, and Suno without a user picker. Google Cloud launched a Gemini workplace agent. The Haiku 5.5 price card is restated: $0.10 / $0.50 under 100,000 tokens, then 5×.

Full text · 4,189 chars
| | | 🧠 OpenAI launches GPT-6 for everyone LINK | OpenAI has released GPT-6 to all ChatGPT users, bringing an "Intelligent UI" that can answer questions with interactive diagrams, side-by-side comparisons, or plain text depending on what fits best. The system runs on a library of streamable components and a compiler that builds the interface piece by piece as the model writes, so results appear progressively instead of waiting for the full response. GPT-6 also handles web searches better, deciding more wisely when to look something up, and resists attempts to bypass its safety training across multi-turn attacks involving cyberattacks, biological threats, and violence. | 💰 Samsung forecasts world record-breaking $80B profit LINK | Samsung expects to post a quarterly operating profit of about 107 trillion won ($80 billion), which would be a record for any technology company, driven by the global memory chip shortage. The profit, up more than nine times from a year ago, would mark Samsung's fourth straight record quarter, as AI demand pushes prices higher for DRAM, NAND flash, and high-bandwidth memory chips. Samsung's full breakdown arrives on October 29, but analysts expect the chip business drove most of the gains, while its smartphone unit suffers from higher component costs forcing price increases. | 🔐 Researcher warns AI may break crypto wallets LINK | Ethereum researcher Justin Drake has warned that artificial intelligence could crack the ECDSA cryptography protecting crypto wallets, urging users to shift their funds to fresh addresses before quantum computers pose the same threat. In an X post on Wednesday, Drake said recent AI math results mean ECDSA could break "in months not years," letting attackers recover private keys from exposed public keys far sooner than people expected. Drake pointed to OpenAI's work, including 10,000 AI agents that cracked the Navier-Stokes equation in 88 hours, and advised large holders to move balances first, though he cautioned that a rushed migration would backfire. | 🤖 Grok will now use rival Claude LINK | Elon Musk said on Tuesday that SpaceXAI's Grok Bot will route jobs to outside AI services, including Anthropic's Claude Opus 5.5, image maker Midjourney and music tool Suno, picking a provider based on each task. Customers can't pick the model themselves, since Grok Bot manages the choice with no picker; usage analytics reveal which model handled each request, and billing follows whichever model actually did the work. Enterprise accounts get a team allowlist of permitted models honored by default, but Grok Bot warns it may not follow the list, and Musk offered no tests proving the routing gives better results. | 💥 Google Cloud unveils Gemini work agent LINK | Google Cloud has launched the Gemini agent, a single assistant for work that handles tasks from answering questions and creating content to coding, all starting from one prompt box. The agent plans the work, picks the right model for each job, and connects to a company's business systems, delivering finished results inside the documents, inboxes, and developer tools employees already use. Built for businesses, the Gemini agent carries each organization's full work context and comes with cost controls plus the security, administration, and governance that enterprise customers expect. | 💲 Anthropic reveals Haiku 5.5 LINK | Anthropic released Claude Haiku 5.5, its fast, low-cost model, dropping the price roughly 90% from the previous Haiku 4.5 to just $0.10 per million input tokens and $0.50 per million output tokens for shorter prompts. The new price matches OpenAI's GPT-6 Luna up to 100,000 tokens, but beyond that it jumps 5x to $0.50/$2.50, and a less generous tokenizer counts about 1.25x more tokens than Haiku 4.5 for the same text. Unlike the older Haiku, the 5.5 model supports reasoning levels and can't fully disable them, defaulting to medium; it draws far better than its predecessor, and a top-effort pelican test cost just over 3 cents and took about five minutes. | |
16:26

Harness Acquires Augment Code Assets to Expand Reach into AI Coding - DevOps.com

DevOps.com covers Harness buying Augment Code assets to reach further into AI coding. It says most DevOps teams already use AI coding tools and it is unclear how far they are toward agentic engineering. No deal terms in the stub.

Full text · 150 chars
It's not clear how far down the path toward agentic engineering DevOps teams are today. Most are making extensive use of AI coding tools. The next ...
17:07

Vinod Khosla bets on ex-DeepMind engineer's agent startup Wajo, citing trust | Dealroom.co

Vinod Khosla is backing Wajo, an agent startup from an ex-DeepMind engineer, with trust as the pitch. The Dealroom stub says when sites block agents the product can prompt the user, open a verification group, or hire a human. Funding size is not in the capture.

Full text · 149 chars
When sites block AI agents, it can prompt the user to act or create a verification group. It can also hire a human to complete a task when needed ...
17:13

Harness acquires Augment Code to automate software from idea to production

Harness is buying Augment Code to automate software from idea toward merge-ready code. Cosmos becomes the Harness Cosmos Software Factory Agent. The Dealroom stub does not include price or close date.

Full text · 151 chars
Cosmos becomes the Harness Cosmos Software Factory Agent , tasked with automating engineering work from idea to merge-ready code. The product: When ...
17:55

GitHub Copilot is going local — but Microsoft won't say what gets sent to the cloud

The New Stack says GitHub Copilot is adding local inference but Microsoft will not say what still goes to the cloud. The stub notes a prompt uses GitHub issue and pull-request metadata. No routing table is in the capture.

Full text · 145 chars
Yet the prompt uses GitHub issue and pull request metadata instead of ... Amanda Caswell is an AI journalist, certified prompt engineer , and ...
18:05

Walking Robotic Hand Uses AI to Learn New Tricks - IEEE Spectrum

IEEE Spectrum covered a walking robotic hand that uses AI to learn new tricks. Columbia’s Matei Ciocarlie calls it a “cool result” that suggests robot hands can pick up extra skills. The paper, hardware, and scores are not in the stub.

Full text · 154 chars
However, Matei Ciocarlie, an associate professor of mechanical engineering at Columbia University, says it's a “cool result” that suggests robot hands ...
18:08

Adobe in Trouble After Someone Reverse Engineered Adobe Creative Suite and Released ...

Futurism says someone reverse-engineered Adobe Creative Suite and released an open-source alternative called ArtCraft. The stub notes prompt windows exist but AI is not required. No author, license, or feature list is in the capture.

Full text · 152 chars
... prompt windows, though AI use is by no means required by ArtCraft's apps. ... More on software: Software Engineer Says AI Is Causing Chaos Among ...
18:29

Dutch firm uses AI to help drone systems 'talk' on Ukraine's battlefield - Defense News

A Dutch firm is testing an AI layer so drone systems can share commands in Ukraine. Defense News caption: an engineer holds an interceptor while Nexus is tested on September 30, 2026. The company name beyond “Dutch firm” and any performance numbers are not in the stub.

Full text · 147 chars
An engineer holds an interceptor drone, while a Dutch company tests commands on the AI -powered Nexus system, Sept. 30, 2026. (Piroschka van de ...
18:32

Keysight gives RF engineers agentic AI keys to automate design and speed up dev time

A second Keysight alert says RF engineering has been slow to adopt agentic AI because it leans on specialist knowledge and schematics that language models handle poorly. Same Fierce Wireless story as the other Keysight item. No product details in the stub.

Full text · 149 chars
Radio frequency (RF) engineering has been slow to adopt agentic AI because it relies on specialist expertise and schematics and layouts that LLMs ...
18:33

Keysight gives RF engineers agentic AI keys to automate design and speed up dev time

Keysight is adding agentic AI so RF engineers can tell software to run design chores. Fierce Network says radio-frequency work has been slow to adopt agents. No product name, price, or time-saved figure is in this stub.

Full text · 149 chars
Keysight is bringing agentic AI to its software, enabling engineers to direct AI agents to automate complex design tasks; RF engineering has been ...
18:40

The Missing Handoff Between Schematic Generation and PCB Layout Automation

An Electronic Design piece argues the missing step is the handoff from schematic generation to PCB layout automation. After an engineer approves system boundaries, several agent and engineer roles can proceed. No tool names or benchmarks are in the stub.

Full text · 154 chars
It becomes a working agreement among the agents and engineers . Once an engineer approves the system boundaries, several roles can proceed at the same ...
19:22

Google launches Gemini AI workplace agent that can write code and run tasks, company says

Google unveiled a Gemini workplace agent that the company says can write code and run tasks. CBS News frames it as competing with other tech giants’ work assistants. Features, pricing, and availability are not in the stub.

Full text · 147 chars
Google on Thursday unveiled a new AI agent tailored for workplace tasks, positioning it to compete with other technology giants racing to offer ...
20:39

Cal AI's 19-year-old founder just raised $10M for his new AI startup | TechCrunch

The teen co-founder of calorie app Cal AI raised $10 million for a new personal AI agent startup. TechCrunch names Zach Yadegari and says the new company competes with existing personal agents. Other terms are not in the stub.

Full text · 142 chars
Zach Yadegari, the teen co-founder of popular Cal AI calorie tracking app, has launched a new personal AI agent startup that competes with ...
21:05

Quoting Carson Gross

A short note argues programming stays a job even when AI writes more of the code. Carson Gross says the work is solving problems with computers and controlling complexity. Simon Willison posted the quote on October 8. There is no extra argument in the capture beyond those two bullets.

Full text · 431 chars
8th October 2026 Computer programming is, fundamentally, about two things: - Problem-solving using computers - Learning to control complexity while solving these problems I have a hard time imagining a future where knowing how to solve problems with computers and how to control the complexity of those solutions is less valuable than it is today, so I think it will continue to be a viable career even with the advent of AI tools.
06:08

Prompt engineering is the design and fine-tuning of text input to guide artificial intelligence models.

A newspaper Facebook post is teaching the dictionary definition of prompt writing. It calls prompt engineering an emerging discipline that crafts instructions or inputs to guide AI systems. This is a primer, not a product launch. No examples or tools are in the stub.

Full text · 148 chars
Prompt engineering is an emerging discipline that focuses on crafting effective instructions or inputs (called "prompts") to guide AI systems in ...
12:23

Anticipatory intelligence in clinical microbiology | Nature Reviews Bioengineering

A Nature Reviews Bioengineering article on anticipatory intelligence in clinical microbiology mentions prompt-engineering-enabled LLMs or MLLMs plus bioinformatics for SARS-CoV-2 antibody work. The stub is a single methods line, not the paper.

Full text · 150 chars
Prompt - engineering -enabled LLM or MLLM and instigative bioinformatics pave the way to identify and characterize significant SARS-CoV-2 antibody ...
13:38

AI transformation starts at the top: Closing the VP leadership gap on AI

HR Brew says AI transformation starts with closing a VP leadership gap. General Assembly training includes prompt engineering plus live sessions and one-on-one coaching via Ezra. No enrollment numbers or prices are in the stub.

Full text · 153 chars
... prompt engineering . The program also includes additional training through live, hands-on sessions and one-on-one executive coaching via Ezra, an ...
15:13

Scaling Agentic AI for Enterprise Success - Tech Mahindra

Tech Mahindra published a view on moving agentic AI from experiments to enterprise impact. The stub lists industry knowledge, engineering, and integration as what they bring. No client metrics are included.

Full text · 152 chars
Tech Mahindra brings industry knowledge, engineering expertise, and enterprise integration capabilities. Collectively, they enable transformation at ...
15:44

Anthropic & Godel: What is the AI-Powered Godel Factory? | AI Magazine

AI Magazine asks what Anthropic’s Godel Factory is. The stub says agentic engineering teams direct workflows toward defined business and technology outcomes and can be “Godel’s engineers.” No product spec is in the capture.

Full text · 151 chars
Agentic engineering teams direct and orchestrate the workflows towards defined business and technology outcomes. They can be Godel's engineers, the ...
16:02

Real IT Solutions Named Finalist for 2026 AI-Powered Award by MSP Titans of the Industry

A managed-service firm is a finalist again for an industry AI award. Real IT Solutions was named for the 2026 AI-Powered Award by MSP Titans of the Industry, its third straight year as a Titans finalist. CEO Matt Kahle mentions a prompt-engineering method called CRAFTED. The rest of the method is not in the stub.

Full text · 150 chars
... prompt - engineering method called CRAFTED. "We're honored to be named a finalist this year," said Matt Kahle, CEO of Real IT Solutions. "This ...
17:04

Scaling Sandboxed Agentic AI on Xeon CPUs

Intel posted about scaling sandboxed agentic AI on Xeon CPUs. The stub is bylines: Jun I Jin, Michael Zhang, Rohan Vaidya. No throughput numbers or Xeon SKU are in the capture.

Full text · 152 chars
Jun I Jin, Principal Engineer - Cloud and System Performance, Intel. Michael Zhang, Cloud Software Architect, Intel. Rohan Vaidya, Software Engineer ...
17:28

IoT device failures threaten enterprise AI deployment - survey - The Engineer

A UK engineering magazine says a survey found IoT device failures threaten enterprise AI deployments. The Engineer stub does not include sample size, failure rates, or the survey sponsor.

Full text · 150 chars
... AI applications, a new survey has found. ... The Engineer is a UK-based monthly magazine covering the latest developments and business news in ...
17:35

What Happens When AI Agents Start Maintaining Legacy Code?

A HackerNoon tease asks what happens when agents start maintaining legacy code. It says older systems look like an obvious place to save engineering time. No case study or repo is in the stub.

Full text · 154 chars
That makes legacy software look like an obvious place where these agents could save considerable engineering time. Many older systems have exactly the ...
17:43

Physician edits to artificial intelligence (AI)-generated messages to patients were associated with ...

A medical brief says doctor edits on AI-drafted patient messages changed how long replies took. The 2-Minute Medicine stub names Poursoltan and stops. No edit types, sample size, or time numbers are in the capture.

Full text · 155 chars
Physician edits to artificial intelligence (AI)-generated messages to patients were associated with different response time burdens · 1. Poursoltan and ...
17:48

Zallpy Digital Expands AI Agents in Transformation Consulting - Markets Insider

A Markets Insider blurb says Zallpy Digital is expanding AI agents inside transformation consulting. Possible offerings listed: applied AI, agents, automation, data, integration, platforms, or software engineering. No client names, headcount, or revenue are in the stub.

Full text · 152 chars
Potential solutions identified may include applied AI, AI agents , automation, data, systems integration, digital platforms, software engineering or ...
17:57

WSU turns to artificial intelligence to assist sports communications - HeraldNet.com

Washington State University is using AI to help athletics communications “expand storytelling.” The HeraldNet stub also mentions Gonzaga staffers turning to AI. No tool names or examples are in the capture.

Full text · 151 chars
... artificial intelligence to “expand storytelling” by its athletics department. ... artificial intelligence for help. Gonzaga University staffers ...
18:02

Peter Thiel slams Obama, Pope Leo in new 'Antichrist' lectures

POLITICO says Peter Thiel is escalating attacks on AI critics, naming former President Barack Obama and Pope Leo in new “Antichrist” lectures. The stub has no lecture date or quotes beyond the headline frame.

Full text · 150 chars
Tech billionaire Peter Thiel is escalating his attacks on critics of artificial intelligence , accusing former President Barack Obama and the pope ...
18:04

In the era of agents , GTM alpha means self-learning - The GTM with Clay Blog

Clay’s GTM blog says the edge in an agent era is self-learning outreach, not more research automation. As GTM engineers automate research and routine mail, reps spend time on the parts that still need a person. No product launch details are in the stub.

Full text · 152 chars
As GTM engineers automate more of the research and routine outreach, reps will spend their time on the parts of selling that depend on a person, and ...
18:04

Recapping the products announced at Sculpt 2026 - The GTM with Clay Blog

Clay recapped products announced at Sculpt 2026. The stub says GTM engineers now work with Clay the way software engineers work with coding agents: set the outcome, approve the plan, let agents build. The product list is not in the capture.

Full text · 149 chars
GTM engineers now work with Clay the way software engineers work with coding agents : set the outcome, approve the plan, and let the agents build ...
18:05

Want to fix the broken university model in the age of AI ? Look to tradition | Fortune

A Fortune commentary says fixing the university model in the AI age means looking back at tradition. The stub identifies the author as an École Polytechnique and Cambridge-trained engineer who helped scale Uber internationally. The argument itself is not in the capture.

Full text · 148 chars
An École Polytechnique and Cambridge-trained engineer , he helped scale Uber in its early days in International Markets and spent time at Oliver ...
18:15

7 Best Resources to Learn About Self-Evolving AI Agents

KDnuggets listed seven resources for learning about self-evolving AI agents. The stub mentions an agent that rewrites its prompt and one that builds a… then cuts. The seven links themselves are not in the capture.

Full text · 144 chars
An agent that rewrites its prompt , one that builds a ... ML Engineer , AI Engineer , or LLM Engineer : Which Role Actually Builds What in 2026?
18:16

AI NSFW 2026: Prompt Engineering , Output Control and Platform Differences - Apps Script

A buyer’s guide stub says it will cover prompt engineering, output control, and platform differences for AI NSFW tools in 2026. The captured body is one sentence. No product list, prices, or methods are in the stub.

Full text · 146 chars
This article is a buyer's guide for navigating AI NSFW tools in 2026. It focuses on prompt engineering , output control, platform differences, ...
18:33

Trump declares anyone who uses the term Artificial Intelligence to be 'The Enemy!'

The Independent says Trump called anyone who uses the phrase “artificial intelligence” the enemy, because “artificial” makes intelligence sound fake. Preferred substitutes listed: Intelligence, Superior Intelligence, or Extreme Intelligence. The stub is a clipped quote, not a speech transcript.

Full text · 152 chars
... intelligence ” because “use of the word artificial makes intelligence fake. ... Intelligence ,” “Superior Intelligence ,” or “Extreme Intelligence .
18:46

Artificial Intelligence : From Technology to Rights

The University of Naples posted a session titled “Constitutional Europe and Artificial Intelligence,” chaired by Francisco Balaguer Callejón of Granada. The stub is a program line, not the talks.

Full text · 155 chars
The morning session, “Constitutional Europe and Artificial Intelligence , ” will be chaired by Francisco Balaguer Callejón ( University of Granada) and ...
18:48

Meet the 2026 Ming Hsieh Institute Scholars - USC Viterbi | School of Engineering

USC Viterbi introduced its 2026 Ming Hsieh Institute Scholars. The quote is that full-stack AI touches every part of the department. No scholar names or project titles are in the stub.

Full text · 150 chars
“When we talk about full-stack AI , every aspect of it is touched by our department, not just in terms of education, but really inventing the next ...
18:49

BBC Staff Concerned About Matt Brittin Praise For AI 'Doctor Who'

BBC staff want Matt Brittin to take back praise for an AI-generated Doctor Who episode. Deadline says employees would like him to travel back in time and scrub the comments. No clip or quote from Brittin is in the stub.

Full text · 147 chars
BBC employees would like Matt Brittin to travel back in time and scrub from the record his comments about an AI -generated episode of 'Doctor Who.'
19:04

Prompt engineering is already becoming obsolete. The next skill is getting AI agents to ...

A LinkedIn post says prompt engineering is already going stale and the next skill is making agents actually operate. It points at Claude Code as the direction. The capture is a truncated teaser with no examples or data.

Full text · 150 chars
Prompt engineering is already becoming obsolete. The next skill is getting AI agents to actually operate. Claude Code is moving in that direction: ...
19:14

How to Become an AI Engineer from Beginner to Job-Ready: Complete Roadmap 2026

A 2026 career roadmap stub lists Python, SQL, machine learning, data analytics, generative AI, and prompt engineering for becoming an AI engineer. It is a course-site outline. No hours, cost, or placement data are in the capture.

Full text · 150 chars
Learn how to become an AI Engineer with the 2026 roadmap covering Python, SQL, Machine Learning, Data Analytics, Generative AI, prompt engineering ...
19:16

How to choose a major (in the age of AI ) - NDSMC Observer

An NDSMC Observer column on choosing a major in the AI age is written by someone who is now an AI researcher and startup investor. They are giving a talk called “Aristotle for Founders.” The advice itself is not in the stub.

Full text · 147 chars
Today, I'm an artificial intelligence researcher and startup investor coming “home” this week to give a talk called “Aristotle for Founders” on ...
19:18

Building confidence through innovation at Compass Data & AI | Hire Waterloo

The University of Waterloo posted that Mehta finished an eight-month Data & AI Engineer co-op at Compass. The stub praises Compass’s student growth. No project details are included.

Full text · 151 chars
Mehta recently completed an eight-month work term as a Data & AI Engineer co-op at Compass. Compass's commitment to student growth, experimentation ...
19:29

Why Your AI Agents Can't Talk to Each Other (Yet) — Vlad Luzin, BAND - BigGo Finance

BAND’s Vlad Luzin argues multi-agent systems still cannot talk to each other cleanly. The BigGo Finance stub says engineers already run a planning session and a review session. No protocol name or demo is in the capture.

Full text · 147 chars
Vlad Luzin, co-founder and CTO of BAND, argues that the multi- agent systems engineers already run — a planning session and a review session in ...
19:33

AI Keeps Pissing Off Mathematicians

Full text · 125 chars
Description: Visit https://brilliant.org/coldfusion to try Brilliant's tutor for free and get 20% off an annual subscription.
19:39

President Trump Remarks at Science and Technology Summit in Washington, D.C. | Video

C-SPAN posted video of President Trump hosting a science and technology summit in Washington. The stub says the event highlights technology, AI, scientific innovation, and an updated science agenda. No transcript is in the capture.

Full text · 149 chars
President Trump hosts a summit to highlight technology, artificial intelligence , scientific innovation, and his administration's updated science ...
19:55

Worries around artificial intelligence cross party lines, with the vast majority of Americans ...

A Scripps News Facebook post says worry about AI crosses party lines. The vast majority of Americans, it says, want the U.S. government to prioritize keeping… and the sentence cuts off. No pollster, date, or percentages are in the stub.

Full text · 150 chars
Worries around artificial intelligence cross party lines, with the vast majority of Americans saying the U.S. government should prioritize keeping ...
20:41

AI policy should not be written by a handful of billionaires who sat in the front row of Donald ...

Senator Elizabeth Warren posted that AI policy should not be written by a handful of billionaires who sat in the front row of Donald Trump’s inauguration. The Facebook stub cuts off at “who make financial…”. No bill text is attached.

Full text · 143 chars
AI policy should not be written by a handful of billionaires who sat in the front row of Donald Trump's inauguration and who make financial ...
20:42

AI Governance: From Strategy to Execution Digital - IAPP Store

IAPP is selling the official textbook for its AI Governance Professional certification. “AI Governance: From Strategy to Execution” is the AIGP® course book. No table of contents or price is in the stub.

Full text · 140 chars
AI Governance: From Strategy to Execution” is the official textbook for the IAPP's AI Governance Professional (AIGP®) certification program.