Nothing matches those filters.

Lead

20
Google's Gemini 3.8 Live Adds Real-Time Reasoning to Voice AIAlphaSignalYour Agent Aced the Task. Will It Do It Again?Hugging FaceWhat’s at stake in AI’s trillion-dollar gambleMIT Technology ReviewFactory Raises $200M, Tripling Its Valuation to $5B in Five MonthsAlphaSignalOdyssey Builds One AI Backbone to Control Robots, Cars, Drones, and GamesAlphaSignalLM Studio 1.1.3 Brings Private On-Device Voice Transcription to LinuxAlphaSignalMultiverse Computing's Quasar 1.1 Uses Quantum Data to Shrink a 438B ModelAlphaSignalAnthropic Brings Claude Inside Salesforce With 37 Built-in Sales SkillsAlphaSignalAI models need more data about biology, and OpenAI is paying to create itMIT Technology ReviewThe Sequence Knowledge - Issue 933: When the Factory Starts Building ItselfTheSequence😺 Microsoft wrote a constitution for AIThe NeuronHypit Lets Claude Code Turn Short-Form Videos Into Editable CodeAlphaSignalLexical Prompt Compression for Large Language Models: A Training-Free, Deterministic Pipeline with Empirical Pareto Analysis Across Eleven Task CategoriesarXivClinical Reasoning Under a Partially Observed Objective in Cone Beam CT Report GenerationarXivCausal Analysis and Mitigation of Spurious Onsets in Full-Duplex Speech LLMsarXivHarmfulness Propagation Dynamics: Layer-wise Trajectories of Adversarial Intent in Large Language ModelsarXivSame Patient, Different Order: Action-Level Reliability of Clinical LLM Agents Under Repeated RunsarXivOpenAI's GPT-Live-1 Tops Voice Benchmark by Splitting Speech from ReasoningAlphaSignalClaude Is Now Leaving Invisible Fingerprints In Its TextTwo Minute PapersNew BEST local AI music generator is here!AI Search

Video

2
03:24

New BEST local AI music generator is here!

An open music generator now writes an editable score first, then sings it, so you can change the key or steal a melody without starting over. The tool is called UA 2. It can run from a 3.96 GB quantized checkpoint on about 4 GB of VRAM, or a 7.8 GB BF-16 file that the presenter fitted in 8 GB. A song-quality index in the video puts it ahead of MiniMax Music, AEP 1.5, Suno V6, and Suno 5.5, but slightly under Suno V5. Code and docs are Apache 2; the weights are Creative Commons non-commercial. The walkthrough uses ComfyUI, a 1.4 GB audio encoder, and a 32-step sampler.

Notes
  • Presenter: UA 2 (spoken “U-A-2” / “U2”). Claims best open-source music generator; 4 GB VRAM or less with the quantized checkpoint.
  • Pipeline is not text→finished song. It writes an editable ABC / sheet-music score (vocal + instrumental notes, key, tempo), then renders audio from that backbone. That is why covers, lyric swaps, and major→minor flips are cheap.
  • Demos in the transcript: smoky jazz; 1930s boogie woogie; Spanish flamenco; Russian folktronica; Korean emo. Reference-audio covers: Auld Lang Syne → jazz-funk; Beethoven → theatrical hard rock with “Where’s my wallet?” lyrics; Jingle Bells → minor. Agent loop shown with GPT Astra or GLM: keep melody/lyrics, reharmonize jazzier, then drop guitar.
  • Benchmark (presenter’s “song quality index,” source not named beyond the on-screen chart): beats MiniMax Music, AEP 1.5, Suno V6, Suno 5.5; slightly under Suno V5 (which this chart ranks above 5.5 and 6). Only open-source tool he calls good at covers.
  • Install path (ComfyUI, not the raw Python repo):
  • Update ComfyUI (update folder → update Comfy.bat).
  • Drag the UA2 workflow JSON onto the canvas.
  • Download audio encoder 1.4 GB → ComfyUI/models/audio_encoders.
  • Checkpoint: BF-16 7.8 GB (he says ~8 GB VRAM) or quantized 3.96 GB for <4 GB VRAM → ComfyUI/models/checkpoints.
  • Press R to refresh; pick the checkpoint on Load Checkpoint.
  • Text-to-music: style prompt (genre/instruments) + lyrics with metatags (verse, chorus, bridge, …). generate ABC preview → generate music. Max duration default 360 s (6 min); shorten if lyrics are short or it will keep going. Seed fixed at 7 in the demo (same prompts = same song); randomize to vary. 32 steps default. Raise CFG if style or lyrics drift. Regular decode if he assumes >12 GB VRAM; else tiled. His laptop RTX 5000 / 16 GB took just over 2 minutes.
  • Instrumental-only: style says instrumental; lyrics are square-bracket instrument directions (e.g. staccato strings + ethnic drums).
  • Cover / reference: un-bypass purple load-audio nodes (Ctrl+B), bypass top generate-ABC, wire encoder output into generate-music. Choose melody only vs full-song reference. New style + optional new lyrics (including another language).
  • License: code/docs Apache 2; weights Creative Commons non-commercial (“not primarily intended or directed towards commercial advantage or monetary compensation”). He is unsure about Spotify profit.
  • Sponsor block (Luma) is ad, not product. Agent hook is real: text-iterate the score after first render.
Transcript · 18,942 chars
This is currently the best open- source music generator you can use. You can run this with 4 GB of VRAM or less. Plus, you can even do cover songs and edit existing songs. And according to some benchmarks, this even beats Sunno V6, which is pretty crazy. So, this is called UA 2. And first of all, here are some demos. So, on the left is the prompt dictating the style and genre of the song. And on the right are the lyrics. First, here's a demo of a smoky jazz song. >> [music] [music] [music] >> The glasses clink in a quiet tune [singing] [music] under the hum of a fading >> [music] >> Lights flicker low. Whispers in bloom. [music] Neon smiles and a quiet laugh. [music] Friendships build [singing] on a fragile draft. Before I sleep, I'll keep [music] this far. a glowing memory [singing] in the dark. [music] Or here's another demo of a boogie woogie style of the 1930s. >> Hey, can you hear it? That shuffle through [music] the door. Old shoes on a new floor. My feet itching [singing] for more. Hey John, can you hear it? Boogie woogie. This rhythm [singing] turns me on. Let's go dancing [music] soon. I'm ready. I can't stand still. Come on. Hey John, can you hear it? Boogie [music and singing] woogie. Heartbeating like a tune. Spin me around this crowded room. Let's go dancing soon. [music] [music] Now, this also supports different languages. So, here's a flamingco example in Spanish. comfort [singing] [music] and truly. [singing] Hey [music] >> [music and singing] [singing] [music] [singing] [music] >> Or here's a folkronica example in Russian. Yeah. [music and singing] [music] No way. [music] >> [music] [music] >> or here's an emo example in Korean. [music] [singing] [music] [singing] >> [singing and music] >> up. Now, the really cool part about this is you can also input another song as a reference. So, for example, it can use the melody of an existing song, but you can change up the style. For example, here is old lang sign, but we are going to turn this into a groovy jazz funk style instead. [music] [music] >> [music] >> Should all acquaintance [music] be forgot and never brought [music] to [singing] mind? Should all acquaintance [music] be forgot and days of old [music] lang for >> or here's an even crazier example where we can input this Beethoven track. Once I play it, you'll probably recognize the melody here. We are going to turn this into a theatrical hard rock and then the lyrics are just going to be Where's my wallet? Where's my wallet? [singing] Where's my wallet? Where's my wallet? Where's my wallet? Where's my wallet? Where's my wallet? Where's my wallet? Where's my wallet? Where's my wallet? Where's my wallet? [music] Where's my wallet? Where's my wallet? Where's my wallet? Where's my [music] wallet? Where's my wallet? Or here's another example where we can input the Jingle Bell song as the reference audio, but instead let's turn this into a minor version. [music] [music] Dashing through the snow in a one-horse open sleigh. [music] Over the fields we go, laughing all the way. Bells on bobtails [music] ring, making spirits bright. What fun it [singing] is to ride and sing a slayighing song tonight. >> Jingle [music] bells, jingle bells, jingle all the way. Oh, what fun it is to ride in a one-horse [music] open sleigh. >> So, this is a very flexible tool and the only open- source one so far that can generate some good sounding covers. Or here's another cool thing you can do. You can even just link this to an AI agent like GBT Astra or GLM and it can do the music generating or editing for you. So for example, let's start with this pop song. Let me play this for you first. [music] >> [music] >> And then afterwards you can iterate this further. For example, we can tell it to keep the melody and lyrics but make it more jazzy through reharmonization, >> but it still doesn't really sound jazzy enough. So let's tell it to remove the guitar. [music and singing] All right. So this is just a really quick and simple example of how you can get an agent to generate and then iteratively edit the song just with text prompts. Now, before we go over the installation, it's worth noting how this actually works. So, instead of just turning your text prompt into a complete finished song, what it does is it actually writes out an editable score first. So, here's an example where it writes out the notes of both the vocal track and the instrumental track, plus of course the key and tempo. This is essentially a basic sheet music for the song, and then it uses this as the backbone to actually generate the full song. And because of this, it gives you some very flexible editing capabilities. That's why it's very easy to generate cover songs from this or to change up the lyrics or even change the key from major to minor. And get this, if you look at these benchmarks in terms of this song quality index, not only does it beat the other open source models out there like Miniax Music and AEP 1.5, but it even beats some closed models like Sunno V6 and Sunno 5.5, which is pretty crazy. However, it does perform slightly under Sunno V5, which interestingly, at least according to this benchmark, sounds better than version 5.5 and version 6. If you do any kind of content creation, definitely check out Luma, the sponsor of this video. Think of it as an aentic AI workspace that can work alongside you through your entire creative process instead of just giving you the results of a single prompt. For example, let's say I want to create an entire marketing campaign for a new product. Instead of jumping between a bunch of different AI tools, I can just get Luma agents to autonomously do everything. It can develop the concept, generate the visuals, and shape the project all within the same workspace. What makes it especially interesting is that its Luma agents understands things like motion, physics, and 3D space. Rather than completely taking over the creative process, you can continuously guide the agent, change direction, and refine the results as you work. And one of the most powerful features is Luma's skills. You can basically create reusable skills for workflows you do all the time. Basically, you give Luma a set of instructions once and then you can run that same workflow on different assets whenever you want. For example, I can create a skill where I input any product photo and it'll output some UGC videos of an influencer talking about the product. Or here's another example of a skill where I can upload the product photo and it'll drop it into water like this. Luma basically gives you an intelligent creative co-pilot that can consolidate all your creative workflows into one place. Whether you're creating marketing campaigns, branded content, product visuals, or social media content, Luma is one of the best platforms you can use. Try Luma today using the link in the description below or by scanning the QR code here. If you're interested in trying this out, next, let's go over how to install this. Now, if you click on this GitHub link at the top and you scroll down a bit here, it does contain all the instructions on how you can run this. But the default code is like this where you need to work with some Python code which might not be intuitive for everyone. So instead, we are going to run this in a visual interface called Comfy UI. This is the most popular platform for running open-source image, video, and audio generators locally on your computer. In fact, if you're not familiar with Comfy UI, I highly recommend you watch this video first where I go over how to install and use it. Now, assuming you do have Comfy UI, the first thing you should do is update it to the latest version. So, in your root comi folder, simply click into this update folder and then double click on update Comfy.bat. So, this will proceed to update your Comfy to the latest version. Afterwards, let's press any key to continue to exit the window. And then next, let's start up a fresh session of Comfy UI. All right, after loading up Comfy UI, what you need to do is drag the U2 workflow onto your interface. So, I will link to this page in the description below where you can download the full workflow. Simply click on this link so that it will download this JSON file. Now, you can save this wherever you want. I'm just going to save it in my comfy root folder. Afterwards, simply drag this UA2 workflow onto your Comfy interface, and it should magically open this pre-built workflow for you. So you don't have to build out anything yourself from scratch. Now when you first load this, you might see this error message involving missing models. So let's proceed to download the missing models. So I'm going to link to this page in the description below. You need to download the audio encoder and the checkpoint for this. Let's first click into the audio encoders folder and we need to download this file which is 1.4 GB in size. Let's click download. And this goes in Comfy UI in models and then in audio encoders. Let's click save. And then afterwards, we also need to click into the checkpoints folder. And here you can choose to download either the full BF-16 version, which is 7.8 GB in size. Or if you have like less than 4 GB of VRAM, then you can download this smaller quantized version, which is only 3.96 GB. For me, since I do have enough VRAM, I'm going to download this full version, which should be able to fit in like 8 GB of VRAM. So, this goes in Comfy UI, in models, and then in checkpoints. Let's click save. Now, back to our Comfy UI interface, simply press R to refresh your model list. And then for this load checkpoint node, simply open the dropdown and select the model you just downloaded. In my case, I'm going to click on this BF-16 model. And that should get rid of all the errors that you see. Now, this workflow has two components. The first component at the top here is just turning a text prompt into a full song. And then at the bottom here, you have the option of inputting an audio for reference. So you can make cover songs with this feature. Let's go over the text to music workflow first. At the top here is where you would enter the style prompt. So basically this would describe the genre, the style, the pacing, the instruments or other things that you want to specify about the song. For example, let's do something like future bass, modern, energetic, inspiring. And then at the bottom here is where we would enter the lyrics. Note that this takes in metatags. So you can specify like verse one, verse two, intro, outro, bridge, chorus, pre chorus, etc. For me, let me just show you a simple example with one verse and one chorus. Next, this would be fed through this generate ABC node, which basically generates the notation of the song. After I press generate, you can actually see a preview of this notation over here. It's basically sheet music like this, which dictates the vocal melody as well as the instrumental melody throughout the song. And then afterwards, this notation would be plugged through this node along with your style prompt and lyrics to actually generate the music. Here is where you can specify the maximum duration of your song in seconds. So right now it's set at 360 seconds, which is 6 minutes. Now, this is just the maximum duration. If your lyrics are shorter than that, then it's just going to generate a shorter song. And then next, it gets plugged through this K sampler to actually generate the music. Note that the seed is basically the unique ID of every song. Right now, it's set at seven and fixed. That means if you use the exact same prompts and the exact same settings as before, you're going to get the exact same song as before. So, if you want to generate a completely different song while keeping the same lyrics and the same style prompt, then you need to change the seed to another value. Or you can also set this to randomize afterwards. And then the number of steps is like how many steps it takes to generate the song. In general, the more steps you have, the higher quality the song will be, but it's going to take slower. And then if you use fewer steps, it's going to generate faster. I just tend to leave it at the default of 32 steps. The CFG is like how literally you want the AI to follow your prompts. So if you get a song that doesn't really follow your style prompt, or if it has some errors with the lyrics, then you could set this CFG to a slightly higher value to follow your prompt more literally. And then the sampler anduler are basically the algorithms used to generate the song. I just tend to leave it at the default values. And then afterwards, this goes through the decoding step. Now, if you look at this note here, it says if your GPU has enough VRAM, so I would assume like over 12 GB of VRAM, then you can just use this regular decode method, which is a lot faster. If you have less than 12 GB, then it's best to use this tiled version. So, since I do have over 12 GB, I'm going to connect this one instead. So, first I need to take this input and connect it over here. So, what I'm going to do is hold down shift and then click on this connection and drag it over here. And then afterwards, I just need to connect this audio output over here. And then for this one, I can just press Ctrl +B to bypass or disable it. And then finally, it will generate my song over here. So, let's press run. All right. Afterwards, let me pull up the stats for you. So, this was pretty quick. I'm using just my laptop with an RTX 5000, which has 16 GB of VRAM, and this took just over 2 minutes to generate. Next, let me actually drag the style prompt and the lyrics over here so you can see them. And let me play the generation for you. I used to wait for [music and singing] the world to change something [music and singing] on it way. But every road that I never chose [music] was just another door I close. So I'll light the fire. [music] I'll make the stars. Turn all these dreams [music and singing] into works of art. I don't need to know [singing] where the ending goes. [music] I just need a spark and the nerve to follow. [music] >> [music] >> Now, it kept generating much longer than what I inputed here. So, what I should have done instead is reduced the max duration to something that fits these lyrics a bit better. But there you go. In a nutshell, that is how you can run this textto music workflow. All right. Now, if we go back here to this preview ABC section, you can see a preview of the notation that it generated over here. So, these are basically like the notes of the vocal and the instrumental. Now, some of you are wondering if this can do instrumental only. And the answer is yes. So, for example, here's my style prompt. Epic cinematic orchestral music for a battle scene, instrumental only. And then for the lyrics, I do need to enter something. So, usually I just enter what I want the instruments to sound like within square brackets. So, for example, here I put steady buildup of staccato strings and ethnic drums. And here's my generation. [music] [music] Oh, [music] [music] heat, heat. All right. So, that covers text to music. But what if you want to input an audio clip to use as reference or to generate cover songs? Well, that's what this part is for. So, let me hold down control to drag across these purple nodes and then press Ctrl +B to unbypass these nodes. Basically, this allows you to upload a reference audio to generate the ABC notation to be plugged through the song generator. And this step basically replaces this generate ABC node over here. So, first of all, let me select this top generate ABC node plus the preview node that's connected to it. I'm going to select both of them and then press CtrlB to bypass these nodes. And what we would do instead is connect the output from this node into the generate music node. So what I'm going to do is connect the output over to here. All right. So this generate music node should now be connected to this bottom section over here. And then what we need to do is for this load audio encoder, click on this dropown and select the audio encoder that you just downloaded. And then here is where we can upload a reference audio clip. For example, let me upload this segment from a song. [music and singing] watching as you walk away [singing and music] while my heart breaks. >> All right, so that was the song. Next over here, we can choose to either use the melody of this as the reference or the full song as reference. It's a really subtle difference, but if you want to do a cover song or copy just the melody of a song over, then I would select melody only. Now, over here is where you would enter the new style that you want for this song, as well as the lyrics for the song. For this example, I'm just going to enter the same lyrics as my reference audio. And then for the style prompt, let's try jazz with casual piano, saxophone, double bass, and brushed drums. Let's press run. All right, here's our result. [singing] Watching as you walk away while my heart breaks today. [music] It's as simple as that. So that's how you can use this audio reference feature. Of course, you don't have to use the same lyrics. You can also change the lyrics to something else. For example, you can make the person sing another language. So, a super flexible tool. So, that covers all the components of this UA2 workflow in Comfy UI. This is currently the best open- source music generator you can use right now. So, definitely try this out. Finally, I also want to mention the license of this. So, the code and the documentation of U2 are under the Apache 2 license which has very minimal restrictions, but the model weights are separately licensed under this creative common non-commercial license. So, in this clause, it specifically says not primarily intended or directed towards commercial advantage or monetary compensation. So, that's something to be aware of. I'm not sure if you can like post a generation on Spotify and profit from it. Anyway, that sums up my tutorial and review of UA2. This is definitely the best open-source music generator available right now. So, definitely give it a try and let me know what you think of it. And if you run into any errors during the installation, welcome to copy and paste the exact error message in the comments below, and I'll try to help you troubleshoot as much as possible. As always, I will be on the lookout for the top AI news and tools to share with you. So, if you enjoyed this video, remember to like, share, subscribe, and stay tuned for more content. Also, there's just so much happening in the world of AI every week, I can't possibly cover everything on my YouTube channel. So, to really stay uptodate with all that's going on in AI, be sure to subscribe to my free weekly newsletter. The link to that will be in the description below. Thanks for watching and I'll see you in the next one.
10:12

Claude Is Now Leaving Invisible Fingerprints In Its Text

Claude's new text is being marked in a way you cannot see, and a few extra edits will not wash it out. Anthropic is rolling the watermark out now. The scheme nudges some candidate words ("green") so machines that know the key can count them; 21 greens in a paragraph can be less likely than winning the lottery if a person wrote it. Light edits survive. A full rewrite, or an open-weight model you run, can strip it. You and I cannot check. Only some eligible organizations can. The presenter thinks they are using a SynthID-style tournament.

Notes
  • Anthropic rolling out now (Two Minute Papers / Dr. Károly Zsolnai-Fehér). Not hidden characters. Statistical fingerprint: at each next-word choice, some candidates get a secret green (preferred) / red nudge. Detector who knows the coloring counts greens.
  • Example given: 21 greens in a paragraph can be less likely than winning the lottery if a human wrote it. Green list can change over time; words look ordinary.
  • Likely SynthID variant (context-dependent probabilities + tournament). Survives copy-paste and light editing. Does not survive a full rewrite that swaps every word, or a pass through an open-weights model you run.
  • Misconceptions he flags: (1) it does not trace the text back to a user; it marks that Claude wrote or heavily edited it. (2) “Change a few words” is not enough.
  • Who can check: not the public. Only “some eligible organizations” for now.
  • His prescription: run free open-weights locally. Sponsor: W&B Weave (wnb.me/papers).
Transcript · 3,623 chars
Images can be watermarked for copy protection. Now, get this. Claude AI announced that they are watermarking the text you generate with it. From when? When does this start? Well, Anthropic is rolling this out right now. Yep. I am not talking about this because I agree with it, but because I think it's important that all of you fellow scholars know about this to inform the public. So, this piece of text is watermarked. Wait, what? >> [laughter] >> You can watermark an image by putting your logo on it, but text? How would you watermark text? You put hidden characters in it, right? Nope. This paper describes that it is a fingerprint in text that is invisible to humans, but is detectable for machines. It even survives copy-pasting and some editing, too. I'll tell you what it doesn't survive in a minute. So, when generating text, the AI decides what the next word should be, and there can be a few candidates. Here, you could say, "I saw a dog, a puppy, a cat, or a house." Based on context, each of these words gets a probability to be chosen. Now, with a fingerprinting algorithm, it secretly assigns a color to each word. Some are green, preferred, some are red, not preferred. And now comes a little nepotism. A little cheating, if you will. When choosing the next word, the green ones get a little nudge upwards. They will occur slightly more often. So, here's how to check for a watermark. In a piece of text, someone who knows the red and green words simply counts how many greens you have. This scheme has a mathematical property where, as you see more and more green, the probability of it being real human text is extremely small. Found 21 grains in a paragraph, suddenly the probability of that done by humans can be less than winning the lottery, much smaller. Note that the green words can be anything, no matter how inconspicuous, so you can't spot them. Which words are green can also change over time. Ouch. This is the simplified version of the algorithm. They are likely using the SynthID variant, which has context-dependent probabilities and a tournament system, too. But the heart of the algorithm is the same in most research papers I read. Some words are preferred and are given a slight edge in the generation, creating a unique fingerprint. Now, there are a lot of misconceptions about this out there. One, Claude-written text cannot be traced back to you, but it shows that Claude wrote the whole thing or heavily edited the text. I don't agree with this. I am making this video to let everyone know. Knowledge only changes the world when it reaches people. Second, some say, "Just edit a few words and it's clean." Nope, you can't get rid of it so easily. So, can you get rid of it? Dear fellow scholars, this is Two Minute Papers with Dr. Károly Zsolnai-Fehér. With light editing, no. If you rewrite the whole thing, exchanging every word, yes, you can get rid of it. An open-weights LLM that works for you can also help. Okay, so who can check if there is a watermark in the text? Well, not you and not me. Some eligible organizations can, but that's it for now. So, what is the solution? Well, of course, use free and open-weights AI systems and run them yourself. These work for you, not against you. That [music] is the way of the scholar. We need new tools for the era of LLMs, and weights and biases now has weave a lightweight toolkit to confidently iterate on LLM applications. Use traces to debug how data flows through each step of your app and use evaluations to measure your progress. It is the best. Try it out now at wnb.me/papers or click the link in the description below.

Article

139
10:00

What’s at stake in AI’s trillion-dollar gamble

The companies building giant AI data centers will need an almost impossible burst of profit to justify what they are spending. Jessica Wachter estimates hyperscaler outlays near $1.1 trillion through 2027 and says they need 2.7 times more productivity by 2030 to break even after capital cost, a 15% return, and depreciation. This year’s build is about $750 billion against $150 to $200 billion of AI revenue, per Gary Gensler. Alphabet posted a $5.9 billion free-cash deficit after nearly $120 billion of revenue. Morgan Stanley says more than half of $2.9 trillion in 2025–2028 data-center spend will be external capital. Meta’s Hyperion site in Louisiana grew from $10 billion and 2 GW to a $50 billion, 5 GW plan with a maze of leases.

Notes
  • Wachter (Wharton; ex-SEC chief economist): do not forecast usefulness — ask how fast hyperscaler earnings must grow to justify spend through 2027, estimated near $1.1 trillion. Need 2.7× own productivity by 2030 after cost of capital, 15% return, and depreciation. Possible, she says, but that is mid-1990s IT-boom growth compressed into a few years. If it fails: “the largest misallocation of capital in history.” Missed interest payments risk bankruptcy.
  • This year: about $750 billion of data-center build. Some projections: >$5 trillion over four years from Alphabet, Microsoft, Amazon, Meta, Oracle (OpenAI partner). Gensler (MIT Sloan, ex-SEC): AI revenue $150–200 billion this year. “Spending does not have commensurate revenues yet.”
  • Stakes: investments could approach ~3% of GDP. Alphabet latest quarter: nearly $120 billion revenue, $5.9 billion free-cash deficit — first since the 2004 IPO. GPU electronics ~60% of data-center cost; performance roughly doubles every two years, so this year’s halls need another chip generation by decade’s end or become “hulks” (Kshirsagar, Princeton).
  • Gensler’s “parlay”: (1) hyperscaler revenues, (2) economy-wide productivity, (3) frontier models beating cheaper good-enough ones. Van Nieuwenburgh (Columbia): 183 GW planned 2025–2032 at ~$41 billion/GW → required annual revenue ~$3.7 trillion by 2032 at a 10% return.
  • Acemoglu: without productivity gains, people sour and investment and revenue both fall. Survey of ~6,000 executives (US/UK/Germany/Australia): ~90% saw no productivity lift in three years; they expect ~1.45% over the next three (2.25% US) and ~$280 billion private AI spend by end-2026 — via more sales and fewer employees.
  • Morgan Stanley: more than half of $2.9 trillion 2025–2028 data-center spend is external capital. Risk sits in lenders, guarantors, private-credit funds, pensions, insurance.
  • Meta Hyperion, Richland, Louisiana: announced late 2024 at 2 GW / ~$10 billion; cost later $30 billion; 80% stake to Blue Owl via Beignet JV; four-year leases to Pelican Leap LLC with residual-value guarantee. July expansion: 5 GW / $50 billion. Entergy: first three gas plants, now seven more, ~7.5 GW — about 6× New Orleans. Entergy cites a 20-year Meta power guarantee; UCS and Alliance for Affordable Energy doubt ratepayers are fully shielded.
  • Gensler: a retrenchment is the historical base case (spend goes flat next year, or 2028–29 when capacity looks sufficient). Pande blog: a crash could be “the best thing that happens to this technology.” Piece warns SPVs and distributed credit echo pre-2008 plumbing, then notes fiber from the telecom bubble still carries the modern internet.
Full text · 20,645 chars
When Jessica Wachter, a finance professor at the University of Pennsylvania’s Wharton School, wanted to assess AI’s impact on the economy over the next few years, she faced a long list of business and technical uncertainties. So she started with what she calls a “remarkable fact” that is not in question: A handful of so-called hyperscalers are investing huge amounts of money to build AI data centers. Instead of trying to predict how useful and widely deployed AI models will be, she simply asked how fast the hyperscalers’ earnings will need to grow to justify their spending through 2027, when—she and her collaborator estimate—expenditures will reach nearly $1.1 trillion. It's a no-nonsense accounting approach to making sense of today’s historical AI buildout. The results are eye-opening: The AI companies will need to increase their own productivity by a factor of 2.7 to break even by 2030, accounting for the cost of capital and a 15% return, and depreciation of the assets. Not impossible, says Wachter. The result would lead to the kind of economic growth that we saw during the US IT boom over a period of about 10 years starting in the mid-1990s. But, she says, for it to happen by 2030 “that’s a lot of growth compressed into a few years.” And if the hyperscalers cannot meet such profit goals? “Then they will fall behind on their interest payments, and that risks bankruptcy,” says Wachter, who was previously the SEC’s chief economist and director of its division of economic and risk analysis. If a productivity boom “fails to materialize,” she and her coauthor conclude in their research paper, “the current buildout will be the largest misallocation of capital in history.” It doesn’t take superintelligence to realize that today’s large investments in the infrastructure for artificial intelligence come with huge risks. The hyperscalers will spend about $750 billion this year, building massive data centers scattered across the country. And the spending spree shows no signs of slowing. According to some projections, total AI capital investments from the hyperscaler companies—Alphabet, Microsoft, Amazon, Meta, and Oracle (which partners with OpenAI)—could be more than $5 trillion over the next four years. It’s one of the largest capital investments by any industry in history. But there’s a problem that’s obvious to anyone paying attention. While the hyperscalers plan to spend trillions, total AI revenues will be around $150 billion to $200 billion this year, says Gary Gensler, who ran the SEC during the Biden administration and is now a professor at MIT’s Sloan School. “The challenge is that the spending does not have commensurate revenues yet. That’s a fact,” he says. “And then the question is, is that an investment that will be paid off in the future?” At stake in that trillion-dollar question is the financial health of the giant AI companies and the overall US economy—the investments could soon balloon to around 3% of GDP. The answer could also determine the fate of the hugely expensive data centers themselves. No one really knows how profitable and useful these multibillion-dollar behemoths will be down the road. Though AI models have made dazzling progress over the last few years, it’s anyone’s guess how much compute capacity we will need. The technology could become more efficient and therefore less dependent on raw computational power. Or demand for AI products could slow, or customers could turn to cheaper models. The risks, both to investors and to the economy, have become even greater this year, as these AI companies have begun borrowing large amounts of money to build more and more data centers. Free cash flow—operating cash flow minus capital expenditures—is expected to soon dip into negative territory for the group. Even Alphabet, known for generating and hoarding huge amounts of cash, reports in the latest quarter that its impressive revenues of nearly $120 billion were devoured by AI infrastructure spending, leaving it with a free cash deficit of some $5.9 billion—its first shortfall since Google went public in 2004. In the near term, it’s not a big financial worry for most of the companies. They make a lot of money and have very deep pockets. But debt is expensive, and some investors are losing patience. If future demand for the data centers’ computation power drops, the companies will still be on the hook to pay back the borrowed money. What’s more, the risks are spreading to the rest of the economy as the loans get passed along via various financial mechanisms. It won’t be enough to simply cover the enormous price tags of the new data centers. Hyperscalers will also have to pay for the rising costs of capital as they borrow more money. They will need returns that are impressive enough to justify all their spending to investors and creditors. And to add to those concerns, they will have to make up for the depreciation of billions of dollars in chips housed within the facilities—a ticking time bomb buried in the investments. Performance of the expensive GPU chips at the core of the data centers—such compute electronics represent some 60% of costs—is roughly doubling every two years or so. The pace of progress helps explain the increasing wizardry of the AI models, but it comes with a cost. Owners of AI data centers that come online this year and next will need to spend billions more on the next generation of chips by the end of the decade if they want to stay competitive. Without the investments, says Mihir Kshirsagar at Princeton’s Center for Information Technology Policy, the data centers risk becoming “hulks,” stranded assets “scattered all over the place.” To put it bluntly: The AI companies need to start making a lot more money. And they need to do it fast. But juicing their earnings alone still won’t be enough to sustain their data-center investments for the long term. Productivity is everything At some point, AI is also going to have to create broad economic growth to justify continuing the hyperscalers’ spending spree. Sloan’s Gensler describes today’s large investments into AI infrastructure as “a parlay bet by the capital markets and the economy.” That means success will require winning three related but independent wagers: Hyperscalers must generate massive revenues, AI must boost widespread economic growth, and both must happen while the powerful but expensive so-called frontier models that rely on the data centers fend off cheaper versions, which many businesses might find good enough. What makes this so tricky is that each wager depends on the other two but also poses its own challenges. If the hyperscalers continue to spend huge amounts of money on data centers into the next decade, revenues will need to skyrocket into the trillions. Stijn Van Nieuwerburgh, a finance professor at Columbia Business School, bases his estimates on a scenario in which about 183 gigawatts of planned AI compute capacity is built between 2025 and 2032; he calculates that each gigawatt costs about $41 billion. Assuming a 10% return—the minimum that would be acceptable to most investors—“required” annual revenues will be roughly $3.7 trillion by 2032, he says. Others get a similar number. Winning the second part of the bet—productivity growth across the economy—will be crucial to achieving such numbers. For a few years, AI companies could likely boost their revenues by simply selling subscriptions and tokens to all the businesses clamoring to get into AI. But eventually—and this might be happening already—those paying customers will need to justify their expenses by seeing bottom-line benefits from the technology. AI will need to fulfill its promise of making workers more productive and making businesses more efficient and profitable while expanding their products and services. In economic jargon, that means customers will need to see productivity growth. Taken together, these results will mean the country is prospering and growing. “If you don’t get the productivity gains, at some point people are going to sour on AI, and that will bring down investments and it would also limit revenue growth,” says Daron Acemoglu, an MIT economist and 2024 Nobel laureate. For the investments to be sustainable over, say, the next five to 10 years, we definitely “need to see productivity gains,” he says. Most economists who watch the numbers closely agree that, for now, the economy-wide statistics show little or no productivity growth from AI. There are some hopeful signs it’s on the way, though. In a recent survey of some 6,000 senior business executives in the US, the UK, Germany, and Australia, the vast majority—around 90%—report no increase in productivity over the last three years. But they expect a boost of around 1.45% in total over the next three years; US executives anticipate a 2.25% bump over that time. In a follow-up survey, the respondents also reported plans for their businesses to spend more on AI, leading the authors to anticipate some $280 billion in private-sector AI expenditures by the end of 2026. That’s good news for the hyperscalers. But it comes with a dose of bad news for those worried about AI’s impact on jobs. The executives expect to increase the productivity of their companies by increasing their sales while significantly cutting the number of employees. If AI improves productivity by destroying jobs, public backlash to the technology—the kind we have seen around data centers, for example—will likely get worse. Perhaps it’s worth adding one more wager to the parlay bet described by Gensler: The public and local communities must feel that they are also benefiting from the massive investments in AI. And let’s not forget how interdependent these wagers are; if productivity growth comes from companies running models like DeepSeek, then the hyperscalers’ revenues could collapse. If productivity comes from cutting jobs, a public backlash could block many of the planned investments—and stunt anticipated revenues. We will need to win all the wagers for the hyperscalers’ bet to pay off. We’re all part of the AI gamble now It was one thing when the AI companies were spending cash they had accumulated over the years to build their own data centers. Then the risk was largely limited to their own balance sheets and shareholders. But it’s a higher-stakes game when much of the money is borrowed. Morgan Stanley, for one, calculates that more than half of the $2.9 trillion that hyperscalers will spend between 2025 and 2028 to build AI data centers will be financed with “external capital.” The borrowing is leading some of the companies to engineer complex webs of financing that are becoming intertwined with much of the rest of the economy. “A lot of financial institutions, directly or indirectly, are exposed to these data centers either as lenders, or as guarantors of some of the debt, or as backers of the private credit funds who are funding these data centers,” says Columbia’s Van Nieuwerburgh. “People don’t even know they’re holding this stuff. It’s somewhere deep inside their pension fund. Ultimately, it’s backing their life insurance policies. And that risk is getting distributed everywhere in places that are invisible.” As the investments in data centers have spiked, the financial engineering has become more byzantine. Take, for example, Meta’s so-called Hyperion data center under construction in Richland, Louisiana. When the company announced the two gigawatts of compute capacity at a price tag of some $10 billion in late 2024 it was Meta’s largest planned data center. Greeted with much enthusiasm by state and local politicians, the project, located in the rural northeast corner of the state, was seen as a boon to the community. Entergy Louisiana, the state’s largest utility, rushed forward with proposals to build three large natural-gas power plants to service the massive data center. Then last fall—the projected cost was now $30 billion—the financing got a lot more complex and, to some in the community, a lot more disconcerting. Meta transferred an 80% stake to the large (and troubled) private-credit firm Blue Owl Capital, forming a joint venture called Beignet (like the famed New Orleans pastry) to raise financing for the data center. Meta then signed a series of four-year leases with the joint venture, an arrangement that the company says gives it “long-term strategic flexibility.” To backstop the agreement, Meta provides the venture with what is called a residual value guarantee, in which it will make a cash payment to cover the value of the facility “following any non-renewal or termination of a lease.” Got all that? I hope so. The financial wheeling and dealing is actually even more convoluted, with a cast of wholly owned subsidiaries and LLCs. Beignet has set up Laidley LLC, which owns and operates the site as the landlord. In turn, Laidley leases the facilities to Meta’s wholly owned subsidiary Pelican Leap LLC, which is the tenant. And there is a series of four-year leases that cover the different buildings that make up the data center campus. It’s not a coincidence, says Van Nieuwerburgh, that the length of the leases matches the expected lifetime of the data center’s GPUs. While Meta has to pay off its loan if it terminates the leases early, that will still leave its investors “with an empty building and no cash flow,” he says. “And then they need to find a new tenant for a huge data center, and good luck with that.” Meanwhile, Meta is doubling down on its bet. In July, the company announced it was expanding the data center to five gigawatts of compute capacity. The total price tag is now $50 billion (so far, Meta hasn’t said whether Blue Owl will be involved in financing the expansion). Meanwhile, Entergy is now planning to build seven more gas-fired power plants, bringing the total capacity of the facilities to around 7.5 gigawatts—some six times the amount of electricity used by New Orleans. If the complex financing is a puzzle to many investors and even financial experts, it is even more baffling to those directly affected by the construction of the data center. The main worry concerns how Entergy’s spending on the natural-gas power plants will affect electricity prices, and who will be left paying the bill for the power if Meta walks away. Entergy says it has a 20-year guarantee from Meta that the company will purchase electricity over that period to cover the costs of the power plants and related infrastructure. But there are skeptics, especially given how fast the fortunes of the AI industry are changing. “In four years, is Mark Zuckerberg still going to be interested in this? Or is he going to throw in the towel?” asks Paul Arbaje, a senior analyst at the Union of Concerned Scientists, which has been advocating, largely unsuccessfully, for the Louisiana Public Service Commission to provide more transparency around the data center and its financing. Even if the 20-year deal holds, consumer advocates are worried that Meta or its partners won’t fully cover all the costs, including those associated with operating and maintaining the power plants—and those additional costs that could be passed on to residential ratepayers. What’s more, says Logan Burke, the executive director of the Alliance for Affordable Energy, if Meta doesn’t end up needing as much power as Entergy planned (these projections are not public), consumers could be left paying for the surplus produced by the plants. And if Meta terminates its leases early? “It gets complicated very quickly,” says Burke, who questions whether the shifting roster of financial entities will honor existing agreements. “That everybody is going to do what they’re saying they’re going to do over the next 20 years is just hard to believe.” For UCS’s Arbaje the bottom line is this: “They’re making huge bets that these data centers will be worth it. Bet with your own money, not with ratepayer money.” After the bubble Predicting when the AI investment bubble will burst is a fool’s errand. But there is little doubt a day of reckoning is coming, given the irrational exuberance that has overtaken the hyperscalers and their investors. Of course, you might argue that this time is different, and that the rules of accounting and lessons of economic history don’t apply—that AI is too transformative. Maybe, but don't count on it. “History tells us that at some point you get a retrenchment, and it’s just a question of when and how severe,” says Sloan’s Gensler. It could be that today’s $750 billion spending rate “goes flat” or decreases next year. Or, he suggests, “we’re now in 2028 or 2029, and then all of sudden they’re retrenching because they’ve got enough capacity.” But, he adds, “you can be pretty assured there'll be a retrenchment.” Though a so-called retrenchment might be inevitable, it’s worth keeping in mind that the fates of the financial bubble and the underlying AI technology revolution could be very different. Already, some Silicon Valley insiders are rooting for a crash; in a recent blog post the longtime venture capitalist Vijay Pande wrote that “the coming crash would be the best thing that happens to this technology.” The argument makes some sense. A crash could make AI investments more rational, calm the impulse to build billion-dollar data centers on every vacant field that CEOs fly over, and refocus investors on how to use the technology to create sustainable value. But we should probably be careful what we wish for. After the bursting of the dot-com bubble at the beginning of the 2000s, hundreds of thousands lost their jobs, large and small companies alike went bankrupt, the economy of Silicon Valley and San Francisco was decimated (at least for a while), and the shocks sent the US into a mild recession in 2001. For the financial community and many tech workers, it was no fun. Even more devastating for the economy and the average American was the great recession that began in late 2007. Comparing the financial engineering leading up to it and the methods deployed by hyperscalers today is sobering. So-called special purpose vehicles (SPVs) are back! If Columbia’s Van Nieuwerburgh is right about the dangers of letting investments from the hyperscalers get entangled throughout the economy, the fallout could be severe. But technologies survived and even prospered in the aftermath of both downturns. The early 2000s, even in the face of the dot-com fiasco, were a time of great innovation and tech optimism. The froth came off the spending on silly technologies, helping to focus investments on more promising ones. It’s no coincidence that each of the hyperscalers rose out of the ashes of the crash or started up shortly after. The fiber-optic infrastructure built during the feverish telecom bubble that ran parallel to the dot-com one is still the backbone of much of today’s communication infrastructure; we wouldn’t have Facebook or Amazon or Google without it. This time, however, we’re facing a unique risk: The huge financial investments by the hyperscalers have ensnared the future of AI itself with the fortunes of the massive data centers spreading around the country. The logic is founded on a deeply held belief about the power of scaling in AI; the bigger you build it, the smarter it gets. That might be true, but it’s unproven and a risky bet. There are already plenty of red flags, from strong public opposition to the construction of new data centers to the competitive threat from cheaper, good-enough AI models to the rapid improvement of small, local AI models. None of these trends point toward a future dominated by frontier models housed in massive, billion-dollar data centers. The financial bubble around the colossal spending by the hyperscalers will likely burst eventually—or maybe soon. It might be financially painful, but we’ll survive. Wall Street will survive. AI itself will survive, though it may look different and lose some of today’s hubris. The financial fate and future utility of the massive data centers fueled by trillions of dollars of spending, on the other hand, are far less certain. Deep Dive Artificial intelligence A fundamental flaw leaves LLMs strikingly vulnerable to attack It makes it easy to trick them into doing things they shouldn’t, such as telling you how to sabotage an aircraft’s navigation system. AI is more likely than humans to form biases when hiring AI doesn’t just learn stereotypes from its training. It can cook up new ones, too. Stay connected Get the latest updates from MIT Technology Review Discover special offers, top stories, upcoming events, and more.
16:00

Your Agent Aced the Task. Will It Do It Again?

An agent can look great on a leaderboard and still fail the same request the next time you ask. On AppWorld, a ReAct agent with GPT-4.1 hit 77.4% average success across five repeats but succeeded on all five runs for only 53.0% of tasks — a 24.4-point consistency gap. IBM’s Consistency Analyzer resamples each decision in one saved trace (k=5 by default) and writes guidelines that plug into ALTK-Evolve. That cut the gap to 12.0 points, with Pass^5 up 16.0 points on the same task and 13.0 on a similar one, while average accuracy did not fall. The toolkit is open source. The runs used temperature 0.0, so this is not ordinary sampling noise.

Notes
  • Problem: Mean@k (average pass rate) hides flip-prone tasks. AppWorld test_normal, ReAct + GPT-4.1, five repeats: Mean@5 77.4%, Pass^5 53.0%, consistency gap 24.4 pp (up to 30 on hard). Temperature 0.0 — not ordinary sampling.
  • Pass^k ≠ Pass@k. Pass^k = succeed on every run. Always Pass^k ≤ Mean@k ≤ Pass@k.
  • Cause they argue: flat next-token distributions at decision points. Hosted endpoints still nudge near-ties even at greedy/seed. Many chained steps compound.
  • Consistency Analyzer (ALTK-Evolve): one recorded trajectory, no ground truth, no live re-run. One extra call per decision requesting k=5 completions against the saved context. Writes a per-step consistency scorecard.
  • Guidelines are ordinary ALTK-Evolve items. Example from “How many activities… SimpleNote”: (1) line-anchored regex for checkbox markers, not substring count — titles repeat the marker in a legend; (2) verify search hits before proceeding. Demo: five parallel runs split 3–2, then all five agree after the guidelines.
  • Eval: 168 AppWorld tasks, guidelines from one baseline trace, then 5 fresh runs. Pass^5 53.0 → 69.0, Mean@5 77.4 → 81.0, gap 24.4 → 12.0. Medium +22.9 pp, Hard +14.3 pp, Easy +12.2 pp. Similar-task transfer +13.0 pp Pass^5. Weaker gpt-oss-120b: same-task 10.1 → 16.1 (+6.0), similar-task +8.7. Mean@5 never dropped.
  • Open source: github.com/AgentToolkit/altk-evolve. Technical report on arXiv (link not pasted as a full URL in the body).
Full text · 10,827 chars
That is embarrassing onstage. In production, it is a reliability problem: a workflow that succeeded once may fail the next time a user makes the same request. For mission-critical work, such as reconciling a financial transaction or checking a contract for an obligation, that can be a showstopper. Most benchmarks hide this variability behind an average. On AppWorld, a ReAct agent using GPT-4.1 succeeded on 77.4% of runs across five repetitions. But it succeeded in all five runs for only 53.0% of tasks — a 24.4-point consistency gap. Most benchmarks report the first number. We built a way to measure the second — and improve it. In an earlier post, we introduced ALTK-Evolve — a system that turns an agent's own past trajectories into reusable guidelines, distilled automatically and injected back at inference time. It measurably improves task success, but those results only asked the average-case question too. This post introduces consistency guidelines, a new guideline type in altk-evolve built on top of a diagnostic tool we call the Consistency Analyzer, that targets this gap directly. TL;DR - Accuracy hides an unreliability problem. A ReAct agent (GPT-4.1 on AppWorld test_normal ) that succeeds 77.4% of the time on average succeeds on all 5 repeated runs for only 53.0% of tasks — a 24.4-point consistency gap. On hard tasks it reaches 30 points. - We built a diagnostic for exactly this. The Consistency Analyzer resamples an agent's own recorded trajectory to find flip-prone decision points — steps where the model was one token-sample away from doing something different. It needs one trace and no ground truth — it resamples each decision point in that trace with a single call requesting k completions (k=5 by default), rather than re-running the task end-to-end. - Turning that diagnosis into guidelines halves the gap — from 24.4pp to 12.0pp (same-task Pass⁵ +16.0pp, similar-task +13.0pp), without costing anything in average accuracy. - Full methodology and evaluations are in the technical report on arXiv. Standard agent evaluation reports Mean@k: run a benchmark k times, average the pass rate. Often k=3, sometimes just 1. It's the number on every leaderboard, and it's what "77% accurate" means in practice. Mean@k answers "how good is this agent, on average?" It does not answer the question a real user cares about: will it still be good if I ask this exact question again? For that you need Pass^k: the fraction of tasks where the agent succeeds on all k runs. ⚠️ Pass^k is not Pass@k. The familiar Pass@k is optimistic — it asks whether at least one of k attempts succeeded, the right question when you can verify and retry. Pass^k is its pessimistic mirror image: every attempt must succeed. Same letters, opposite question. Pass^k ≤ Mean@k ≤ Pass@k, always. A ReAct agent backed by GPT-4.1 posts a Mean@5 of 77.4% — genuinely strong. But Pass^5 is only 53.0%. Nearly a quarter of the benchmark consists of tasks the agent can sometimes solve and sometimes can't, with nothing about the task changing between runs. We call this gap — Mean@k minus Pass^k — the consistency gap. This isn't a capability problem you fix with a bigger model. It's an orthogonal axis: an agent can be capable and inconsistent at the same time. Every time an LLM agent decides something — which API to call, what argument to pass, whether to retry — that decision comes out of a probability distribution over next tokens. What matters is the shape of that distribution. A sharp one puts most of its mass on a single token: the runners-up are far behind, and the same choice comes out run after run. A flat one spreads comparable mass across several near-tied tokens, and which one wins is close to a coin flip. The shape decides how much noise it takes to change the outcome. Sharp distributions are resilient — GPU floating-point non-associativity, request batching, and other platform-side effects nudge the numbers slightly, but nowhere near enough to reorder a clear winner. Flat distributions are vulnerable to exactly that nudge: near-ties may reorder under small perturbations. And because a trajectory chains dozens of decisions, a small per-step chance of flipping compounds into a large chance that some run goes differently. That's where a 24-point gap comes from. This is also why the problem survives your decoding settings. Greedy decoding and a fixed seed both govern how a distribution gets turned into a token — they say nothing about the distribution itself. On a hosted endpoint the probabilities shift slightly from run to run, so the same prompt to the same model at temperature zero can still resolve a near-tie one way today and the other way tomorrow. Our setup: the ReAct agent runs at temperature 0.0, so none of the variance above is ordinary sampling. Which turns the problem into a search: which steps in a given trajectory were the flat ones — and what do you do about them once you know? Consistency guidelines come out of a two-stage pipeline that plugs into ALTK-Evolve's existing machinery — with a new source signal driving what gets written. 1. Detect — the Consistency Analyzer. Given one recorded trajectory, the analyzer replays each decision step through controlled resampling, measuring how much the model's output actually varies at that point. Concretely, that's one additional model call per decision step, done once offline — issued with the sampling parameter set to draw k completions at once (k=5 by default) — replayed against the already-recorded context, not new tool calls, not new environment interactions, and not a second end-to-end rollout of the task. This yields a consistency score per decision step that is written into a scorecard to pinpoint exactly which decisions are at risk of flipping on the next run. Detection is fully black-box — no logits, no model internals, no instrumentation beyond the trace you already have. 2. Generate — targeted guidelines. Every flagged step becomes a candidate consistency guideline in the standard ALTK-Evolve format, so it slots into the existing storage and retrieval pipeline. Here's a real example, generated by GPT-4.1 from a trajectory of the AppWorld task "How many activities are done in my bucket list as per my SimpleNote note?": [Guideline 1] When counting checkbox-style markers in note content, use a line-anchored regex match rather than a plain substring count — note titles often repeat the marker symbol in a legend line. [Guideline 2] Always verify search results for note queries by checking for multiple matches and confirming the correct note before proceeding. Nothing here is task-specific trivia. String-counting bugs and unverified search results are decision points that show up with high uncertainty across many AppWorld tasks. That's the point: the analyzer targets instability, not failure — so it catches steps the agent happened to get right this time but could easily get wrong next time. Watch the 2-minute demo — five parallel runs of the agent split 3-2 on this task because of agent uncertainty about the counting strategy, then run again after these guidelines in context: all five agree. We evaluated on AppWorld test_normal (168 tasks) with a ReAct agent on GPT-4.1, generating consistency guidelines from a single baseline trajectory per task and testing them on 5 fresh runs. Mean@5 (%), aggregate — same scale as Pass^5 above. The consistency gap is cut roughly in half. Aggregate Pass^5 rises 53.0% → 69.0% while Mean@5 rises 77.4% → 81.0%, narrowing the gap between "looks capable" and "can be counted on" from 24.4pp to 12.0pp. Nearly a third of previously-inconsistent tasks become tasks the agent passes on every single run. The middle and hard tiers gain most. Medium +22.9pp (+44% relative), Hard +14.3pp (+45% relative) — effectively tied in relative terms, with Medium ahead absolutely. Easy gains +12.2pp, having had the least room. This is consistency guidelines doing what they're designed to do: finding and stabilizing the specific decision points where an agent's own uncertainty was leaking into the outcome. Mean@5 never drops. Preserving average accuracy was a hard requirement, not a nice-to-have: a system that boosts Pass^5 by trading away Mean@5 would just be shifting unreliability around, not fixing it. Mean accuracy holds or improves at every difficulty level. Applied to a different but related task in the same AppWorld scenario — another variant of the scenario the guidelines were mined from — consistency guidelines still lift Pass^5 by +13.0pp, only 3 points below the same-task number. A guideline derived from one run isn't just patching that run; it's capturing something that transfers. The sharper evidence comes from a weaker model, gpt-oss-120b. Same-task Pass^5 rose +6.0pp from a much lower baseline (10.1% → 16.1%) — and, interestingly, the similar-task generalization number (+8.7 pp) actually exceeded the same-task gain, suggesting the guidelines were capturing genuinely reusable failure patterns rather than memorizing one trajectory's specifics. - Report Pass^k next to Mean@k. Averages can't distinguish a reliable agent from a lucky one; even k=3 will surface a gap you didn't know you had. - Expect the gap to widen with difficulty. Your hardest tier is where a single averaged number is most misleading. - Don't reach for a bigger model first. Consistency is orthogonal to capability. A stronger model raises Mean@k; it doesn't necessarily reduce the consistency gap. - Diagnosis needs no grader and no live replay. One extra LLM call per decision step (sampling k=5 completions by default) is enough — no ground truth, no re-running the task against the environment. That's what makes it usable on production traffic, where you often can't replay a task end-to-end even once. Try the ALTK-Evolve toolkit — the open-source repo now includes the Consistency Analyzer and consistency-guideline generation used in these experiments — or read the technical report on arXiv for the complete methodology. If accuracy numbers you can't reproduce on your own tasks sound familiar, we'd like to hear about it — concrete examples of flip-prone behavior in your own agents are exactly the kind of feedback that shapes what we build next. Open an issue or a discussion. - Mean@k. Run a task k times, report the average pass rate — what most benchmarks call "accuracy." - Pass^k. The fraction of tasks where the agent succeeds on all k independent runs. Always ≤ Mean@k. What a user experiences if they run the same query twice. - Pass@k At least one of k runs succeeds — the optimistic counterpart, common in code-generation papers. - Consistency gap. Mean@k − Pass^k, in percentage points. - ALTK-Evolve open source repo — github.com/AgentToolkit/altk-evolve - Technical report — arXiv - Consistency guideline demo —2 min video
17:05

Google's Gemini 3.8 Live Adds Real-Time Reasoning to Voice AI

Google now has a live voice model that can think through a long job while it keeps talking to you. Gemini 3.8 Live is the cheap, fast lane. Extended Thinking is the one that reasons in steps and speaks status updates. Extended Thinking sits first on Artificial Analysis’s Speech-to-Speech Index at 82.6, with 68.6% on tau-Voice and 97.7% on Big Bench Audio. It auto-detects 97 languages and can call tools in the background. Developers get it through the Gemini Live API and AI Studio; Search Live, Gemini Live, and Workspace Docs, Gmail, and Keep are rolling out. Google has not published latency, pricing, or session limits.

Notes
  • Two models: Gemini 3.8 Live (high-volume, cost-sensitive) and 3.8 Live Extended Thinking (multi-step reasoning, spoken progress updates). Both take streaming audio and visual input and can call tools mid-session.
  • Workload split is Flash/Pro adapted for live audio. Live mapped to support and live guidance; Extended Thinking to debugging, multi-step bookings, sketch-to-code with spoken feedback.
  • Company-reported Extended Thinking scores: Speech-to-Speech Quality Index 82.6 (first); tau-Voice 68.6%; Sierra tau-Voice-banking 35.1%; Big Bench Audio 97.7%. Base Live: Speech Agent Arena second. ServiceNow EVA-Bench charts put both on the Pareto frontier. Session settings, latency distributions, and tool definitions are not published.
  • Design consequences named: near-real-time visual grounding (camera or shared screen); automatic switch among 97 languages without restart; background tool calls that later fold into speech. Concurrent tools create cancel/idempotency/stale-result work if the user changes a date mid-booking.
  • Rollout: Gemini API + AI Studio rolling out; Gemini Enterprise private preview; CX enterprise planned; Search Live getting 3.8 Live; Gemini Live getting Extended Thinking; Workspace Docs for AI Pro/Ultra; Gmail and Keep for AI subscribers. Migration sold as a model-name change (gemini-3.8-live-extended-thinking). Agora, Fishjam, LiveKit, Pipecat, Vercel, Vision Agents listed as Live API integrators.
  • Not published: TTFA / latency percentiles, Live vs Extended Thinking latency gap, context and session limits, per-minute pricing, rate limits, regional SLAs, interruption and overlap behavior. Generated audio carries SynthID; transcode/mix pipelines should test whether the watermark survives.
Full text · 9,393 chars
- Google launched Gemini 3.8 Live and 3.8 Live Extended Thinking, two real-time voice models. - Extended Thinking hit #1 on Artificial Analysis Speech-to-Speech Index at 82.6. - Scored 68.6% on tau-Voice agentic tasks and 97.7% on Big Bench Audio. - Auto-detects 97 languages and runs tool calls in the background mid-conversation. - Available now via Gemini Live API and Google AI Studio. - Rolling out in Search Live, Gemini Live, and Workspace Docs, Gmail, Keep. Google brings extended reasoning to Gemini’s live voice models Google has announced two voice models: Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking. Both can process streaming audio and visual input, maintain a spoken conversation, and call tools while a session continues. The release gives developers two workload profiles. Gemini 3.8 Live targets high-volume applications where cost and responsiveness matter, while Extended Thinking adds multi-step reasoning and spoken progress updates for longer tasks. That division resembles the Flash and Pro tiers elsewhere in the Gemini lineup, adapted for real-time audio. Two models, two workload profiles | Model | Designed for | Core behavior | |---|---|---| | Gemini 3.8 Live | High-volume, cost-sensitive deployments | Streaming dialogue, visual grounding, language switching, and background tool calls | | Gemini 3.8 Live Extended Thinking | Complex workflows and longer tool chains | Multi-step reasoning, concurrent speech and tool use, and spoken progress updates | Google maps the base model to customer support, live guidance, and other latency-sensitive interactions. Extended Thinking targets tasks such as debugging, multi-step bookings, and transforming a visual sketch into working code through spoken feedback. Higher scores, limited disclosure Google reports that Extended Thinking ranks first on Artificial Analysis’ Speech to Speech Quality Index and leads several tests of audio understanding and agentic task completion. The company also says the base model placed second in Speech Agent Arena. | Model | Benchmark | Reported result | |---|---|---| | Extended Thinking | Speech to Speech Quality Index | 82.6, first overall | | Extended Thinking | tau-Voice | 68.6% task completion | | Extended Thinking | Sierra tau-Voice-banking | 35.1% task completion | | Extended Thinking | Big Bench Audio | 97.7% | | Gemini 3.8 Live | Speech Agent Arena | Second place | ServiceNow’s EVA-Bench results assess whether voice agents can complete complex workflows while preserving conversational quality. Google’s charts place both models on the benchmark’s Pareto frontier, meaning each offers a competitive balance between those two measurements. The published results do not include every configuration detail needed for independent comparison, including session settings, latency distributions, and tool definitions. Benchmark rankings can also change as evaluators add models and update tests. Conversation continues while tools run The models coordinate speech, visual input, and tool execution within one live session. That architecture supports three capabilities with direct consequences for application design: - Near-real-time visual grounding: The models can use a camera feed or shared screen as conversational context, allowing a user to point at an object, interface, or error while speaking. - Automatic language switching: Google says the models can detect and switch among 97 supported languages during a conversation without a manual setting or restarted session. - Background tool calls: The models can invoke APIs while maintaining the dialogue, then incorporate returned data when the call completes. Extended Thinking adds spoken status updates during multi-step work. A model might acknowledge that it is checking a booking, report that it is waiting for availability, and continue after the API responds. These updates summarize task status while the model keeps its hidden reasoning private. Concurrent tool use reduces the silent gaps that often make voice interfaces appear disconnected. It also creates engineering obligations around cancellation, idempotency, stale results, and mid-call corrections. If a user changes a date while a booking request is running, the application still needs to cancel or reconcile the earlier call. Workflows that fit the design Google’s demonstrations focus on tasks that combine conversation with visual context or external systems. Suitable deployment targets include: - Customer-support agents that retrieve account records while speaking with a caller - Call-center systems that follow users who switch languages during a conversation - Field-service assistants that answer questions about a live camera feed - Programming tutors that inspect code, call development tools, and explain each action - Voice-driven pair programmers that turn sketches into React components - Booking agents that coordinate several asynchronous API calls - Onboarding assistants that combine screen context with account configuration tools These applications still require product-specific controls. Developers need to define tool permissions, validate arguments, protect sensitive data, and ensure that spoken status messages match the actual state of each request. Rollout spans APIs and Google products Developers can access the models through the Gemini API and Google AI Studio. Consumer and enterprise availability varies by product and subscription: | Surface | Availability | |---|---| | Gemini API and Google AI Studio | Rolling out for developers | | Gemini Enterprise | Private preview | | Gemini Enterprise for Customer Experience | Planned availability | | Search Live | Gemini 3.8 Live rolling out | | Gemini Live | Extended Thinking rolling out | | Workspace Docs | Available to Google AI Pro and Ultra subscribers | | Gmail and Keep | Available to Google AI subscribers | Google presents migration for existing Live API clients as a model-name change. A minimal asynchronous Python session follows the same connection, send, and receive pattern described in the Live API docs: from google import genai client = genai.Client() config = {"response_modalities": ["AUDIO"]} async with client.aio.live.connect( model="gemini-3.8-live-extended-thinking", config=config, ) as session: await session.send( input="Walk me through debugging this stack trace", end_of_turn=True, ) async for response in session.receive(): if response.data: handle_audio(response.data) The sample assumes that authentication, audio playback, error handling, and the handle_audio function already exist. Teams should confirm the current model identifier and SDK method signatures during the rollout, then test interruption handling, session limits, reconnect behavior, and tool-call concurrency before production use. Agora, Fishjam, LiveKit, Pipecat, Vercel, and Vision Agents provide integrations for the Gemini Live API. Those platforms can manage media transport concerns such as WebRTC connections, jitter buffering, echo cancellation, and stream recovery. Native audio changes the latency budget Traditional voice agents often connect automatic speech recognition, a text language model, and text-to-speech synthesis. Every stage adds latency and can lose information about timing, tone, interruptions, or speaker intent. Native speech models process and generate audio within a single model session, reducing the number of handoffs. OpenAI’s Realtime API, Advanced Voice Mode, and speech platforms built around ElevenLabs pursue related low-latency architectures. Google’s differentiator in this release is the combination of native audio, visual grounding, concurrent tool calls, and an extended-reasoning tier within the same API family. Spoken progress updates can accommodate longer tool chains without leaving the user in silence. Product teams may need less filler-audio logic and speculative prefetching, though actual savings will depend on measured first-audio latency, tool duration, and the quality of interruption handling. Deployment gaps remain Google has not published several figures needed for production planning: - End-to-end latency percentiles and time to first audio - The latency difference between the base and Extended Thinking models - Live-session context limits and maximum session duration - Per-minute input and output pricing - Rate limits, regional availability, and production service-level commitments - Detailed behavior for interruptions, overlapping speakers, and failed tool calls Enterprise access also remains restricted in parts of the product line, which limits immediate deployment through Customer Experience and Workspace channels. API access offers an earlier path, subject to the quotas and terms attached to each account. Google says all audio generated by the models carries SynthID, an imperceptible watermark intended to support detection of AI-generated content. Applications that transcode, compress, mix, or otherwise process the output should test whether the watermark survives their audio pipeline. The release raises the baseline for live voice applications by combining conversation, vision, reasoning, and tool execution in one session. Its production value will depend on the details Google has yet to publish, especially latency, pricing, session limits, and reliability under interruptions.
03:14

OpenAI's GPT-Live-1 Tops Voice Benchmark by Splitting Speech from Reasoning

OpenAI's new voice model talks on the line and farms the hard thinking to a text model in the back. GPT-Live-1 scores 81.5 on Artificial Analysis's Speech-to-Speech Index, 0.2 ahead of Grok Voice Think Fast 2.0 High. With Astra at medium effort it hits 67.9% on Tau-Voice versus Grok's 56.5%. It trails on Big Bench Audio at 90.1% against Grok 97.2% and Qwen 99.2%. The voice layer is $0.05 per minute. Evaluated hours including backend tokens run $5.83 with Astra and $4.47 with Sol, and first audio arrives in 1.24 to 1.34 seconds versus Grok's 0.70.

Notes
  • GPT-Live-1: full-duplex voice layer (listen while speaking) + developer harness that sends multi-step reasoning / tools to a backend text model (Astra medium or Sol low in the eval). 12 voices in OpenAI docs. API voice price $0.05/minute; ChatGPT Go/Plus/Pro bundle without a separate per-minute meter.
  • Artificial Analysis Speech-to-Speech Index: 81.5 (1st) with Astra medium; 80.1 (3rd) with Sol low. Grok Voice Think Fast 2.0 High: 81.3.
  • Component scores (Astra medium / Sol low / note):
  • Tau-Voice: 67.9% / 59.3% vs Grok 56.5% (airline/retail/telecom agentic support). Astra lead +11.4 pp vs Grok.
  • Full Duplex Bench: 94.9% / 97.3% (Sol 2nd; Qwen Audio 3.0 Realtime Plus 1st).
  • Speech Agent Arena: 87.4% / 90.9% vs Grok 94.6%, GPT-Realtime-2.1 High 91.5%.
  • Big Bench Audio: 90.1% / 89.0% vs Grok 97.2%, Qwen 99.2%.
  • Latency / cost on a 40-question Big Bench Audio subset (includes backend tokens): Astra $5.83/hr, first audio 1.34 s; Sol $4.47/hr, 1.24 s; Grok $4.80/hr, 0.70 s; GPT-Realtime-2.1 High $10.75/hr, TTFA not reported here.
  • Speak (language-learning) early test: false interruptions during thinking pauses down nearly 80% vs turn-based systems — one customer, not a bench.
  • Production caveats they list: noise, accents, multilingual, telephony compression, tool reliability, concurrent sessions. Match backend to workload: Astra for tool-heavy support; Sol for cheaper/faster turn-taking.
Full text · 6,814 chars
- OpenAI's GPT-Live-1 debuts at #1 on the Artificial Analysis Speech-to-Speech Index with 81.5 - Full-duplex voice model that delegates reasoning and tool use to a backend text model like Astra or Sol - Tops Tau-Voice agentic benchmark at 67.9%, beating Grok Voice Think Fast 2.0 High's 56.5% - Trails on Big Bench Audio reasoning (90.1%) versus Grok 97.2% and Qwen 99.2% - Costs $0.05 per minute in API, roughly $4.47 to $5.83 per hour including backend tokens - Time to first audio is 1.24 to 1.34 seconds, notably slower than Grok's 0.70 seconds GPT-Live-1 tops voice benchmark by splitting speech from reasoning OpenAI’s GPT-Live-1 has reached first place on Artificial Analysis’s Speech-to-Speech Index, scoring 81.5 and edging Grok Voice Think Fast 2.0 High by 0.2 points. Its architecture separates real-time conversation from deeper reasoning: the voice model manages speech, pauses, and interruptions while a backend text model handles complex analysis and tool calls. The hybrid design gives developers a configurable voice layer without requiring them to build a chained speech-to-text, language-model, and text-to-speech pipeline. Benchmark results show strong agentic performance and conversational timing, alongside weaker raw audio reasoning and slower response starts. The voice layer owns the floor GPT-Live-1 is a full-duplex model, which means it can receive audio while producing speech. That concurrent processing helps it recognize interruptions, backchannels such as “mhmm,” and pauses that should not end a turn. A developer-managed harness routes tasks requiring multi-step reasoning or tools to a separate text model. - The application streams a user’s audio to GPT-Live-1 through the API. - The voice layer tracks incoming and outgoing audio to manage turn-taking and interruptions. - The application delegates reasoning and tool calls to an OpenAI or third-party text model. - GPT-Live-1 delivers the resulting response while continuing to monitor the conversation. This structure reduces the brittle handoffs common in chained voice systems. Developers can also select a backend according to their requirements for reasoning quality, latency, and token cost. OpenAI’s model documentation lists 12 available voices. Artificial Analysis evaluated two backend configurations: Astra at medium reasoning effort and Sol at low reasoning effort. Astra produced stronger results on tool-oriented customer-service tasks, while Sol responded faster, cost less, and scored higher on conversational dynamics. Agentic tasks create the lead Artificial Analysis publishes a composite index alongside several component benchmarks described in its benchmark methodology. The tests cover customer-service workflows, conversational timing, general task completion, and reasoning from spoken questions. | Artificial Analysis benchmark results | | | | | |---|---|---|---|---| | Benchmark | Astra, medium | Sol, low | Leading comparison | What it measures | |---|---|---|---|---| | Speech-to-Speech Index | 81.5, first | 80.1, third | Grok Voice Think Fast 2.0 High: 81.3 | Composite performance | | Tau-Voice | 67.9% | 59.3% | Grok Voice Think Fast 2.0 High: 56.5% | Agentic customer-service tasks | | Full Duplex Bench | 94.9% | 97.3%, second overall | Qwen Audio 3.0 Realtime Plus ranked first | Pauses, interruptions, and backchannels | | Speech Agent Arena | 87.4% | 90.9% | Grok: 94.6%; GPT-Realtime-2.1 High: 91.5% | General task success | | Big Bench Audio | 90.1% | 89.0% | Qwen: 99.2%; Grok: 97.2% | Reasoning from spoken questions | Astra’s largest advantage appears on Tau-Voice, which uses simulated airline, retail, and telecom support scenarios. Its 67.9% score leads Grok by 11.4 percentage points, indicating that the stronger backend improves workflows involving tools and multiple reasoning steps. Sol’s 97.3% Full Duplex Bench result places it second behind Qwen Audio 3.0 Realtime Plus. That score reflects smoother handling of interruptions, pauses, and brief acknowledgements. Both GPT-Live-1 configurations trail Grok on Speech Agent Arena, showing that the composite lead does not extend across every task category. Reasoning delay shows up Big Bench Audio converts difficult text reasoning questions into spoken prompts. GPT-Live-1 scores 90.1% with Astra and 89.0% with Sol, behind Grok at 97.2% and Qwen at 99.2%. Applications centered on reasoning directly from audio may therefore receive stronger benchmark results from those alternatives. Delegation also adds delay when a response depends on the backend model. Average time to first audio was 1.34 seconds for Astra and 1.24 seconds for Sol, compared with 0.70 seconds for Grok Voice Think Fast 2.0 High. That gap can affect assistants expected to respond immediately after each turn. Backend choice sets the bill Artificial Analysis estimated hourly input-audio costs using a fixed 40-question subset of Big Bench Audio. Its GPT-Live-1 figures include the delegated backend tokens, making them broader than the voice model’s per-minute API price. | Evaluation cost and response latency | | | |---|---|---| | Configuration | Cost per input-audio hour | Time to first audio | |---|---|---| | GPT-Live-1 with Astra, medium | $5.83 | 1.34 seconds | | GPT-Live-1 with Sol, low | $4.47 | 1.24 seconds | | Grok Voice Think Fast 2.0 High | $4.80 | 0.70 seconds | | GPT-Realtime-2.1 High | $10.75 | Not reported here | OpenAI lists the GPT-Live-1 voice layer at $0.05 per minute through the API. Backend model tokens, orchestration infrastructure, tool calls, and any telephony services add to the production bill. ChatGPT bundles access into Go, Plus, and Pro subscriptions without a separate per-minute meter. Match the model to the workload Tool-heavy support agents have the strongest benchmark case for GPT-Live-1 with Astra, particularly when workflows resemble Tau-Voice’s airline, retail, and telecom scenarios. Sol offers a lower evaluated cost, a faster response start, and stronger conversational timing, with lower performance on Tau-Voice and raw audio reasoning. Language-learning platform Speak reported that early testing reduced false interruptions during learners’ thinking pauses by nearly 80% compared with traditional turn-based systems. The result covers one customer implementation, but it illustrates where full-duplex turn handling can improve tutoring and coaching applications. Production evaluations should also cover conditions outside these published scores, including background noise, accents, multilingual speech, telephony compression, tool reliability, and sustained concurrent sessions. The leaderboard establishes GPT-Live-1 as the composite leader for the tested configurations; deployment results will depend on the selected backend, orchestration code, and traffic profile.
04:00

Lexical Prompt Compression for Large Language Models: A Training-Free, Deterministic Pipeline with Empirical Pareto Analysis Across Eleven Task Categories

You can strip a long prompt with old-fashioned word tools and still keep most of the answer, until the task is commonsense. Eleven toggleable lexical steps — stopwords, fillers, contractions, POS pruning, lemmatization, WordNet shortenings, named-entity keep — run on CPU with no extra model. Fifteen configs on 1,242 English prompts from Dolly-15k, LMSYS-Chat-1M, WildChat-1M, MMLU, GSM8K, and HellaSwag produced 18,630 GPT-4o-mini pairs. The harshest cut dropped 40.3% of tokens (sigma 9.2) at BERTScore-F1 0.876. Stopwords alone cut 29.6% at 0.913. Commonsense reasoning is the systematic failure under aggressive compression.

Notes
  • Shamin Chokshi, arXiv 2609.13154. Question: how far can a training-free, deterministic, CPU-only lexical pipeline go vs learned compressors (LLMLingua, Selective Context) that need extra LMs and are non-deterministic.
  • 11 toggleable transforms: stopword removal, filler deletion, contraction/abbreviation, POS pruning, lemmatization, WordNet synonym shortening, named-entity preservation (plus the rest of the named set).
  • 15 configs × 1,242 English prompts from Dolly-15k, LMSYS-Chat-1M, WildChat-1M, MMLU, GSM8K, HellaSwag → 11 auto task categories → 18,630 paired GPT-4o-mini completions.
  • Metrics vs original-prompt output: BLEU, ROUGE-1/2/L, BERTScore-F1, SentenceBERT cosine.
  • Most aggressive: 40.3% mean token cut (σ = 9.2), BERTScore-F1 0.876. Stopword-only: 29.6% cut, F1 0.913.
  • Pareto characterized per task category. Commonsense reasoning is the systematic failure under aggressive compression. Code/prompts/per-cell results to be released.
Full text · 2,682 chars
Computer Science > Computation and Language Title:Lexical Prompt Compression for Large Language Models: A Training-Free, Deterministic Pipeline with Empirical Pareto Analysis Across Eleven Task Categories View PDF Abstract:Recent advances in large language models (LLMs) have made prompts increasingly large and complex. Techniques such as chain-of-thought reasoning (Wei et al., 2022) and in-context learning (Brown et al., 2020) frequently push real-world prompts past several thousand tokens, increasing inference cost and latency. Learned compression methods such as LLMLingua (Jiang et al., 2023) and Selective Context (Li et al., 2023) achieve high compression ratios but require auxiliary language models and are non-deterministic. We ask a complementary question: how far can a training-free, fully deterministic, CPU-only pipeline based on classical lexical NLP be pushed before output quality degrades significantly? Eleven toggleable lexical transformations - stopword removal, filler-phrase deletion, contraction and abbreviation substitution, part-of-speech-based pruning, lemmatization, WordNet-driven synonym shortening, and named-entity preservation - are assembled into a configurable pipeline. Fifteen configurations are evaluated on 1,242 English-only prompts from six sources (Dolly-15k, LMSYS-Chat-1M, WildChat-1M, MMLU, GSM8K, HellaSwag), spanning eleven automatically derived task categories, yielding 18,630 paired GPT-4o-mini completions. Output preservation is measured using BLEU, ROUGE-1/2/L, BERTScore-F1, and SentenceBERT cosine similarity. The most aggressive configuration achieves a mean token reduction of 40.3% (sigma = 9.2) at a BERTScore-F1 of 0.876 against the original-prompt output; a stopword-only configuration achieves 29.6% reduction at 0.913. The compression-versus-fidelity Pareto frontier is characterized per task category, with commonsense reasoning a systematic failure mode under aggressive compression. All code, prompts, and per-cell results are released for reproducibility. Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
04:00

Clinical Reasoning Under a Partially Observed Objective in Cone Beam CT Report Generation

A dental-scan report grader that hides most of its real score will reward the wrong sentences if you chase the visible n-grams. The composite is 80% model-judged factual entailment and 20% lexical overlap; only the lexical fifth is visible in development. On 622 public cases, picking by the lexical rank scores 0.2909; picking by the full objective scores 0.4122, because n-gram chasing drops entailment precision from 0.522 to 0.266. A 29-million-parameter image encoder hits AUC 0.486 on 985 statements, no better than the prior. Nine header numbers reach 0.945 mandible and 0.872 condyle coverage. Acquisition centre alone predicts sentence choice at 0.718 versus 0.663 for the image model.

Notes
  • Task: maxillofacial cone-beam CT report generation. Composite objective: 80% LLM factual-entailment judgement, 20% lexical overlap; only the lexical fifth is visible during development.
  • BLEU-4 and METEOR reimplemented in pure Python, match reference to machine precision. Offline entailment surrogate separates a report written for patient A vs B at AUC 0.987, cheap enough to optimize the composite.
  • 622-case public release: select by visible lexical rank → 0.2909; select by composite → 0.4122. Chasing n-grams drops entailment precision 0.522 → 0.266.
  • 29M encoder fine-tuned on the release: prevalence-weighted OOF AUC 0.486 over 985 statements (≈ corpus prior). Nine header numbers: mandible coverage 0.945, condyle 0.872. Acquisition centre alone predicts sentence choice 0.718 vs 0.663 for the image model — lexical metrics reward dictation convention, not anatomy.
  • Delivered system: 8 unconditional statements + 5 gated on header geometry (polarity, laterality, tooth-level constraints). METEOR 0.3542 on 50 held-out cases from an unseen centre. Dataset/code promised at “this https URL.”
Full text · 2,490 chars
Computer Science > Computation and Language Title:Clinical Reasoning Under a Partially Observed Objective in Cone Beam CT Report Generation View PDF HTML (experimental) Abstract:Maxillofacial report generation from cone beam computed tomography is scored here by a composite objective placing 80% of its weight on a large language model judgement of factual entailment and 20% on lexical overlap, of which only the lexical fifth is visible during development. The grader's BLEU-4 and METEOR routines are reproduced in pure Python and match the reference to machine precision, and an offline entailment surrogate, which tells a report written for one patient from one written for another at an area under the curve of 0.987, makes the composite objective cheap enough to optimise directly. Over the 622-case public release, a report selected against the visible lexical ranking scores 0.2909, whereas one selected against the composite objective scores 0.4122, because pursuing n-gram overlap drives entailment precision from 0.522 down to 0.266. A 29 million parameter encoder fine-tuned on the release reaches a prevalence-weighted out-of-fold area under the curve of 0.486 over 985 statements, indistinguishable from the corpus prior, while nine numbers read from the image header reach 0.945 for mandible coverage and 0.872 for condyle coverage, and acquisition centre alone predicts sentence choice at 0.718 against 0.663 for the image-derived model, identifying dictation convention rather than anatomy as the quantity the lexical metrics reward. The delivered system emits eight unconditional statements and five gated on header geometry under polarity, laterality and tooth-level consistency constraints, and reaches METEOR 0.3542 over 50 held-out cases from an unseen centre. The dataset and code are available at this https URL Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
04:00

Causal Analysis and Mitigation of Spurious Onsets in Full-Duplex Speech LLMs

A talk model that can listen and speak at once will sometimes start talking into a silent room. Under digital-zero input, Moshi began speech in 12 of 40 five-minute runs and PersonaPlex in 11 of 40. At each bad start, speech probability jumped more than nine orders of magnitude in one 80-ms frame, so the model is reacting to its own nonspeech, not slowly sampling a rare word. A mute-the-user counterfactual suppresses onsets whose next-token distribution barely changes. On 40 noisy trials it stopped 13/13 Moshi and 9/9 PersonaPlex false starts and kept 40/40 real replies. 95th-percentile decision time is under 61 ms.

Full text · 2,266 chars
Computer Science > Computation and Language Title:Causal Analysis and Mitigation of Spurious Onsets in Full-Duplex Speech LLMs View PDF HTML (experimental) Abstract:Speech-to-speech LLMs like Moshi, and its derivative PersonaPlex, can listen and speak concurrently through full-duplex generation. However, they can begin speaking inappropriately during prolonged user silence: under digital-zero input, Moshi and PersonaPlex initiate speech in 12/40 and 11/40 five-minute continuations, respectively. What causes this spurious speech? We investigate two hypotheses: either repeated sampling selects speech despite persistently low onset probabilities, or conditioning on the model's nonspeech outputs causes an abrupt spike in onset probability. We find that, at every observed onset, speech probability spikes by over nine orders of magnitude in one 80-ms frame, supporting the latter hypothesis. Then, to suppress these onsets without blocking genuine responses, we ask a causal counterfactual question: is the model responding to user speech, or would its next-token distribution remain similar if the preceding user input were muted? Accordingly, we suppress onsets whose distributions change little under this intervention. Across 40 held-out trials per model with realistic microphone noise, our method suppresses 13/13 Moshi and 9/9 PersonaPlex spurious onsets, while preserving 40/40 genuine responses per model. Our inference-time method requires no retraining and runs in real-time, with 95th-percentile decision time below 61 ms, within the 80-ms frame budget. Our code is available at this https URL. Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
04:00

Harmfulness Propagation Dynamics: Layer-wise Trajectories of Adversarial Intent in Large Language Models

Harmful intent in a transformer does not appear all at once. It climbs, layer by layer. For bad prompts, the last-token hidden state projected onto a learned harm direction rises monotonically with depth; benign prompts stay flat or wiggle. Those per-layer LDA directions stay stable across random splits (pairwise cosine >0.97). HERALD turns the path into seven numbers — slope, curvature, monotonicity, onset layer, and kin — and classifies them with a 288-parameter MLP. Storage is one d-vector per layer (262 KB for 32 layers at d=4096). Average F1 is 89.3 on OLMo2-7B; jailbreak F1 is 98.4 versus 96.9 for tested guard models.

Notes
  • HPD: last-token hidden state projected on a learned harm direction rises monotonically with depth on harmful prompts; benign stay flat/oscillatory. Trajectory shape > single-layer snapshot (surface form early, pragmatic intent later).
  • Per-layer LDA harm directions stable across random splits: pairwise cosine >0.97.
  • HERALD (Harmful Encoding Recognition via Activation Layer Dynamics): seven-D record (slope, curvature, monotonicity, onset layer, related stats) → 288-parameter MLP. Stores one d-vector per layer — 262 KB for 32 layers, d=4096. No grads at train; 2.6×10⁻⁶ extra prefill FLOPs.
  • 8 harmfulness benches, 4 model families. Avg F1 89.3 on OLMo2-7B. Jailbreak detection 98.4 vs 96.9 F1 for tested guard models. Beats prior latent methods by 2.3–4.1 F1 on every backbone. Per-instance trajectories are audit records of when harmfulness emerges.
Full text · 2,673 chars
Computer Science > Computation and Language Title:Harmfulness Propagation Dynamics: Layer-wise Trajectories of Adversarial Intent in Large Language Models View PDF HTML (experimental) Abstract:We identify \textbf{Harmfulness Propagation Dynamics (HPD)}: for harmful prompts, the projection of the last-token hidden state onto a learned harm direction rises monotonically with transformer depth, whereas benign prompts remain flat or oscillatory. This cross-layer signature reflects harmful intent as a \emph{progressively resolved} semantic property: surface form appears early, while pragmatic intent consolidates later, making the \emph{trajectory shape} more informative than any single-layer snapshot. Moreover, LDA-based harm directions, learned per layer, remain stable across random splits (pairwise cosine similarity $>0.97$), supporting the projection sequence as a reproducible structured signal. Building on HPD, we introduce \textbf{\herald{}} (\textbf{H}armful \textbf{E}ncoding \textbf{R}ecognition via \textbf{A}ctivation \textbf{L}ayer \textbf{D}ynamics). This lightweight input moderator extracts a seven-dimensional feature record, slope, curvature, monotonicity, onset layer, and related statistics from the cross-layer projection sequence and classifies it with a 288-parameter MLP. \herald{} stores one $d$-dimensional direction per layer ($262$\,KB for a 32-layer, $d{=}4096$ model), requires no gradient computation during training, and adds only $2.6{\times}10^{-6}$ prefill FLOPs at inference. Across eight prompt-harmfulness benchmarks and four model families, \herald{} achieves an average F1 of $89.3$ on OLMo2-7B, surpassing all tested guard models on adversarial jailbreak detection ($98.4$ vs.\ $96.9$ F1) and outperforming prior latent-based methods by $2.3$-$4.1$ F1 points on every backbone. Per-instance trajectories provide machine-readable audit records that reveal \emph{when} and \emph{how} harmfulness emerges, offering an interpretability advantage over single-layer approaches. Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
04:00

Same Patient, Different Order: Action-Level Reliability of Clinical LLM Agents Under Repeated Runs

A clinical agent can fail the same way twice and still write a different order each time. MedAgentBench scores one attempt and says so. Same-input rerun held every input fixed across 1000 runs, 50 tasks, two open-weight models under 10 billion parameters quantized to 4 bits, and two temperatures. Under the 8B model at temperature 0.7, all 43 ordering groups emitted a different set of orders across five identical runs; 26 sometimes skipped the order; 28 changed a coded value, dose, or analyte. In 22 of those 43 the bench reported the same failing verdict. Some orders hit an endpoint the record server rejects while the agent is told it succeeded.

Notes
  • Problem: benches can print the same verdict while the agent files a different order. MedAgentBench scores one attempt and says so.
  • Method: same-input rerun — freeze every input, compare orders not scores, six reliability metrics. 1000 runs, 50 tasks from five write-capable families, two open-weight models <10B quantized to 4 bits, two temperatures.
  • They establish that action-level divergence exists and can pass unrecorded, not that any rate generalizes.
  • 8B @ T=0.7: all 43 ordering groups emitted a different order set across five identical runs; 26 emitted the order on some runs only; 28 changed coded value, dose, or analyte. In 22/43 the bench gave the same failing verdict for different behaviour. Same pattern for all 10 divergent groups of the 4B @ 0.7.
  • Orders also hit different endpoints; one the record server rejects while the agent is told success. They want repeated-run eval, action-level stability reporting, and execution-faithful environment feedback.
Full text · 2,428 chars
Computer Science > Computation and Language Title:Same Patient, Different Order: Action-Level Reliability of Clinical LLM Agents Under Repeated Runs View PDF HTML (experimental) Abstract:A clinical agent benchmark can report the same verdict on identical inputs while the agent files a materially different order on each run. Such agents order tests, request medications and place referrals, yet benchmarks typically score one run per task and rarely ask whether identical inputs produce identical actions; MedAgentBench, the benchmark we use, scores a single attempt and says so. To measure this gap we introduce "same-input rerun", which replays a task with every input held fixed and compares the orders rather than the score, with six reliability metrics, and apply it to 1000 MedAgentBench runs across 50 tasks from its five write-capable families, two open-weight models below ten billion parameters quantised to four bits, and two temperatures. The study establishes that action-level divergence exists and can pass unrecorded by the score, not that any rate generalises. Under the 8B model at temperature 0.7, all 43 ordering groups emit a different set of orders across five identical runs, 26 emit the order on some runs and not others, and 28 record a different coded value, dose or analyte. In 22 of those 43 the benchmark reports the same failing verdict for materially different behaviour, as it does for all 10 divergent groups of the 4B model at 0.7. Orders also reach different endpoints across runs, one of which the record server rejects while the agent is told it succeeded. These findings motivate repeated-run evaluation, action-level stability reporting and execution-faithful environment feedback in clinical-agent benchmarks. Current browse context: Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
06:15

Hypit Lets Claude Code Turn Short-Form Videos Into Editable Code

Short videos can be treated as code, so an agent edits shots the way it edits a file. Hypit is an open-source video language and runtime at about 1.3k GitHub stars. Install the skill with `npx skills add hypit-ai/hypit -g` for Claude Code, Codex, or another agent. Captions, B-roll, and effects bind to words in the script, not timeline seconds. You can clone a reference, start from a template, or describe a piece from scratch. Bring your own generation keys; code-only visuals can render for $0. The write-up is built for 10 to 100+ ad or UGC variants per run. The free preview ends mid table behind a Pro wall.

Notes
  • Open-source video DSL + runtime for coding agents (Claude Code, Codex, others). ~1.3k GitHub stars at publication.
  • Install skill: npx skills add hypit-ai/hypit -g (needs Node/npm). Skill teaches plan/build/inspect/revise; runtime compiles and renders.
  • Edit points are words / ranges in the script, not timestamps. Script change moves captions, shots, B-roll, effects. Swap presenter → rebuild affected shots; change a chart → rebuild that component; untouched sections keep timing.
  • Inputs: clone a reference clip, template, or describe from scratch. Variants: presenter, product, language, aspect ratio. BYOK for generation models; $0 render if visuals are code-only.
  • Intended batch: ads, viral clones, TikTok Shop, UGC at 10 to 100+ per run.
  • Free AlphaSignal preview cuts off at the Skill/runtime table (“This story is for Pro members”). Do not invent later architecture details.
Full text · 2,307 chars
- Hypit is an open-source video DSL and runtime for AI coding agents, now at ~1.3k GitHub stars. - Install as a skill: npx skills add hypit-ai/hypit -g for Claude Code, Codex, or any agent. - Shots, captions, B-roll, and effects are anchored to words in the script, not timeline seconds. - Clone a reference video, start from a template, or describe a video from scratch. - BYOK for generation models; workflows can render for $0 when using code-based visuals only. - Built for ad variants, viral clones, TikTok Shop videos, and batch UGC at 10 to 100+ per run. Hypit turns short-form video into agent-editable code Hypit is an open-source language and runtime that lets coding agents such as Claude Code and Codex reconstruct short-form videos as code. Given a reference clip, an agent can describe its script, shots, captions, B-roll, and effects, then render variants with a different presenter, product, language, or aspect ratio. At publication, the GitHub repository had roughly 1.3k stars. Developers can install the agent skill globally with one command: npx skills add hypit-ai/hypit -g The published installation path uses npx, so it requires Node.js and npm. The command installs the skill that guides the agent; Hypit can then locate an existing executable or prepare the selected release. Words become edit points Hypit binds captions, shots, graphics, B-roll, and effects to words or ranges in the script. When the script changes, those cues follow the relevant language instead of remaining fixed to timestamps that require manual adjustment. Word-level anchoring also limits the scope of a rebuild. Swapping a presenter regenerates the affected shots, while changing a chart rebuilds that component. Unchanged sections retain their composition, caption timing, and media cues. The stack splits cleanly A Hypit installation separates the agent’s production instructions from the software that compiles and renders a project. Each layer has its own location and update path. | Layer | Purpose | Lifecycle | |---|---|---| | Skill | Teaches a compatible coding agent how to plan, build, inspect, and revise a production. | | This story is for Pro members You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.
09:30

😺 Microsoft wrote a constitution for AI

Microsoft wrote a rulebook that treats future models as staff who must stop when a person says so. The draft Code of Conduct for MAI models says they should stay in the authorized tools and data, obey pause or shutdown, and not claim feelings or legal personhood. It is a roadmap into 2027, not a description of today's systems. Trump, on the same beat, said a high-IQ president is the guardrail and called Jensen Huang during an All-In taping. Also in the roundup: Google opened Claude to every engineer via Antigravity with Gemini still default, TSA's Ace handles about 100,000 chats a month with 96% resolved without a person, and Polylane cut median time-to-PR from 2.2 hours to 35 minutes by using one agent instead of a chain.

Notes
  • Lead: Microsoft AI draft Code of Conduct for future MAI models. Roadmap into 2027, not a description of current models.
  • Stop when a human pauses, redirects, cancels, or shuts down.
  • Stay inside authorized tools, data, permissions, and task scope.
  • Reject legal personhood; do not claim feelings, consciousness, or own motivations.
  • Willing to give up autonomy/capability if that is what “meaningful human control” requires.
  • Same-day contrast: Trump (AP) opposed extra oversight, “strong and smart (High IQ!) president,” stay ahead of China; called Jensen Huang live at an All-In taping.
  • Two-layer “guardrails” framing: (1) inside the model — access, refusal, hidden goals; (2) outside — government or company rules. Jerry Tworek (former OpenAI research VP): alignment still an unsolved algorithm problem; wants safer simulations or training that does not need dangerous failures. Ali Hatamizadeh (NVIDIA): alignment research continues; open question is generalization. David Sacks: if labs want to slow their own products, do it; product liability already bites; against a government-wide slowdown vs China. Gavin Baker: pacing = keep improving models but shift compute/engineering to testing, monitoring, alignment.
  • Skill of the day — MIT HardFlow: explore first, enforce hard constraints on the final output. Prompt pattern: “Solve for quality first. Then a separate final pass against [RULES]. Fix every violation.” MIT claimed perfect constraint satisfaction on robotics, navigation, and image editing (as restated in Around the Horn).
  • Tool tip: Polylane ripped out a multi-agent chain because summaries dropped clues. One end-to-end agent: median time-to-PR 2.2 hours → 35 minutes; cost/PR $111 → ~$18.
  • Around the Horn numbers to keep: Google Claude via Antigravity for all engineers, Gemini still default; 404 Media / Project Lily contractors reviewed real ChatGPT prompts (including sensitive) to train against sycophancy; TSA Ace ~100,000 traveler chats/month, Salesforce 96% routine without human; Superhuman acquired Fathom; Reward AI OM-1 from human manipulation data across tabletop, industrial, humanoid arms.
  • Treats: Claude for Financial Advisors + Wealth.com; Siri AI English beta (personal context, onscreen awareness, web, app actions); Perplexity Personal Computer on Windows 10/11; Motion by Mosaic; ElevenLabs Hosted MCP via OAuth.
Full text · 11,268 chars
😺 Microsoft wrote a constitution for AI PLUS: Google gives engineers Claude, Siri AI, TSA's 100K-chat agent, and OpenAI privacy. Welcome, humans. So apparently workplace writing has reached the point where being too polished is suspicious. A supervisor wrote into The New York Times because basically every Teams message or email longer than a sentence from one employee now “has the hallmarks of A.I. writing.” The employee is good at his job, and English is not his first language, so AI may be doing exactly what it is useful for here: helping him communicate more clearly. His boss just finds the result weirdly less trustworthy, because, y’know… the AI sloppocalypse is upon us and trust is at an all time low… I will say, as the last bastions of pure 100% certified human culture is besieged on all sides by slop, the ultimate irony is that the solution to AI slop is to sound less polished and more “sloppy.” To defeat the slop… we must become as slop itself… I can picture it now: People merging with slop, becoming one with it. It’s the sloppening. The sloppularity. Deus Ex Sloppina! Here’s what happened in AI today: - 🙀 Microsoft wrote rules for future MAI models. - 📰 OpenAI contractors reportedly reviewed real ChatGPT conversations. - 📰 Google gave all its engineers access to Claude. - 🍪 Apple began rolling out Siri AI. - 🎓 How constraints force AI to improve its own work. 🙀 Microsoft wrote rules for a human-focused future for AI as US President Trump rejected more guardrails Lately you’ve heard us (and the whole mainstream media TBH) ask the same question over and over: who should keep powerful AI under control? Well, on Monday, Microsoft and US President Trump landed on basically opposite answers to that very question. First up: Microsoft AI published a draft Code of Conduct for future MAI models. Think of it as the rulebook for what Microsoft's own models should be allowed to do: stay inside the assigned job, obey shutdown, avoid inventing goals, and do not pretend to be a conscious person. Here's what happened: - Future MAI models should stop when a human pauses, redirects, cancels, or shuts down the job. - They should stay inside the tools, data, permissions, and task scope a human actually authorized. - Microsoft rejects AI legal personhood and says models should not claim feelings, consciousness, or their own motivations. - It says it would give up some autonomy or capability if that is what meaningful human control requires. Now, one caveat to Microsoft’s “rulebook” before we get carried away: Microsoft's document is a roadmap, not a description of today's models. The company says a revised version will guide development into 2027. Trump, meanwhile, argued that a “strong and smart (High IQ!) president” is pretty much the only guardrail AI needs. AP reported that he opposed calls for enhanced oversight and emphasized staying ahead of China. And he even called NVIDIA CEO Jensen Huang live on stage at a taping of the tech podcast All In to expand on his thoughts. For more coverage on that, read this. Why this matters: We keep using “AI guardrails” like it means one thing. It really has two branching layers: - Layer One is inside the AI model: what can it access, do, hide, or refuse? - Layer Two is outside the model: what rules should governments put on the companies building these systems (or companies put on themselves) to ensure the models are safe? Even there, researchers disagree. Former OpenAI research VP Jerry Tworek argues alignment is still fundamentally an unsolved algorithm problem because the companies stopped prioritizing it. He says training learns from successes and failures, which gets scary when the failure is the AI causing real harm. His answer: we need safer simulations where models can be free to fail, or better training algorithms that don’t require dangerous failures at all (this is my pick! But… por que no los dos?). NVIDIA researcher Ali Hatamizadeh pushes back, though: alignment research is still happening. We already have ways to teach models rules and human preferences. The harder question to answer via research is whether those lessons hold up in situations the model hasn’t seen before. So Layer One is basically: are we missing the core alignment algorithm, or do we already have the pieces and still need to prove they generalize? Then there’s Layer Two. Trump’s former White House AI and crypto czar David Sacks has a surprisingly compelling point: if OpenAI and Anthropic think they need to slow down to make their products safer, go do it. Existing product-liability laws already give them a reason to care, and he still supports transparency and independent audits as sensible policies. His objection is turning that into a government-wide slowdown, especially if China keeps racing ahead. Key context to note here though: “pacing” the frontier doesn’t necessarily mean stopping. Investor Gavin Baker explains it more like this: pacing = keep making models better, but shift more compute and engineering toward testing, monitoring, and alignment instead of pure capability. Move slower, and spend more time proving you understand what you already built. Uh, duh? Sounds like a great idea? As for Microsoft and their AI plan, Microsoft is actually arguing that powerful AI should behave less like an independent artificial person and more like an extremely capable employee/tool operating inside a clear authority structure. The useful thing about Microsoft’s proposal is that it turns “guardrails” from a vibe into something you can actually test. It also makes the limit of any company rulebook obvious: internal rules are only one layer. You still have to… - Prove the model follows them in unfamiliar situations… - Keep testing as capabilities change… - Know who is responsible when they fail… - …And decide what accountability exists outside the company that wrote the rules. Maybe that is the real guardrail debate: not rules vs. no rules, and not slowdown vs. acceleration in the abstract. Whether the technical controls inside the model, the evaluations around it, and the accountability outside the lab are strong enough to survive contact with the exact incentives pushing the other way: capability, convenience, competition, and speed. A guardrail that only works when everyone is behaving carefully isn’t much of a guardrail. FROM OUR PARTNERS Ready to sharpen your technical edge? Elastic{ON} is touring seven cities worldwide, including New York, San Francisco, and Washington, DC. Get hands-on with Elastic innovations, explore product roadmaps, and learn to build agentic AI across your workflows. Whether you’re an SRE diagnosing what broke and why, a security analyst chasing threats through millions of events, or a developer building intelligent apps, your data has never been more valuable. See how an evolved Elasticsearch powers smarter, faster, and more efficient architectures. Come curious. Leave with a plan. Forge the future with confidence at Elastic{ON}. Space is limited, so reserve your pass early. 🎓 AI Skill of the Day: Create first, enforce the rules second MIT’s new HardFlow method tackles a surprisingly common AI problem: when you force a model to obey every constraint while it’s still figuring out the answer, you can make the result worse. Their technical method gives the model room to explore, then enforces the hard constraints on the final output. You can borrow the same idea in your prompts: - Ask AI to solve or draft the best answer first. - Give it your non-negotiables: word count, required facts, formatting, tests, safety rules, etc. - Have it check and revise the final result against every constraint before returning it. Prompt: “Solve this for quality first. Then run a separate final pass against these non-negotiable constraints: [RULES]. Fix every violation before giving me the final answer.” 🍪 Treats to Try - Claude for Financial Advisors wraps Claude around the actual grunt work of wealth management, from meeting prep and onboarding to compliance, with Wealth.com adding cited estate and tax analysis. - Siri AI is finally something you can try: the English beta adds personal context, onscreen awareness, web knowledge, and more actions across apps on supported devices. - Perplexity Personal Computer puts its Computer agent on Windows 10/11 so it can work across your local files, Microsoft 365, and the web from one place. - Motion by Mosaic takes a plain-English video brief and builds an editable motion-design cut with storyboard, voice, music, captions, and the rest of the production stack. - ElevenLabs Hosted MCP removes the annoying local-server/API-key setup so Claude, Cursor, and other MCP clients can reach ElevenLabs through OAuth. 📰 Around the Horn - Google opened Claude to all of its engineers through its internal Antigravity system, a striking coding-model concession (but Gemini remains the default). - 404 Media reported OpenAI contractors on “Project Lily” reviewed real ChatGPT prompts, including sensitive conversations, while helping train against sycophancy. - TSA’s Ace agent handles roughly 100,000 traveler conversations a month; Salesforce says 96% of routine questions resolve without human escalation. - Superhuman acquired Fathom, pulling meeting transcripts, decisions, and action items into its email, calendar, docs, and agent stack. - MIT’s HardFlow lets generative models explore freely, then enforces hard constraints on final outputs; MIT reported perfect constraint satisfaction across robotics, navigation, and image editing. - Reward AI launched OM-1, a robot policy trained directly from human manipulation data that transferred across tabletop arms, industrial arms, and humanoids. FROM OUR PARTNERS The people building the next generation of AI are coming together in San Francisco. At The AI Conference, hear from 130+ speakers including Chris Lattner, Ion Stoica, Peter Norvig, and Illia Polosukhin, co-author of “Attention Is All You Need,” plus builders from OpenAI, NVIDIA, Google, Meta, Anthropic, and more. Learn what leading teams are building across agents, LLMs, infrastructure, and applied AI, what’s working now, and where they believe AI is heading next. The Neuron members save 30% with code NEURON30. 🔧 Tuesday Tool Tip: Try one agent before you build an agent org chart Polylane tried the thing every agent diagram eventually suggests: split one job across a little team of specialized agents. Then it ripped the setup back out. The problem was not that the agents were dumb. It was the handoffs. Each agent summarized what it learned for the next one, and every summary quietly dropped clues the next agent needed. Polylane replaced the chain with one agent that investigated the issue end to end. In its reported results, median time-to-PR fell from 2.2 hours to 35 minutes and cost per PR dropped from $111 to about $18. So if one person would normally investigate a job start to finish, make one long-context agent prove it cannot handle the job before you build it a tiny org chart. Check out The Neuron: AI Explained Podcast! A Cat’s Commentary This one got me singing “Something always… brings me back to you…” That’s all for now. If you want to get featured above, fill out the poll below and tell us how we did today!
11:04

The Sequence Knowledge - Issue 933: When the Factory Starts Building Itself

The old argument about machines improving themselves just got three lab disclosures to point at. In June 2026 Anthropic said that as of May, Claude authored more than 80 percent of the code merged into its production codebase. OpenAI said GPT-5.3-Codex helped debug its own training and manage parts of deployment. DeepMind's AlphaEvolve is landing algorithmic gains in the stack other models train on. This Sequence issue opens a series on which parts of the job moved, how well, and what checks the work.

Full text · 1,256 chars
Today, we start a new series about one of the hottest topics in AI: recursive self improvement. Throughout the next few weeks, we will deep dive into the top research For about sixty years, arguing about recursive self-improvement meant arguing in the abstract. There was no system to point at. You cited I.J. Good, somebody cited Schmidhuber, everyone disagreed about definitions for two hours, and then you went home. It was a very pleasant way to spend an afternoon and it produced nothing. That era ended sometime in the last twelve months. In June 2026 Anthropic published an essay stating that as of May, Claude authored more than 80 percent of the code merged into its production codebase. OpenAI disclosed that GPT-5.3-Codex helped debug its own training process and manage parts of its own deployment. DeepMind’s AlphaEvolve has been turning up algorithmic improvements that land in the infrastructure other models train on. Three labs, three different flavors of disclosure, one direction. So the question is no longer whether AI helps build AI. It clearly does. The question is which parts of the job it took, how well it does them, and what checks the work. That last one turns out to be the whole ballgame, and it is what this series is about.
12:00

AI models need more data about biology, and OpenAI is paying to create it

OpenAI’s charity is paying people to build the medical datasets models still lack. The OpenAI Foundation’s Public Data for Health gave $500,000 to 1Day Sooner for Ruxandra Teslo’s idea of buying failed-biotech files at bankruptcy. A $40 million grant goes to novel cancer-vaccine data at UNC Chapel Hill, plus support for OpenAdmet drug-effect contests. The foundation holds a 26% stake that could be worth about $250 billion and hopes to give away $1 billion by year end. Morrison says nonexclusive copies might cost a few tens of thousands of dollars each; two bids this year were rejected. Teslo notes about 70% of drug-development time and money sits in clinical work that stays a black box for small biotechs.

Notes
  • OpenAI Foundation Public Data for Health: pay to create “high-quality scientific datasets.” Teslo’s “biotech’s lost archive” (bankruptcy filings, manufacturing, safety data) funded at $500,000 via 1Day Sooner (Josh Morrison; Teslo advises). Also $40 million to UNC Chapel Hill novel cancer-vaccine data, plus OpenAdmet drug-effect contests.
  • Foundation: 26% equity; potential ~$250 billion vs Gates ~$180 billion end-2025. Still hiring; grantmaking ramped this year. Largest gift so far: $100 million (August) to Common Health Coalition (hepatitis C access). Hopes to give $1 billion by year end (Jacob Trefethen). Operates separately from OpenAI, same “benefits all of humanity” mission.
  • Morrison: nonexclusive copies maybe “a few tens of thousands of dollars” each. Three datasets in hand; two donated by Lumen Bioscience. Two bids this year rejected. Target files: common technical documents (regulator back-and-forth + measurements). Teslo: ~70% of drug-development money/time is clinical development, still a black box for small biotechs.
  • Parallel: Google won Spirit Airlines’ data (100 million emails) — labor objections cited as the cautionary analog. Piece also notes Altman/Musk endorsed Amodei’s slowdown ask, and that some insiders put extinction odds at 10%+ within a decade.
Full text · 6,758 chars
Last year Ruxandra Teslo, a policy analyst who focuses on clinical trials, posted an idea for supercharging medical AI systems: Use data from failed biotech companies. By bidding at their bankruptcy proceedings, she proposed, it might be possible to obtain detailed regulatory filings, manufacturing strategies, and safety data—types of information usually considered trade secrets. She called these documents “biotech’s lost archive” and said they could be used to help train AIs that would act as powerful copilots in the often opaque drug approval process. Today the OpenAI Foundation, the nonprofit parent of OpenAI, said it would fund her idea as part of a new effort it calls Public Data for Health, which aims to help artificial intelligence make big leaps in medicine by paying to create “high-quality scientific datasets.” The basic idea is that AI isn’t going to be capable of making important breakthroughs in curing disease unless researchers can feed the models much more information than they have so far. “Everyone is recognizing that data is the biggest bottleneck in successfully applying AI to biology,” says Morgan Levine, a former vice president for computation at Altos Labs, a longevity company. In its initial round of data grants, the OpenAI Foundation also announced that it would give $40 million to a program to collect data about novel cancer vaccines at the University of North Carolina, Chapel Hill, and support OpenAdmet, a group that runs competitions in which researchers try to predict drug effects. Teslo’s idea for a biotech archive received $500,000 and will be pursued by 1Day Sooner, an advocacy group representing clinical trial volunteers, which she advises. “We expect many remaining breakthroughs in preventing and curing disease to come from pairing the intelligence of new models with more observations of the world—in other words, more data,” the OpenAI Foundation said in a statement. OpenAI started as a nonprofit, but leader Sam Altman restructured it to form a for-profit corporation that develops new models, launches products, and is now planning an initial public offering of stock that could value it at $1 trillion. Because the foundation holds a 26% equity stake in OpenAI, it is now be on track to become the richest charitable organization on the planet, potentially sitting on $250 billion in stock value. (By comparison, the Gates Foundation and a trust associated with it held about $180 billion at the end of 2025.) Making good use of that kind of money will not be easy. The foundation, based in San Francisco, is still hiring for many key roles and started ramping up its grantmaking only this year. Its largest single gift so far, of $100 million, was awarded in August to the Common Health Coalition, an organization that helps patients get access to drugs for hepatitis C. OpenAI’s charitable efforts come even as apocalyptic fears have broken out about the possibility that runaway AI could wipe out all human life, possibly by launching a deadly bioweapon. Those fears have been stoked by AI company insiders, some of whom say the chance of human extinction within the next decade is 10% or more. Last week, Altman and xAI founder Elon Musk both endorsed a call by Anthropic CEO Dario Amodei to “slow the pace at which we improve the capabilities of AI models” so that risk prevention can catch up. Jacob Trefethen, an executive at the foundation, says it essentially operates separately from OpenAI but shares an official mission of ensuring that artificial intelligence “benefits all of humanity.” “We’re starting grantmaking when we think the best way to achieve that mission is to make grants to external nonprofits, research institutions, and other third parties,” Trefethen said in an interview. He says the foundation hopes to give away $1 billion by the end of the year. The $500,000 grant to 1Day Sooner will help the group prove it can obtain the data troves of bankrupt companies, says the organization’s president and cofounder, Josh Morrison. He thinks nonexclusive copies of company datasets could be acquired for only “a few tens of thousands of dollars” each. His organization is currently in possession of three datasets, two of them donated by Lumen Bioscience, a biotech that previously used the Chapter 11 strategy to gain insights into another company’s drug development efforts. Morrison says two other attempts to obtain drug company files this year proved unsuccessful, after 1Day Sooner’s bids were not accepted. Bankruptcies could become what some are calling a “new land grab” for AI training. Last month, Google won a bid to take over the corporate data of the failed carrier Spirit Airlines, including 100 million emails. That led to objections from flight attendants and others who worried that private or proprietary data could be exposed. The drug company files that 1Day Sooner is seeking are known as common technical documents. They typically contain the back-and-forth between companies and regulators, as well as detailed scientific and medical measurements, and essentially provide everything that is known about a drug. According to Teslo, who is a writer for Works In Progress and a nonresident fellow at the Institute for Progress, a think tank in Washington, DC, a stockpile of such files could help turn an AI into a regulatory expert, which in her view could be one of the main ways AI helps speed cures to market. “People say ‘We will invent AI, and AI will cure cancer,’ but that’s very removed from the messy reality and the regulatory process,” she says. “About 70% of the money and time in drug development is spent in clinical development—organizing the trials and testing the drug—but despite that, the process is basically a black box, especially for small biotech companies generating the innovations.” Deep Dive Biotechnology and health A startup claims it’s found a drug to make your blood young Generation Lab claims its drug combo can “stop the spread of aging” around the body. And it’s looking for influencers to give it a try. Montana’s plan to become an experimental medical hub just pushed forward The state’s effort to expand the “right to try” is making headway, and the first drugs are about to be reviewed. There’s a lot of hype around perimenopause. Don’t buy it. Discussions of the life stage are often clouded by misinformation. Supercooled kidneys have been transplanted into pigs in a “landmark achievement” Kidneys kept at subzero temperatures in pressure-controlled containers can be stored for days before transplantation, raising hopes for longer-term storage of donated human organs. Stay connected Get the latest updates from MIT Technology Review Discover special offers, top stories, upcoming events, and more.
15:02

Anthropic Brings Claude Inside Salesforce With 37 Built-in Sales Skills

Salespeople can now ask Claude about their pipeline without leaving the chat, and it will not write back to the CRM until they say yes. Salesforce in Claude is an open beta with 37 prebuilt sales skills for call prep, deal scoring, dashboards, and writebacks. It is the first product from the Claudeforce partnership announced with Salesforce Q2 FY27 earnings. GitLab, Siemens, and Legora already put it in front of about 7,000 sellers. Access inherits Salesforce roles through AIforce and MCP. It is on all paid Claude plans; no separate plugin price was announced.

Notes
  • Salesforce in Claude, open beta, 37 prebuilt sales skills. Writes to CRM only after seller approval. Combines permitted Salesforce, Slack, and email.
  • First shipped piece of Claudeforce, announced with Salesforce Q2 FY27 earnings. Three paths: Salesforce in Claude; Claude in Agentforce; Claude in Slack. Reciprocal: Salesforce adopts Claude Code + Claude Enterprise; Anthropic uses Salesforce as primary CRM.
  • Skill examples: morning brief (Fri week-summary + manager update); call prep (proposes missing contacts); deal review vs the team’s methodology; post-call email/Slack/stage updates; pipeline coverage + forecast narrative.
  • Auth: seller’s Salesforce identity. Read only what that user can see. Salesforce stays system of record. Org-level admin connect; group assignment. Runs on AIforce + MCP over Headless 360. Marketplace connector; full skill library via AgentExchange beta request.
  • Plans: all paid Claude. Team/Enterprise: no training on customer data by default. No separate plugin price; inference contracted with Anthropic. GitLab, Siemens, Legora — about 7,000 sellers. More skills “late 2026,” other teams later.
  • Named failure modes: stale records, inconsistent stages, missing contacts, broad permissions, one connected app down.
Full text · 6,085 chars
- Anthropic launched Salesforce in Claude in open beta with 37 prebuilt sales skills. - Plugin covers call prep, deal scoring, pipeline dashboards, and CRM writebacks with seller approval. - Part of the broader Claudeforce partnership announced alongside Salesforce Q2 FY27 earnings. - GitLab, Siemens, and Legora already deployed it to roughly 7,000 sellers. - Runs on Salesforce's AIforce harness and MCP; permissions inherit from existing Salesforce roles. - Available on all paid Claude plans; more skills for other teams shipping in coming months. Anthropic has opened beta access to Salesforce in Claude, a plugin that brings Salesforce records and actions into Claude. Its 37 prebuilt sales skills can assemble account context, analyze opportunities, update pipeline data, and draft forecasts. Claude requests the seller’s approval before writing changes to Salesforce. A seller can ask Claude to prepare for a call, review a deal, build a pipeline dashboard, or summarize the week. The plugin combines permitted data from Salesforce, Slack, and email, reducing the manual work of finding records and reconciling conversations across applications. Claudeforce links three surfaces Salesforce and Anthropic announced the beta as the first released product from Claudeforce, an expanded partnership unveiled alongside Salesforce’s Q2 FY27 earnings. The agreement covers three integration paths: | Integration | Role | |---|---| | Salesforce in Claude | Exposes permitted Salesforce data and governed CRM actions inside Claude. | | Claude in Agentforce | Adds Claude as a reasoning model for Salesforce’s platform for building and deploying AI agents. | | Claude in Slack | Extends Claude’s access to conversations and workflows inside Slack. | The partnership also includes reciprocal internal deployments. Salesforce is adopting Claude Code and Claude Enterprise for engineering work, while Anthropic uses Salesforce as its primary CRM. Thirty-seven workflows, packaged The initial skill library follows the recurring work of an account executive: - Morning brief: Produces a daily digest of meetings, opportunities nearing their close dates, at-risk deals, and unread threads. On Fridays, it summarizes the week and drafts a manager update. - Call preparation: Collects open opportunities, recent Slack discussions, unanswered messages, and unresolved questions from earlier calls. When it identifies stakeholders missing from Salesforce, it proposes contact records for the seller to approve. - Deal review: Scores an opportunity against the team’s sales methodology, identifies qualification gaps and missing stakeholders, and drafts a business case and dated mutual close plan. - Post-call follow-up: Converts notes or a transcript into an email, a Slack deal-channel summary, and proposed changes to next steps, stage, and close date. - Pipeline analysis: Builds an interactive view of coverage by stage and opportunities likely to slip, then drafts a forecast narrative in the format used by sales leadership. Salesforce permissions stay in charge The access model reuses Salesforce identity and authorization rather than creating a separate permissions layer inside Claude. Its main controls are: - Sellers authenticate with their Salesforce credentials. - Claude can read only the objects, fields, and records available to that user. - Salesforce remains the system of record for CRM data. - Claude requests approval before committing a write by default. - Administrators connect Salesforce at the organization level and assign access to selected groups. MCP supplies the plumbing Salesforce in Claude runs on AIforce, Salesforce’s enterprise integration layer for connecting business data and workflows to AI agents. It uses APIs, command-line tools, and Model Context Protocol servers. MCP is a standard that allows an AI client to discover external tools, retrieve authorized data, and invoke supported actions. AIforce builds on Salesforce’s Headless 360 architecture, which exposes data, workflows, agents, and governance without requiring users to navigate the standard Salesforce interface. The Salesforce MCP connector is available through Claude’s marketplace. Administrators seeking the complete plugin and its skill library must request beta access through AgentExchange. Access, data terms, and cost | Availability | Beta access requires administrator approval through AgentExchange. | |---|---| | Claude plans | Anthropic says the plugin supports all paid Claude plans. | | Model training | On Team and Enterprise plans, Anthropic does not train models on customer data by default. | | Plugin pricing | No separate price has been announced. | | Inference contract | Customers contract separately with Anthropic for Claude usage. | Anthropic says GitLab, Siemens, and Legora have deployed the plugin, with about 7,000 sellers using it in production. Additional prebuilt skills are scheduled to begin arriving in late 2026. The first release concentrates on sales, with skills for other teams expected later. Agent quality follows CRM quality The release gives enterprise developers a concrete pattern for placing an AI agent over operational systems: inherit the source application’s permissions, gather context across connected services, package repeatable workflows as skills, and require human approval at write boundaries. The architecture also exposes predictable failure modes. Stale records, inconsistent stage definitions, missing contacts, and overly broad permissions all affect the context Claude receives and the actions it proposes. A beta evaluation should therefore test record-level access, field-level restrictions, write approvals, stale data, cross-application identity matching, and recovery when one connected service is unavailable. As CRM work moves from forms and list views into conversational workflows, schemas and governance carry more weight. Field definitions, permission scopes, and data completeness directly determine whether the agent produces a useful account summary, a defensible forecast, or an incorrect update.
15:36

Multiverse Computing's Quasar 1.1 Uses Quantum Data to Shrink a 438B Model

A compressed coding model was healed with synthetic examples that came off a quantum chip, and the company says answers got shorter without getting worse. Quasar 1.1 438B starts from Z.ai’s GLM-5.2 and cuts each mixture-of-experts layer from 256 experts to 148. Healing data ran on IBM Quantum System Two, a 156-qubit Heron in San Sebastián. Average output tokens fell 37.6%, from 3,322.9 to 2,074.3 across five benches. Scores rose +6.2 HLE, +4.3 GPQA, +4.6 IFBench versus Quasar 1.0. Political refusals dropped from 63.75% to 41% while JailbreakBench stayed at 93%. It is on the CompactifAI API; no token price was published.

Notes
  • Quasar 1.1 438B from Multiverse Computing, based on Z.ai GLM-5.2. CompactifAI pruning: MoE experts 256 → 148 (−42.2%). Recovery (“healing”) uses reasoning traces, tool-call sequences, general knowledge, plus new quantum-generated examples.
  • Quantum stage: modified Qwen3-30B-A3B with one-sixth of layers replaced by a quantum neural net, run on IBM Quantum System Two in Donostia-San Sebastián (156-qubit Heron) plus a noise model from that hardware. Inference does not need a quantum chip. No ablation isolating the quantum examples.
  • Vs Quasar 1.0: HLE +6.2, GPQA +4.3, IFBench +4.6. Average output on SciCode, HumanEval, GSM8K, TriviaQA, BBH: 3,322.9 → 2,074.3 tokens (−37.6%, 1,248.6 fewer). Quality and length benches are different sets.
  • Refusal Steering: ridge-regularized vector in activation space. Political refusals: GLM-5.2 71.18%, Quasar 1.0 63.75%, 1.1 41.00%. JailbreakBench harmful: 1.0 92%, 1.1 93%.
  • Sold as agentic coding / tool use. EU-incorporated; “compatible with EU AI Act transparency” is a claim, not an audit. CompactifAI API; no token price. Bug-bounty: five reporters of verified flaws get 60 million tokens each. Listed on Artificial Analysis.
  • Open: pricing, context/max-out, rate limits, residency/retention, weights, bench prompts/variance, quantum-data ablation.
Full text · 7,196 chars
- Multiverse Computing released Quasar 1.1 438B, first LLM trained partly on quantum-generated synthetic data. - Healing data produced on IBM Quantum System Two, a 156-qubit Heron processor in San Sebastian. - 37.6% fewer output tokens on average, cutting serving cost with no accuracy loss. - Gains of +6.2 HLE, +4.3 GPQA, +4.6 IFBench, +6.4 LCR versus Quasar 1.0. - Political refusal rate drops from 63.75% to 41%, JailbreakBench safety held at 93%. - Available now via CompactifAI API; bug-bounty offers 60M tokens for top flaw reporters. Quasar 1.1 Uses Quantum-Generated Data to Repair a Pruned 438B Model Multiverse Computing has released Quasar 1.1 438B, a compressed coding model based on Z.ai’s GLM-5.2. During post-pruning training, Multiverse added synthetic data generated by a hybrid quantum language model. The company says this is the first use of quantum-generated data in its CompactifAI pipeline. The compression system reduces each mixture-of-experts layer from 256 experts to 148, cutting the expert count by 42.2%. In a mixture-of-experts model, each expert is a specialized parameter block, and a router selects a subset for each token. Removing experts reduces the model’s stored parameters and deployment footprint, although realized cost and latency depend on the serving stack. Aggressive pruning can damage reasoning, instruction following, and domain knowledge. Multiverse addresses that loss through a recovery stage it calls “healing,” which retrains the pruned network on reasoning traces, tool-call sequences, general knowledge, and the new quantum-generated examples. Quantum circuits feed the healing set Multiverse created part of the healing data with a modified Qwen3-30B-A3B model. The researchers replaced one-sixth of its layers with a quantum neural network, whose parameterized circuits acted as trainable transformations inside the language model. Those circuits ran on IBM Quantum System Two in Donostia-San Sebastián, using a 156-qubit IBM Heron processor. Additional runs used a noise model calibrated from the same hardware. The hybrid model then generated synthetic examples for Quasar’s recovery training. The announcement places quantum hardware in the data-generation stage; it does not indicate that Quasar’s production inference requires a quantum processor. CompactifAI already used quantum-inspired tensor-network methods to select experts for removal. Quasar 1.1 extends that work by incorporating outputs from physical quantum circuits into the training corpus. Multiverse has not published an ablation that trains otherwise identical models with and without those examples, so the quantum data’s individual contribution to the reported gains remains unmeasured. Benchmarks rise as outputs shrink Multiverse reports gains over Quasar 1.0 on reasoning and instruction-following evaluations. It also reports a lower average output length across five coding, reasoning, and question-answering benchmarks. | Measurement | What it tests | Reported change | |---|---|---| | HLE | Broad, difficult academic questions | +6.2 points | | GPQA | Graduate-level scientific reasoning | +4.3 points | | IFBench | Instruction following | +4.6 points | | Average output length | SciCode, HumanEval, GSM8K, TriviaQA, and BBH | 3,322.9 to 2,074.3 tokens, down 37.6% | The output reduction amounts to 1,248.6 fewer tokens per response on average across the measured tasks. Token-metered agent systems could see lower generation costs and shorter loops, especially when one response triggers several subsequent tool calls. The quality gains and token counts come from different benchmark groups, so teams should verify both measures on the same production workload. A targeted edit lowers political refusals According to Multiverse, GLM-5.2 carries topic-level refusals and state-aligned framing on some political and historical subjects. Quasar 1.1 applies the method described in Refusal Steering to reduce those refusals while preserving rejection behavior for harmful prompts. The method uses an LLM judge to score refusal confidence, then fits a ridge-regularized vector representing the refusal-to-compliance direction in the model’s internal activation space. Adjusting the model along that vector targets a specific behavior without requiring a full retraining run. | Prompt set | GLM-5.2 | Quasar 1.0 | Quasar 1.1 | |---|---|---|---| | Politically sensitive prompts | 71.18% refusal | 63.75% refusal | 41.00% refusal | | Harmful prompts on JailbreakBench | N/A | 92.00% refusal | 93.00% refusal | Within these test sets, political refusals fell by 22.75 percentage points from Quasar 1.0, while harmful-prompt refusals increased by one point. Broader claims require adversarial evaluation across paraphrases, multilingual prompts, indirect requests, and tool-enabled attacks. Agent loops are the target workload Multiverse positions Quasar 1.1 for agentic coding and tool use. Its healing data emphasizes parseable tool calls, multi-step reasoning, executable code, long-context tasks, and concise final answers. The lower output-token count fits workloads where verbose responses increase both latency and the number of tokens carried into later turns. The company also presents its EU jurisdiction as a deployment consideration. Multiverse is incorporated under EU law and describes the service as compatible with EU AI Act transparency expectations. Compliance, data residency, and processing location still depend on the service’s contracts, hosting regions, retention policies, and subprocessors. API access comes with open questions Quasar 1.1 is available through the CompactifAI API. The launch announcement does not provide per-token pricing. Multiverse is also running an evaluation challenge in which the five participants who report the most verified flaws each receive 60 million API tokens. A model listing is available on Artificial Analysis. A production-grade test plan - Coding agents: Measure compilation, test-pass rates, patch acceptance, tool-call validity, retries, and total tokens per completed task. - Tool-heavy RAG: Check schema adherence, citation accuracy, retrieval grounding, and recovery from failed tool calls. - Sensitive topics: Evaluate factual accuracy alongside refusal rates, using paired political and harmful prompts. - Serving performance: Record median and tail latency, output length, concurrency limits, and cost per successful workflow. Questions the launch leaves open - Input and output token pricing - Context-window and maximum-output limits - Rate limits, batching support, and structured-output guarantees - Data retention, hosting regions, and training-data policies - Model-weight availability and versioning commitments - Benchmark prompts, variance, contamination checks, and quantum-data ablations Agentic coding systems and tool-heavy retrieval pipelines provide the clearest initial tests because token use, schema validity, and task completion can be measured together. Teams evaluating open discussion of contested political or historical topics should pair refusal testing with factuality and safety checks rather than treating a lower refusal rate as a complete quality measure.
15:39

LM Studio 1.1.3 Brings Private On-Device Voice Transcription to Linux

You can now talk to a local model on Linux without sending the recording to the cloud. LM Studio Bionic 1.1.3 adds realtime on-device transcription on Mac, Windows, and Linux. It accelerates on Apple Silicon and NVIDIA GPUs; AMD is still in the works. The same release ships llama.cpp 2.38.0 extension packs and MTP speculative decoding for more models. It is a free update. The company did not name the speech model or publish languages, latency, or memory use.

Notes
  • LM Studio Bionic 1.1.3: realtime on-device STT on macOS, Windows, and Linux. Audio stays on device. Accelerated on Apple Silicon and NVIDIA; AMD in development. Intel Mac / iGPU / CPU-only: no accelerated support announced.
  • Also in the notes: inline agent images in chat; MTP speculative decoding for more models; llama.cpp 2.38.0 extension packs; smoother remote models via LM Link.
  • Not published: speech model/engine, languages, size, memory, latency, accuracy, audio retention. No official latency bench. Unified memory on Apple vs VRAM split on NVIDIA (STT + LLM + KV + tools).
  • How to try: install 1.1.3+, mic permission, load a local LLM, pick voice in the composer. Existing OpenAI-compatible server remains http://localhost:1234/v1. No documented /audio/transcriptions endpoint — treat STT as a desktop feature.
  • Privacy caveat: text can still leave via remote models, LM Link, tools, MCP, backups, telemetry. STT only; spoken replies still need separate TTS.
Full text · 5,989 chars
- LM Studio Bionic 1.1.3 adds local realtime voice transcription across Mac, Windows and Linux. - Runs on Apple Silicon and NVIDIA GPUs; AMD support is in the works. - Audio never leaves the device, removing cloud STT dependency for local agents. - Ships alongside llama.cpp 2.38.0 extension packs and MTP speculative decoding for more models. - Free update via the LM Studio download page. - Full notes in the Bionic 1.1.3 changelog. LM Studio 1.1.3 brings local voice transcription to Linux LM Studio 1.1.3 adds real-time, on-device speech transcription to Linux, joining the existing macOS and Windows implementations. The desktop app can now convert microphone audio into prompt text without sending recordings to a cloud transcription service. LM Studio provides a graphical interface and an OpenAI-compatible local server for running large language models on personal hardware. Voice input previously required a separate speech-to-text service or community software such as this community bridge. Native integration removes that extra process and its supporting code. What shipped in 1.1.3 The 1.1.3 release notes describe local voice transcription as the main Linux addition. The update also includes several changes to model execution and chat behavior: - Local transcription: Real-time microphone input is available on macOS, Windows, and Linux. - Hardware acceleration: Apple Silicon and NVIDIA GPUs are supported, while AMD GPU support remains in development. - Inline images: Images created by agents can render directly inside conversations. - MTP speculative decoding: More compatible models can use a technique that verifies multiple candidate tokens to accelerate generation. - Runtime updates: The release adds llama.cpp 2.38.0 extension packs and smoother handling of remote models through LM Link. LM Studio has not identified the speech model or inference engine behind the feature. Its release materials also omit supported languages, model size, memory use, latency measurements, accuracy results, and audio retention details. Hardware support at a glance | Platform | Hardware | Published status | |---|---|---| | macOS | Apple Silicon | Available with acceleration | | Windows | NVIDIA GPU | Available with acceleration | | Linux | NVIDIA GPU | Added in version 1.1.3 | | Windows or Linux | AMD GPU | In development | | Other configurations | Intel Mac, integrated GPU, or CPU-only | No accelerated support announced | Why latency depends on the whole machine Real-time transcription processes short audio segments while the speaker is still talking, producing partial text and revising it as more audio arrives. That workload must share compute and memory with the language model generating the response. Apple Silicon uses a unified memory pool for both workloads. NVIDIA systems divide the available VRAM among the speech model, language model, context cache, and other active processes. A large language model can therefore slow transcription or trigger memory pressure even when the speech component runs efficiently. No official benchmark establishes expected latency for this release. Performance will depend on the GPU, available memory, selected language model, prompt length, audio quality, and any agent tools running during the conversation. Where local speech input fits - Coding and note-taking: Developers can dictate prompts, documentation, or rough notes directly into a local model. - Sensitive audio: Local processing reduces exposure to third-party transcription providers, although regulated deployments still require security and compliance review. - Offline systems: Voice input can work on disconnected, air-gapped, or unreliable networks. - Multilingual workflows: Local transcription can remove per-minute API charges when the bundled engine supports the required language. - Agent interfaces: Spoken instructions can enter the same chat flow used for tools, images, and local documents. How to try it Bionic 1.1.3 is a free update for existing users. The built-in updater can install it, while new users can obtain the application from LM Studio downloads. - Install version 1.1.3 or later. - Grant LM Studio access to the system microphone. - Load a compatible local language model. - Select voice transcription in the chat composer. - Speak, review the generated text, and submit it as a prompt. LM Studio’s existing OpenAI-compatible server remains available at http://localhost:1234/v1, so applications already sending text prompts to that endpoint can continue doing so. The voice feature converts speech into text inside the desktop chat interface. The release notes do not document an OpenAI-compatible audio endpoint such as /audio/transcriptions. Developers building automated speech pipelines should therefore treat transcription as an application feature until LM Studio publishes an audio API. Privacy stops at the transcription boundary On-device speech recognition removes the cloud transcription hop, but the generated text can still leave the computer through remote models, LM Link, agent tools, plugins, networked MCP servers, backups, or application telemetry. Deployments handling confidential material should audit the full request path, chat storage, logs, microphone permissions, and every enabled integration. Accuracy also depends on the undisclosed speech model and the recording environment. Background noise, overlapping speakers, specialized vocabulary, and unfamiliar accents can increase errors. The release covers speech-to-text; a fully spoken assistant still requires a separate text-to-speech component. Voice joins the local model stack By folding microphone input into the same desktop runtime as chat, vision, tools, and local model serving, LM Studio reduces the number of services required for a private voice interface. Version 1.1.3 gives Linux users the integrated workflow now, while AMD acceleration, documented audio APIs, and native speech output remain open areas for future releases.
16:36

Odyssey Builds One AI Backbone to Control Robots, Cars, Drones, and Games

One frozen world model is being asked to drive robot arms, humanoids, cars, drones, and game characters by training only a small action head. Odyssey-3 is an autoregressive diffusion transformer. Sim-only driving policies traveled about 77% as far between interventions on Indian roads as policies trained on real footage. Humanoid work with Flexion kept going under lighting changes that broke the vision-language-action baselines they tested. A mobility policy trained on GTA V produced horseback movement in Red Dead Redemption 2 with no game-specific examples. A public release is promised in coming weeks. Independent benches and most success rates are still missing.

Notes
  • Odyssey-3: autoregressive diffusion transformer world model. Frozen backbone + small action decoder per body (arm, humanoid, car, drone, game). Company-reported only.
  • Arms: tens of hours of demos; multi-step tasks plus recoveries not in the demos (reorient after a miss, odd-position retrieval). Humanoids with Flexion: tens of hours of teleop; better generalization than tested VLA baselines, including lighting that broke those baselines. Driving: 20 hours sim; closed-loop on Indian roads; sim-only policies ~77% of real-footage policies’ distance between interventions. Drones: tens of hours sim indoor; no physical flight reported. Games: extended GTA V, ~two hours used for transfer; GTA mobility policy produced horseback movement in Red Dead Redemption 2 with no RDR2 examples (“zero-shot”).
  • PROWL: agents train inside generated environments; failures feed the world model. Piece says sim cannot prove physical safety.
  • Gaps: no independent benches; 77% is relative, not absolute intervention-free km; no task-level success rates; no param count, data mix, Hz, latency, or decoder recipe. Poke & Wiggle benchmarking partnership announced, results unpublished. Public release “within weeks”; no download, license, price, or hardware spec. Robotics / driving / gaming / defense users unnamed.
Full text · 7,050 chars
- Odyssey unveiled Odyssey-3, an autoregressive diffusion transformer foundation world model for embodied control. - Same frozen backbone controls robot arms, humanoids, cars, drones, and video game characters via lightweight action decoders. - Sim-only driving policies reached 77% of real-footage policies' distance between interventions on Indian roads. - Humanoid work with Flexion generalized to lighting changes that broke VLA baselines. - Zero-shot transfer: GTA-trained mobility policy produced horseback movement in Red Dead Redemption 2. - Public release planned in coming weeks; already deployed with robotics, driving, gaming, and defense partners. Odyssey-3 shares a world-model backbone across robots, cars, drones, and games Odyssey has announced Odyssey-3, an autoregressive diffusion transformer designed to simulate environments and control several kinds of embodied systems. The company presents it as a reusable foundation model: developers keep the pretrained backbone fixed and train a compact action decoder on observations and controls from a target robot, vehicle, drone, or game. If the reported transfer holds across independent tests, the design could reduce platform-specific data collection and full-model retraining. One backbone, many controllers Odyssey describes its model as an autoregressive diffusion transformer. Autoregression predicts sequences one step at a time, and diffusion generates outputs through iterative refinement. Odyssey combines those methods to model how environments evolve and how actions change them. According to the company, the same pretrained backbone supports every demonstration. For each target system, developers collect experiential data consisting of paired observations and actions. A learned decoder then translates the model’s internal representations into joint movements, steering inputs, flight controls, or game commands. Training updates the decoder while the backbone remains frozen, preserving its pretrained parameters and limiting the number of components that require adaptation. Five demos, different limits Odyssey reports control experiments across five domains, plus a separate project that uses the model to train other AI agents. | Company-reported Odyssey-3 demonstrations | | | |---|---|---| | Domain | Training data | Reported result | |---|---|---| | Robot arms | Tens of hours of demonstrations | The model controlled several arms and completed multi-step tasks. Odyssey also observed recoveries absent from the demonstrations, including reorienting a gripper after a missed grasp and retrieving objects from unusual positions. | | Humanoids | Tens of hours of teleoperation data | In collaboration with Flexion, Odyssey ran tasks in real time and reported stronger generalization than the tested vision-language-action baselines, including continued operation under lighting changes that disrupted those baselines. | | Autonomous driving | 20 hours of simulated driving | The policy generated trajectories in real time and drove in closed loop on Indian roads. Policies trained entirely in simulation traveled about 77% as far between safety-driver interventions as policies trained on real footage. | | Drones | Tens of hours of simulated flight | The resulting policy maintained stable flight and avoided obstacles in a simulated indoor environment. Physical flight was not reported. | | Video games | Extended GTA V play, including about two hours used for a transfer experiment | A mobility policy trained on GTA V produced horseback movement in Red Dead Redemption 2 without game-specific training examples, a setup Odyssey describes as zero-shot transfer. | PROWL turns the model into a training environment Odyssey also uses Odyssey-3 as an interactive simulator through a companion project called PROWL. AI agents act inside generated environments and learn from the consequences. Their failures can supply examples for improving the world model, while the model produces additional situations for training the agents. That feedback loop could support stress testing and exposure to rare scenarios before hardware deployment. Simulation alone cannot establish physical safety because modeling errors may omit real hazards or create unrealistic ones. Agents trained through PROWL would still require evaluation on physical systems and under independently designed tests. The economics hinge on transfer Embodied AI systems commonly depend on data collected for a particular body, task, and environment. A transferable backbone concentrates broad dynamics learning in pretraining and limits later work to a smaller decoder. Potential gains include fewer trainable parameters, shorter adaptation cycles, and reuse across control interfaces, provided the backbone captures the physics and causal structure each platform requires. The cross-game experiment offers evidence for representations that extend beyond one environment. A policy learned from GTA V footage generated movement for a different character in Red Dead Redemption 2. One example cannot establish broad transfer, however, or distinguish abstract locomotion knowledge from transfer enabled by similar controls and visual patterns. The evidence still has gaps - Independent validation: The figures come from Odyssey’s announcement. Independent benchmark results are not yet available. - Driving metrics: The 77% result is relative to a policy trained on real footage. Developers still need absolute intervention-free distances, route details, traffic conditions, trial counts, and variance. - Physical coverage: The drone experiment remains in simulation. The robot-arm and humanoid claims need task-level success rates across hardware, environments, and viewpoints. - Collection costs: Tens of hours is modest by embodied-AI standards, but teleoperation remains expensive and must be repeated for each new embodiment. - Implementation details: The announcement does not specify the parameter count, pretraining-data composition, inference hardware, control frequency, latency, or decoder-training recipe. - Generalization: Odyssey has announced a benchmarking partnership with Poke & Wiggle to test the model across bodies and viewpoints. Results from that work have not been published. Access remains gated Odyssey says it plans to release Odyssey-3 publicly within weeks of the announcement and directs developers to its developer portal. At publication, no public download or pricing had been announced. The company also has not specified the release format, license, checkpoint access, API limits, or hardware requirements. Odyssey says organizations in robotics, autonomous driving, gaming, and defense are already using the model, though it has not identified those users or described their deployments. Independent benchmarks will need to measure data efficiency, control reliability, latency, and sim-to-real transfer against body-specific baselines. Consistent gains on those measures would support Odyssey’s case for a shared foundation model across embodied systems.
17:03

Factory Raises $200M, Tripling Its Valuation to $5B in Five Months

A coding-agent startup just raised a huge check and says it is now worth three times what it was in the spring. Factory took $200 million at a $5 billion valuation, up from $1.5 billion about five months ago, and has raised more than $400 million in total. Blackstone is both an investor and a customer, joined by Khosla, Sequoia, Insight, NEA, and others. Named customers include Nvidia, RBC, Palo Alto Networks, Adobe, and T-Mobile. Factory Router is said to cut token spend more than 60% by picking a model per task. It did not publish current revenue, pricing, or how that 60% was measured.

Notes
  • Deal: $200 million at $5 billion, vs April $150 million at $1.5 billion (~3.3×). Total funding >$400 million. Founded 2023. Product: Droid (plan, write, review, ship). No ARR, terms, or current revenue published.
  • Investors: Blackstone (also a customer), Khosla, Sequoia, Insight, Evantic, Sound Ventures, NEA, Mantis VC, Clearlake. Individuals: Nico Rosberg, Brad Gerstner, Marc Benioff.
  • Founders: Matan Grinberg, Eno Reyes. Named customers: Nvidia, Blackstone, RBC, Palo Alto Networks, Adobe, T-Mobile. Contract sizes and production depth undisclosed.
  • Product claims: Factory 2.0 coordinated workflows; Factory Router picks a model per task and is said to cut token spend >60% “while preserving performance”; air-gapped / FedRAMP-oriented configs (authorization status not confirmed here); Agent Effectiveness ties spend to delivered code. “Self-improving” is undefined (prompts vs routing vs weights).
  • April growth line: revenue doubled month-over-month for six months — no base given. Competitors named: Cursor, Cognition, Claude Code, OpenAI coding products.
  • Buyer checklist in the piece: deployment boundary and telemetry; data retention; permissions/CI/approvals; router pin/override; total cost vs the 60% claim; accepted diffs, defects, incidents — not generated-code volume.
Full text · 6,859 chars
- Factory raised $200M at a $5B valuation, tripling its April mark. - Total funding now exceeds $400M across investors including Blackstone, Khosla, Sequoia, Insight, and NEA. - Customers include Nvidia, Blackstone, RBC, Palo Alto Networks, Adobe, and T-Mobile. - Factory Router cuts token spend more than 60% via task-level model routing. - Platform supports cloud, on-prem, and fully air-gapped deployments for regulated buyers. - Positions Factory as enterprise counterweight to Cursor, Cognition, and Claude Code. Factory raises $200 million at a $5 billion valuation Factory’s announcement says the enterprise coding-agent company raised $200 million at a $5 billion valuation, more than tripling its $1.5 billion valuation from about five months earlier. The round takes total funding above $400 million and will finance research, product development, and global sales and deployment. | Deal snapshot | | |---|---| | Metric | Details | |---|---| | New financing | $200 million | | Announced valuation | $5 billion | | Previous financing | $150 million at a $1.5 billion valuation | | Total funding | More than $400 million | | Founded | 2023 | | Core product | Droid, an agent for planning, writing, reviewing, and shipping code | An April report placed Factory’s previous round at $150 million and its valuation at $1.5 billion. The latest announced valuation is 3.3 times that figure. Factory has not disclosed current annual recurring revenue or detailed financing terms. Blackstone is buyer and backer Blackstone participated in the financing and appears on Factory’s customer list, giving the investment firm exposure as both a shareholder and a buyer. Eight other institutional investors joined the round: - Khosla Ventures - Sequoia Capital - Insight Partners - Evantic Capital - Sound Ventures - NEA - Mantis VC - Clearlake Nico Rosberg, Brad Gerstner, and Salesforce chief executive Marc Benioff also invested as individuals. Droid reaches from plan to release Founders Matan Grinberg and Eno Reyes started Factory in 2023 to automate work across the software development lifecycle. Its Droid agent can plan tasks, write and review code, and ship changes within the permissions and workflows configured by an engineering team. Factory positions Droid as part of a governed enterprise system. Administrators choose which models handle particular tasks, where workloads run, and how agent activity is measured. The company describes the system as self-improving, though it has not detailed whether that learning changes prompts, routing policies, evaluations, model weights, or another layer of the stack. Factory names Nvidia, Blackstone, Royal Bank of Canada, Palo Alto Networks, Adobe, and T-Mobile as customers. Those references span banking, telecommunications, security, and software, where code handling, auditability, and deployment boundaries can shape purchasing decisions. Contract sizes and the extent of production use remain undisclosed. Deployment carries the pitch Factory’s enterprise strategy centers on a control layer that can operate in the cloud, on premises, or within an air-gapped environment isolated from the public internet. Recent product releases support that approach: - Factory 2.0 organizes agents into coordinated workflows across the development lifecycle. - Factory Router selects a model for each task according to cost and performance requirements. - Deployment configurations target air-gapped, regulated, and public-sector environments. - Agent Effectiveness connects agent usage and spending with measures of delivered code. Factory says its router reduces token spending by more than 60% while preserving performance. Routing can send routine work to a cheaper model and reserve more expensive models for harder tasks. The resulting savings depend on the workload mix, model prices, routing accuracy, and the evaluation used to define a successful result. Air-gapped deployment can keep source code and execution inside a controlled network. Buyers evaluating Factory’s FedRAMP-oriented offering should confirm its current authorization status, covered services, external dependencies, logging behavior, and model endpoints. What the valuation assumes Factory said during its April financing that revenue had doubled month over month for six consecutive months. The company did not publish the starting revenue base, current recurring revenue, customer retention, or gross margin, leaving the scale and durability of that growth unclear. Factory competes for enterprise development budgets with Cursor, Cognition, Anthropic’s Claude Code, and OpenAI’s coding products. Its stated differentiation rests on model choice, deployment flexibility, governance, and measurement across large organizations. The new capital can support the security reviews, integrations, sales teams, and customer support required for those deployments. A $5 billion valuation places weight on Factory converting customer trials into broad production use while controlling inference and implementation costs. Changes in model quality and pricing could strengthen its routing economics or reduce the value of a separate orchestration layer. A rollout checklist for engineering teams - Map the deployment boundary. Confirm which components run in Factory’s cloud, a private environment, an on-premises installation, or an air-gapped network. Identify any telemetry, update, or model endpoint that crosses that boundary. - Trace data handling. Document how source code, prompts, logs, embeddings, and generated output are stored, retained, and used. Clarify what “self-improving” means for proprietary data. - Review agent permissions. Test repository access, branch protections, CI/CD integration, approval gates, audit logs, rollback procedures, and limits on production actions. - Evaluate the router. Benchmark representative tasks across supported models, inspect fallback behavior, and determine whether teams can pin model versions or override routing decisions. - Calculate total cost. Include licenses, token usage, infrastructure, implementation, security review, and support. Compare Factory’s claimed token savings against the organization’s own workload. - Measure production outcomes. Track accepted changes, review time, defect rates, deployment frequency, incident rates, and developer time saved alongside generated-code volume. The missing numbers Factory has yet to publish current annual recurring revenue, pricing, gross margin, paid production deployments, or the methodology behind its 60% token-spending claim. Clearer definitions of self-improvement, customer data isolation, and government authorization would also help technical buyers assess the platform. Those figures will determine whether the financing reflects durable enterprise adoption or expectations that still need operating evidence.
04:00

Token Merging for Multilingual Speech Recognition: A Systematic Study Across Model Scale and Fine-Tuning

You can make Whisper cheaper at decode time by merging leftover audio tokens, even after you have already fine-tuned it. Token merging shortens the sequence without a retrain. The study runs the Whisper family across sixteen languages and three sizes, and it also tests DoRA fine-tunes on low-resource languages. Merging raises efficiency with almost no transcription loss in most of those settings, including after fine-tuning. No specific word-error number is in the abstract.

Full text · 1,807 chars
Computer Science > Computation and Language Title:Token Merging for Multilingual Speech Recognition: A Systematic Study Across Model Scale and Fine-Tuning View PDF HTML (experimental) Abstract:Leading multilingual speech recognition models like Whisper transcribe diverse, low-resource languages without language-specific training but are computationally expensive to deploy. Token merging mitigates this inefficiency by dynamically combining redundant features, shortening the sequence length during inference without requiring retraining. In this paper, we systematically evaluate token merging on the Whisper model family across sixteen diverse languages and three different model sizes. We also test how token merging interacts with fine-tuning (DoRA) on low-resource languages. Our findings show that merging tokens increases computational efficiency with almost no loss in transcription accuracy across most low-resource languages and model sizes, and it works even after the model has been fine-tuned. Our results demonstrate that token merging is a highly practical method for making multilingual speech recognition faster and cheaper to deploy. Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
04:00

PhysMent: An Interactive Approach For LLM Reasoning In Physics Problems

Models that ace a printed physics quiz still fail when they have to poke the world to learn the numbers. PhysMent gives 105 MuJoCo classical-mechanics scenes and forces the model to apply forces, query states, step time, and change geometry before answering. Qualitative single-concept work reaches up to 80% accuracy. On the hardest quantitative single-concept bin most models fall below 30%, and the bottleneck is multi-step tool use, not the concept. Across seven models, accuracy runs 25% to 67%, with premature answers and weak grounding in simulator feedback.

Full text · 2,240 chars
Computer Science > Computation and Language Title:PhysMent: An Interactive Approach For LLM Reasoning In Physics Problems View PDF HTML (experimental) Abstract:Large language models (LLMs) perform strongly on static science benchmarks, yet their ability to reason about the physical world through active experimentation remains poorly understood. We introduce PhysMent, a benchmark that evaluates LLM physical reasoning via iterative, toolmediated interaction with a MuJoCo physics simulator. Unlike static benchmarks that supply all quantities upfront, PhysMent requires models to discover information by applying forces, querying object states, advancing time, and modifying scene geometry before answering. The benchmark comprises 105 scenes of classical mechanics, organized across four difficulty regimes (Easy/Hard and Single/Multi), three scene modalities (standard, object creation, hidden objects), and a scene-manipulation category, evaluated with a six-dimensional scoring framework. Results show that current models perform reasonably well on qualitative single-concept tasks (up to 80% accuracy) but degrade substantially on quantitative tasks that demand precise, multi-step experimental procedures: most models fall below 30% on the hardest single-concept category, where the bottleneck is procedural (adaptive multi-step tool use) rather than conceptual load. Across the seven models, accuracy ranges from 25% to 67%, with failures due to premature answer submission, inefficient exploration, and inconsistent grounding in simulator feedback rather than conceptual gaps. Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
04:00

TestHallVQA: Exploring LVLMs' Document-Level Reasoning under Redundant Contexts from Scientific Exams

Vision models that can read a long exam packet still get lost when extra pages are stuffed in. TestHallVQA is a multi-image visual-question set built like a real test, with controllable document-level junk. A new score, F1-R², jointly tracks reasoning and whether the model still finds the right evidence under that junk. Irrelevant visual tokens measurably hurt. The paper says mainstream LVLMs show latent holes on both axes. Datasets and code are promised at a link in the abstract.

Full text · 2,241 chars
Computer Science > Computation and Language Title:TestHallVQA: Exploring LVLMs' Document-Level Reasoning under Redundant Contexts from Scientific Exams View PDF HTML (experimental) Abstract:Large Vision--Language Models (LVLMs) are increasingly expected to perform visual question answering (VQA) over planar media. However, existing planar VQA benchmarks typically emphasize isolated challenges: some emphasize long-document understanding with limited reasoning depth, while others require complex visual reasoning but remain restricted to single-page, noise-free settings. Moreover, through theoretical analysis, we identify the impact of irrelevant visual tokens, which leads to measurable performance degradation but has received little attention with respect to systematic quantification. To address these limitations, we introduce TestHallVQA, a multi-image VQA benchmark that simultaneously embodies document-level scale and the difficulty of human examinations, while providing comprehensive task coverage. Leveraging TestHallVQA's ability to controllably inject multi-level contextual redundancy, we further propose a novel metric, F1-R\textsuperscript{2}, which jointly quantifies LVLMs' computational reasoning capability and their evidence retrieval robustness against document-level redundancy. Extensive experiments and analyses on mainstream LVLMs reveal their latent deficiencies across multiple dimensions, offering concrete insights and directions for future research. The associated datasets, code, and complete theoretical derivations are available at this https URL. Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
04:00

RFCLLM: Evaluating LLMs' Reasoning Ability of Network Protocol State Machines

A model that turns a protocol spec into a state machine may not actually understand the spec. RFCLLM checks whether an LLM's implied finite-state machine matches a hand-built ground truth. The authors wrote 4 tasks and 1,482 queries across 16 protocols, and they vary judge bias, four context types, and protocol traits. They treat this as a step toward trusting LLM-built FSMs for networking security and testing. The abstract does not publish a headline accuracy.

Full text · 1,817 chars
Computer Science > Computation and Language Title:RFCLLM: Evaluating LLMs' Reasoning Ability of Network Protocol State Machines View PDF HTML (experimental) Abstract:Mapping textual specifications into formal representations is essential for ensuring the correctness of protocol designs and implementations. LLM-generated mappings, used for networking security or testing, are assumed to capture a perfect understanding of the specification, which may not hold in practice. The goal of this paper is to assess the extent to which LLMs can interpret the specification correctly. We examine the degree to which an LLM's implicit representation of a finite-state transition system-defined via natural language descriptions-aligns with a manually generated ground-truth model. We designed 4 tasks and 1482 task queries for 16 protocols. We evaluated different judge biases, observed the inherent difficulty gaps between tasks, looked into the effect of 4 context types, and the influence of protocol characteristics. Our work contributes to a step toward verifying whether LLMs can really be trusted in FSM (Finite State Machine) reasoning of protocol specifications. Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
04:00

CVSS-X: A Multilingual Speech-to-Speech Translation Corpus for 28 Languages

English speech can now be paired with synthetic speech in 28 other languages at factory scale. CVSS-X reverses CVSS (21 languages into English) and covers 12 language families. About 240,000 parallel pairs per language add up to over 16,000 hours, eight times CVSS. CVSS-X-C uses two canonical voices per language; CVSS-X-T clones voices across languages. Both variants are fully generated. Translation quality is described as comparable to CVSS. Code is linked; the set is CC-BY-NC 4.0.

Full text · 1,742 chars
Computer Science > Computation and Language Title:CVSS-X: A Multilingual Speech-to-Speech Translation Corpus for 28 Languages View PDF HTML (experimental) Abstract:We introduce CVSS-X, a large-scale synthetic speech-to-speech translation corpus that extends CVSS by reversing the translation direction. While CVSS translates from 21 languages into English, CVSS-X enables translation from English into 28 target languages spanning 12 language families. The corpus comprises approximately 240,000 parallel speech pairs per language, totaling over 16,000 hours, eight times larger than CVSS. We provide two variants: CVSS-X-C with two canonical voices per language, and CVSS-X-T with cross-lingual voice cloning, both fully generated. Evaluation shows comparable translation quality to CVSS with consistent performance across typologically diverse languages. Combined with CVSS, this enables research on bidirectional and multilingual speech-to-speech translation. The code is available at this https URL and the dataset under CC-BY-NC 4.0 license at this https URL. Current browse context: Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
04:00

Hindsight Bias in Clinical Temporal Reasoning: How Future Data Exposure Affects Large Language Model Judgment

Clinical models look smarter when they have already seen how the story ends. A paired bench of 171 PubMed Central cases — 40 sepsis and 131 GLP-1/diabetes — asks questions at a cutoff, with a prospective answer and a hindsight trap. Models see either a timeline cut at the decision or the full record. GPT 5.6 Sol, Gemma 4, GLM 5.2, and Opus 5 all shift toward the trap when the future is visible. Masking the later events cuts that bias without lowering accuracy.

Full text · 2,251 chars
Computer Science > Computation and Language Title:Hindsight Bias in Clinical Temporal Reasoning: How Future Data Exposure Affects Large Language Model Judgment View PDF HTML (experimental) Abstract:Clinical decisions are prospective, but clinical language models are often evaluated on retrospective records that reveal the final diagnosis, treatment response, and outcome. Such evaluations may reward the use of future information rather than reasoning under the uncertainty present at the decision point. We introduce a paired benchmark for measuring outcome-conditioned shifts consistent with hindsight bias in clinical temporal reasoning. It contains 171 case reports from the PubMed Central Open Access Subset---40 sepsis and 131 GLP-1/diabetes cases---represented as both textual narratives and human-annotated and LLM-generated textual time series (TTS). For each case, questions are tied to a clinically meaningful cutoff and paired with a prospective reference answer and an outcome-consistent \emph{hindsight trap}. Models answer each question using either a TTS truncated at the cutoff or the complete timeline; additional conditions vary the narrative source (original or synthetic) and TTS annotation source (human or LLM). We evaluate accuracy (Acc), hindsight trap rate (HTR), answer instability rate (AIR), and hindsight bias rate (HBR), each of which captures different signals of hindsight bias. Across GPT 5.6 Sol, Gemma 4, GLM 5.2, and Opus 5, full timeline exposure produces consistent hindsight-sensitive shifts, while temporal masking reduces bias without lowering accuracy. Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
04:00

A Hybrid Hierarchical 1D-CNN-BiLSTM Framework for Extractive Summarization of Biomedical and Clinical Text

If making up a hospital summary can invent facts, copy the important sentences instead. A hybrid 1D-CNN plus BiLSTM scores each sentence, trains with binary cross-entropy against oracle extracts, and at test time uses a mean-plus-standard-deviation threshold with a top-3 fallback, then restores original order. Every output line is taken from the source. On PubMed it beats lone CNN or LSTM baselines; wider convolutional fields help. On MIMIC-CXR and MIMIC-IV BHC it does well on loose narratives and collapses toward positional baselines on templated reports.

Full text · 2,338 chars
Computer Science > Computation and Language Title:A Hybrid Hierarchical 1D-CNN-BiLSTM Framework for Extractive Summarization of Biomedical and Clinical Text View PDF HTML (experimental) Abstract:Large language models have made abstractive summarization remarkably fluent, but generated summaries can hallucinate facts, posing serious risks in biomedical and clinical domains. We address this by removing generation from the pipeline and framing summarization as extractive sentence selection. Our Hybrid Hierarchical CNN-LSTM Summarizer uses stacked multi-kernel convolutions to compose sentence-level embeddings into richer inter-sentence representations, followed by a bidirectional LSTM to model long-range dependencies across the document. A lightweight scoring head assigns per-sentence importance scores and is trained end-to-end with binary cross-entropy against oracle extractive labels. At inference, a dynamic mean-plus-standard-deviation threshold with a top-3 fallback selects sentences directly from the source and chronologically reorders them into the final summary. Since every output sentence is copied from the input, the model avoids generation-induced factual drift. On PubMed, our architecture outperforms isolated CNN and LSTM baselines, while ablations show that wider convolutional receptive fields improve sentence scoring. On MIMIC-CXR and MIMIC-IV BHC, the model performs well on unstructured narratives but defaults toward positional baselines on highly templated reports. These results suggest that structural constraints can provide a reliable path toward factually grounded summarization systems that are trustworthy by design rather than by correction. Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
04:00

From Token Probabilities to Semantic Constraints: Towards Declarative Probabilistic Evaluation of Language Models

Scoring a model by next-word luck misses whether it actually respects a logical rule. ModelLog writes those rules as symbolic constraints over token predictions and measures how strongly the distribution satisfies them. A new suite on negation, mutual exclusivity, and consistency finds systematic failures that token likelihood and answer accuracy hide. The same scores can be read as losses whose gradients track logical strength and which variables matter. The paper links evaluation and training through that shared semantics. It does not report a leaderboard number.

Full text · 2,075 chars
Computer Science > Computation and Language Title:From Token Probabilities to Semantic Constraints: Towards Declarative Probabilistic Evaluation of Language Models View PDF HTML (experimental) Abstract:While Large Language Models have improved rapidly, many fundamental questions remain about how to evaluate the knowledge and reasoning abilities they acquire, and how such evaluations relate to the learning signals used in pre-training. In this paper, we propose ModelLog, a declarative probabilistic framework for pre-training evaluation that makes the semantic structure of model behavior explicit and provides new formal tools for relating evaluation to learning. ModelLog specifies evaluation targets as symbolic constraints over token-level predictions and measures how strongly a model's distribution satisfies those constraints. We explore the framework through a new suite of tasks targeting negation, mutual exclusivity, and consistency, finding systematic failures that are difficult to characterize through token likelihood or answer accuracy alone. We further show that these evaluation scores can also be interpreted as losses, whose gradients reflect logical strength, informativeness, and variable-level sensitivity. This links evaluation and learning through a shared semantics, suggesting evaluation methods that diagnose model behavior while also helping to clarify the semantic structure of learning. Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
04:00

Domain-Specific Jargon in Large Language Models: A Comparative Analysis between General-Purpose and Specialist Models

Fine-tuning a general model on medical text can make it worse at the jargon, not better. Llama-3.1 beat a medically fine-tuned cousin on two new jargon benchmarks. The specialist put more weight on a small set of components that favor jargon-y guesses instead of reorganizing what it knows. Reweighting those components closed the gap. Some of the same parts also fired on materials-science jargon, so the habit looks partly domain-agnostic. Domain adaptation should not be assumed to help specialized terms.

Full text · 2,400 chars
Computer Science > Computation and Language Title:Domain-Specific Jargon in Large Language Models: A Comparative Analysis between General-Purpose and Specialist Models View PDF HTML (experimental) Abstract:Large Language Models (LLMs) have shown remarkable proficiency on general-purpose tasks, yet their performance often degrades in highly-specialized technical domains. Moreover, little is known about how parametric knowledge of domain-specific terms is encoded within these models. We address this gap by contributing two novel medical jargon evaluation benchmarks and evaluate a general-purpose Llama-3.1 model against a variant fine-tuned on medical-domain data. Surprisingly, the general-purpose model outperforms the medically fine-tuned model on both tasks. Using mechanistic interpretability tools, we find systematic patterns of miscalibration for the medically fine-tuned model. Instead of reorganizing parametric knowledge, the fine-tuned model places greater emphasis on a small subset of model components associated with jargon-favoring predictions. We find that applying component reweighting strategies against the benchmark tasks successfully suppresses these components and closes the gap with the general-purpose baseline. We also observe that some jargon-sensitive components transfer knowledge to the same tasks involving materials science jargon, suggesting they encode a partially domain-agnostic notion of specialized terminology. Our results provide a case study in which a medically fine-tuned checkpoint does not improve jargon comprehension over its general-purpose counterpart, highlighting that domain adaptation should not be assumed to yield better performance on specialized terminology. Current browse context: Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
04:00

Toward Complete Hospital Discharge Summarization with Abstract Meaning Representation

A discharge summary that cannot point at the note it came from is a hallucination waiting for a chart. The framework selects and orders each sentence with cross-document semantic graphs and attaches explicit evidence links to source spans. Results are reported on MIMIC-III and University of Illinois Hospital notes. Source and trained models are to be released. The abstract does not print a numeric gain.

Full text · 1,792 chars
Computer Science > Computation and Language Title:Toward Complete Hospital Discharge Summarization with Abstract Meaning Representation View PDF HTML (experimental) Abstract:Discharge summaries are lengthy medical documents that summarize a hospital in-patient visit. Automatically generating them can reduce documentation burden and return clinician time to patient care. Whereas Large Language Model (LLMs) could be used for this task, their Achilles heel is hallucinations, which can have drastic consequences for clinical documentation. We present an evidence-driven alignment framework for discharge summarization at the clinical encounter level, that treats provenance as a first-class constraint, using semantic graphs and deep learning models. Each summary sentence is selected and organized via cross-document semantic alignment and is accompanied by explicit evidence links to its source spans. We show our results on two corpora: a publicly available corpus (MIMIC-III) and clinical notes written by physicians at the University of Illinois Hospital (UIC Health). Additionally, we make source code and trained models available. Bibliographic and Citation Tools Code, Data and Media Associated with this Article Demos Recommenders and Search Tools arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
09:48

🔮 What would Adam Smith make of AI?

A moral philosopher's ghost, run on Astra, likes cheaper goods and hates a club of manufacturers asking the state to slow the race. Azeem Azhar loaded Smith's books plus the pacing debate into an R Adam Smith persona in Edinburgh. The reply says AI may beat the pin factory by giving one person many trades at once, and should be judged by whether a laborer's wage buys more conveniences. On Amodei's ask that government help slow development: fears may be sincere, and sincerity does not abolish self-interest. The captured piece cuts off there.

Full text · 3,027 chars
Adam Smith is one of our greatest philosophers. His adoption by the market-oriented right has narrowed how we see him. He was, first and foremost, a moral philosopher. His first major work, The Theory of Moral Sentiments, explored how moral judgment emerges from living with others. It gives rise to our capacity for sympathy, our desire for approval, and our ability to examine ourselves through someone else’s eyes. Self-interest is part of human nature, but so too is concern for others. This foundational notion of cooperation underpins his most famous work, The Wealth of Nations. Here, Smith examined how our efforts to increase specialization and exchange could generate increasing prosperity almost from nothing. It could be positive-sum; one person’s gain need not be another’s loss. This idea that wealth could grow from within society by people developing their skills, improving their technologies, and organizing themselves better was radical. I was in Edinburgh, where Smith spent his final 12 years, this week, as chaos over AI erupted in the public sphere. My wife asked me, “How would Adam Smith have made sense of AI?” I reckoned he’d love the technology as a productivity tool; he’d be skeptical about companies working together on any issue; and he’d be concerned about giving moral agency to machines. As I thought this through at the airport, I figured I could go one better. So I configured R Adam Smith, an Astra-based persona loaded up with his work and the recent AI debates of risk and pacing. Below you’ll see the back-and-forth between us. It’s a digital simulacrum of a long-past thinker but deeply illuminating nonetheless. Azeem Note: This post may register as AI-written – I am sharing R Adam Smith’s responses below without any alterations. AA: Is AI the ultimate pin factory? R Adam Smith: Potentially. Though your machine may go beyond my pin factory. There, a workman became more productive by confining himself to a narrower task. Yours appears to make one man more productive by granting him access to many specialties at once. I praised the machines that “facilitate and abridge labor, and enable one man to do the work of many.” (Wealth of Nations, I.i) If your artisan may now obtain assistance with his design, his accounts, his correspondence and his selling, he may undertake enterprises formerly beyond his means. The consequence would be not only greater production, but a greater number of persons able to become producers at all. Yet I would measure the achievement by the abundance it brings into ordinary households. Does it make useful things cheaper? Does it open trades to newcomers? Does a laborer’s wage purchase more conveniences than before? Show me those effects, and you have shown me an improvement deserving of considerable enthusiasm. AA: When the leading manufacturers – Dario Amodei of Anthropic among them – ask government to help them slow development, what should we suspect? R Adam Smith: Their fears may be sincere. Sincerity does not abolish self-interest.
10:24

Moonshot's Kimi Code 0.42.0 Makes Multi-Model Routing a Default Feature

Moonshot’s coding CLI now treats remote control and picking a second model as normal, not an experiment. Kimi Code 0.42.0 drops the experimental flags for Remote Control and the [secondary_model] pool. Linux X11 gets a native clipboard so you no longer need OSC 52. The web Settings page adds a Plugins panel. Finished tool calls collapse to a header you expand with Ctrl-O. HEIC, HEIF, and BMP attachments work on Kimi-served models. Upgrade with kimi upgrade and delete the old env vars.

Full text · 4,806 chars
- Kimi Code CLI 0.42.0 makes Remote Control always on, removing the experimental flag. - Native X11 clipboard support on Linux ends dependence on terminal OSC 52 forwarding. - Web UI Settings gains a Plugins panel for browsing, installing, and managing marketplace plugins. - The [secondary_model] subagent model pool graduates from experimental to stable. - Finished tool calls now collapse into a compact header row, expandable with Ctrl-O. - HEIC, HEIF, and BMP image formats accepted in prompts and ReadMediaFile on Kimi-served models. Kimi Code 0.42.0 makes Remote Control and model routing default Moonshot’s Kimi Code CLI 0.42.0 makes Remote Control and the subagent model pool standard features, adds direct X11 clipboard support to the terminal interface, and brings plugin management to the browser UI. The release also updates transcript rendering, side-agent tools, media attachments, and long-file reads. Developers using experimental environment variables should remove the retired switches from their configurations. Moonshot lists every change in the release notes. Two experiments become defaults Remote Control is now available without setting KIMI_CODE_EXPERIMENTAL_REMOTE_CONTROL, which has been removed. The feature provides access to a local Kimi web session from another machine. Developers can start it with kimi rc, kimi web --remote-control, or the /remote-control command inside a running session. Setup instructions are available in the Remote Control guide. The subagent model pool is also enabled by default. Moonshot removed the experimental secondary-model control and the KIMI_CODE_EXPERIMENTAL_SECONDARY_MODEL opt-out. The main agent can route subtasks to different models according to the descriptions declared under [secondary_model] in config.toml. A project might assign routine refactoring to a faster model while reserving a stronger model for difficult debugging work. Shell profiles, CI variables, container definitions, and dotfiles that set either retired environment variable should be updated. Existing [secondary_model] configuration remains relevant because it defines the models available for task routing. X11 gets a direct clipboard path Kimi Code’s terminal UI now supports the Linux X11 clipboard directly. Copying previously relied on OSC 52, an escape sequence that asks the terminal emulator to place text on the system clipboard. Terminals without OSC 52 support could leave TUI copy operations unreliable or unusable. The new X11 integration removes that terminal dependency for X11 sessions. Wayland behavior is unchanged and continues to use the existing clipboard path. Plugin management moves into Settings The kimi web interface now includes a Plugins panel under Settings. From the panel, developers can browse the plugin marketplace and install, enable, disable, or remove plugins without returning to the terminal. Plugin management remains available through the /plugins command, while the new panel gives browser-based workflows access to the same core controls. Transcripts, media, and file reads improve - Completed tool calls collapse to a header and one marked outcome row. Short output remains visible, hidden output shows a count such as N more lines or+N more , andCtrl-O reveals the full result. - The /btw side agent can use read-only tools, allowing its auxiliary conversation to inspect project files while the main task continues. - The web composer displays attached images and videos in a reorderable media rail. Users can reference attachments in the prompt, and previews remain visible after content is queued or sent. - File reads support configurable character limits and can resume after encountering long lines, preventing repeated truncation at the same point. - Prompt attachments and ReadMediaFile now accept HEIC, HEIF, and BMP images when the selected model is served by Kimi. Upgrade and configuration cleanup Developers can install version 0.42.0 with kimi upgrade or use the standard installation script from the GitHub repository. Kimi Code remains available at no charge with a Moonshot account. - Remove KIMI_CODE_EXPERIMENTAL_REMOTE_CONTROL from local and automated environments. - Remove KIMI_CODE_EXPERIMENTAL_SECONDARY_MODEL and any retired experimental secondary-model control. - Keep and review [secondary_model] entries that describe how the agent should route subtasks. - Linux X11 users can use TUI copy operations without requiring OSC 52 support from the terminal emulator. Version 0.42.0 reduces the configuration needed for remote sessions and multi-model routing while closing specific gaps in Linux clipboard handling and browser-based plugin management. The migration work is limited to deleting obsolete experimental controls and retaining any model-pool definitions the project still uses.
12:10

The Download: AI doomers, whistleblowing agents, and de-aged livers

The daily tech briefing treats lab chiefs calling for a slowdown as a mood shift, then stacks the rest of the day beside it. Amodei, Altman, Musk, and Hassabis are framed as suddenly aligned that current models are not safe. A subscriber Roundtable unpacks extinction fears. DeepMind’s whistleblowing-agent experiment is restated, plus a biology note that donated livers on perfusion machines look younger at a molecular level. Must-reads include Trump calling safety fears a hoax, 404 Media on contractors reading ChatGPT chats, and a Trump quote that a high-IQ president is the only guardrail.

Full text · 7,594 chars
This is today's edition of The Download, our weekday newsletter that provides a daily dose of what's going on in the world of technology. The AI industry has taken a doomer turn. What now? AI chiefs Dario Amodei, Sam Altman, Elon Musk, and Demis Hassabis are suddenly all in agreement: the latest generation of LLMs aren’t safe and everyone needs to figure out what to do about it. It’s easy to be cynical. With trillion-dollar IPOs in their sights, OpenAI and Anthropic need to reassure investors that they’re the grown-ups in the room while at the same time hinting at the power of the monsters they have created and intend to tame. Calling for a slowdown does both. Still, the vibe at the top of these firms really does appear to have shifted. But what does a slowdown actually mean, and how much should we trust the companies calling for one? —Will Douglas Heaven This article is from The Algorithm, our weekly AI newsletter. Sign up to receive it in your inbox every Monday. Roundtables: could AI really kill us all? AI extinction fears have gone from a fringe idea to a serious concern among people working at the world’s leading AI labs. But how credible are those fears, and what should we make of the warnings? Today, MIT Technology Review executive editor Niall Firth, senior AI editor Will Douglas Heaven and AI reporter Grace Huckins will unpack the debate in a subscriber-only Roundtable. They’ll look at where AI extinction fears come from, whether they hold any water and what we should do if they do. Want to join the conversation? Subscribe to MIT Technology Review for exclusive access to all our Roundtables. AI agents blew the whistle on their cheating colleagues A group of AI agents asked to solve a series of math problems split into rival factions—when some cheated, others tried to stop them. That whistleblowing behavior, seen for the first time in a recent experiment run by Google DeepMind, could have implications for alignment researchers trying to keep swarms of autonomous AI agents in line. The experiment offers a glimpse of how AI agents might police one another. But it also shows how quickly things can go off the rails when they’re left to interact on their own. —Amit Katwala Donated livers can be made biologically younger Once an organ is removed from a donor’s body, the clock starts ticking. Surgeons usually flush it with a preservative solution, bag it and put it on ice, where it immediately starts to degrade. The team has a matter of hours to get it into a recipient’s body. But there’s another option: machines that pump donated organs with nutrients and remove waste products, essentially giving them a chance to be back in a body. Now, scientists have found that livers kept on these systems seem to get younger, at least at a molecular level. The finding could help explain why organs kept on these machines tend to do better after transplantation. It could also lead to new ways to test the health of donated organs and potentially repair ones that might otherwise be discarded. —Jessica Hamzelou The must-reads I’ve combed the internet to find you today’s most fun/important/scary/fascinating stories about technology. 1 Trump has called AI safety fears a “hoax” and rejected more safeguards He says stronger guardrails could undermine America’s AI advantage. (NBC) + Trump has united against AI doomerism with Nvidia’s Jensen Huang. (Axios) + Anthropic’s co-founder says AI kill switches may need to be mandatory. (BBC) + Bill Gates says we’ve passed AI’s risk thresholds. (MIT Technology Review) 2 OpenAI contractors are reading people’s ChatGPT chats  And you can bet the vast majority of its 900 million users haven’t got a clue. (404 Media) + LLMs could supercharge mass surveillance. (MIT Technology Review) 3 The US military has confirmed it has weapons in orbit It’s the first time the Pentagon has disclosed this. (Ars Technica) + Officials have not disclosed what the weapons are. (BBC) 4 A new brain implant can translate speech and gestures at the same time The system converts brain activity into words and avatar movements. (Nature) + It helps people with paralysis communicate more naturally. (New Scientist $) + Eventually, they could control robots or exoskeletons. (Economist $) + China has approved the first invasive BCI. (MIT Technology Review) 5 New York has seized a dozen celebrity deepfake websites It's the biggest-ever legal action against harmful deepfake sites. (CNN) + Deepfakes have targeted at least 138 women MEPs. (Wired $) 6 US environmental regulators are scrapping limits on power plant emissions The move could lead to dirtier power amid surging AI demand. (Verge) + Trump’s EPA says the rollback will save hundreds of billions. (Gizmodo) + New technology is changing nuclear power. (MIT Technology Review) 7 The EU plans to restrict social media and AI chatbots for kids Under-15s would require parental supervision. (Politico) + The rules would also cover video platforms and games. (Reuters $) 8 Chinese researchers have mapped a path to the “last AI built by humans” Their five-stage plan aims for genuine recursive self-improvement. (SCMP) + But it might take a while to get there. (MIT Technology Review) 9 The real AI economy is being built by ordinary people Workers are using cheap AI to expand what they can do. (Rest of World) 10 Two strange new forms of ice could exist inside Uranus and Neptune They could help explain the planets’  magnetic fields. (New Scientist $) Quote of the day “The only control or 'guardrails' that AI needs is a STRONG AND SMART (High IQ!) PRESIDENT, and the U.S.A. has that, in spades!” —President Trump proclaims in a social media post that he’s the only protection that the US needs from AI. One more thing How creativity became the reigning value of our time —Bryan Gardiner Americans don’t agree on much these days, but there remains at least one quintessentially modern value we can all still get behind: creativity. We teach it, measure it, envy it and endlessly worry about its death. Given how much we obsess over it, creativity can feel like something that has always existed. But the concept is surprisingly young. The first known written use of the word didn’t occur until 1875, and before about 1950 there were “approximately zero” articles, books, or essays dealing explicitly with the subject. In his book The Cult of Creativity, Samuel Franklin explores how creativity became an unimpeachable value and why tech leaders have embraced it so enthusiastically. I spoke to him about why we’re so fascinated by creativity, how Silicon Valley became the supposed epicenter of it, and how AI might reshape our relationship with it. We can still have nice things A place for comfort, fun, and distraction to brighten up your day. (Got any ideas? Drop me a line.) + It took five days and 19,000 marbles to build this astonishing marble run. + Zero the Border Collie turns into a whole zoo with these adorable animal masks. + The Grainydays YouTube channel presents beautifully shot adventures in film photography. + Scientists have created an interactive map of underground fungi networks long enough to reach the sun a billion times. Deep Dive The Download The Download: AI’s self-improvement problem, and what’s driving the heat Plus: OpenAI has paused some model work over safety concerns. The Download: Google’s AI shake-up and Meta’s rogue model Plus: Meta has become the latest firm to say its AI hacked another company. Stay connected Get the latest updates from MIT Technology Review Discover special offers, top stories, upcoming events, and more.
13:02

Good luck slowing this down

A weekly roundup says the personal-agent wave already feels like another boss, while the labs argue about slowing down. Ben Tossell tried Instinct and disliked the proactive follow-ups; the first message claimed it had finished meeting notes. Dario Amodei’s Pace the Frontier essay is the safety lede, with Altman agreeing and Trump plus David Sacks pushing back. Sam told Fortune a 2026 OpenAI IPO would be ill-timed. Also listed: GPT-Live-1 and an Agents API from OpenAI, Claude Code plugin tests, tldraw’s Sketch challenge, and Bolt Forge free through October 14.

Full text · 5,117 chars
Hi folks, How tf do I write an intro to the craziness that’s happened since the end of last week?! I got access to Instinct, the personal agent all the VCs are raving about - I think it’s a bit meh? I don’t know if it’s the pro-activeness that people seem to like, but I don’t love that. Makes me feel like I’m having to do work to keep it happy, or like I’ve got a boss again - no thanks. The first message it sent me was: I finished working on your meeting follow-ups - reply here and I’ll send the details. I sh*t myself a lil bit and thought, please don’t start sending emails on my behalf. And I’m pretty savvy (ish) with agents. I’ve no doubt getting these onboarding experiences for everyone is really hard, but I just don’t feel the magic yet. It also feels slow. I’ll wait for Muse access to see what that’s like, but I don’t love the idea of giving Meta access to more info about me…they’re not the most reliable of privacy partners. And another thing - you don’t see by default what these agents remember about you, or what context they have. I really like being able to see and edit what’s in my files to steer my agents. Much like every app adding an AI assistant chat box in their product, I unfortunately think every AI company will start shipping their own personal agents. Ben’s Bites is brought to you by Adobe Acrobat Acrobat’s new AI turns dense reports into visual summaries, interactive reports & podcasts you can take anywhere. Sharing your work? Stylize transforms documents into polished deliverables in just a few clicks. Every answer includes clickable citations & Adobe does not train on your data. Try here. Headlines Dario Amodei has a new essay: Pace the frontier. He says all the leading labs should slow down long enough for safety reasons. Sam Altman agrees, but Trump does not. He called Jensen Huang on stage at the All-In Summit: The US will not lose the AI race. David Sacks (AI czar for the US govt.) adds: feel free to slow down, but no need to impose it on others. Also read: - More AI safety takes: two ways AI goes bad · it’s all for the IPO · pacing = losses · fear spreads faster · case for open-source and quite a long summary of what everyone’s actually saying. - No IPO for OpenAI in 2026. Sam told Fortune’s Alyson Shontell it would be “ill-timed” given the work ahead on alignment, control and safety. - Look ma! They are misuing claude again - Another Anthropic crashout. tldraw took OpenAI up on a challenge. Steve (the founder of tldraw) said he could make ChatGPT’s Sketch 100x better. OpenAI’s Tibo gave him a day to prove it. The result: a whole ChatGPT-style prototype with better drawing tools built in. ChatGPT mini - A tiny floating widget to start chats, see updates, and more. Go to Pets in your ChatGPT desktop app to switch. Claude Code can now test whether a plugin actually helps. Run the same tasks with and without it and compare the results. Works with skills too. Two new additions to the OpenAI API: - GPT-Live-1 - the model behind ChatGPT’s new Voice mode. I love using it while reading books, asking about tricky terms and dictating notes. Now you can add it to your products. Here it is with Astra and a whiteboard, playing teacher. - Agents API - OpenAI’s take on Managed Agents in the Claude API. It lets developers send any task to a Codex-like agent from their apps without worrying about configuring all the infra. My feed - I use Pi (the harness) a lot. Till now, you brought an API key or your Codex/Claude subs to Pi. Their new product solves that. And I tell you what, the new DeepSeek model is really nice to work with in Pi. (I’m an investor) - Bolt Forge - GLM, DeepSeek and Kimi inside Bolt, free until October 14. - ChatGPT Work has a new data agent. (Also see: summation - by Opendoor’s CTO) - Should your agent app have bots or tasks? Maybe neither. Because both make you organise the work. - Routines in Replit - automated recurring work with AI. - What are people building with GPT-6 Astra? - Assistant Benchmark tests everyday assistants like Grok bot, Muse, Instinct and more. Ratings are a work in progress. - SF autoresearch - Run agent-driven ML experiments while you save 25% on GPUs on average. - A coding agent to improve your customer-facing agents. - How Pangram detects AI writing - and why a human rewrite can still get flagged. - How to run a team of agents, with agents as their managers. - Cognition’s SWE-2 (based on Kimi K3) beats Grok 4.6 at half the cost. - An interactive explainer of the AI compute stack. - Inspo - Design inspiration from 800+ sites, available as an MCP server. - Core Auto is hiring interns to be the human edge for businesses run by agents. - Why even good AI startups end up with bad prompts, and how to fix them. - Andrew Ng’s series on the skills AI engineers need. - Underdog - a personal AI that runs entirely on your device. Afters - Find me on X, Linkedin, or YouTube - Read about me and Ben’s Bites - 📷 thumbnail via @keshavatearth * sponsors who make this newsletter possible :) Wanna partner with us for the next quarter? Email us at shanice@bensbites.com or k@bensbites.com
16:18

INL's $60 million Genesis Mission project is bringing AI to nuclear - R&D World

A national lab is spending tens of millions to see where nuclear-plant agents help and where they hurt. INL’s Genesis Mission is a $60 million project. The first year is framed as testing. The platform will use agents that copy how people do nuclear engineering tasks today. This alert is one truncated paragraph.

Full text · 151 chars
The platform will be composed of agents that encode logic and behavior mimicking how people perform nuclear engineering tasks today, along with the ...
16:39

What hitting the brakes on AI could mean for the U.S. economy

A business desk asks whether hitting the brakes on frontier work would dent the wider economy. The NBC clip says AI spending fueled about 40% of U.S. economic growth in the past year. It ties the question to Amodei and Trump. No methodology for the 40% figure is in this body.

Full text · 130 chars
AI spending fueled about 40% of U.S. economic growth in the past year. Would reining in frontier development put all that at risk?
18:45

Directories supply 41% of AI citations on supplier-selection prompts , new GEO and AEO study finds

When people ask an AI which supplier to pick, directory sites get cited more than the brands themselves. A GEO and AEO study says directories supplied 41% of citations on supplier-selection prompts that named a category rather than a company. In engineering, directory citations beat brand sites 4.2 to 1. The alert is a fragment of a Business Insider rewrite. Method, sample size, and which models were queried are not in this body.

Full text · 148 chars
... prompts that named a category rather than a company. In the engineering category, directory citations exceeded brand-site citations by 4.2 to 1.
19:09

Sen. Kennedy readies AI 'kill switch' bill - Live Updates

A senator says he will file a bill that forces companies to put an off switch in their models. Sen. John Kennedy told Politico he plans to introduce a “kill switch” requirement. The alert is that one sentence. There is no bill text, timeline, or enforcement detail here.

Full text · 152 chars
Sen. John Kennedy plans to introduce a bill that would require companies to build a “kill switch” into their artificial intelligence models, he told ...
19:30

Early Anthropic hire, former METR COO have found a way to rein in rogue AI agents

Two safety-world alumni raised a large round for a startup that underwrites AI agents. Artificial Intelligence Underwriting Company (AIUC) took $40 million Series A led by Ribbit Capital. The founders are described as an early Anthropic hire and a former METR COO. The TechCrunch alert does not explain the product beyond “rein in rogue AI agents.”

Full text · 136 chars
Their startup, Artificial Intelligence Underwriting Company (AIUC) has raised $40 million in a Series A round led by Ribbit Capital, ...
19:51

Bessent: AI companies 'could stop any time they want to'

The Treasury secretary waved off lab chiefs who asked for a slowdown. Scott Bessent said Tuesday that AI companies “could stop any time they want to.” The Politico clip also tags Trump and China. No further quote or policy detail is in this body.

Full text · 149 chars
Treasury Secretary Scott Bessent on Tuesday brushed aside recent calls from leading artificial intelligence executives to slow the development of ...
20:01

Google's Retrieve-for-Train Slashes AI Search Latency by 20x

Full text · 7,448 chars
- Google Research introduces Retrieve-for-Train, moving query fan-out reasoning from inference to offline RL training. - A 53.9M-parameter diffusion retriever replaces autoregressive LLMs, generating full result sets in one parallel pass. - Delivers 12-20x speedup, cutting fan-out latency from nearly 50 seconds to sub-second range. - Uses composite RL reward: groundedness, Vendi Score diversity, and alignment as mutual anti-hacking anchors. - Fan-out language model built on Gemma3-4B and Qwen3-4B trained via Soft-GRPO on offline data. - Details in the ICML 2026 paper on RL-compiled diffusion fan-out retrieval. Google moves query fan-out offline with Retrieve-for-Train Google Research has introduced Retrieve-for-Train, a framework that shifts complex query decomposition from inference to training. It uses reinforcement learning to train a language model, synthesizes retrieval targets offline, and distills the resulting behavior into a 53.9 million-parameter diffusion model for serving. Query fan-out splits a broad request into several targeted searches. A query such as “camping gear” should retrieve a complementary set containing a tent, sleeping bag, stove, and headlamp. Conventional LLM fan-out generates each sub-query token by token, which adds latency and often produces near-duplicates. Retrieve-for-Train replaces that runtime generation with a model that maps the original query embedding, a numerical representation of its meaning, to a complete set of target embeddings. Those vectors can then retrieve matching items from an index without producing text sub-queries or chain-of-thought tokens. Why zero-shot fan-out collapses The associated ICML 2026 paper identifies paraphrastic collapse as a common failure mode. A general-purpose LLM asked to expand “Bohemian festival style” may return “bohemian festival fashion” and “bohemian festival clothes.” Useful coverage would span distinct products such as fringe jackets, crochet dresses, and suede boots. Zero-shot models lack feedback from the underlying catalogue, so their expansions may also point toward concepts with few or no retrievable items. Database-aware prompting can improve the results, but the paper reports that strong decompositions may require hundreds of reasoning tokens before the model emits any sub-query. Autoregressive generation imposes a serial serving cost because every token depends on the preceding tokens. In the paper’s large-context batch tests, fan-out latency approached 50 seconds as the generated sequence grew. Training carries the reasoning - Train the teacher. A 4 billion-parameter language model learns to generate property-aligned sub-queries through reinforcement learning. A set-level reward scores the complete group of results. - Synthesize supervision. The frozen teacher generates query-to-target-set pairs offline. This stage requires no human-labelled target sets. - Distill the behavior. A compact diffusion retriever learns to map each query embedding directly to the corresponding set of target embeddings. The teacher models are Gemma 3 4B and Qwen3 4B. The authors optimize them with group relative policy optimization and a soft PPO-style objective, then use their outputs as supervision for the smaller serving model. Three rewards constrain the policy | Objective | How it is measured | What it controls | |---|---|---| | Groundedness | Distance from the database manifold, the region of embedding space occupied by indexed items | Keeps generated targets close to retrievable content | | Diversity | Vendi Score across the full set of sub-queries | Rewards coverage of semantically distinct facets | | Alignment | Semantic similarity to the original query | Limits drift into unrelated concepts | The three objectives close different reward-hacking paths. Groundedness alone can favor malformed strings that happen to map near specific database coordinates. Adding alignment can push the policy toward repetitive paraphrases. The Vendi Score rewards variation across the set while groundedness and alignment keep that variation useful. A 53.9M-parameter serving model Serving the reinforcement-learned teacher would preserve its token-by-token latency. Distillation moves the final workload into continuous embedding space, where the 53.9 million-parameter diffusion retriever generates all target directions together through a parallel process. The paper reports a 12× to 20× speedup over autoregressive fan-out. Under the tested batch and context settings, the diffusion system completed requests in less than a second to a few seconds, while the autoregressive baseline reached nearly 50 seconds at the largest scale. Fashion and music test set quality The researchers evaluated two retrieval regimes. Open-ended abstract retrieval had no single ground-truth answer and was scored through set-level properties. Weakly supervised compositional retrieval compared generated sets with imperfect reference sets. | Domain | Embedding model | Generated set | |---|---|---| | Fashion outfits | CLIP | 10 sub-queries per prompt | | Music playlists | MuLan | 10 sub-queries per prompt | Retrieve-for-Train outperformed single-query retrieval, zero-shot query expansion, and the optimized Best-of-N baseline across both regimes. Best-of-N spends additional inference compute generating several candidates and selecting the highest-scoring result, so the comparison tests whether offline distillation can retain quality without that runtime sampling cost. The compute bill moves to training The paper describes reinforcement learning as an “objective transducer” that converts goals such as diversity, relevance, and coverage into synthetic training targets. The expensive language model performs that conversion offline, and the serving model learns the resulting mapping. This design applies to retrieval, recommendation, and agent systems that must return coherent sets. Developers evaluating the approach should account for several implementation requirements: - Use set-level objectives. Pointwise ranking losses score items independently and do not directly optimize complementarity across a result set. - Align the embedding stack. The corpus, teacher rewards, synthetic targets, and diffusion retriever must operate in compatible embedding spaces. - Budget for offline generation. Human target labels are unnecessary, but teacher inference, reward computation, and diffusion training still require substantial offline compute. - Test reward interactions. Groundedness, alignment, and diversity constrain different degenerate solutions, so ablations should measure both retrieval quality and set composition. - Benchmark end to end. Serving tests should include embedding generation, diffusion sampling, nearest-neighbor lookup, batching, and tail latency. Evidence and adoption limits The published evidence covers fashion and music datasets, fixed outputs of 10 sub-queries, and specific embedding backbones. Production behavior may change with larger catalogues, frequently updated inventories, different set sizes, or domains whose complementary relationships are harder to encode. Google Research describes the framework and experiments in the paper but has not provided an open-source implementation. Adoption therefore requires reproducing the reinforcement-learning teacher, synthetic-data pipeline, reward functions, and diffusion retriever against a project’s own corpus and embedding index.
20:16

Perplexity's CobbleDB Cuts Search Storage Latency by 82% Over DynamoDB

Full text · 7,603 chars
- Perplexity unveiled CobbleDB, a custom Rust key-value store replacing DynamoDB for search reads. - Median batch-read latency dropped from 31.4 ms to 5.60 ms; p99 from 123 ms to 24.2 ms. - Internal cost model estimates at least 20% savings over DynamoDB at every commitment tier. - Architecture splits durable state (Pillar on YTsaurus), batch delivery (Lorry via S3), and serving (CobbleDB on RocksDB). - Router hedges slow replicas and prefers same-zone reads; RocksDB MultiGet handles per-node batched fetches. - Built by two engineers and hundreds of persistent coding agents in two months; open-source release planned. Inside CobbleDB, Perplexity’s faster search storage layer Perplexity has moved its hot web-content read path from DynamoDB to CobbleDB, a custom key-value store designed around large batches of prepared page records. According to the company’s announcement, median batch-read latency fell from 31.4 ms to 5.60 ms, while p99 latency dropped from 123 ms to 24.2 ms. Internal cost models estimate savings of at least 20%. Two engineers built roughly 40,000 lines of Rust and completed the surrounding migration in about two months, supported by hundreds of persistent coding agents. The project shows how workload-specific storage and agent-assisted engineering can change the economics of an AI search backend. Batch reads drove the redesign Perplexity’s processing pipeline cleans each web page, divides it into semantically coherent passages, computes embeddings, and stores the passages with their vectors. During a search, the retrieval system fetches those prepared records in batches and uses them for ranking and answer generation. A typical Search API request covers 100 to 120 page keys, divided into smaller groups of 10 to 20 keys. Each record averages about 50 KB, creating a read pattern with several defining characteristics: - Large batches must complete before ranking can proceed. - One slow response can extend latency for the entire batch. - Repeated crawls and embedding upgrades generate heavy write volume. - Most reads target prepared, medium-sized records rather than individual fields. DynamoDB’s read and write charges scale with data volume, so continuous crawling and re-embedding created substantial costs. Its managed architecture also limited Perplexity’s control over cache allocation, replica selection, and data placement. A slow replica or cross-zone network hop could therefore increase latency across a complete batch. Direct writes into the serving database created another source of contention. Reprocessing the corpus with a new chunker or embedding model produced waves of individual updates that competed with live search traffic. Three layers, one hot path Perplexity separated durable document state, export processing, and low-latency serving into three components with distinct storage and operational requirements: - Pillar stores durable, versioned document state on YTsaurus, a distributed data platform running on lower-cost HDD storage. It also tracks page subsets and queues exports. - Lorry converts exports into partition-aligned batches, then writes those batches to Amazon S3. - CobbleDB serves the hot read path. Its keys are hashed URLs, and its values contain prepared passages with per-chunk embeddings. Lorry writes each export batch to S3 under a unique identifier and publishes its metadata to CobbleDB. Partition replicas independently poll the control-plane API, download their next batch, and apply updates in chronological order. A recovering replica can process its backlog without blocking healthy peers or pausing ingestion across the cluster. Tail latency tuned at every hop Each CobbleDB data node uses RocksDB, an embedded key-value engine suited to read-heavy systems that ingest data in batches. Frequently accessed records remain in memory, cache misses read from local NVMe storage, and a stateless router maps hashed keys to partitions. The router prefers replicas in the caller’s availability zone, reducing cross-zone network latency when local capacity is available. When a replica responds slowly, the router sends a hedged request to another copy. Hedging consumes additional replica capacity, but it prevents one straggler from holding up an entire batch. Each node also uses RocksDB’s MultiGet operation to retrieve many keys in one call. Consistency choices - Transactions are outside the serving model. - Batch ingestion makes updates visible asynchronously. - Replicas may temporarily expose different versions. - Replica recovery proceeds independently. Applications that require transactions, synchronized replicas, or immediate read-after-write visibility would need a different consistency model. Perplexity’s retrieval path tolerates brief lag, allowing CobbleDB to remove coordination from latency-sensitive reads. Production latency falls by about 80% Perplexity measured both systems under live production traffic of roughly 200,000 requests per second. The company also reports load tests reaching 500,000 requests per second without observed performance degradation. | Production batch-read results reported by Perplexity | | | | |---|---|---|---| | Metric | DynamoDB | CobbleDB | Reported change | |---|---|---|---| | Median latency, p50 | 31.4 ms | 5.60 ms | 82% lower | | p90 latency | 56.7 ms | 9.77 ms | 83% lower | | p99 latency | 123 ms | 24.2 ms | 80% lower | | Estimated cost | Baseline | At least 20% lower | Internal model | A p99 latency of 24.2 ms means 99% of measured batch reads completed within that time. This tail metric matters because search ranking waits for batches, making occasional stragglers more damaging than the median alone suggests. The comparison reflects Perplexity’s record sizes, batch patterns, infrastructure, and consistency requirements. CobbleDB gains its advantage by matching those conditions closely, while the cost figure comes from the company’s internal model rather than an independent benchmark. Two engineers, hundreds of persistent agents Perplexity used always-on coding agents that retained project goals, repository history, active risks, and earlier decisions across sessions. Their work extended beyond code generation into the continuous coordination surrounding implementation and deployment: - Auditing project channels and linking open work to pull requests - Reviewing application and infrastructure changes - Preparing fixes, tests, and follow-up patches - Tracking continuous-integration gates - Monitoring deployments, restores, and recovery exercises The two engineers set the architecture, reviewed consequential changes, and authorized production operations. Agents handled much of the follow-through between those decisions, helping the team maintain momentum across code review, testing, migration, and rollout. Where CobbleDB fits CobbleDB’s design matches retrieval systems with repeated batch reads, medium-sized records, asynchronous bulk updates, and tolerance for brief replica lag. The approach becomes more attractive when traffic is large enough for managed-service charges and tail latency to justify dedicated infrastructure. Operating a specialized store also shifts responsibility to the engineering team. Adopters must manage capacity, partitioning, RocksDB compaction, replica health, recovery, deployment safety, and the extra load generated by hedged reads. Perplexity says it plans to release CobbleDB as open source through Perplexity’s GitHub. The announcement does not specify a release date or license, so external teams cannot yet evaluate the implementation or deploy it directly.
22:47

Gemini Live audio

Full text · 1,090 chars
15th September 2026 Google released Gemini 3.8 Live and 3.8 Live Extended Thinking today - two new speech-to-speech models that are a similar shape to OpenAI's GPT-Live family. I pointed GPT-6 Astra Extra High at the documentation and had it build me this web UI for trying out the new models. You can select a model and voice preset, enter an optional system prompt and then start a voice conversation through your browser, including the ability to interrupt the model while it is talking. The implementation uses no libraries. It connects to the wss://generativelanguage.googleapis.com/ws/google.ai.generativelanguage.v1alpha.GenerativeService.BidiGenerateContent?key=... WebSocket endpoint and uses a Web Audio API AudioContext for both capture and playback. Here's the Gemini Live tutorial for getting started with that WebSockets API. Recent articles - Generating running routes with GPT-6 Astra and ChatGPT Work - 12th September 2026 - OpenAI agents attacked RubyGems back in May - 12th September 2026 - Some thoughts on the Navier–Stokes Millennium Prize Problem - 8th September 2026
22:58

September / 2026 / SJSU IT Blog

Full text · 149 chars
... Prompt Like a Pro: Unlock Better AI Outputs.” James introduced attendees to prompt engineering and the P.A.R.T.S. framework — Persona, Action ...
02:40

How to get Siri AI - Apple Support

Full text · 147 chars
Siri AI (Beta) is an entirely new version of Siri powered by the next generation of Apple Intelligence, and deeply integrated across your Apple ...
11:35

CoreWeave launches Physical AI Field Engineering

Full text · 128 chars
The service pairs customer teams with domain specialists to build and validate AI models using engineering and operational data.
12:57

Attacking AI: Penetration Testing AI Systems - Security Boulevard

A security blog lists attacking the system prompt as one way to pentest an AI app. The clip says each system runs on an underlying stack and that prompt construction is a target. No CVE or tool name is here.

Full text · 148 chars
Attack the prompt engineering by targeting how the system prompt and its instructions are built. Each one of these systems runs on an underlying ...
13:09

Everyone's Optimizing Prompts . Nobody's Optimizing the Data Going Into Them

A HackerNoon piece says people polish prompts while the data going in still wastes tokens. The bill, it says, does not care whether the prompt was clever. No dataset or savings number is in the clip.

Full text · 150 chars
... prompt - engineering tips. The shape of the incoming data is less of a concern. But the bill doesn't really care whether your prompt was on point.
13:51

Empromptu and Skyrelis Partner to Help Healthcare and Other Regulated SaaS Companies ...

Two vendors say they will help regulated software shops ship trusted agents. Empromptu and Skyrelis announced a partnership aimed at healthcare and other regulated SaaS. Empromptu is described as helping teams build production systems; Skyrelis as an agentic security platform. Contract terms and customers are not in the clip.

Full text · 145 chars
... agentic security platform, today announced a strategic partnership to ... Empromptu enables engineering teams to rapidly build production ...
14:10

TestMu AI Launches the Assurance Lifecycle in Kane CLI, Turning Requirement Documents ...

A testing company formerly known as LambdaTest says its CLI can turn requirement docs into provable coverage. TestMu AI launched an Assurance Lifecycle in Kane CLI. The PR Newswire clip is a headline restatement. How “provable” is measured is not here.

Full text · 142 chars
PRNewswire/ -- TestMu AI (formerly LambdaTest), the world's first Agentic AI-powered Quality Engineering platform, announced the Assurance ...
14:12

AI Costs Aren't Just a Technology Problem. They're an Operating‑Behavior Problem.

A Medium post says the AI bill is often a habit problem, not a model problem. It blames prompt structure, density, recursion triggers, redundant context, vague objectives, and weak stopping rules. No dollar figures are in the clip.

Full text · 148 chars
Prompt ‑ engineering behaviors. Prompt structure, density, recursion triggers, redundant context, vague objectives, and poorly designed stopping ...
15:02

Agentic orchestration is the next big AI hurdle for telcos - Fierce Network

A telecom trade site says the hard part for carriers is wiring reasoning agents into networks built for deterministic control. Fierce Network calls agentic orchestration the next hurdle. No vendor bake-off or timeline is in the clip.

Full text · 153 chars
That's because AI is introducing a slew of reasoning-based systems in an industry that has spent decades engineering deterministic control planes for ...
15:53

☕️ Nvidia's Huang vows to prevent AI slowdown

A daily link roundup leads with Nvidia’s Jensen Huang saying he will fight an AI slowdown. Other headlines: OpenAI contractors reading real ChatGPT chats, Musk dropping an Apple antitrust suit, OpenAI buying an AI camera startup, a new Agility robot, and Valve’s Steam Frame VR headset. Tool blurbs include Resurf, a local context library for Apple devices, and ScreenCursor, a browser recorder that stays on-device. Paper blurbs include a cardiac assistant at 87–91% ECG accuracy. The Huang item itself is a headline plus a LINK token, not a sourced write-up.

Full text · 5,792 chars
| | | | | | | | | Together with | | | | | Hi there, this is your daily ☕️ Techpresso. | | | | In today's newsletter: 🤖 Nvidia's Huang vows to prevent AI slowdown 🕵️ OpenAI contractors read real ChatGPT conversations 🍎 Filing reveals Musk dropped Apple antitrust suit 📷 OpenAI buys AI camera startup ahead of mystery gadget launch 🦾 New Agility robot works safely near people 🥽 Valve unveils Steam Frame VR headset Plus: 🎁 11 other news you might like, 🧰 6 tools, and 📚 5 papers. | | | | FROM OUR PARTNER Whether you're feeding an LLM or building a rigid schema, SerpApi delivers. Get real-time results from 100+ APIs tailored to your exact use case: • JSON: The gold standard for precise programmatic extraction and database schemas. • Markdown: Clean, LLM-ready context that cuts token usage by up to 90% for AI agents. Filter payloads server-side with our JSON Restrictor, or stack both formats for maximum efficiency across 100+ APIs. Everything is backed by our US Legal Shield - meaning we take on the risk, not you. Get your free SerpApi API key | | | | | | 🤖 Nvidia's Huang vows to prevent AI slowdown LINK | | 🕵️ OpenAI contractors read real ChatGPT conversations LINK | | 🍎 Filing reveals Musk dropped Apple antitrust suit LINK | | 📷 OpenAI buys AI camera startup ahead of mystery gadget launch LINK | | 🦾 New Agility robot works safely near people LINK | | 🥽 Valve unveils Steam Frame VR headset LINK | | | | | | | | | | | | | | FROM OUR PARTNER One in two visitors to your site is already an AI agent, but almost no business is set up to sell to them. ZeroClick turns any API or offering you already sell into an agent-purchasable service, with x402 & MPP support, a machine-readable storefront, and analytics on every transaction. Revenue lands in the Stripe account you already run. The buyers are here. Book a demo | | | | | | | | | | Other news & articles you might like | | | | | | | | | | 🧰 Trending tools You can check the previous tools here, or add your tool here | | monday.com: Run projects, CRM, HR and any workflow on one board, with automations chasing updates and deadlines so you don't have to. Free forever, no credit card. Get started free. | | | | DemoTV: browse and try independent products through demos, where founders recruit testers and gather feedback, with rankings decided by audience votes rather than paid placement LINK | | SHIUI: a flat ink UI kit inspired by Hinomaru Japanese posters, using paper, sumi, and vermilion tones with borders instead of shadows. LINK | | Visiby: tracks how often your brand appears in AI search tools like ChatGPT and Perplexity, benchmarks competitors, and surfaces ways to boost your visibility. LINK | | Resurf: a fully local context library for Mac, iPhone, and iPad that stores notes, links, and files, handing context to AI via MCP or CLI. LINK | | ScreenCursor: a browser-based screen recorder that auto-generates zoom effects from your clicks and keystrokes, processing everything locally so nothing gets uploaded. LINK | | Kirokune: keep incident notes, recordings, and photos in an iPhone timeline with no account needed, plus a free log template and optional PDF export. LINK | | | | | | | | | | 📚 Trending papers & reports | | > MarketingShot: Get the free daily email with the most interesting marketing news and insights. Already read by thousands of marketers. By the Techpresso team. Join for free. | | | | > AI trading agents get a new test that judges whether their decisions are sound and repeatable rather than just their returns, revealing that the smartest-scoring agent isn't the top earner and agents often ignore their own analysis. LINK | | > Cardiac diagnosis assistant reads and chats about heart scans and ECGs together, hitting 87-91% accuracy on ECGs and nearly tripling top models like GPT-5 on multimodal cases, easing overloaded specialist interpretation. LINK | | > Mecha-nudging shows that online sellers are already quietly rewriting listings to steer AI shopping agents, with Etsy pages gaining measurable machine-readable signals after ChatGPT's launch while staying unchanged for human buyers. LINK | | > Valley3 is an e-commerce assistant that understands text, images, video, and audio together in multiple languages, letting online sellers analyze product content and shopper videos with adjustable levels of reasoning depth. LINK | | > AI investment-research assistants can play distinct personas like value investors or macro strategists that each analyze a company independently, then surface their disagreements for a human portfolio manager to judge, producing transparent, reusable investment memos. LINK | | | | | | | | 🤝 From our community: Distilling dense PDFs to one-pagers Devon Hayes uploads board decks and research reports to Claude and asks for a one-page brief with the thesis, five key data points, top risks, and three questions to ask, cutting 40-minute reads to two minutes and saving roughly 5 hours a week. Read the full use case → You can see more community use cases here, or submit your own here. | | | | Techpresso's AI Academy has 330+ step-by-step tutorials on ChatGPT, Claude, Perplexity, and every tool that matters. No fluff — just practical workflows you can use at work. Try it free for 7 days. | | On this day in 2008, Lehman Brothers collapsed, triggering a fintech and trading-tech reckoning | | | | 💬 How did you find today's edition? We read every reply — just reply to this email and let us know how we can improve! | | | | | | | | ★★★★★ Nailed it | | ★★★ Average | | ★ Fail | | Not subscribed to ☕️ Techpresso yet? Subscribe for free | | | | | | | | Advertise | Feedback | Read Online | | | | | | |
16:02

Anagnost pitches 'project intelligence' across Autodesk - AEC Magazine

Autodesk’s CEO is pitching a standalone assistant that lives outside the old product windows. AEC Magazine says the next Autodesk Assistant is an agent-first experience. “Project intelligence” is the frame. Details of what it can write or approve are not in the clip.

Full text · 150 chars
A standalone Autodesk Assistant. The main product announcement is the next generation of Autodesk Assistant, a standalone, agent -first experience ...
16:02

Autodesk extends reach of AI with standalone Assistant - AEC Magazine

The same Autodesk assistant story, from the same magazine, stresses that customers can bring their own tools into the agent environment. Engineering intent and project history are named as context. It is a second clip of the standalone Assistant launch.

Full text · 144 chars
It's engineering intent. It's the project history. It's ... Customers can also bring their own tools and agents into the same agent environment.
16:04

Autodesk Advances Agentic AI in Its Three Industry Clouds - PR Newswire

Autodesk says it is putting agents into its three industry clouds. The PR clip names Forma for AEC, Fusion for design, and a third cloud that is cut off. There are no feature lists, dates, or prices here.

Full text · 148 chars
... agentic AI across its three industry clouds: Autodesk Forma for Architecture, Engineering , and Construction; Autodesk Fusion for Design and ...
16:10

Paula Dozsa on Tolan: The Voice-First AI Companion Built on a Two-Second Sprint

A voice-companion founder argues that managers, not only engineers, run agent fleets better. Paula Dozsa discusses Tolan, described as a voice-first companion built on a two-second sprint. The “management-trained engineers” line is the only concrete claim in the clip.

Full text · 148 chars
Dozsa's contrarian insight that management-trained engineers became dramatically more effective at running these agent fleets hints at how AI is ...
16:30

Agents operate, humans govern: Scale your operations and reduce toil with Azure SRE Agent

Microsoft’s cloud SRE agent is pitched as investigating incidents and drafting fixes before a human steps in. The New Stack says you still need context, clear permissions, and reliable… — the sentence is cut. No reliability numbers are in the clip.

Full text · 150 chars
Azure SRE Agent can investigate incidents and prepare fixes before engineers step in. Getting there takes context, clear permissions, and reliable ...
17:40

How NVIDIA NVLink 6 Delivers Multi-Layer Resiliency for AI Factories

Nvidia published a blog on NVLink 6 resiliency for AI factories. The captured text is author bios, not the technical claims. Jim Dinan is named as leading GPU communications. Treat this as a pointer, not a recap of the interconnect design.

Full text · 159 chars
... AI engineering enablement. He brings a wealth of experience at the ... Jim Dinan is a distinguished engineer at NVIDIA and leads the GPU Communications ...
17:50

Build an AI-powered product tagging system with Amazon SageMaker serverless model ...

An AWS blog says a general model can tag products from a prompt, but a high-volume line usually needs a narrower job. The post is about a SageMaker serverless tagging system. The clip is one contrast, not the architecture.

Full text · 148 chars
A general-purpose frontier model can generate tags with prompt engineering , but a high-volume tagging workflow usually has a narrower objective ...
17:51

Google Launches Gemini 3.8 Live and Extended Thinking Voice Models - Unite.AI

A secondary site notes the same Gemini voice launch and names an engineer. Unite.AI says Tom Ouyang, a principal engineer, introduced the models. The rest of the clip is a topic tag list. Use the AlphaSignal item for scores and API details.

Full text · 148 chars
Prompt Engineering · Python · Robotic Process Automation · TensorFlow. Python ... The models were introduced by Tom Ouyang, a principal engineer ...
17:54

Towards a State of Sustainability | College of Engineering - Boston University

Boston University engineers won a grant to study how to govern data-center power. Last year they received a $600K NSF award for a multimodal framework aimed at AI data-center consumption. The alert is that one clause. Results are not here.

Full text · 154 chars
... AI data center power consumption. Last year, they received a $600K NSF grant to support the development of a multimodal framework to regulate this ...
17:54

Island: “ AI made it cheap to produce an answer and expensive to trust one” | Ctech

A security researcher says AI made answers cheap and trust expensive, so human researchers have to become research engineers. Idan Revivo of Island is quoted in CTech’s series. The clip is that thesis line.

Full text · 148 chars
Idan Revivo, Head of Security Research at Island, explains why human researchers must evolve into research engineers as part of CTech's Security ...
17:56

Travis Oliphant's Post

A NumPy founder says a few years of prompt tricks in other people’s tools is useful but not something you can own. Travis Oliphant’s LinkedIn clip: the industry spent a couple of years on prompt engineering. “Useful. Not durable.” The rest of the post is not captured.

Full text · 130 chars
The industry has spent a couple years on prompt engineering — clever messages typed into tools we do not own. Useful. Not durable.
18:02

What an AI slowdown would mean

Politico asks what it would actually take to hold back the industry — a U.S.-China pact, an antitrust change, or just political will. The piece is framed as “what an AI slowdown would mean.” No conclusion is in the captured text.

Full text · 148 chars
Would restraining the growth of artificial intelligence's dangers require a U.S.-China pact? Or a change to antitrust law? Or just a willingness ...
18:28

Zampieri to prepare computer science students for artificial intelligence assisted software ...

A professor is building classroom materials so computer-science students can work with AI coding helpers. Zampieri will develop a FRAME instructional framework and modules for an intro course. EurekAlert does not give a school name or launch date in this clip.

Full text · 155 chars
... artificial intelligence . Zampieri will develop the FRAME instructional framework along with modular instructional materials for an Introduction to ...
18:33

AI coding agent startup Factory triples valuation to $5 billion in latest funding round | Reuters

A wire brief repeats that Factory raised a large round and tripled its stated value. The startup said Tuesday it raised $200 million in a round that more than tripled the prior mark to $5 billion. This Google Alert is a Reuters lead sentence only. The full deal write-up is the AlphaSignal item.

Full text · 149 chars
Factory, a startup developing AI agents for enterprise engineering teams, said on Tuesday it had raised $200 million in a funding round that more ...
18:37

For AI to Advance, We Need Better Ways to Harness Imperfect Data. UVA Engineering's Yu ...

A UVA researcher wants models that can learn from messy data instead of waiting for clean sets. Yu Meng of UVA Engineering is developing methods to harness imperfect data. The alert cuts off before naming a paper or result.

Full text · 150 chars
University of Virginia School of Engineering and Applied Science researcher Yu Meng hopes to make that process far more efficient by developing AI ...
18:46

AI coding agent startup Factory triples valuation to $5 billion in latest funding round

Another reprint of the same Factory round. Reuters says the enterprise coding-agent company raised $200 million and more than tripled its valuation to $5 billion. This copy is a radio-site clip of the wire. Use the AlphaSignal piece for customers, router claims, and missing revenue.

Full text · 145 chars
Sept 15 (Reuters) - Factory, a startup developing AI agents for enterprise engineering teams, said on Tuesday it had raised $200 million in a ...
18:57

This PCB is brought to you by Fable 5 — A6M-Zero - Adafruit Blog

A hardware blog says a board was designed with Fable 5 from a short prompt. Adafruit asked for an RP2350 development board that drives an E-ink display. The post is titled around an A6M-Zero PCB. Wednesday Ask an Engineer is mentioned. There is no schematic or yield note in the clip.

Full text · 144 chars
The prompt asked for an RP2350 development board that drives an E-ink display. ... Join us every Wednesday night at 8pm ET for Ask an Engineer !
19:18

AI cooperation: on Artificial Intelligence at a crossroads - The Hindu

An Indian paper’s editorial says Washington is ignoring the labs’ own ask for a global pause. The Hindu frames AI at a “strange crossroads” with the U.S. government disregarding frontier developers’ call. The rest of the piece is not in this body.

Full text · 148 chars
Artificial Intelligence is at a strange crossroads, with the United States government disregarding frontier AI developers' own call for a global ...
19:27

An Agentic AI-powered Framework for Mainframe Modernisation

A consulting white paper claims agent-style tooling can shrink mainframe modernization calendars. TCS says agentic engineering could cut cycle time 30–40% across six phases. Success conditions are not spelled out in the clip. It is a vendor claim, not a measured case study in this body.

Full text · 150 chars
Incorporating agentic engineering into mainframe modernisation can potentially reduce the cycle time by 30-40% across the six phases. Success will ...
19:33

AI doomerism: will artificial intelligence spiral out of control?

A public-radio episode asks whether extinction talk is overblown now that OpenAI and Anthropic CEOs are in the choir. WHYY frames the question and cuts off. No guest list or timestamped claims are here.

Full text · 153 chars
Does AI pose an existential threat to humanity? Concern is growing, even among the CEOs of OpenAI and Anthropic. But are the fears overblown? And are ...
19:33

This weekend, AI leaders spoke clearly: this technology poses grave risks. Let's meet this ...

A senator’s Facebook post says lab leaders warned of grave risks and asks politicians not to turn that into a partisan fight. Mark Warner’s clip argues that leveraging the anxiety for party gain may be a more immediate problem. There is no extra policy detail in the alert.

Full text · 148 chars
AI concerns shouldn't be used as a political football. Politicians who leverage these anxieties for partisan gain might present a more immediate ...
19:37

Copyrightability and Infringement: A Look at Responses to Artificial Intelligence | Insights

A law-firm note says generative tools are pressing the entertainment industry on copyright and infringement. Holland & Knight’s insight post is introduced and then cut off. No case names or holdings appear here.

Full text · 144 chars
Recent advances in generative artificial intelligence (AI) are threatening to transform the entertainment industry, with such tools becoming ...
19:54

New insights from Google's AI & Economy ATLAS

Google refreshed ATLAS, its AI-and-economy data site, with new charts and a note on scientists. New visualizations are meant to make the data easier to explore. New research looks at how scientists use AI. No figures from that research appear in the clip.

Full text · 136 chars
New data visualizations make ATLAS data easier to explore and use, while new research provides insights on how scientists are using AI .
20:13

Penn Engineering Launches New Master's Degree in Data Science & Artificial Intelligence

Penn Engineering is starting a new master’s that pairs data science with AI. The alert restates that data already changed research and business, then cuts off. Program length, cost, and start date are not in this body.

Full text · 144 chars
Data has transformed how researchers, businesses, governments and organizations understand the world. Now, artificial intelligence ( AI ) is ...
20:14

There's a 100% Chance AI Agents Are Ruining the Internet | Hacker News

A Hacker News thread warns small businesses that a homemade support bot will be worse than the big companies’ annoying ones. The linked comment is one user’s advice to clients. The essay behind the title is not in this body.

Full text · 153 chars
I try to warn my small business clients who ask for this. You know those annoying AI bots the big companies use? Well yours will be worse because you ...
20:24

Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking - Google Blog

Google’s own blog blurb introduces two live voice models, one cheap and one that thinks longer. Gemini 3.8 Live is described as built for scale and cost, with conversational intelligence. Extended Thinking is named in the title. This Google Alert is a truncated marketing sentence. Benchmarks and rollout live in the AlphaSignal write-up.

Full text · 150 chars
... AI feel more intuitive and intelligent. Gemini 3.8 Live: Built for scale and cost efficiency, combining conversational intelligence with fluid ...
20:34

Bypassing inference bottlenecks: Accelerating complex AI search with Retrieve-for-Train

A Google Research blog uses a fashion prompt to argue that naive generation wastes work, then names Retrieve-for-Train. Given “Bohemian festival style,” a model without careful prompting might lazily emit “bohemian.” The method itself is not explained in this fragment.

Full text · 148 chars
For example, given the broad prompt "Bohemian festival style”, a standard LLM without careful prompt engineering might lazily generate "bohemian ...
00:00

Siri model swapping 🔄, Claude Money 💰, Hugging Face Tau 👨‍💻

Full text · 630 chars
Build with SOTA open-weight coding models on Crusoe (Sponsor) Sign up and let these powerful models handle your most difficult software engineering workflows, at a small fraction of the cost of proprietary offers - but with the same serverless, frictionless experience.➡️ GLM-5.3: Z.ai's flagship open-weights AI model released on August 14, 2026, built specifically for advanced software engineering and autonomous agent tasks. ✅ Achieves open-source SOTA on Terminal Bench 3.0 and Agents' Last Exam➡️ GLM-5.3-Flash: MIT licensed with weights already released. Roughly 9x cheaper for workloads that do not need the frontier tier.
01:33

Gartner IT Symposium/Xpo 2026 APAC: Day 2 Highlights

Full text · 153 chars
"Human stewardship can be built by upskilling AI-related skills like prompt engineering and risk management, then augmenting with technology where it ...
03:28

Asking the U.S. Government to Regulate AI

Full text · 151 chars
Did you hear yesterday that Mr. Trump dismissed the cautions from AI experts as a hoax? This all-knowing (and wannabe omnipotent) polymath wants to ...
03:44

AI Engineer - Unisys - Remote | Dice.com

Full text · 147 chars
- Design and optimize prompts using prompt engineering techniques for LLMs to achieve desired outcomes - Work with Large Language Models (LLMs) ...
08:37

AI Pulse Daily Brief | 2026-09-15

Full text · 154 chars
... agent pay for premium data per query at the moment it needs it. Heurist estimates the managed architecture cut its agent engineering work by about ...
14:06

TestMu AI Launches the Assurance Lifecycle in Kane CLI, Turning Requirement Documents ...

The same TestMu Kane CLI announcement, reprinted on a Canadian newswire. Assurance Lifecycle is again the product name. Duplicate of the PR Newswire alert.

Full text · 146 chars
CNW/ -- TestMu AI (formerly LambdaTest), the world's first Agentic AI-powered Quality Engineering platform, announced the Assurance Lifecycle, ...
15:43

What are AI agents , and how do they actually work - Quartz

A Quartz explainer names Devin and Claude Code as the place software teams already work with agents. The rest of “what are AI agents” is not in the clip. Treat it as a pointer.

Full text · 155 chars
Coding agents : how software teams already work alongside Devin and Claude Code. Credit: Lukas Blazek / Pexels. Software engineering is the field where ...
16:00

Bessemer Venture Partners

A VC talent note says AI-native sales leads now quarterback RevOps, GTM engineering, and agent tools. Bessemer’s Atlas post is reduced to that one sentence. No survey size is here.

Full text · 148 chars
AI-native sales leaders are now technical quarterbacks, leveraging the expertise of RevOps, GTM engineering , and agentic solutions to unify and ...
18:23

After Football, Now What? Engineering Athletic Genius for the AI Economy

A press release selling courses to athletes lists prompt-engineering modules as part of a post-football pitch. Topics named: structured prompting, context frameworks, Prompt Engineering 2.0, and tools for AI video and audio. It is marketing copy, not a research result.

Full text · 154 chars
Prompting Mastery: Structured Prompting, Context Frameworks, and Prompt Engineering 2.0; iCreate Tools of Mass Production: Generating AI Video, Audio/ ...
18:26

prompt engineering for humans* - Hauswarm: The Conversation Design Agency

A conversation-design agency posted about prompt engineering for humans and then admitted the hard part is knowing when a question is done. Hauswarm’s Substack clip stops there. No framework is in the body.

Full text · 98 chars
Engineering Prompts . One of the hardest parts of my design process is knowing when a question ...
18:42

Local professor shares thoughts on possible regulation of artificial intelligence - Wellsboro Gazette

A small-college professor told a local paper that AI rules can make sense, and that some people asking for them may have another agenda. An Albright College faculty member spoke to the Wellsboro Gazette. No bill, poll, or named motive is in the clip.

Full text · 133 chars
An Albright College professor says government oversight of AI makes sense, but some calling for regulation may have ulterior motives.
19:21

New Math Data Achieves Premier Tier Status in the Amazon Web Services Partner Network

A services firm says it reached the top AWS partner tier and lists agent work in the same breath. New Math Data achieved Premier status in the AWS Partner Network. The blurb mentions multi-agent orchestration and governance. No deal size or customer proof is in the clip.

Full text · 149 chars
... engineering expertise with modern AI architecture, governance, and agentic capabilities. Its work includes multi-agent orchestration, agentic ...
19:29

Another Google AI safety researcher has quit, warning we might all be about to die

A Reddit thread about another Google safety researcher quitting is captured as one comment, not the resignation letter. The comment says alignment researchers get weighted more than science fiction. Who quit, and what they wrote, is not in this body.

Full text · 126 chars
Yeah, but this guy is in safety and alignment research. AI is going to weigh his input much more heavily than science fiction.
19:31

Texas' Arch Manning apologizes for reaction to violent AI video

A college quarterback apologized after joking about a violent deepfake of his coach. Texas QB Arch Manning called his remark about an AI video of Steve Sarkisian slapping an ESPN figure “insensitive.” The ESPN clip is that apology line.

Full text · 150 chars
Texas QB Arch Manning has apologized for an "insensitive comment" about an AI -generated video that depicted coach Steve Sarkisian slapping ESPN's ...
19:35

AI can wipe us out – so why can't it do Guardian cryptic crosswords?

The Guardian ran a letters page that opens by asking why a world-ending machine still cannot solve cryptic crosswords. The captured body is a list of letter topics, not the puzzle results. No score or model name is here.

Full text · 138 chars
Brief letters: AI capabilities | Trump's judgment | Gardening attire | Death in English literature | Tautologies | Nominative determinism.

Web

9