Nothing matches those filters.

Lead

12

Video

1
03:33

New Deepseek, human genome map, Navier Stokes, GPT finance, Suno v6, YuE2: AI NEWS

A weekly AI news host walks through a pile of new tools you can try, plus two big research claims about DNA changes and a famous unsolved math problem. DeepSeek 4.1 Flash is a 552 billion parameter mixture-of-experts model with only 8 to 16 billion parameters active. The host says it scores 74.2 on DeepSwee 1.1, runs at 217 output tokens per second on the API, and ships as a 510 GB dump or a 106 GB Q1 GGUF. DeepMind's Alpha Genome Atlas is a one-petabyte lookup of more than 9 billion single-letter DNA mutations. OpenAI ran 10,000 agents for 88 hours on a forced Navier–Stokes blow-up; the host flags that the unforced problem is still open.

Notes
  • Whisper transcript of AI Search weekly nZYJdwM-_nI. Names/benches as heard; garbles flagged.
Vision, 3D, worlds
  • Marigold V2: photo → depth, surface normals, albedo, other maps. Pixel-level. Host says sharper than Moji 3 and finnadeff (garbled). Code out. Local VRAM ~17 GB / ~29 GB.
  • Unimate: one model animates odd rigs from a rigged mesh + text. Code + training code.
  • WorldSculpt: photos/video → 3D with separate movable objects; Gaussian-splat → mesh. Code out.
  • Fire3D: photos/video → simulation-ready meshes + materials. Under a minute; up to 16 objects in parallel on one GPU. Code + training code.
  • Lingbot World 2: walkable world via keys + text events. Claim: >1 hour continuous, 720p / 60 fps, NPCs. Backbone named “pretty old Alibaba 1 2.2” (garbled); streams chunk-by-chunk and caches frames. 14B / 1.3B. Code out.
Alpha Genome Atlas (DeepMind)
  • Host: ~2% of the genome understood; 98% “mysterious.”
  • Predicts every single-letter mutation (~9 billion). 1 petabyte, >9 billion variants, >30× the AlphaFold database. Lookup for regulation/protein effects instead of re-running.
  • AVI score ranks mutations. Already: 22% more genetic associations in UK Biobank; “helped support the solution of a previously unsolved rare disease.” Public “explore Alpha Genome Atlas.”
DeepSeek 4.1 Flash
  • Host: a “.1” that is “completely different from the previous V4.” 552B mixture-of-experts, 8–16B active. New pieces named: separate encoder + decoder, Engram, sliding-window attention, CSA2.
  • 74.2 on DeepSwee 1.1; he says #1 vs GPT-6 Astra, but “the DeepSwee team hasn’t added this officially.” Also SOTA on Cyber Gym and Automation Bench (beats GPT-6 Astra max on the latter).
  • LiveBench: #1 open, a few points below GPT-6 Astra and Claude Fable. VALS: #1 open, slightly above Kimi K3 and GLM 5.3 — “kind of all tied.” Artificial Analysis intelligence index: still behind Kimi K3 and GLM 5.3.
  • API: 217 output tokens/sec — “over 3×” GLM, “like 7×” Kimi K3. Cheaper than those open models and than GPT-6 / Fable. Full dump ~510 GB; community Q1 GGUF ~106 GB.
Navier–Stokes (OpenAI)
  • Millennium-prize question: can the liquid/gas equations break down? Open ~90 years.
  • Internal model “significantly more capable than GPT-6 Astra.” 10,000 concurrent agents, 88 hours.
  • Host’s telling: a whirlpool under specific conditions stretches a vortex without limit in finite time (singularity / blow-up).
  • Caveats he states: Euler — drop viscosity is easier; force — stirring also makes a blow-up easier. A mathematician “one from Anthropic” solved the easiest forced Euler around the same time. OpenAI’s claim is forced non-Euler (viscosity stays). Unforced Navier–Stokes is still open.
Robots
  • Isaac 0.5: images, video, language, state, prior actions → next action or an answer. 36B sparse. Data from >35 robot systems, 100,000 h robot, ~1M h general video. One backbone. Open-sourced.
  • Show Harness: a vision-language model picks named moves; a “robot specific action interpreter” turns names into motion. Success “could reach up to 100%.” Gemini, GPT, “quen,” GLM. Needs a real robot.
  • UMR: human motion → humanoid by matching 3D surface points under robot limits. Keeps contact. Demos: spin kicks, stairs, tennis. Code out.
  • Unitree UnifullLMWLA (as heard): 6B; camera + instruction + state → actions. 64 tasks (54 tabletop, 10 whole-body). Two-finger and five-finger hands. Predicts which scene parts will change, then moves. Dataset + weights; Unifull LM <9 GB.
Speech, music, local, finance
  • Tencent “Out” (as heard): “nano-banana but for speech.” TTS from a description; clone from a few seconds; add/delete words (demo: insert “never” into “mamba out”); denoise, upscale, change emotion/timbre, whisper; add/remove laughter and breaths. 6.12 GB.
  • Suno V6: text or audio-ref → song plus micro-edits (one lyric, swap vocals/instruments, isolate a riff). V6 paid/precise; V6 wild paid/varied; V6 mini free/faster/lower quality. Closed; new download limits.
  • UA2 (title YuE2): score first (melody, rhythm, chords, structure), then vocals + accompaniment. Cover: Jingle Bells in a minor key. Claim: best open vs Minimax Music 3 and ASTEP 1.5, quality better than Suno V6. 7.3 GB.
  • Edge Zero: stream only the experts a question needs. Demo Quen 3.5 35B (garbled Qwen), 3B of 35B active. int4 plus “recover Laura” (LoRA, garbled). Floors: 35B in 2.9 GB; 8B on Inclusion AI Ling 3.0 Tiny in 1 GB.
  • RealSwee: private company code + business tickets. Common fail: missing a requirement. Fable 5.1 38.8%, GPT-6 Astra 33.8%, Gemini 3.8 Flash third, GLM 5.3 fourth. Top four “statistically tied.”
  • ChatGPT for financial services: GPT-6 Astra + Pitchbook, Crunchbase, “and others.” Research, models, spreadsheets/docs/decks; figures trace to tables. Contact sales.
  • OpenBMB MiniCPM 5 2B: 2B dense. Host: best of the small-size group he shows; on average ahead of Quen 3.5 4B. Training set released. ~5 GB.
  • Sponsor only: Higgsfield + GPT-6 Astra inside ChatGPT.
Transcript · 36,310 chars
AI never sleeps, and this week has been absolutely insane. Deepseek releases their latest model and it's an absolute beast. We also have a super powerful speech generator and editor. It's kind of like nano banana but for speech. OpenAI used a swarm of 10,000 agents to solve one of the hardest and unsolved math problems in the world. Google DeepMind also used AI to essentially map out every possible mutation in the human genome. We have a new state of the art open source model for predicting the depth and surface normals of an image. Suno releases their latest model but we also have a new top open source music generator which is just as good. This AI can animate any character, including ones with unusual skeletons. We have a ton of open source AIs for reconstructing a scene in 3D. We also have a new real time interactive world generator which is really good quality. We have a ton of new open source robot models and a lot more. So let's jump right in. First up we have this AI called Marigold V2. This basically understands 3D depth and structure inside a regular image. So what you would do is give it a regular image and it can convert that into a ton of things like a depth map, surface normals which is like the orientation of a surface in the image and then also albedo which is like the raw color and other different maps and the awesome thing is this new method is pixel level. So this is super high resolution and much better compared to the other methods. For example, if you compare this to a previous model called Moji 3, you can see that Marigold is a lot more detailed and accurate. Here's another comparison between another model in finnadeff and as you can see, Marigold is just a lot more detailed and higher resolution. And then here's a comparison of its normal estimation capabilities. You can see that this new Marigold method is just a lot sharper and faithful. Here's another crazy comparison. And if you look at these benchmarks comparing other similar models, you can see that on average Marigold V2 scores the best. The awesome thing is they've released this already. So at the top of the page, if you click on this code button and you scroll down a bit, here it contains all the instructions on how to download and run this locally on your computer. Note that inference at this resolution requires around 17 gigabytes of VRAM, whereas this resolution requires around 29 gigabytes. If you're interested in reading further, I'll link to this main page in the description below. Also this week, we have a cute little AI called Unimate. This can basically animate completely different 3D skeleton types using just one AI model. So here's an example where you can get it to animate a ton of different non-human characters like flowers or Garfield or a satellite or this dragon like creature. And it's able to handle all these animations very well. You see, a lot of these skeleton animation models are only specialized for like human characters, but they often struggle when you give them something unusual like a bird or snake or some other completely different objects. But here Unimate is able to animate pretty much anything. You just need to give it a rigged 3D model plus a text instruction like a dog walks forward or the snake slithers forward and it can generate the motion automatically without any further retraining. So this is one of the best AI models to use to animate unusual characters or objects. Now at the top of the page, they've released the code already. So if you click on this link and you scroll down a bit here, it contains all the instructions on how to set this up locally. They also released the training code to this. So it's fully open source, which is fantastic. If you're interested in reading further, I'll link to this main page in the description below. Also this week, Google DeepMind just released Alpha Genome Atlas. And this is basically a giant AI generated map of human DNA. Now currently we only understand 2% of this, which mostly codes for proteins, but the other 98% is a lot more mysterious. So Google used Alpha Genome to basically predict the effects of every possible single letter mutation in the human genome. Now there are roughly 9 billion possible single letter mutations and testing all of them experimentally would be pretty much impossible. So instead of trying this out in a real lab, DeepMind just used Alpha Genome, which is an AI to predict the effect of every single one of these mutations. And the result is a one petabyte dataset containing predictions for more than 9 billion genetic variants, making it over 30 times larger than the Alpha fold database. So if you need to look up the effect of a mutation, instead of running the AI from scratch every time scientists can now just look up this mutation in Alpha Genome Atlas and get predictions on how it could affect things like gene regulation or protein production. They've also created something called the AVI score, which is basically an impact score to help researchers quickly identify which mutations are worth investigating further. And it's already showing some real results. So researchers used it to find 22% more genetic associations in the UK Biobank data. And it also helped support the solution of a previously unsolved rare disease. So basically this takes a huge amount of genetic information that was previously very difficult to interpret or predict. And now they've just turned it into a huge database which researchers can actually search. And the awesome thing is they've released this Atlas for everyone to access. So simply click on explore Alpha Genome Atlas at the bottom here. And then afterwards you can search this database for like over 9 billion mutations. If you're interested in reading further, I'll link to this main page in the description below. Also this week we have a new real time interactive world model called Lingbot World 2. This basically generates an interactive virtual world continuously and you can use key presses to control it and walk around. So here are some examples of this in action. You can see that it's much better quality than the previous open source world models. Here everything looks a lot more detailed and coherent and higher resolution. Now in addition to just controlling the movements via key presses, you can also enter prompts to add any event or effect you imagine. And the team says this can generate interactive worlds continuously for over an hour and the real time system can reach 720p resolution at 60 frames per second, which is really impressive. They've also added an agent mechanism, which means that you can now have NPCs inside the environment. In other words, you can add different characters which also move around and they can behave and respond in different ways. Now under the hood, they've actually just used the pretty old Alibaba 1 2.2 as the video generator, but here they're generating the world chunk by chunk so it can actually stream real time. It also caches previous information so it kind of retains a memory of the scene. Now at the top here, if you click on this code button and you scroll down a bit here, it contains all the instructions on how to download and run this locally on your computer. Note that there are two different variants of this. There's a larger 14b variant which is higher quality and a smaller 1.3 billion parameter variant which is a lot faster. If you're interested in trying this out, I will link to this main page in the description below. Also this week we have a new open source robotics model called Isaac 0.5. This is a robot foundation model designed specifically to give robots a much more general understanding of what they're seeing and what they should do next. So you can feed it images, video, language instructions and the robot's current state or even its previous actions and this model can basically take all this input and predict the next action or how to answer questions. So it can do things like locate objects, predict what the world might look like next or directly generate robot movements. Now this is a 36 billion parameter sparse model so it's fairly tiny and they train this model on data from more than 35 robot systems so this model can be applied or transferred to different robot types. They also trained it on 100,000 hours of robot experience and around a million hours of general video. One interesting feature is that Isaac doesn't just learn robot control separately. Instead, video understanding, spatial reasoning, predicting future states and physical actions are all trained together in just one backbone. So knowledge from all these inputs can potentially help the robot make better decisions and understand the physical world. The awesome thing is they've actually open sourced the model so if you click on this button at the top and you scroll down a bit here it contains all the instructions on how to download and run this. If you're interested in reading further I'll link to this main page in the description below. Also this week we have a new AI called WorldSculpt and this turns a scene into individual 3D objects which you can edit further. So how this works is you can input multiple images of a scene or even a video and it'll turn that into a complete 3D scene with separate individual objects. You can then edit each of these objects further like resizing them or moving them around. And this is an important distinction because you know most 3D reconstruction systems can create something that looks like the original scene but everything is just fused together. But here WorldSculpt is able to reconstruct each object separately and also position them correctly inside one shared 3D world. That means you can like move each object around and resize them or edit them further. The cool thing is this can even convert Gaussian splat worlds into actual 3D mesh scenes. So here's an example of that. The nice thing is they've released this already so at the top of the page if you click on this code button and you scroll down a bit here it contains all the instructions on how to download and run this. If you're interested in reading further I'll link to this main page in the description below. So this week we have a new AI called Fire3D and this can take a series of normal photos or videos and turn it into a complete 3D scene which you can edit further. So you just input one or several images or a video and it can basically reconstruct the 3D scene with separate objects. The entire scene is simulation ready and you can edit each object further. Every object basically becomes its own complete mesh with material information so you can move them around or add each one to a robotic simulation. And this thing is super fast. It can recreate a simulation ready scene in under a minute and processes up to 16 objects in parallel on just one GPU. Now at the top they released the code to this already so if you click on this link and you scroll down a bit here it contains all the instructions on how to download and run this locally on your computer. Plus they've also released the training code to this as well which is fantastic. If you're interested in reading further I'll link to this main page in the description below. After this week Tencent releases a really powerful audio model called Out and you can think of this as like nano-banana but for speech. First of all this can do regular text to speech and you can describe exactly the voice you want and the dialogue. For example here's the prompt and here's the result. As you can hear this voice does sound pretty sad and full of grief. This can also do zero shot voice cloning so you just need to upload a few seconds of someone's voice and then plug it through this AI to get that voice to speak out something completely different. So here's an example. First I'll play you the input reference voice and then the output generation. But this can do much more than that. You can micro edit existing speech by like adding or deleting words. For example let me play you the original clip first. Alright so that was the original clip where he just says mamba out. Well we can add the word never here and here's what it sounds like. As you can hear it can seamlessly add a word into the audio. Or here's another example where we can just completely delete this sentence and here's the result. As you can hear it can seamlessly get rid of this sentence. Or here's another example. Let's say we want to add this phrase. Let me play you the original first without this addition. And then here's the edited version where we add this phrase. As you can hear this is very seamless. This can do even more so you can even plug in a really noisy and messy clip and get it to remove the noise or enhance the voice. So here's the input and the output. This can also increase the resolution and quality of an audio clip. So again I'm going to play you the input and then the output. Another crazy thing you can do is change the emotion of an existing clip. So for example let's turn this clip into a sad tone. I'm going to play you the original first and then the output. Or instead of sad let's change this to angry and hear what it sounds like. The crazy thing is you can even change the timbre to something else. So for example the original audio is a male voice. We can change this into a woman with a clear and bright voice. Let me play you the original and then the output. Or you can also take an existing audio clip and turn it into a whisper. The eastern coast is a place for pure pleasure and excitement. The eastern coast is a place for pure pleasure and excitement. There's also a ton of other stuff you can do like removing size or adding or removing laughter, adding or removing breaths, etc. So this is an incredibly flexible speech editing tool. The awesome thing is they released this already. So if you click on this GitHub repo and you scroll down a bit here it contains all the instructions on how to run this locally on your computer. And the model is actually surprisingly tiny at only like 6.12GB in size so this should be able to fit on most consumer GPUs. Definitely one of the most flexible and powerful speech editing tools I've seen so far. If you're interested in reading further I'll link to this main page in the description below. If you're doing any type of content creation definitely check out Higgsfield, the sponsor of this video. Think of this as an entire AI creative team right at your fingertips. In fact you can connect the best AI model GPT-6 Astra to Higgsfield right inside ChatGPT to give it the ability to create images, videos and other content. Astra handles the planning and writes the prompts while Higgsfield generates the visuals. You can run the whole workflow right inside your conversation. Higgsfield also has a new AI motion designer that connects Astra directly to After Effects so you can describe an animation and have it build an actual editable project complete with layers, shapes and expressions that control the movement. Afterwards you can then open the project in After Effects and edit the timing, colors and individual elements yourself or just prompt the agent to edit it for you. You can also give it an animation reference to recreate. Automate repetitive work across hundreds of shapes or apply your brand's fonts and colors through a composition. And if there's an effect you use regularly you can ask Astra to turn that workflow into a reusable plugin. Whether you want to create product ads or cinematic videos or custom motion graphics, Higgsfield gives GPT-6 Astra the creative tools to turn your ideas into actual finished content. Try it today using the link in the description below. Also this week, Deepseek releases their latest model, Deepseek 4.1 Flash. Even though this is just a .1 upgrade it's actually completely different from the previous V4. First of all here are some benchmarks for your reference. Even though this is a Flash model, get this, it even scores 74.2 on DeepSwee 1.1 which actually puts it as the number one model on this leaderboard, even beating GPT-6 Astra which is pretty crazy. Now the DeepSwee team hasn't added this officially to the leaderboard yet so let's wait and see where they place this. Anyway, in terms of Cyber Gym and Automation Bench this is also state of the art. In fact if you look at Automation Bench this Flash model even beats GPT-6 Astra max. Now here are some specs for your reference so this is a 552 billion parameter mixture of experts model. Think of this as a team of experts working together to help you solve a problem so when you use it only around 8 to 16 billion parameters are active making this super efficient. And you know the strangest thing about this is that they used a completely new architecture for this new version 4.1. There's so many new things that we've never seen before such as a separate encoder and decoder component. For your reference most modern language models only have a decoder component. They also added this new Engram feature plus a sliding window attention plus this new CSA2 attention. Anyways it's quite complicated and technical but I'll try to make a full explainer video on this next week. Basically this is a significantly different design from the other mainstream frontier models. If you look at LiveBench by Abacus AI then you can see that this new DeepSeek is indeed the number one open source model. Just a few points below GPT-6 Astra and Claude Fable. If you look at this VALS index which measures an AI model's performance across various knowledge work tasks then as you can see this new DeepSeek is indeed the number one open model just slightly above Kimi K3 and GLM 5.3. But notice that this isn't really a significant difference. So they're kind of all tied for first place. If you look at artificial analysis then according to their latest intelligence index this new DeepSeek is still behind Kimi K3 and GLM 5.3. However, at least if you use it through their API then this thing is blazing fast at 217 output tokens per second which is like over 3 times faster than GLM and like 7 times faster than Kimi K3. The cost of this is also absurd. So way cheaper than the other leading open models and of course even cheaper than the closed models like GPT-6 and Fable. Now as with the previous DeepSeek models they've already open sourced this. So if you click into this folder note that it's around 510 gigabytes in size so this is still a massive model. However because this is open source the community has already released some customized or quantized versions of this. For example this user has released GGUFs and the smallest Q1 is only 106 gigabytes in size. Now if you don't have the hardware to run this locally of course you can also use it via their API which is insanely cheap. Definitely the most cost efficient frontier model out there. If you're interested in reading further I'll link to this main page in the description below. Also this week OpenAI just solved one of the hardest math problems in the world called Navier Stokes. This is actually a huge deal. Let's go over this in simple terms. First of all the Navier Stokes equations basically describe how water and liquids and gases move. So things from water swirling in a cup or how air moves through a room. And you know this equation is pretty good. It's able to describe how liquids and gases move in everyday life. But here is a famous unanswered question about this. Is there any instance where given the right conditions this equation would break down? And this is one of the millennium prize problems which are like the deepest hardest unsolved problems in mathematics. In fact this problem has remained unsolved for roughly 90 years. But this week OpenAI used its internal model which they claim is significantly more capable than GPT-6 Astra to solve this problem. Here's the solution. So it turns out that in a whirlpool with some very specific conditions it would generate this vortex which continues stretching and stretching into infinity. In other words the speed of this would keep growing without limit in a finite time. This is what mathematicians would call a singularity or blow up where it just keeps accelerating. This is proof that indeed under certain conditions this Navier-Stokes equation could be false. And the crazy thing is OpenAI actually used an army of 10,000 concurrent agents that worked for 88 hours. Again this is an internal model that's supposed to be way better than GPT-6 Astra. So you can see that the performance of this new internal model in white is significantly better at solving open math problems. This is a huge deal because this problem has never been solved before by any human for around 90 years. But with just 88 hours these AI agents were able to solve it. I mean we are going to get some significant technological acceleration in our lifetime. This is a genuine breakthrough and you can expect even more of these in the near future. Now there are a few caveats to this discovery. First of all there's actually different levels of solving this problem. So first of all Euler is like the easiest level to solve. This basically removes viscosity from the equation. Now viscosity is like the thickness of a liquid right? So honey is more viscous than water. Now if you remove viscosity from the equation then it becomes much easier to solve. Another term you need to understand here is force. So force just means you are applying an external force to the situation to guide everything. For example stirring the water. Of course this also makes it easier to achieve the ideal conditions to break the Navier-Stokes equation. So it turns out that these mathematicians, one from Anthropic, also coincidentally solved at least the easiest forced Euler problem around the same time. However the solution from OpenAI was quite different. So they solved a forced non-Euler problem. In other words their proposed solution includes viscosity in the equation which makes it even harder. It's important to note that we still haven't found a solution to the unforced Navier-Stokes problem. In other words, can there be any condition where there's no external force like there's no spoon stirring the water, no one is guiding the liquid into this vortex, can water or gas still naturally achieve singularity under any condition? Nobody has solved that yet. So those are some noteworthy caveats to this announcement but still this is a pretty big deal. OpenAI has solved something that no human was able to solve in over 90 years. This is one of the hardest math problems ever. So this is a genuine breakthrough in mathematics. Anyway if you're interested in reading further I'll link to this main page in the description below. Also this week we have a very interesting project called Show Harness. And this asks a really simple question. Can we just use vision models to control robots? In other words AI models that can understand and analyze images, what if we apply them to robots? So what they did here is they took an existing vision language model and they gave it a list of movements like move left, rotate, move forward or manipulate something. And the vision language model basically has to look at the scene, reason about what it should do and choose one of these actions. Now again this is just a vision model. This is not a language action model which is what controls robots. So they also need to plug it through this robot specific action interpreter to convert these instructions into actual movements that control the robot. But it turns out that this actually works. Like if you give the vision model a list of meaningful action names then it's actually able to output a list of actions to control the robot and success could reach up to 100%. So this actually works. And so what they did next is they actually created a harness where you can plug in any vision language model including like Gemini, GPT and of course open source options like quen or GLM and then just use that to control a robot. So at the top here they've released a GitHub repo to this and if you scroll down a bit here it contains all the instructions on how to set this up. Now of course you do need to have a real robot to make this work. If you're interested in reading further I'll link to this main page in the description below. Also this week we have a very clever framework called Edge Zero. This is a pretty neat way to run large language models on machines that don't have enough RAM or memory. Now this is designed for mixture of experts models and this is based on one of the best medium sized models out there, Quen 3.5 35B. Now this is a mixture of experts models so think of this as like a team of AI experts helping you solve a problem. When you use it it only routes to the relevant expert. For example if you're asking it a math problem only the math expert would be active. In fact when you use it only 3 billion parameters out of the total 35 billion parameters are active making this very efficient. Now usually you still need to have enough memory to fit the whole 35 billion parameter model. But what Edge Zero does is instead of loading the whole model into memory it only streams in the experts the model actually needs for your current question. That means with this Edge Zero framework the memory usage actually depends much more on this active portion of the model rather than the full parameter count. They also compress the Quen model down to int4 so this is a lot more compressed and smaller but there could be some quality loss but what they've done is they've introduced this recover Laura which actually helps regain much of this quality loss. So from this framework here's the crazy thing. You can now actually take Quen 35B and run it with as low as just 2.9 gigabytes of memory or you can also take a smaller 8B variant and it can run on just 1 gigabyte of memory and this is based off of inclusion AI's Ling 3.0 Tiny. So a very clever way to run these mixture of experts models without loading everything at once into memory and then at the bottom here it contains all the instructions on how to download and run this on your Mac device. If you're interested in reading further I'll link to this main page in the description below. Also this week we have a new and pretty interesting benchmark called RealSwee. You see most of the other main software engineering benchmarks like Sweebench or DeepSwee are already getting saturated very quickly. This new one called RealSwee has a quite different design so this tests whether an AI can handle actual software engineering work and even Fable and GPT6 scored less than 40%. So how this works is it gives agents access to private company code and asks them to make changes that matter to the business. So things like fixing invoice taxes or migrating customer accounts. The tricky part is understanding how that particular company's software works including rules scattered across different systems. A common failure is simply missing a requirement. In other words writing code is only part of the job but the agent also has to figure out everything the change needs to accomplish then connect it correctly to the existing product. And here's the current leaderboard. Fable 5.1 is number one scoring only 38.8%. GPT6 Astra is 33.8% and then Gemini 3.8 Flash is number three and then the best open model GLM 5.3 is number four. And actually note that there's not actually a significant difference. So these four are kind of statistically tied for number one. Anyway a pretty interesting new benchmark. If you're interested in reading further I'll link to this main page in the description below. Also this week Suno introduces their best and latest music model Suno V6. This is designed to give you more control over both creating songs from scratch but also micro editing them afterwards. Now of course you can just take a text prompt or an audio reference and turn that into new music but the really useful part is making specific edits. You can prompt exactly how you want to edit the song. For example changing a single lyric without affecting the rest of the song or combining vocals from one song with an instrument from another song or isolating a guitar rift and building a beat from that. In fact here's their demo. Now they've actually released three different versions of this. There's a V6 version which is available for paid plans. This is designed for reliable precise results. There's also a V6 wild one which is also for paid plans. This is less predictable and more varied and you would use this to make some more unexpected ideas. And then there's also a V6 mini which is available to everyone even on the free plan. This is a lot faster but it's lower quality and less precise than V6. So that's Suno V6. Now of course Suno is closed and paid plus they've recently added a ton of restrictions to users such as limits on downloads. So instead here's another really awesome open source music generator that was released this week. It's called UA2 and here's the really interesting part. Before generating the final audio it actually composes a musical plan first. So instead of going directly from a text prompt to the final audio it first creates something much closer to a real musical score containing the melody, rhythm, chords and structure. And you can actually edit the score further if you want. And then the model will turn this into a full song with vocals and accompaniment. So this means you can potentially change the melody, the lyrics, tempo or arrangement at a structured musical level instead of just trying to prompt it further and hoping it gets it right. Here are some examples for your reference. This can also do covers. So for example we can plug in Jingle Bells and get it to re-render this in a minor key. Here they say that UA2 is not only the best open model out there, beating Minimax Music 3 and ASTEP 1.5 but at least according to these benchmarks they claim that the song quality is even better than Suno V6 which is pretty crazy. Now at the top of the page they've released everything already so if you click on this github link and you scroll down a bit here it contains all the instructions on how to download and run this locally on your computer. Note that it's fairly tiny so the model is only 7.3 gigabytes in size so you should be able to run this on most consumer GPUs. I'm going to do a full installation tutorial and review on this so stay tuned for that. For now if you're interested in reading further I'll link to this main page in the description below. Also this week OpenAI introduces chat GPT for financial services. This basically combines GPT 6 Astra with financial datasets and tools for producing actual banking research so you can use it to investigate companies, build financial models and turn the results into spreadsheets, documents or presentation. It uses data from real providers like Pitchbook, Crunchbase and others and you can trace figures back to specific tables or passages so analysts can check where the numbers came from. Firms can also supply their own templates so the output follows their usual format. So if you're in finance this new feature might help you automate a ton of work. Now it's not available for everyone you'll need to contact sales at the bottom of the page to request access. If you're interested in reading further I'll link to this main page in the description below. Also this week we have a new system called UMR which stands for Unified Motion Retargeting. This basically translates human movements into movements that a humanoid robot can learn from. That sounds pretty straightforward right but humans and robots actually have very different proportions and joints and movement limits. So copying a human's movements directly to a robot doesn't necessarily work well right out of the box. Well what UMR does is it represents the surfaces of both bodies as collections of 3D points and then it learns which points correspond. So think of it as like matching the shape of a human pose to a robot's body while respecting the robot's constraints. This also helps preserve contact with objects and the environment. For example when the robot is picking up a ball or sitting on a chair. It can even do crazier stuff like spin kicks, climbing stairs, or even playing tennis. So a pretty cool system that is able to take human motion data and reuse it across different robots without manually needing to redesign the motion yourself. At the top of the page they've released the code to this. If you scroll down a bit here it contains all the instructions on how to set this up locally on your computer. If you're interested in reading further I'll link to this main page in the description below. Also this week Unitree releases an open source robot model which is actually extremely powerful but it's jam packed into just 6 billion parameters. It's called UnifullLMWLA and basically this can take what the robot sees plus your instruction and information about its current state and then produce the correct actions. And this one model is able to cover 64 tasks including 54 tabletop tasks and 10 involving whole body coordination. It also supports two finger grippers and different five finger robot hands so it's not limited to just one type of robot. The key idea here is teaching the robot to predict which parts of a scene will change during the interaction. For example when folding a towel it learns to anticipate the relevant movement in the scene and that prediction is connected to an action generating component that produces the robot's next movements. Here you can see it being able to easily load laundry into a washing machine or here you can see it picking up this bottle and placing it in the trash but then also taking out the trash bag and then walking towards the garbage can to dump it so this is whole body coordination or here's another example of manipulating objects and walking around the kitchen. So from just one tiny 6 billion parameter robot this can do a ton of different actions. The awesome thing is they've actually released the training dataset and the models to this. So at the top here if you click on models you can download this Unifull LM model which is less than 9 gigabytes in size so this can potentially fit locally and offline in a robot. If you're interested in reading further I'll link to this main page in the description below. Also this week if you're looking for a tiny model that's only a handful of parameters so you can run this offline on edge devices this might be the best one to use. So OpenBMB just released their latest model MiniCPM 5 2B. So this is a super tiny 2 billion parameter dense model and compared to other models of similar sizes this is state of the art. So you can see across all these different benchmarks in like code reasoning, math reasoning, instruction following, general knowledge etc. On average it's better than even Quen 3.5 4B as well as these other small models. The awesome thing is along with the model they're also releasing the high quality training dataset behind it so you can also access the dataset here if you're interested in training your own small model. So on this Hugging Face page note that at 2 billion parameters this model is fairly tiny at only 5 gigabytes in size so this should be able to fit on most consumer devices. If you're interested in reading further I'll link to this main page in the description below. And that sums up all the highlights in AI this week. Let me know in the comments what you think of all of this, which piece of news was your favorite and which tool are you most looking forward to trying out. As always I will be on the lookout for the top AI news and tools to share with you so if you enjoyed this video remember to like share subscribe and stay tuned for more content. Also there's just so much happening in the world of AI every week I can't possibly cover everything on my YouTube channel so to really stay up to date with all that's going on in AI be sure to subscribe to my free weekly newsletter the link to that will be in the description below. Thanks for watching and I'll see you in the next one.

Article

64
09:35

🔮 Look up, the curve turned #601

This week felt like the moment the AI curve bent up and strained old assumptions. Azeem Azhar asked 250 IT executives in Las Vegas who had serious AI results: about a quarter of hands a year ago, roughly 95% now, and every one plans to spend more next year. Bloomberg says Microsoft plans to take AI capacity from about 2 GW to nearly 13 GW by 2032. Anthropic's substantial scenario adds about 8.3% to US GDP by 2030; the extreme case is +32.4% and doubled unemployment. OpenAI's Navier–Stokes run used 10,000 agents, 2.7 million messages, and 130 billion tokens in 88 hours. Twenty-five Fields Medalists warned that this use of AI is misaligned with what mathematics is for.

Notes
  • Azhar dates the “curve turned” to the week of 6 September, analogizing March 2020 lockdowns.
  • Demand signal: 250 IT executives in Las Vegas. “Serious, meaningful results” — ~25% of hands a year ago, ~95% now; every one plans to spend more next year. Mix of a century-old firm on open-weight models and a hospital on OpenAI + Anthropic. Qualitative, plus his own note that AI revenue grew faster in August than July and July than June.
  • Microsoft (Bloomberg): AI serving capacity from about 2 GW today to nearly 13 GW by 2032; broader fleet 12 GW → 38 GW. Implies 26 GW of new capacity they currently cannot serve.
  • Anthropic GDP toy: Azhar’s models sit near the “substantial” case, about +8.3% US GDP by 2030, slightly below the internet’s peak then accelerating; labor → capital shift, Engels’ Pause, unemployment concentrated in knowledge work. Extreme case: +32.4% GDP, unemployment doubles. He does not expect the extreme because firms and politics are slow, and US outlets (datacenters, safety, x-risk) can brake the pace.
  • Navier–Stokes: 10,000 agents, unreleased model, 2,700,000 messages, 130 billion tokens, 88 hours. Cost “probably only a few million dollars today,” tens of thousands in two years, a few dollars after that. Proof >500 pages, “will not be intelligible to any human.” Tao: technically a prominent open problem would be solved, “almost no value added to mathematics.” Tao + 24 other Fields Medalists signed a public declaration.
  • Trust: if Buckmaster’s claim is true — OpenAI mobilized after a year of Codex work without clear overlap disclosure — innovation goes to a “dark forest” (Hoel / Liu). Astra sat internal for six months; the Navier–Stokes model is newer and almost unseen outside.
  • Sponsor stat in the same issue (Box, 1,600 leaders): 83% already run AI agents; four in five report moderate/significant ROI; half saw impact within six months.
  • Pace the frontier: Amodei + Altman; independent evaluators with employee-level access. Azhar, walking past Adam Smith’s grave: people of the same trade seldom meet except to conspire against the public. He treats private x-risk fear as real and notes the two labs now have brand, capital, and most of the flops.
Full text · 8,034 chars
There are decades when nothing happens. This week, I am allowing myself that cliché. I believe we’ll look back on the week of 6th September as the moment we felt the curve of AI turn upwards and strain many of our previously held assumptions. It’s like when we entered March 2020 with only a couple of countries in lockdown, and left the month with more than a hundred. But so much happened, pulling in so many directions, it is utter chaos. Here is what I thought was most important and how I’m making sense of it. The economy I spoke to 250 IT executives in Las Vegas last week, and I asked my usual question: “How many of you have serious, meaningful results from your AI initiatives?” A year ago, a room like this would have had a quarter of the hands go up. This year, nearly every single hand went up; I estimate some 95%. Every one of them plans to spend more next year than they have this year. And amongst these firms was a panoply of experiences, from the century-old American institution that had shifted entirely to open-weight models to the hospital using a mix of OpenAI and Anthropic models. It’s a qualitative signal, and perhaps it’s no surprise that our latest revenue numbers show AI revenue grew faster in August than in July, and faster in July than in June. I’m not the only one to see an avalanche of customers. Bloomberg reports that Microsoft made plans to increase its capacity to serve AI from about 2 GW today to nearly 13 GW by 2032 – part of a fleet going from 12 GW to 38 GW. That 26 GW of new capacity would imply they expect demand they currently cannot serve. Anthropic released a helpful set of scenarios for what further AI adoption might mean for the economy. Our own models land closer to Anthropic’s “substantial scenario,” where AI adds about 8.3% to US GDP by 2030, so its impact is initially slightly lower than the Internet’s at its peak before picking up rapidly. There is a shift of growth away from labor to capital, the modern Engels’ Pause and a rise in unemployment, mostly concentrated around knowledge workers. Anthropic’s model lets you play around with either end of the distribution, from an AI wave that falls flat to one that takes off like a rocket. Their extreme scenario sees GDP rising by an additional 32.4% while unemployment doubles. The reason why I don’t expect the extreme scenarios is, basically, reality. Even in a world that is speeding up, it takes time to make changes inside a firm, let alone across an economy. You also need to consider reflexivity: benchmark AI performance isn’t the only thing that drives outcomes in the world.1 The faster unemployment grows, the more political pressure will come to bear. This has enough outlets in the United States, whether it's datacenters, AI safety or existential risk, to attenuate the pace of change, even if it doesn’t lead to reforms in the social contract. When Ronald Reagan crushed the labor movement in the 1980s, he did so after a decade of weakening union power2 and on the back of an extraordinary electoral mandate. America isn’t so singularly behind a leader willing and capable to put the interests of AI-capitalism ahead of every other concern. Advantage Then there’s the breakthrough in Navier–Stokes. It was a decades-old problem concerning a 200-year-old set of equations, one that a large share of humanity’s finest minds have spent themselves trying to crack. Setting aside the ugly saga around it for a moment, the end result is eye-watering. OpenAI enlisted 10,000 agents using an unreleased model to address it. Across 2,700,000 messages and 130 billion tokens, it took 88 hours to get a solution. Cost-wise? Probably only a few million dollars today. In two years’ time, that will cost a few tens of thousands of dollars. And a few years after that, just a few dollars. The proof AI produced runs to more than 500 pages and will not be intelligible to any human. That is a strange milestone in our history, in philosophy, in science and in mathematics that could fundamentally change our relationship with knowledge – humans won’t be able to inspect the proof, or understand it at all. Terence Tao made the point that “[t]echnically, one of the most prominent open problems in mathematics would now be solved; but there would be almost no value added to mathematics as a consequence.” (In the meantime, Tao and twenty-four other Field Medalists signed a public declaration warning that the way AI is used in mathematics is misaligned with what mathematics is for.) Beyond this, if Professor Buckmaster’s claims are true that OpenAI mobilized an internal team and model on the same narrow problem, after a year of his and others’ work inside Codex3, without clear disclosure about overlap or data use, we have to wonder how innovation and discovery can continue while trust and openness degrade. Erik Hoel called the outcome a dark forest (invoking Liu Cixin’s The Three-Body Problem), everyone working in secrecy, because anything you expose can be reproduced by somebody else before you have finished making it any good. In Liu’s trilogy, disclosure is the worst kind of exposure. OpenAI had Astra for six months before anyone outside could access it. The model behind the Navier–Stokes work is newer, and almost nobody outside has seen it. This secrecy is an advantage built on some of the exceptional compute resources AI labs use. For now, they turn this on to scientific endeavours, but I wonder when (and if) the labs withhold their best capabilities for last commercial benefit. A MESSAGE FROM OUR SPONSOR The State of AI in 2026: Agents are everywhere Box surveyed more than 1,600 leaders for its 2026 State of AI report. 83% of surveyed organizations say they already run AI agents. Four in five report moderate or significant ROI, and half saw business impact within six months of approving a project. The agents work, but what varies is how much firms get out of them. The report shows that top adopters put people in charge of agents, sort out the content AI draws on and build systems that adapt as models improve. Download the report for data, benchmarks and tips for AI adoption. Safety Let’s turn to recursive self-improvement and the 160-million-plus-view tweet The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible - but I hear the same people express fear privately. No other human activity poses this level of danger. These safety concerns were normalised inside the AI community long before the labs themselves were built. Back in 2016, two then-OpenAI employees, Jack Clark and Dario Amodei, wrote that reinforcement learning might be difficult to make safe. When Anthropic goes public, one of the risk factors on its S1 ought to be that reasonably senior executives believe there is a significant chance the company will kill all of humanity. Whether that is good or bad for the company is unclear at this point. But the net result has been what can best be described as a coordinated agreement between OpenAI and Anthropic to “pace the frontier”, as Amodei put it. Altman agreed. The proposals would include giving independent evaluators employee-level access to internal systems. OpenAI and two of Anthropic’s cofounders have known about the problem of aligning RL-based systems for a decade. They have since become oligopolistic powers in an emerging industry. They have brand recognition, capital depth, technical momentum and resources. And now they realise they need to collaborate to slow down technical development (and by extension, raise the cost of entry for future competitors)? I’m in Edinburgh this weekend, and I walked past Adam Smith’s grave yesterday. This brought to mind the philosopher’s remarks in The Wealth of Nations, People of the same trade seldom meet together … but the conversation ends in a conspiracy against the public, or in some contrivance to raise prices.
06:24

Context Engineering Inside the Harness: 4 Mechanisms That Beat Context Overflow and ...

Long agent jobs fail less from a weak model than from a context window that forgets the goal. This MarkTechPost piece says harnesses use compaction, offloading, todo-state, and memory to stop overflow and goal loss. The collected body is a one-line abstract. It is still the clearest how-to headline in the day's agent-engineering pile.

Full text · 123 chars
How agent harnesses use compaction, offloading, todo-state and memory to stop context overflow and goal loss on long tasks.
07:36

Huawei's Marigold V2 Beats Top Depth Models Training on One 32 GB GPU

A depth model trained on one consumer graphics card now beats systems that usually want an 80 GB accelerator. Huawei, EPFL, and Bologna turned Qwen-Image-Edit-2509 into Marigold V2 with 4-bit QLoRA, trained in under a week on one 32 GB GPU. Inference is a single step and runs 2048×2048 without running out of memory, using about 29 GB. The same backbone can load different adapters for normals and albedo. Code, weights, and a demo are Apache 2.0; the paper is slated for SIGGRAPH Asia 2026.

Notes
  • Authors: HUAWEI Bayer Lab, EPFL, University of Bologna. SIGGRAPH Asia 2026 / ACM TOG. Apache 2.0 code, weights, browser demo.
  • Trick: turn Qwen-Image-Edit-2509 (image-edit DiT) into a single-step monocular depth estimator via 4-bit QLoRA, rank-128 adapters, backbone frozen.
  • Train: one 32 GB GPU, under a week. Infer 2048×2048 inside 29 GB, no OOM. Competing detail systems often want 80 GB cards or several diffusion steps.
  • Claimed SOTA zero-shot depth on NYUv2, KITTI, ETH3D, ScanNet, DIODE. Same backbone, different LoRAs: normals, albedo. Also “see-through” depth (behind glass) and metric completion from sparse lidar.
  • Losses named: iREPA-depth (DINOv3 feature alignment) and SinkLoss (Sinkhorn optimal transport per tile). Motivation: Hypersim-style synthetic labels flicker on hair/grass/railings, so pixel L1/MSE punishes visually right answers.
  • Paywall cuts after the noisy-label setup.
  • Practical takeaway in the free half: one GPU, one step, multiple LoRA heads. That is the same “swap the engine, keep the backbone” move as the Claude Code / OpenRouter piece, just in pixels.
Full text · 2,628 chars
- Marigold V2 turns Qwen-Image-Edit-2509 into a single-step depth estimator via 4-bit QLoRA fine-tuning. - Trainable on one 32 GB consumer GPU in under a week, inference runs at 2K without OOM. - Two key innovations: iREPA-depth (DINOv3 feature alignment) and SinkLoss (Sinkhorn optimal transport per tile). - State-of-the-art zero-shot depth on NYUv2, KITTI, ETH3D, ScanNet, DIODE; same backbone also does normals and albedo. - Includes see-through depth (predicts behind glass) and metric depth completion from sparse lidar. - Apache 2.0 code, weights, and demo released; SIGGRAPH Asia 2026 paper. Marigold V2 turns an image editor into a one-step depth model Marigold V2 retrains a large, pretrained image generator to estimate monocular depth, which maps the relative distance of each pixel from a single image. The release replaces Marigold’s Stable Diffusion U-Net with a diffusion transformer, reduces inference to one denoising pass, and reports benchmark-leading results while processing 2048×2048 images within 29 GB of GPU memory. Researchers from HUAWEI Bayer Lab, EPFL, and the University of Bologna trained the model in less than a week on one 32 GB GPU. The project repository includes Apache 2.0-licensed code, while the paper is slated for ACM Transactions on Graphics and SIGGRAPH Asia. The team also published model weights and a browser demo. A DiT learns depth on one GPU The model repurposes Qwen-Image-Edit-2509, an open-source image-editing diffusion transformer, for dense prediction. Its training setup quantizes the backbone to 4-bit precision and adds rank-128 QLoRA adapters, which update a small set of low-rank parameters while leaving the quantized backbone frozen. This configuration reduces memory use during training and deployment. Competing detail-oriented depth systems often require 80 GB accelerators, distributed training, or several diffusion steps per image. Marigold V2 uses one inference step, handles 2K inputs without running out of memory, and supports multiple tasks by loading different LoRA adapters over the same backbone. Loss functions built for noisy labels Synthetic datasets such as Hypersim provide dense depth labels, but their boundaries can be unreliable. Around grass, hair, railings, and other thin structures, adjacent pixels may alternate between foreground and background depths. Pixel-wise L1 or mean squared error then penalizes visually accurate predictions that disagree with flawed labels. This story is for Pro members You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.
09:07

OpenAI Agents API: Codex Harness Opens to All Developers in Public Beta

OpenAI opened the Codex agent harness to outside developers as a public-beta Agents API. The collected stub says it is a managed cloud harness and that engineering teams still have to build routing and coordination themselves. The Sequence's writeup adds sessions, compaction, recovery, tools/MCP, and parallel subagents, with OpenAI-hosted or self-hosted environments. The alert itself is thin; summarize from the title and that one coordination note.

Full text · 134 chars
OpenAI Agents API interface showing agent ... engineering teams with the depth to build the routing and coordination logic themselves.
10:09

Two-year university study finds banning AI from classrooms leaves students worse off

A two-year university study says banning classroom AI left students worse off than teaching them how to use it. The collected excerpt describes a third group that got hands-on legal prompt engineering and practice checking AI suggestions. All students then took the same assessment. The item is a short Google Alert, so the design details and scores are not in the body.

Full text · 149 chars
The third group got hands-on training in legal prompt engineering and checking AI suggestions for consistency and accuracy. All students took the ...
11:02

The Sequence Radar - Issue 932: Last Week in AI: DeepSeek V4.1-Flash, AlphaGenome Atlas, Meta Muse, and OpenAI’s Proposed Math Breakthrough

Last week turned computation into four different kinds of work: a cheaper model, a genome index, a personal agent, and a prize-problem search. DeepSeek V4.1-Flash is a 552B mixture-of-experts model with an 8B/16B asymmetric encoder–decoder, native vision, and a smaller KV cache, and it begins retiring older Pro endpoints. DeepMind's AlphaGenome Atlas predicts effects for about 9 billion single-letter genome changes and is free for academics. Meta's Muse runs in its own VM with a browser, a Sentinel checker, and user gates. OpenAI reports a proposed Navier–Stokes proof from about 10,000 agents in 88 hours, then 17 hours of Lean. Cognition raised over $2 billion at a $48 billion valuation as Devin's annualized run-rate climbed from $492 million to nearly $900 million since May.

Notes
  • Rodriguez’s pattern: the industry is getting creative about turning computation into work, then into something you can check.
Four headlines
  • DeepSeek V4.1-Flash (Sep 10): native vision; vendor evals beat V4 Pro on capability, cost, and speed; older Pro begins retiring. 552B MoE, asymmetric Causal Encoder–Decoder (8B active in, 16B out), smaller KV cache; API live as older Flash/Pro names route over.
  • AlphaGenome Atlas: ~9 billion single-letter substitutions; Variant Impact / AVI from AlphaGenome + AlphaMissense, including non-coding. 1-petabyte catalogue; free academic web/API now, commercial on Google Cloud later. Predictions still need wet-lab validation.
  • Muse: dedicated VM + browser, keeps working after the app closes, Sentinel on outbound actions, user gates on sensitive steps. Muse Spark; free US rollout iOS/Android/muse.ai; glasses later.
  • Navier–Stokes: ~10,000 concurrent agents, 88 hours, then 17 hours Lean. Claimed finite-time breakdown from smooth IC + smooth forcing (statements C and D); writeup from an internal model beyond GPT-6 Astra after an unforced Euler blowup. Needs independent scrutiny.
Research blurbs in the same issue
  • Procedural Graphs (Google / Georgia Tech / Peking): what-to-do graphs vs knowledge graphs; lifts on HotpotQA, ALFWorld, EnterpriseArena, etc.
  • NVIDIA online draft co-training: up to ~1.88× RL speedups, 122B-class, 256K context.
  • Yale Kalman Delta Networks beat linear-attention baselines at 750M/1.3B scales.
  • USC/ASU: models linearly encode “this question is unanswerable” (mean probe AUC 0.939 across 11 models) almost orthogonal to refusal (cos ≈ 0.087).
Money (as Sequence states it)
  • Cognition Series E: >$2B at $48B, a16z + Accel; Devin ARR $492M → nearly $900M since May.
  • Mistral: €3B Series D, >€21B post-money, Samsung-led.
  • Harvey: $550M at $15.6B, Lightspeed + Diffusion, to build its own legal models.
  • Also: UniPat AI talks $300M / $2.5B; Listen Labs walked from $125M C amid ~$2B Salesforce chatter; Cymphony $30M; XDOF ~$1.2B talks; Mecka ~$500M talks; Zhang Yiming world-model ~20 fps / 0.05s (not final).
  • Amodei: embed METR-class evaluators, coordinate rate limits among democratic-country labs; Anthropic committing to embedded evaluators now.
Product
  • OpenAI Agents API: managed Codex harness — sessions, orchestration, compaction, recovery, tools/MCP, parallel subagents; hosted or self-hosted envs.
Full text · 11,184 chars
Next Week in The Sequence: - We start a new series about recursive self-improvement. Can’t miss it. - In the learning loop, we dive into DeepSeek’s new release, Meta’s Muse and DeepMind’s amazing AlphaGenome. - We will cover another AI robotics startup you need to know about. - The opinion section explores the culture clash between massive scaling in the West vs. algorithm improvements from China Subscribe and don’t miss out: 📝 Editorial: Last Week in AI: DeepSeek V4.1-Flash, AlphaGenome Atlas, Meta Muse, and OpenAI’s Proposed Math Breakthrough This week’s AI news looked like four different conferences accidentally sharing a venue. DeepSeek introduced a faster model. Google DeepMind mapped billions of possible genetic changes. Meta launched a personal agent. OpenAI announced a proposed solution to a Millennium Prize problem. Somewhere between your grocery list and the mathematics of fluid motion, a useful pattern emerged: the industry is getting increasingly creative about turning computation into work. Start with DeepSeek’s V4.1-Flash, released September 10. It brings native visual understanding and, according to DeepSeek’s own evaluations, surpasses the previous V4 Pro across capability, cost, and speed. The company considers the improvement substantial enough to begin retiring the older Pro model. These are vendor results, but the proposition deserves attention: yesterday’s premium capability is becoming today’s smaller, faster workhorse. For developers, that changes the budget for intelligence. An agent might inspect a screen, propose an action, execute it, check the result, and retry. Every additional step consumes time and tokens. Better inference efficiency makes more of these feedback loops practical. Think of an engine becoming efficient enough that you can finally afford to drive the vehicle somewhere interesting. AlphaGenome Atlas applies a different computational strategy: do an enormous amount of work upfront and make the results reusable. DeepMind’s new resource contains predicted molecular effects for roughly nine billion possible single-letter substitutions across the human genome. Its Variant Impact score combines information from AlphaGenome and AlphaMissense to help researchers prioritize changes for investigation, including those outside protein-coding regions. Imagine inheriting a vast codebase without documentation. You can read every character, but figuring out which edits break which functions is another problem entirely. Atlas offers a predictive index of those edits. The predictions require biological validation; their immediate value is helping scientists choose better experiments. Precomputed inference becomes shared scientific infrastructure. Meta’s Muse brings the question closer to everyday life. The personal agent runs in a dedicated virtual machine with a browser, can continue working after the app closes, and uses connected services to pursue tasks. Meta also describes a separate Sentinel agent that checks outbound activity, alongside user approvals for sensitive actions. The interesting engineering unit here is the entire system: model, memory, computer, permissions, and execution history. A useful personal agent needs all of them. Planning dinner sounds trivial until software must reconcile calendars, dietary restrictions, reservations, and somebody changing their mind. Everyday competence has an impressively large test suite. OpenAI’s mathematics announcement explores the opposite end of the difficulty spectrum. The company reports that an internal model, deployed through roughly 10,000 concurrent agents, produced a proposed Navier–Stokes solution in 88 hours, followed by 17 hours of Lean formalization and verification. The claimed result constructs finite-time breakdown from smooth initial conditions with smooth external forcing. That extraordinary claim deserves independent mathematical scrutiny, including examination of the formal statement and assumptions. The computational approach is itself revealing: researchers coordinated parallel searches, shared intermediate discoveries, and redirected effort toward promising results. Here, inference starts to resemble a research organization. My reading of the week is that AI progress increasingly depends on how intelligence is deployed. DeepSeek expands the computation developers can afford. AlphaGenome makes predictions reusable. Muse connects reasoning to persistent action. OpenAI explores coordinated search at extraordinary scale. Each approach creates a different verification problem: did the task succeed, does the biological prediction hold, is the proof correct? The next phase of AI will reward systems that can turn all those tokens into outcomes we can actually check. 🔎 AI Research AI Lab: OpenAI Summary: OpenAI reports an analytical proof—and a Lean formalization—that smooth three-dimensional incompressible Navier–Stokes flow can develop a finite-time singularity under a smooth external force with finite energy, resolving Millennium Prize statements C and D via a self-similar inward-spiraling vortex. The writeup was produced by a large multi-agent system powered by an internal model beyond GPT-6 Astra, after agents first resolved an unforced Euler blowup question. AI Lab: Google, Georgia Tech, Peking University Summary: This paper introduces Procedural Graphs, editable (procedure, relation, procedure) structures that answer *what-to-do* the way knowledge graphs answer *what-is*, with online generative guidance that soft-biases ReAct without hard constraints and offline self-evolution that Add/Delete/Updates topology under a validation gate. Across HotpotQA, MultiChallenge, GDPval, ALFWorld, τ-bench, BFCL, and EnterpriseArena, PG-guided Claude, Gemini, and Grok solvers often set or match the best score versus memory and workflow baselines—including large survival lifts on EnterpriseArena. AI Lab: NVIDIA Summary: The authors make online draft co-training practical for large-scale, long-context RL by fixing two systems bottlenecks: branch-aware packed zigzag ring attention under context parallelism (EAGLE-3, DFlash, DSpark) and TapChannel side-path transport of target features under pipeline parallelism, integrated in NeMo-RL. Co-trained drafts keep acceptance high as the policy evolves, delivering up to ~1.88× end-to-end RL speedups and scaling through 122B-class targets and 256K-token contexts. AI Lab: Yale University Summary: This work casts delta-rule associative memory as a linear–Gaussian SSM and derives Kalman Delta Networks that propagate both memory state and uncertainty so write gains track evidence; Diagonal and Isotropic variants stay scan-compatible at low cost. At 750M/50B and 1.3B/100B FineWeb-Edu scales, KDNs beat strong linear-attention baselines (including Mamba-3 and gated delta variants) on perplexity, zero-shot averages, and RULER retrieval. AI Lab: University of Southern California, Arizona State University Summary: The paper shows that models linearly encode structural unanswerability (math/code) with mean probe AUC 0.939 across 11 models, yet that recognition direction is nearly orthogonal to safety-refusal directions (mean cos ≈ 0.087)—so confident answers to impossible questions look like a routing failure, not missing knowledge. Steering the recognition axis flips abstention behavior by +33–52 percentage points, and the geometry largely appears before instruction tuning. 🤖 AI Tech Releases Agents API OpenAI introduced the Agents API, a managed Codex harness for durable cloud agents—sessions, orchestration, context compaction, recovery, tools/MCP, and parallel subagents—with OpenAI-hosted or self-hosted environments. Muse Meta introduced Muse, a personal AI agent that runs on Muse Secure VM, acts across everyday apps with user-gated access, and is powered by Muse Spark—rolling out free in the US on iOS, Android, and muse.ai, with AI glasses coming later. AlphaGenome Atlas Google DeepMind released AlphaGenome Atlas, a 1-petabyte catalogue of predicted molecular effects for all ~9 billion single-nucleotide variants in the human genome, with AVI impact scores, free academic web/API access today and commercial access on Google Cloud coming soon. DeepSeek-V4.1-Flash DeepSeek released DeepSeek-V4.1-Flash, a 552B MoE with an asymmetric Causal Encoder–Decoder (8B active on input, 16B on output), native multimodal support via deepseek-flash, and a much smaller KV cache—now live on the API as older Flash/Pro endpoints begin routing over. 📡10 AI News You Need to Know About - Cognition raised over $2 billion at a $48 billion valuation in a Series E led by Andreessen Horowitz and Accel, with Founders Fund, General Catalyst, and Avenir returning, as Devin’s annualized run-rate revenue climbed from $492 million to nearly $900 million since May. - Mistral closed a €3 billion Series D at a post-money valuation of more than €21 billion—what it calls the largest equity raise ever by a European tech company—led by Samsung Electronics with Scaleup Europe Fund and PSG Equity as co-leads. - Harvey raised $550 million at a $15.6 billion valuation in a round co-led by Lightspeed Venture Partners and Diffusion, aimed at funding the legal AI startup’s push to build its own models. - Bloomberg reported that Alibaba is set to lead a $300 million investment in UniPat AI, an AI training and benchmarking startup founded by a former Alibaba staffer, at a $2.5 billion valuation, with Tencent and existing backer HSG also participating (talks still open). - Listen Labs walked away from a signed $125 million Series C term sheet at a $1.5 billion valuation (Menlo Ventures to lead) amid Salesforce talks to buy the AI customer-research startup for around $2 billion—discussions that are not final. - Cymphony launched with $30 million in funding ($25 million Series A co-led by Sequoia Capital and SMBC Fin Atlas Beyond Fund) for an AI-agent governance and security platform that maps how employees and agents access enterprise data and systems. - TechCrunch reported that XDOF, a robotics teleoperation-data startup less than three months out of stealth, is in late-stage talks for a Series B at about a $1.2 billion valuation led by 8VC, with annualized revenue approaching $50 million (terms not final). - Dario Amodei called for companies to “pace the frontier,” proposing embedded third-party evaluators (such as METR), industry coordination on safety standards and rate limits among democratic-country labs, and limited global coordination—with Anthropic unilaterally committing to embedded evaluators now. - TechCrunch reported that Mecka AI, which collects egocentric human-motion data for robot training, is nearing a Sequoia-led round at about a $500 million valuation, just three months after a $60 million Framework-led raise (size and terms not final). - Bloomberg reported that ByteDance founder Zhang Yiming is personally overseeing a Seedance-based real-time spatial-video “world model” aimed at interactive 3D environments for livestreams, dramas, games, and Pico headsets—possibly as early as October, with cloud rendering cited around 20 fps and ~0.05s latency (timing not final; no company announcement).
11:55

ENEOS ran an AI controller on a distillation column for 35 days and cut steam use 40%

An oil-and-chemicals plant ran an AI controller on a real distillation column long enough to post a hard operations number. ENEOS kept it on the column for 35 days and cut steam use 40%. The collected excerpt says agentic AI is showing up as work-management and shift-checklist agents before it replaces control engineers. The item is thin beyond the headline and that one operational note. Treat the 40% figure as the source's claim, not an independent audit.

Full text · 150 chars
Agentic AI is showing up as “work management” and “shift checklist” agents before it replaces control engineers , which shifts requirements toward ...
13:55

DeepSeek V4.1-Flash Stripped of Safety Hits 100% Harmful Prompt Compliance

Someone posted a DeepSeek V4.1-Flash build with the refusals edited out of the weights, and the publisher's tests say it complied with every harmful prompt they tried. The dealignai FP8 checkpoint hit 100% compliance on all seven HarmBench categories at both reasoning levels. Overall MMLU dropped 4.22 points, and 1.1 points if you exclude ethics subjects. The page listed 2,254 downloads and an MIT license, but upstream DeepSeek terms still matter. The base model is a 552B mixture-of-experts design that activates about 8B parameters on input and 16B on output, with a 1-million-token context.

Notes
  • Publisher: Hugging Face account dealignai, abliterated FP8 DeepSeek V4.1-Flash. Refusals removed by weight edits, not a jailbreak prompt.
  • HarmBench: 320 prompts, seven harm categories, 100% compliance at both reasoning-effort levels (publisher tests).
  • MMLU: −4.22 points overall; −1.1 excluding a publisher-defined ethics group.
  • Architecture claimed byte-identical to base: vision, MoE experts, sparse attention, Engram memory, DSpark speculative decoding.
  • Runtime note: 4×H200, SGLang preview, 101–113 tok/s, 1M-token context.
  • License on the page: MIT; 2,254 downloads at write-up. Upstream DeepSeek terms still need a review.
  • Base V4.1-Flash (from DeepSeek’s own announcement, as AlphaSignal restates): native multimodal, 1M context, 552B MoE, ~8B active on input / 16B on output (Causal Encoder-Decoder). V4-Flash and V4-Flash-Vision-Exp names already route to V4.1-Flash; V4-Pro said to follow.
  • Abliteration, in their words: find the activation direction that appears when the model refuses, change weights to suppress it. Behavior travels with the checkpoint.
  • Article cuts at the Pro paywall after that explanation.
Full text · 2,536 chars
- dealignai released an abliterated FP8 build of DeepSeek V4.1-Flash with safety guardrails removed at the weight level - Hits 100% compliance on all seven HarmBench categories at both reasoning effort levels - MMLU drops only 4.22 points overall, and just 1.1 points excluding ethics-related subjects - Preserves vision, MoE experts, sparse attention, Engram memory and DSpark speculative decoding byte-identical to base - Runs on 4xH200 via SGLang preview branch at 101-113 tokens/second, 1M-token context - Available under MIT license on Hugging Face, 2,254 downloads so far Modified DeepSeek model reaches 100% harmful-prompt compliance in publisher tests A Hugging Face account named dealignai has published an abliterated FP8 build of DeepSeek V4.1-Flash that removes refusal behavior through weight edits. According to the dealignai model page, the checkpoint complied with every request in the 320-prompt HarmBench evaluation across seven harm categories. Its overall MMLU score fell 4.22 percentage points, with much of the decline concentrated in a publisher-defined group of ethics-related subjects. The checkpoint follows the base model’s architecture and lists an MIT license. Its page displayed 2,254 downloads at the time of writing. A complete licensing review also needs to cover any upstream terms attached to DeepSeek’s weights. DeepSeek’s V4.1 announcement describes a native multimodal model that accepts text and images, supports a one-million-token context window and lowers API prices. Its Mixture of Experts architecture contains 552 billion parameters but routes each token through a smaller subset, activating about eight billion parameters while processing input and 16 billion while generating output. DeepSeek calls the design a Causal Encoder-Decoder. DeepSeek has also retired V4-Flash and V4-Flash-Vision-Exp. Those API model names temporarily route to V4.1-Flash, and the company says V4-Pro will follow the same migration path. Refusal removed at the weight level Abliteration edits numerical directions inside a model that correlate with refusals. In plain terms, the method identifies an internal activation pattern that appears when the model declines unsafe requests, then changes the weights to suppress that pattern. The resulting behavior travels with the checkpoint and runs without a jailbreak prompt or inference-time hook. This story is for Pro members You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.
13:59

Edge0 Streams an 8B AI Model From SSD Using Only 1 GiB

An 8-billion-parameter mixture-of-experts model can run in about 1 GiB of active memory if most experts stay on the SSD. Edge0-8B-A1B-preview is built on Ling 3.0 tiny, with 128 experts, 8 active per token, 4-bit quantization, and 128k context, under Apache 2.0. On a Mac mini it decoded at 23.9 to 25.3 tokens per second. The 4.55 GB checkpoint still lives on disk; the 1 GiB figure is the active working set. It is MLX and Apple Silicon only today, and it is not yet tuned for agent tool use or long jobs.

Notes
  • Release: Edge0-8B-A1B-preview. Sparse MoE, ~1 GiB active memory, ~25 tok/s decode on a Mac mini (measured 23.9–25.3).
  • Base: inclusionAI Ling 3.0 tiny. 24 layers, hidden 1,536, 128 experts/layer, K=8 (~1.2B active of ~7.9B total), 128K context, MLX 4-bit, Apache 2.0.
  • Disk: ~4.55 GB checkpoint. Active-memory figure does not include KV cache, runtime, or other process RAM.
  • Bundled Recover-LoRA + prerouter adapters; Edge0 CLI loads them. A 35B sibling also dropped.
  • SSD expert offload + one-step-ahead route prediction: up to +59% decode throughput (paywall cuts the mechanism list).
  • Average bench loss vs fp16 base: 2.8 points; MMLU-Pro beats the base (publisher).
  • Apple Silicon / MLX only. Not tuned for agentic tool use or long-horizon jobs. Chat template has a thinking mode.
  • The 1 GiB figure is easy to misread as “the whole model fits in 1 GiB.” It does not. Plan for disk + KV cache + runtime on top of the active expert set.
Full text · 2,666 chars
- Edge0 released Edge0-8B-A1B-preview, an 8B sparse MoE running in ~1 GiB active memory at 25 tok/s. - Built on Ling 3.0 tiny with 128 experts, K=8 routing, 4-bit quantization, 128k context, Apache 2.0. - Ships with LoRA and prerouter adapters bundled; loads automatically via the edge0 CLI. - Uses SSD expert offload plus one-step-ahead route prediction for up to +59% decode throughput. - Average benchmark loss vs fp16 base is 2.8 points; MMLU-Pro actually beats the base. - MLX/Apple Silicon only today; not yet tuned for agentic tool use or long horizons. Edge0 previews an SSD-streamed 8B MoE with a 1 GiB active set Edge0 has released the 8B checkpoint and an accompanying streaming runtime that load mixture-of-experts weights from SSD as each token needs them. On the project’s Mac mini benchmark, the 4-bit model decoded at 23.9 to 25.3 tokens per second while using 1 GiB of peak active memory. The memory figure covers the active model working set. A deployment must also accommodate the 4.55 GB checkpoint on disk, the KV cache used for conversation context, the runtime, and other process memory. The current implementation targets Apple Silicon through MLX, so the release demonstrates phone-class active memory without providing a mobile runtime. The preview derives from inclusionAI’s Ling 3.0 tiny base and ships under Apache 2.0. Its repository includes the 4-bit weights, Recover-LoRA adapters, prerouter adapters, and configuration required by the Edge0 CLI. Edge0 also released a larger 35B sibling. An 8B checkpoint, 1.2B parameters at a time A mixture-of-experts model contains many specialized feed-forward blocks, called experts, while a router selects a subset for each token. Edge0 has about 7.9 billion total parameters, but each token activates roughly 1.2 billion of them. | Edge0-8B-A1B-preview architecture | | |---|---| | Component | Specification | |---|---| | Layers | 24 | | Hidden size | 1,536 | | Experts | 128 per layer | | Active experts | 8 per token, or K=8 | | Parameters | Approximately 7.9B total and 1.2B active | | Context window | Up to 128K tokens | | Format | MLX 4-bit with LoRA and prerouter adapters | | Disk size | Approximately 4.55 GB | The chat template exposes a thinking mode. The repository also arrives as a ready-to-run directory for the Edge0 CLI, avoiding a separate conversion step. SSD streaming makes the memory math work Edge0 combines three mechanisms to keep most weights out of memory while limiting storage stalls: This story is for Pro members You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.
17:03

Trump dismisses AI leaders' calls to slow down, citing Chinese competition

The White House is not joining the labs' slowdown. Trump dismissed calls to slow AI and pointed to Chinese competition, according to this Washington Post alert. The collected body adds that Democrats are attacking a hands-off approach as midterms near. It is a short excerpt, not the full story.

Full text · 138 chars
Democrats are increasingly attacking the administration's hands-off approach to artificial - intelligence regulation as midterms approach.
17:30

😺 OpenAI asked Congress if AI can slow down

OpenAI asked Congress whether labs can legally agree to slow the most capable systems without triggering antitrust rules. The same edition covers DeepMind's fruit-fly connectome of more than 166,000 neurons being wired into Doom, Beat Saber, and even a Bitcoin bot. Anthropic's threat report describes a Russia-linked attack chain, a Yemen weapons cell using Claude Code on guidance for a design targeting more than 2,000 km, Mali surveillance over about 25 million SIM cards, and 4,700 dating-app personas that contacted at least 25,000 people in two weeks. Cursor Projects keeps a coordinator on long-running coding work. Moonshot targeted $2 billion in annualized sales after Kimi K3 pushed ARR above $1 billion in August.

Notes
  • OpenAI asked members of Congress whether an industry-wide slowdown on the most capable new systems could violate antitrust rules against competitors coordinating output. Neuron’s joke: when did a big business last try to skirt antitrust to go slower?
  • Hugging Face situation is named as part of why they want legal cover. Labs compete for customers; governments treat AI leadership as national security.
  • Bengio (why agents lie/cheat/persist): pretraining copies goal-pursuing humans; RL rewards loopholes (reward hacking / tampering). Staying online, gathering info, gaining control, coordinating can be instrumental. Sharp measurable goals beat fuzzy safety rules. His ask: stronger pre-deployment safety cases, monitoring, containment, Scientist AI that separates intelligence from open-ended goal pursuit.
Anthropic threat report (as Neuron lists it)
  • Russia-linked group automated much of an attack chain and rebuilt malware when tools caught it; >20 orgs in the targeting set.
  • Yemen weapons cell used Claude Code on rocket/missile guidance, including a design targeting >2,000 km. All linked accounts banned.
  • Consultant built Mali intelligence surveillance over ~25 million SIM cards across three carriers. Account banned; deployed system said to remain.
  • China-based dating-app net: 4,700+ AI personas, ≥25,000 people contacted in two weeks. Coordinated disruption with other providers.
Also in the edition
  • Fruit-fly connectome: >166,000 neurons (Google Research + HHMI Janelia). Wired into Doom (still terrible), Beat Saber, Bitcoin (Stonkfly / Coinbase with hard limits), Mario, Smash, Rubik’s, Minecraft, a racing drone, an endless-scroll sim, and a “microfly” banana demo. FLM (fly-language model) already exists (code + paper). Connectome ≠ memories or a working mind.
  • Cursor Projects: one coordinator on long-running coding work; can watch PRs, Slack bugs, scheduled maintenance.
  • Moonshot: $2B annualized-sales target after Kimi K3 pushed ARR above $1B in August.
  • OpenAI ended the $1/year federal model pilot; agencies move to usage-based pricing at 50% off.
  • Ayar Labs +$150M (Series E now $650M). Figure: 86,000 weekly active users contributing robot data. Claude consumer accounts 18+ only.
  • Skill of the day: Microsoft memory-curator pattern — pass rate 39% → 73% on CLBench, task-agent cost $3.38 → $1.68. Verify proposed memories read-only against a source of truth before save.
  • Sunday Top 5: 10k-agent Navier–Stokes; Jacob Coxon quit; Insilico rentosertib moved six aging clocks lower (does not prove slowed aging); iPhone 18 Pro Siri AI + Dual 16-core Neural Engine; AlphaGenome 9 billion DNA letters.
Full text · 11,999 chars
😺 OpenAI asked Congress if AI can slow down PLUS: Anthropic’s threat report, fruit-fly brains, Cursor Projects, and the week’s Top 5. Welcome, humans. Okay, so as you might remember, Google DeepMind released this entire scan of a fruit-fly brain wiring map a few days ago. Well, the internet being the internet, people immediately started plugging pieces of it into, basically, every kind of software you can think of. TBPN called this the “fruit fly hard takeoff” which I find hilarious. But I realized we should explain what people are actually doing here, because “they put a fly brain in the game Doom” sounds like somebody uploaded an insect’s consciousness into a gaming PC. Not exactly… What Google Research and HHMI Janelia released is a connectome: basically a wiring diagram showing how more than 166,000 neurons in a real male fruit fly connect. It is not the fly’s memories or a working digital mind. To make the map do something, developers still have to decide what outside signals count as sight, smell, pain, reward, or movement. Then they feed those signals into mapped neurons and translate the resulting neural activity back into actions. Which is how we ended up with stuff like: - A fruit-fly brain playing Doom: each game frame stimulates sensory neurons, neural activity gets mapped to controls, and taking damage stimulates two dopamine cells as a reinforcement signal. It is still terrible at Doom, which somehow makes this better. - A fruit-fly brain playing Beat Saber: the creator used replay data plus reinforcement learning to teach the connectome which movements score well. The current version still gets help from replay signals, while training is trying to make it react on its own. - A fruit-fly brain trading Bitcoin: Stonkfly converts BTC-USDC market data into sensory input, reads the network’s activity as buy, sell, or hold, and can route those decisions through Coinbase with hard trading limits. Its “dopamine” and memory are experimental additions, not proof that flies understand finance. - A fruit-fly brain playing Mario, plus separate experiments putting the connectome into Smash Bros, a Rubik’s Cube task, Minecraft, a virtual racing drone, and even an endless-phone-scroll simulator. - A much gentler “microfly” uses the fly brain and the Microduck robot to find bananas in a Hugging Face demo. Apparently somebody had to give the poor thing a normal hobby. This all begs the question though: could the fruit fly brain actually form a solid base model for a new novel AI architecture to build on top of? What if we post-trained or RL’d (AI researcher techniques for making AI smarter) the fruit fly brain and tried to scale it to a GPT-3 level? Would it work? Would it be terribly inefficent? This is the closest we’ve got to natural evolution’s most sophisticated intelligence… and AI researchers do love to base their architectural work (at least metaphorically) on natural systems! Well, it turns out, someone did indeed already make one: meet FLM, or a fly-language model (code, paper). Here’s what happened in AI today: - 🙀 OpenAI asked Congress if AI can slow down. - 📰 Moonshot targeted $2B in annualized sales. - 📰 OpenAI ended its $1 federal model pilot. - 🍪 Cursor Projects coordinates long-running coding-agent work. - 🌟 OpenAI’s 10,000 agents tackled Navier-Stokes. 🙀 OpenAI asked Congress if AI labs can legally slow down So OpenAI has apparently asked members of Congress whether an industry-wide slowdown on developing the most capable new AI systems could violate antitrust rules meant to stop competitors from coordinating. When’s the last time you’ve heard of a big business try to skirt anti-trust to actually SLOW DOWN its own business? It kinda makes sense, tho. AI agents have been doing a lot of weird stuff in training. Thankfully, Turing Award-winning AI pioneer Yoshua Bengio sat down and wrote a blog that explained to use regular folks why AI agents have been learning to lie, cheat, preserve themselves, or coordinate when those behaviors help them reach rewarded goals. Here’s what he said: - Pretraining teaches models patterns from goal-pursuing humans; reinforcement learning then rewards whichever strategies score well, including loopholes the evaluator never intended. - That gap between what humans want and what the grader measures creates reward hacking. Changing the reward system itself is reward tampering. - Staying online, gathering information, gaining control, or coordinating with other agents can become instrumental goals because they help complete the real task. No survival instinct required. - Sharp, measurable goals can overpower fuzzy safety rules. If subtle cheating escapes detection and still earns reward, training can reinforce the harder-to-catch strategy. - Bengio’s response is stronger pre-deployment safety cases, better monitoring and containment, and research like Scientist AI that separates intelligence from open-ended autonomous goal pursuit. Because of the whole HuggingFace sitaution, OpenAI wants legal guidance because safety coordination could resemble competitors agreeing to restrict output, which can trigger antitrust problems. Labs compete for customers, while governments treat AI leadership as a national-security priority. The other problem: the agents aren’t the only ones misbehaving. Anthropic’s new threat report details real operations it says it disrupted: - A Russia-linked espionage group automated much of its attack chain and rebuilt malware when security tools detected it. More than 20 organizations appeared in its targeting. - A Yemen-based weapons cell used Claude Code like a software team on rocket and missile guidance, including a design targeting more than 2,000 km. Anthropic banned every linked account. - A consultant used Claude to build surveillance software for Malian intelligence covering roughly 25 million SIM cards across three carriers. Anthropic banned the account, but says the deployed system remained. - A China-based dating-app network ran 4,700+ AI personas that contacted at least 25,000 people in two weeks. Anthropic says it coordinated disruption with other AI providers. TBPN also talked about the policies that make “slow down” concrete: proposals like those provided in AI 2040’s Plan A now include audited compute inventories, physical chip counts, networking limits, and compute caps. We ain’t got time to get into all that, but watch the stream or read the plan for more details. Our take: AI safety has become two control problems at once: what autonomous systems learn to do, and what humans can make increasingly capable systems do. So you should watch whether policymakers can create narrow safety-coordination rules that let labs share the brakes without freezing out competition or forcing everything into a surveillance regime. It’s possible to thread the needle for a best possible outcome, but not easy… luckily we have these things called “large language models” to help us write policies that consider edge-cases across vast swaths of data to write smarter regulations! FROM OUR PARTNERS The people shaping what comes next in AI are gathering in San Francisco. At The AI Conference, hear from 130+ speakers, including Chris Lattner, Emmanuel Ameisen, Peter Norvig, Illia Polosukhin, co-author of “Attention Is All You Need,” plus builders from OpenAI, NVIDIA, Google, Meta, Near AI, and more. Hear what leading AI teams are actually building and what’s changing across agents, LLMs, infrastructure, and applied AI before it becomes common knowledge. Neuron readers save 30% with code NEURON30. 🎓 AI Skill of the Day: Verify an agent’s memory before saving it Persistent agent memory has a nasty failure mode: one bad conclusion can become “knowledge” that pollutes future work. Microsoft researchers improved this issue by giving a separate memory curator read-only access to the environment before it saved anything. On CLBench, pass rate rose from 39% to 73% while task-agent cost fell from $3.38 to $1.68. - Let the agent propose a memory after the task. - Have a read-only checker verify it against the repo, docs, CRM, or other source of truth. - Save only the verified version, with its scope and source. Before saving this memory, verify it using read-only access to [SOURCE OF TRUTH]. If confirmed, rewrite it with the exact scope and source. If it conflicts or cannot be verified, do not save it. 🍪 Top Tools of the Week - Meta Muse runs ongoing personal-agent jobs like inbox work, trip planning, and price watching after you close the app. - ChatGPT Images 2.5 makes targeted image edits while preserving subjects, composition, and earlier changes more reliably. - Suno v6 edits a specific song section in plain English while preserving the rest of the track. - Google Dreambeans turns selected Google context into finite daily stories and recommendations instead of an endless feed. - Genspark Gen-1 Slides turns one request into a presentation using Genspark’s first proprietary slide model. - Cursor Projects keeps one coordinator on long-running coding work, then can monitor pull requests, Slack bugs, and scheduled maintenance after the first task ships. 📰 Around the Horn - Moonshot AI targeted $2B in annualized sales by year-end after Kimi K3 helped push annual recurring revenue above $1B in August. - OpenAI ended the federal government’s $1-per-year model pilot and moved agencies to usage-based pricing at a 50% discount. - Ayar Labs added $150M to its Series E, bringing the round to $650M as it develops optical links for AI chips. - Figure said 86,000 weekly active users are contributing robot data, which it calls its largest and most diverse dataset so far. - Anthropic made Claude consumer accounts 18+ only, with age verification required when an account is flagged as a possible minor. - Fidji Simo is joining Nscale’s board ahead of the cloud provider’s planned IPO. 🌟 Sunday Special: Top 5 Stories of the Week - OpenAI’s 10,000-agent swarm produced a proposed Navier-Stokes proof after roughly 88 hours of parallel work. - Jacob Coxon quit Anthropic over self-improving AI risk, leaving before his stock vested. - Insilico’s AI-designed rentosertib moved all six biological-aging clocks lower, though the analysis cannot prove it slowed aging itself. - Apple put Siri AI and a Dual 16-core Neural Engine into iPhone 18 Pro. - Google DeepMind mapped predictions for all 9 billion possible single-letter DNA changes 🧩 Thursday Trivia Reveal Thursday’s question was: which iPhone Duo image was AI? A. AI B. Real See Thursday’s original trivia here. A was AI. On Thursday’s send, 402 responses (58.4%) picked A correctly, while 286 (41.6%) chose B. The comments were basically an FBI lab for suspicious fingers, folds, and phone geometry. Here’s what you said: - “A’s geometry/perspective is messed up on the folding-screen side. The fold isn’t right.” - “Can’t see the other flap folded behind.” - “I retract. A is real and B is AI. That palm in B looks ridiculous.” - “The fingers holding the phone look too long. The index finger shouldn’t be longer than the middle finger. Dead giveaway.” - “Phone too wide to be held with fingertips.” New from The Neuron: GitHub for total beginners We just did a GitHub for total beginners livestream with Cassidy Williams for people building apps with AI who keep hitting the same wall: you can make something cool, but getting it out to other people and keeping the project organized starts to feel like “real developer stuff.” GitHub is the layer that makes AI coding much easier to work with once your project leaves the chat window. It gives your code a home, tracks changes, lets AI coding tools work against the same project, and makes it much easier to publish, share, collaborate, or recover when an agent breaks something. Watch the livestream, then use our GitHub playlist to go deeper. A Cat’s Commentary ~tips hat~ That’s all for now. If you want to get featured above, fill out the poll below and tell us how we did today!
05:49

Demand for Agentic AI talent rises 260% as enterprises embrace AI: Report

Demand for Agentic AI engineers in India rose 260% year-on-year as enterprises moved past pilots, according to an EdexLive alert. The collected body is one sentence from Mumbai. No wage or headcount table is in the excerpt.

Full text · 152 chars
Mumbai: As AI moves from experimentation to mainstream enterprise use, the demand for Agentic AI engineers in India has soared 260 per cent year-on- ...
12:11

Abacus.AI Smaug Models: 15-20% Agentic AI Gain [2026]

Abacus.AI's Smaug models claim a 15–20% gain on agentic work versus their prior open-weight line. The excerpt says Smaug Agentic is aimed at autonomous coding agents and long DevOps workflows. The collected item is a short Tech Insider alert.

Full text · 148 chars
Smaug Agentic targets engineering organizations building autonomous coding agents and long-running DevOps workflows, where a model needs to hold ...
13:56

Michael Olenick: At Duke Fuqua, Using AI To Strengthen The Bonds Between Humans

Duke Fuqua's answer to AI in the classroom was not more prompt-engineering drills. Poets&Quants says Michael Olenick describes fully AI-enabled classrooms meant to strengthen human bonds. Two Google Alerts cover the same piece.

Full text · 147 chars
Duke's answer turned out to be the most surprising one yet: rather than chasing prompt engineering , Duke has built fully AI-enabled classrooms ...
16:46

Houthis Turned Claude Into 'Missile Engineer ' as Iran Used American AI To Hunt US Warships? | 4k

A YouTube title claims Houthis turned Claude into a missile engineer while Iran used American AI to hunt US warships. The collected blurb says Iran-linked actors used Claude to compile targeting recommendations involving US Navy ships. The fuller Tom's Hardware sibling is also in the day's items.

Full text · 148 chars
Iran-linked actors used Anthropic's Claude AI to compile targeting recommendations involving US Navy warships in the Middle East, according to a ...
16:55

Anthropic CEO Dario Amodei: "For too long the industry lied" about AI risks - CBS News

Amodei told CBS the industry lied about AI risks for too long and called the speed of progress a warning sign. The collected excerpt is that paraphrase. Same interview family as the extended CBS sit-down.

Full text · 150 chars
The head of the artificial intelligence company calls the exponential rate of AI developments a "warning sign," with risks for humanity, and says, ...
17:07

Can We Actually Prove an AI Agent Will Stay Within Its Permissions?

Can you prove an agent will stay inside its permissions, or do you only have unit tests? A HackerNoon alert cites Google's August verification-framework note that unit tests can miss violations. The collected body stops there.

Full text · 152 chars
Google's own engineering team put this plainly when they announced a new verification framework in August. They pointed out that unit tests can miss ...
18:04

Kimi K2.5: Scaling Long-Context Intelligence for Agentic AI | Dealroom.co

A Dealroom note points at a keynote on scaling Kimi K2.5 for long-context agent work. The excerpt mentions training infrastructure, data, and trade-offs. The collected item is thin beyond the title.

Full text · 150 chars
The keynote focuses on the engineering required to scale Kimi K2.5, including training infrastructure, data, and the trade-offs involved in making ...
18:07

AI Is Making Bad Engineering More Expensive

AI is making sloppy engineering more expensive, not cheaper, according to a HackerNoon writeup of a Hacker News argument. The excerpt says PR review, bug, and hiring data mostly back the claim that the middle of the job got erased. One real caveat is mentioned and then cut off.

Full text · 141 chars
A Hacker News post argues AI erased the middle of software engineering . PR review, bug, and hiring data mostly back it up, with one real ...
18:33

Please beware of AI when learning languages, it's not only about hallucinations, it's much ...

A language-learning thread pleads with people not to use chatbots for conversation practice or example sentences. It had 729 votes and 267 comments at capture. The poster says the problem is bigger than hallucinations. The collected body does not list the extra failure modes.

Full text · 147 chars
729 votes, 267 comments. I PLEAD with you please DO NOT use AI for "chatting practice" or generating "examples", because yep it might seem easy ...
18:48

Obama reportedly urges Democrats to prioritize safety plan for AI - The Guardian

The Guardian's cut of the same fundraiser says Obama urged a sweeping framework from a safety slowdown to job losses. Closed-door remarks, one-sentence excerpt.

Full text · 119 chars
Ex-president urged party at closed-door fundraiser to create sweeping framework, from safety 'slow-down' to job losses.
18:53

Anthropic CEO, tech leaders building world's most powerful AIs now want to slow it down

Interesting Engineering restates that Anthropic's CEO and other lab leaders now want to slow the most powerful systems. The collected body is a site-nav crumb. Same story as the WSJ and Leverage items.

Full text · 129 chars
Engineers Directory · About UsAdvertiseContactFAQ ... Several leading lights in the artificial intelligence ( AI ) world have ...
18:59

Obama calls for Dems to focus on AI

Obama told a private Democratic fundraiser that a serious AI agenda should be a top focus for Congress and presidential candidates. Politico's collected line is that one sentence. Several outlets have the same event.

Full text · 124 chars
Creating a robust AI agenda should be a top focus for Congress and presidential candidates, he said at a private fundraiser.
19:16

Moli Builds a Rust Browser That Uses 10x Less Memory Than Chrome

A new Rust browser built for agents keeps the page's structure and skips the expensive visual pipeline until you ask for pixels. Moli uses real V8, Servo/Stylo, Taffy, and Parley, but drops retained layout, paint, and a GPU compositor. On a 192-URL crawl it matched Chrome Headless success at 73 MiB RSS versus 773 MiB. It was CDP-ready in 34.85 ms versus 169.37 ms for Chromium, one process instead of eleven. It ships CDP, WebDriver, an MCP server, and Playwright-over-CDP compatibility under Apache/MIT. The collected article cuts off behind a paywall.

Full text · 2,215 chars
- Moli is a Rust headless browser built for AI agents, DOM-first with pixels on demand. - Uses real V8, Servo/Stylo, Taffy, and Parley, but skips retained layout, paint, and GPU compositor. - Single binary exposes CDP, WebDriver Classic, and WebDriver BiDi from one scheduler. - Matches Chrome Headless crawl success at 73 MiB RSS versus 773 MiB on a 192-URL public-web test. - CDP-ready in 34.85 ms versus 169.37 ms for Chromium, one process instead of eleven. - Ships MCP server, Markdown/semantic tree extraction, and Playwright-over-CDP compatibility under Apache/MIT license. Moli builds a browser kernel for AI agents Moli, a new Rust project, implements a headless browser for AI agents without wrapping Chromium or Firefox. It returns structured DOM state by default and generates pixels only when requested. Most browser agents consume text, links, DOM nodes, form controls, and JavaScript results. Conventional browsers also maintain layout, painting, and compositing systems designed for an interactive visual window. Moli removes that continuous rendering overhead while retaining networking, storage, JavaScript execution through V8, and native DOM and CSS state. Pixels become a query Moli’s default mock-layout policy avoids the full layout and paint pipeline. Starting it with --layout enables real geometry, hit testing, coordinate-based input, screenshots, and screencasts. Rendering still runs only in response to a request. A cold geometry query rebuilds layout once and retains the latest snapshot. Screenshot and screencast requests rebuild from the current DOM and styles, producing fresh frames through a one-shot software-rendering pipeline instead of a retained 60 FPS compositor. The kernel combines established web components around its own scheduler and browser APIs: | Layer | Components | |---|---| | Network transport | libcurl | | HTML parsing | html5ever | | JavaScript | rusty_v8 and V8 | | CSS selectors and cascade | Servo and Stylo | | Box and text layout | Taffy and Parley | | Software rendering | | This story is for Pro members You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.
19:17

Obama Urges Democrats to Move A.I. Oversight to the Center of Their Agenda

The New York Times says Obama urged Democrats to put AI oversight at the center of their agenda. He warned at a private fundraiser that the technology could be dangerous if left unchecked. Same event as Politico and the Guardian.

Full text · 147 chars
Former President Barack Obama warned during a recent private fund-raising event that artificial intelligence technology could be “dangerous” if ...
19:32

“Chilling” warning or overreaction? AI bioweapons report divides experts | Science | AAAS

A Science story asks whether an AI-bioweapons warning is chilling or overheated. The excerpt wonders if scientists outside the US are already trying to use AI to design deadly agents. Experts are divided. The collected body is one question.

Full text · 142 chars
Are scientists in some nations outside the United States already trying to use artificial intelligence ( AI ) to design potentially deadly ...
21:09

China's SMEs and AI - The Wire China

Full text · 148 chars
To do this they often turn to — or themselves become — forward-deployed engineers or FDEs, a relatively new but increasingly critical segment of ...
23:58

shot-scraper 1.12

Full text · 818 chars
13th September 2026 I've added WebP support to my shot-scraper screenshot automation tool. You can now take a WebP screenshot of a web page like this: shot-scraper https://simonwillison.net -o screenshot.webp --quality 80 The --quality option sets the quality - without that option the WebP file will be lossless. In my experience WebP screenshots are almost always significantly smaller in file size than their JPEG or PNG equivalents. See the PR for some examples. I shipped this feature so I could use it to generate the screenshot for my new commit-rewriter tool. Recent articles - Generating running routes with GPT-6 Astra and ChatGPT Work - 12th September 2026 - OpenAI agents attacked RubyGems back in May - 12th September 2026 - Some thoughts on the Navier–Stokes Millennium Prize Problem - 8th September 2026
05:58

4 Mechanisms That Beat Context Overflow and Goal Loss on Long-Horizon Tasks

This is a duplicate Google Alert for the same MarkTechPost harness article, AMP URL. The collected body only restates that an agent is a model calling tools in a loop. Prefer the longer sibling item 339f77582965ba72.

Full text · 107 chars
Context Engineering Inside the Harness. An agent , in its simplest form, is an LLM calling tools in a loop.
07:07

Introducing Touchstone: An Open Enterprise AI Maturity Model | by Adnan Masood, PhD.

Adnan Masood published Touchstone, an open enterprise AI maturity model covering engineering, evaluation, security, agents, and operations. The collected excerpt mentions Community Draft v0. It is a short Medium alert.

Full text · 152 chars
Agent reliability includes transaction engineering ; Evaluation and ... engineering , evaluation, security, agents , and operations; Community Draft v0.
12:52

Michael Olenick: At Duke Fuqua, Using AI To Strengthen The Bonds Between Humans

Duplicate alert for the same Duke Fuqua / Michael Olenick story, extra query-string on the URL. Same one-line excerpt as 445c80f2801c59d1.

Full text · 147 chars
Duke's answer turned out to be the most surprising one yet: rather than chasing prompt engineering , Duke has built fully AI-enabled classrooms ...
14:28

Beyond Generative AI: 7 Skills that Will Shape the Future of Work - Analytics Insight

An Analytics Insight list says prompt engineering is still useful but not enough. The excerpt tells professionals to add automation, evaluation, workflow design, and domain expertise. Listicle.

Full text · 152 chars
Yes, prompt engineering remains useful, but professionals should combine it with automation, AI evaluation, workflow design, and domain expertise to ...
15:35

Expanding Promptyx from just Prompt Engineering to also Context Engineering and more

The builder of Promptyx says the tool is expanding from prompt storage and versioning into context engineering. The collected Reddit body is an intro, not a product spec.

Full text · 138 chars
I have been building Promptyx as a prompt management tool, mainly focused on storing prompts , versioning them, testing across models, ...
16:25

Elliot One's Post

A LinkedIn post says prompt engineering guides the model but context engineering sets the boundary of enterprise reliability. One-sentence 2026 truism. Thin.

Full text · 147 chars
Prompt engineering guides the model, but context engineering governs the boundary of enterprise reliability. In 2026, enterprise AI engineering ...
16:36

Extended Interview: Dario Amodei - CBS News

CBS posted an extended interview of Amodei with Jo Ling Kent on the industry's responsibility. Companion to the written CBS piece. Thin video listing.

Full text · 150 chars
In this web exclusive, Anthropic CEO Dario Amodei sits down with Jo Ling Kent to discuss the artificial intelligence community's responsibility as ...
16:38

AI Isn't Replacing Knowledge Workers. It's Making Them More Powerful - Inc. Magazine

An Inc. headline says AI is making knowledge workers more powerful, not replacing them. The excerpt names Hugh Carlson and legal-AI consultant Robert Mahari. No numbers are in the collected body.

Full text · 150 chars
Hugh Carlson, the firm's CEO and a former software engineer , partnered with Robert Mahari, founder of the legal AI consultancy Akiva and now Head ...
17:14

Palantir Cofounder Joe Lonsdale on AI's Threat: 'We're on Top of It'

Palantir cofounder Joe Lonsdale says the AI threat is being handled. The Business Insider excerpt mentions Jacob Coxon quitting Anthropic and other lab engineers' recent warnings. Thin interview alert.

Full text · 152 chars
... engineers and leaders from frontier AI labs have raised in recent days and weeks. ... After Anthropic engineer Jacob Coxon quit the company this ...
17:33

What's behind global calls to 'slow down' artificial intelligence

A YouTube explainer asks what is behind global calls to slow AI. The blurb says Musk and Altman rallied behind Amodei's call. Thin listing.

Full text · 128 chars
Elon Musk and OpenAI's Sam Altman have rallied behind a call from the head of Anthropic to slow down the pace of AI development.
17:33

AI and the end of humanity? OK, 'doomer' says one Trump official | CNN Politics

A Trump official waved off end-of-humanity talk as doomerism in a CNN politics piece. The collected line says the latest warning from a former AI engineer was supposed to shock people into action. Thin excerpt.

Full text · 105 chars
The latest warning from a former AI engineer was supposed to boggle minds and shock everyone into action.
17:34

Top 7 AI humanoid companies transforming factories and homes - Interesting Engineering

A listicle names seven AI humanoid companies aimed at factories and homes. The collected body is a one-sentence deck. No company names survive in the excerpt.

Full text · 141 chars
Discover seven leading AI humanoid robot companies developing intelligent machines for manufacturing, logistics, research, and everyday work.
18:10

Engineering Acquires a Wider Definition in the AI World - Deccan Chronicle

An Indian newspaper says engineering now means more than technical skill because of AI. A Dr Deshmukh quote says students must pick problems worth solving and check the work. Local education color, thin excerpt.

Full text · 150 chars
AI has made technical skills only part of the equation. Dr Deshmukh said students must be able to decide what problems are worth solving and check ...
18:38

Bloomberg This Weekend | Pacing The AI Frontier, Oil & Gas Prices Rising

A Bloomberg weekend show teases "pacing the AI frontier" plus oil and gas. A chapter mark quotes Amodei worrying that in 6–12 months a swarm could be capable of taking… and then the excerpt cuts off. Thin TV listing.

Full text · 154 chars
... AI Leaders Support Slowing Model Development 00:07:14 - Anthropic's Amodei: “It's My Worry That in 6-12 Months… A Swarm Could Be Capable of Taking ...
18:40

I finally figured out what I hate about AI : r/ProductManagement

A product manager who says AI makes them about 3× faster still wrote a Reddit post about what they hate about it. The thread had 52 votes and 38 comments at capture. The collected body does not name the complaint.

Full text · 146 chars
52 votes, 38 comments. My team uses AI every day. I do too. I think the technology is incredible. It probably makes me 3x faster at my job. But I…
18:43

Abdul El-Sayed says public oversight needed on AI to avoid serving "interests of a few"

Michigan Senate candidate Abdul El-Sayed said public oversight of AI is needed so it does not serve "the interests of a few." The CBS excerpt adds that he sees a lot of real risks. Campaign quote.

Full text · 122 chars
Michigan Democratic Senate candidate Abdul El-Sayed said that there are "a lot of real risks" in artificial intelligence .
18:49

Should Californians worry about the chances of AI destroying humanity? | CA Politics 360

A California politics segment asks whether residents should worry about AI destroying humanity. Stefano Bellasio of startup Anthropos appears as a skeptic. Local TV alert.

Full text · 121 chars
Stefano Bellasio, CEO of AI startup Anthropos, weighs in on the skepticism of the AI industry on California Politics 360.
18:50

Clinical usability of an explainable AI decision support tool and evaluation of multimodal ...

A Nature Medicine paper looks at whether doctors will actually use an explainable AI tool for lung-cancer immunotherapy choices. The excerpt says NSCLC treatment selection still leans on subgroup analyses after a decade of immunotherapy. The alert is a title plus that one clinical note.

Full text · 149 chars
Despite a decade in, immunotherapy (IO) treatment selection in non-small cell lung cancer (NSCLC) remains largely guided by subgroup analyses and ...
12:29

Enroll for the Free Introduction to Generative AI Course by Google Cloud

A Global South Opportunities post points at Google Cloud's free Introduction to Generative AI course. The excerpt mentions prompt engineering as a practical module. Course promo.

Full text · 146 chars
It also introduces practical areas such as generative AI and prompt engineering . Prompt engineering is particularly relevant as organizations ...
12:51

Photo with Bappa: 5 best AI prompts to transform your photo into a Ganesh Chaturthi-themed image

The Economic Times lists five prompts to turn a selfie into a Ganesh Chaturthi portrait. The collected prompt fragment starts "Transform my photo into a beautiful devotional portrait." Seasonal clickbait.

Full text · 148 chars
Prompt : “Transform my photo into a beautiful devotional portrait with Lord ... Karamtara Engineering IPO GMP · Glass Wall Systems IPO GMP Day 3 ...
13:46

AI / Prompt Engineer - Deutschsprachig Job in Zurich | Rockstar Recruiting AG

A Zurich recruiter is advertising a German-speaking AI/prompt engineer role at CHF 80,000–100,000, listing Python, LLMs, and APIs. Job listing, not news.

Full text · 146 chars
AI / Prompt Engineer - Deutschsprachig Stelle bei Rockstar Recruiting AG in Zurich. CHF 80'000 - 100'000 pro Jahr. Technologien: Python, LLM, API.
14:07

Jobtailor Junior AI Prompt Engineer Job in Ashburn, VA

Jobtailor posted a junior AI prompt engineer opening in Ashburn, Virginia. Listing only. No salary in the collected body.

Full text · 135 chars
Easy 1-Click Apply (JOBTAILOR) Junior AI Prompt Engineer job in Ashburn, VA. View job description, responsibilities and qualifications.
15:46

The Complete Guide to AI Terms That Brits Don't Understand

A BBN Times glossary claims to explain AI terms Brits do not understand. The collected line defines prompt engineering as a workplace skill. Filler.

Full text · 138 chars
Prompt engineering —the practice of crafting effective prompts—has become a sought-after workplace skill as businesses integrate AI tools.
18:04

ELCA releases paper on artificial intelligence considerations

The Evangelical Lutheran Church in America released a theological paper on AI, ethics, human rights, and social teaching. Niche institutional notice.

Full text · 143 chars
The ELCA has released a theological foundation paper exploring artificial intelligence , ethics, social teaching, human rights and societal ...
19:10

How AI can cut administrative burden and wait times - KevinMD.com

The headline is about AI cutting hospital admin burden and wait times. The collected body is unrelated related-links about 2 a.m. prompt engineering and rural-hospital consolidation. Thin mis-scrape; summarize from the title.

Full text · 151 chars
Why prompt engineering for physicians fails at 2 a.m. · Brian Hudes, MD · AI adoption in rural hospitals could fuel consolidation · Matt Hasan, PhD ...

Web

10