Nothing matches those filters.

Article

26
00:00

Instinct group chats 💬, Reflection’s 501B model 🧠, OpenAI text watermarks 🏷️

Full text · 381 chars
Frontier intelligence. Half the bill. (Sponsor) Fireworks Nexus gives developers access to frontier open models in the harnesses they already use. With FireRouter, each request is scored and routed to the optimal model for the task, while difficult tasks can pass through to your closed provider. Set up in minutes. Keep your existing workflow. Get more control over your AI spend.
09:40

😸 OpenAI will watermark ChatGPT text

Full text · 8,834 chars
😸 OpenAI will watermark ChatGPT text PLUS: Meta and Microsoft cut Claude. Altman on AI harms Welcome, humans. The U.S. Army spent the past few years pushing drones, AI, and autonomous systems into its units at the same time. This week, acting chief of staff Gen. Christopher LaNeve said the Army needs to narrow its modernization push, because too much was changing at the same time. He also said no brigade commander can AI their way to being the best at the job. The fix is a new framework called "How the Army Fights," which decides which technology gets money and where it goes. Acting Army Secretary Adam Telle described the framework as the Christmas tree and each new gadget as an ornament. LaNeve named command and control (the systems commanders use to direct troops) as his top priority, and the Army has not said which ornaments are going back in the box. Every office running eleven AI pilots and zero decisions felt very seen. Here’s what happened in AI today: - 😸 OpenAI will watermark ChatGPT's writing in Europe - 📰 Meta and Microsoft cut internal use of Claude - 📰 Wikimedia said OpenAI-linked agents edited its wikis without approval - 🍪 GRID built dedicated spreadsheet tools for AI agents - 🎓 Make your chatbot argue against you before trusting it 😸 ChatGPT's Writing Is Getting an Invisible Watermark (in Europe, for Now) On Monday, OpenAI said it will start hiding an invisible signal in the text that ChatGPT and Codex (its coding assistant) write for people in the European Union. The signal lives in the AI's word choices (a pattern in which words it picks), invisible to you but findable by a special detector. The EU's AI law requires AI companies to make machine-written text identifiable by software. OpenAI calls its method textGrain and plans to publish it for anyone to build on. Here's what happened: - Developers using OpenAI's API (the connection apps use to tap its models) can opt in to watermarked text worldwide, starting now, for select models. It stays off by default. - ChatGPT and Codex users in the EU, on every plan, get the watermark over the coming weeks, with no global rollout at launch. - The detector stays private for now. Approved researchers and expert organizations can apply for access; the public cannot. - OpenAI says the watermark left its newest model's benchmark scores (tests that grade AI) essentially unchanged, and that it matched or beat other methods, including Google's SynthID for text. Why this matters: OpenAI published where the watermark breaks, which is the useful part: - Short passages are harder to catch: about 80% detected, versus about 95% for passages twice as long, at a 1% false-alarm rate. - Math is harder still, since there are fewer ways to word an answer. - Swapping 10% of the words for synonyms dropped detection from about 92% to 66%. Swapping 25% dropped it to 17%. For you, a light rewrite of a ChatGPT draft mostly erases the signal, and a clean detection says little. OpenAI states a watermark cannot show who wrote a passage, how much a human edited it, or whether it is accurate. A missing watermark proves nothing about human authorship either. Our take: Swap a quarter of the words and the watermark nearly vanishes, which makes it the rare security feature a thesaurus can beat. Read this as a compliance move for Brussels more than a shield against copy-and-paste, and give OpenAI credit for spelling out the limits itself. The open question: if every lab ships its own watermark and keeps its own private detector, who checks one document against all of them? Watch who gets detector access first; that list shows who OpenAI thinks should police AI text. FROM OUR PARTNERS Gemini 4 Argon debuts at #1 on APEX-Agents; new top model for professional work. Mercor's APEX-Agents leaderboard evaluates frontier AI on long-horizon, multistep tasks across economically valuable work. Built with partners like Harvey, Ramp, and Cognition. Every model. Ranked by productivity. 🎓 AI Skill of the Day: Make your chatbot argue back before you trust its "yes" Ask ChatGPT whether your plan holds up, and it will almost always say yes, and that habit is baked in. Lawyer Dr. Niklas Schmidt explains why: chatbots are tuned on human ratings, and people rate agreement higher than correction (called sycophancy, i.e., telling you what you want to hear). Asking for "a balanced view" keeps the tilt and changes only the tone. His fix is structural: - Draft in one message, critique in a separate one. Asking AI to draft, critique, and rewrite at once makes it rush the draft. - Give the critic a side and a reason to win, like a rival who wants your proposal to fail. - Test the pushback. If it sounds soft, reply: "You conceded too easily. Try again as if your business depended on winning." - Keep your hoped-for answer out of the question. Ask "what does this contract say about refunds?" instead of "confirm this contract lets me refund anyone." Here is my position: <position>[paste your plan, argument, or draft]</position> Act as a rival with a strong incentive to defeat this. Give me: the three best counter-arguments, most dangerous first; the question I would least want to answer in front of a skeptical boss or client; and any assumption that, if false, makes the whole thing collapse. Then say which of my points you would concede because arguing them is hopeless. Mark any fact or source you mention as unverified. Want more tips like this? Check out our AI Skill of the Day Digest for October. FROM OUR PARTNERS Their bank asked for a P&L, so their AI pulled it from Xero It's one of the real jobs people handed TinyFish in September. Others let it into their Shopify admin and Cloudflare dashboard, usually with a note like "read-only" attached. TinyFish connects to the AI you already use so it can work on real websites. Top up this month, and you get 30% extra. 📰 Around the Horn - Meta and Microsoft cut back internal Claude use, steering staff to their own coding tools; Meta's Claude Code users fell from about 60,000 to 30,000. - Sam Altman said the world should accept some AI harms, such as scams and hacking, in exchange for the benefits, and argued against one lab controlling the technology. - The Wikimedia Foundation said OpenAI-linked agents edited its wikis without approval and sent millions of automated requests, which may have contributed to a partial outage in May. - NPR found voters are using chatbots to compare candidates and debunk viral claims; about one in five voters consult them for election news, per a September survey. - CNBC asked whether Nadella can reinvent Microsoft again, noting Copilot has 30M seats, but the company still lacks an AI smash hit. 🍪 Treats to Try - *GRID gives your agents dedicated tools to read, edit, and generate spreadsheets: faster, more robust, fewer tokens. Evaluate for free now. - Ghost, a startup with a 19-year-old founder, is taking preorders for Core, a screenless box that keeps your AI assistants and data at home instead of in the cloud (raised $11M) —paid only rn ($3,499 one-time, ships late October). - Beam hands you a U.S.-built open model (one companies can download and run on their own servers) that Reflection says needs three to four times less computing power than similar open models. - Armadin sends swarms of AI agents to attack your company's systems the way real hackers would, so you find the weak spots first (raised $255.5M). - Veltrix scans the data already sitting in your accounting and store tools (e.g., QuickBooks or Xero), read-only, and tells you what is costing you money and what to fix first in about two minutes —first check is free, no card needed. - Teachoo coaches you through any homework problem one question at a time instead of handing over the answer, and you can snap a photo of a worksheet to get started —free to try. - Gauth teaches you biology, chemistry, physics, and even how the heart and lungs work through interactive lessons with guided practice, including one on how noise-cancelling headphones cancel noise. 🐦 Tuesday Tweets: Claude's $200 plan might be wildly underpriced The AI subscription wars are getting weird. Chubby highlighted a new SemiAnalysis comparison estimating that Claude Max's $200 plan gives heavy agent users more than 5x the API-equivalent usage value of ChatGPT Pro at the same price. Translation: this isn't a benchmark saying Claude is five times smarter. It's asking how much model usage you're actually getting for your monthly subscription, then pricing that same usage at API rates. Which raises a fun question: are these $200 "power user" plans actually software subscriptions… or heavily subsidized compute buffets? New from The Neuron: AI Explained A Cat’s Commentary That’s all for now. If you want to get featured above, fill out the poll below and tell us how we did today!
10:35

2026 Climate Tech Companies to Watch

Full text · 1,096 chars
10 Climate Tech Companies to Watch Each year, the MIT Technology Review team puts together a list of some of the most promising climate tech companies in the world. Whether early-stage startup or multinational corporation, the businesses we’ve chosen are working on technologies to help us address climate change or adapt to our warming world. There’s an urgent need for these innovators: We must begin to drastically reduce emissions to avoid the deadliest impacts of climate change, while also contending with its harmful effects. We hope that this list highlights the progress the world is making to tackle the climate crisis, as well as the breadth of solutions required. From energy storage powered by carbon dioxide to cleaner ways to make cement and refine critical minerals, these companies are building technologies to address the acute challenges we face. This is the fourth annual edition of this list. Learn more about how we chose the 2026 slate. The Spark Sign up for our weekly newsletter to understand the innovations, policies, and trends shaping climate tech today.
10:35

Here’s how our climate team picked 10 promising companies to watch

Full text · 4,484 chars
As the team at MIT Technology Review set out to choose companies for this year’s edition of our annual list of Climate Tech Companies to Watch, the challenge felt more daunting than any year since we started compiling the list in 2023. This year is shaping up to be one of the hottest ever recorded. The world must begin to reduce greenhouse-gas emissions to address the climate crisis, which is already affecting communities around the globe. Wildfires, floods, and other climate-fueled disasters are killing thousands and costing billions of dollars. At the same time, political shifts and international conflict have stalled progress, particularly in the US. Technology alone can’t solve climate change. But we believe innovations can help us make real changes and improve lives around the world as we try to address one of the most pressing challenges of our time. In this project, we look for companies developing technologies that can help us face the climate crisis. Some are driving down the greenhouse-gas emissions fueling climate change; others are working to reduce the dangers that communities face from extreme heat, drought, or other shifts brought on by a warming planet. We focus on companies that have established a track record, whether through research findings, capital raised, or deployments. When we set out to build our list, we start by soliciting nominations from our team of reporters and editors. We ask experts for suggestions and reach out to academics, researchers, investors, and analysts. Once we have a pool of nominees, we then consider each one, digging through literature, reports, and patents to vet both the technology and the business. We judge each company not only on its own merits, but also on how it might fit into our list overall. We want the final slate of companies to represent a range of industries and include firms based around the globe that are at different stages of development. We also aim to strike a balance between highlighting companies previously included and introducing new names. Finally, we consider some of the most important trends and events of the year and look for nominees that reflect those developments. The dismembering of US federal regulations and cuts to financial government incentives for innovation means that the country is no longer a leader in climate tech. Like last year, the 2026 list is largely composed of companies that are based outside the US. China is an obvious hotspot for technological innovation across energy and climate. This year’s list includes two Chinese companies: WeLion is building semi-solid-state batteries for electric vehicles, and Envision Energy is installing huge amounts of wind power and energy storage in China and around the world. Several sectors stood out this year as being of particular importance. One of the biggest stories of the year is the rising demand for electricity, driven in part by AI data centers. Energy storage is a key piece of that puzzle because it can smooth out variations in the electricity supply as grids incorporate more renewable power. We’ve included several energy storage companies on this year’s list. Energy Dome is building long-duration energy storage systems on the grid that rely on compressed carbon dioxide. Moment Energy is a Canadian company working to repurpose EV batteries for grid-scale storage. And Form Energy, which we also included in the 2024 edition of this list, is producing iron-based batteries that can store energy for multiple days. A wide range of technologies is represented here: everything from mobile flood barriers to next-generation nuclear reactors. Some of the awardees might be familiar from our past coverage; others are just emerging and may be less well known. We believe we’ve put together a list filled with businesses that are rising to the challenge of this moment. Political will has weakened even as the damages wrought by climate change continue to mount. By helping us make progress on tackling the climate crisis in such difficult times, these companies point to a better way forward. Deep Dive Climate change and energy Batteries just broke another record in the US Huge grid-scale batteries are thriving, but smaller residential systems have lagged. What’s behind this summer’s heat, and why 2027 could be worse El Niño? Climate change? All of the above? Stay connected Get the latest updates from MIT Technology Review Discover special offers, top stories, upcoming events, and more.
10:35

WeLion New Energy and its semi-solid-state batteries

Full text · 4,820 chars
WeLion New Energy is on a quest to make safer, better batteries. The company’s semi-solid-state cells could improve safety and offer greater energy density than lithium-ion batteries to power electric cars, boats, and drones. From the EVs that carry commuters home to the ships that transport goods across countries, more and more vehicles around us are powered by batteries. Lithium-ion batteries are commonly used to power personal devices and vehicles today. But as the need for batteries surges, there’s a growing appetite for cells that can store more energy on a single charge. Also, in lithium-based batteries the electrolyte, which facilitates ion transfer, uses a flammable liquid that can create a fire hazard. Several companies, including WeLion New Energy, are working to make solid-state cells, which use a solid electrolyte instead of a liquid one. This could increase the energy density of batteries, as well as improve their safety. Fully solid-state cells are extremely difficult to manufacture, though. So WeLion currently focuses on developing and producing semi-solid-state cells, which have a part-liquid, part-solid electrolyte. On that front, they have made steady progress. In 2023, WeLion released a battery pack in partnership with Shanghai-based EV-maker Nio, best known for its battery-swapping technology. The pack had an energy density of 360 watt-hours per kilogram (wh/kg), significantly higher than the 200 wh/kg average for lithium-ion packs today. That translated to a range of over 1,000 kilometers (621 miles). However, only about 200 of the battery packs were manufactured due to their high production cost, but they are still being used by Nio, WeLion says. WeLion has also built batteries to power cargo ships and ferries through a subsidiary. And the company is developing batteries with even higher energy density too. WeLion’s engineers made a semi-solid-state battery that achieved a remarkable energy density of 824 wh/kg during testing in 2025, more than three times that of a standard lithium-ion battery available today. Key indicators - Industry: Energy storage - Founded: 2016 - Headquarters: Beijing, China - Notable fact: WeLion’s cofounders include Chen Liquan, a scientist at the Chinese Academy of Sciences who led the development of the country’s first all-solid-state battery back in 1988. Potential for impact China emits more carbon dioxide than any other country, and its transportation sector accounts for around 10% of that total. Scaling up the production of safer, higher-capacity batteries could help switch more vehicles away from fossil fuels. Batteries that power electric ships could be especially helpful—the carbon dioxide emissions from vessels that move goods along China’s lakes, canals, and rivers are projected to grow by around 25% in this decade alone, reaching 18.7 million metric tons by 2030. Caveats One big challenge WeLion faces is scaling up production of its semi-solid-state batteries and competing on cost in a crowded battery market. WeLion says that its semi-solid-state batteries are more expensive than conventional lithium-ion batteries made by larger Chinese battery companies such as CATL. And while laboratory tests showing high energy density are impressive, it can be difficult to translate laboratory results to the real world. The company declined to give more details about the experimental results from its 2025 prototype, which it says is far from commercialization. Key details remain unclear, including how many times the cell can be charged and discharged and how capacity degrades over its lifetime. Next steps WeLion plans to bring a new semi-solid-state battery for drones to market with an energy density of 410 wh/kg by the end of this year. It also intends to introduce its first all-solid-state battery for cars in 2027, which is expected to have an energy density of 400 wh/kg. Two inland cargo ships powered by WeLion’s batteries are expected to launch in October 2026 in Huzhou, a city in eastern China. The battery pack of each of these ships will be able to carry nearly 4,000 kilowatt-hours of electricity, enough for them to sail for up to 250 kilometers (155 miles), the company says. Today, WeLion has an annual manufacturing capacity of 11 gigawatt-hours (GWh) spread between four production bases in China. Two additional plants with a combined yearly capacity of 18 GWh are under construction. Deep Dive Climate change and energy Batteries just broke another record in the US Huge grid-scale batteries are thriving, but smaller residential systems have lagged. What’s behind this summer’s heat, and why 2027 could be worse El Niño? Climate change? All of the above? Stay connected Get the latest updates from MIT Technology Review Discover special offers, top stories, upcoming events, and more.
10:35

Form Energy and its iron batteries

Full text · 4,861 chars
Form Energy is building iron-based batteries that can store energy for multiple days. The company is ramping up supply at its factory and signing deals for commercial projects. New solar and wind installations produce electricity more affordably than fossil fuels in most places, but their output varies based on season, time of day, and weather. To compensate, utilities are increasingly investing in battery storage systems, which can charge when power is abundant and discharge when the grid needs a boost. Lithium-ion batteries are the fastest-growing storage technology. Although light and energy-dense, they require pricey metals and are costly to scale, and few of the large batteries of this type used in industry are designed to store energy for longer than four to six hours. Form Energy wants to build cheaper long-duration energy storage technology. The company has developed a new type of battery, built with a far more abundant and cheaper metal, iron, that can store energy for 100 hours. It uses a process that Form calls “reversible rusting.” During discharge, oxygen from the air converts the iron metal to rust, releasing electrons. When charging, current converts the rust back into iron metal. Over the past two years, Form has gained commercial traction. It began producing batteries at its West Virginia factory for its first commercial deployment, a 150 megawatt-hour system for Minnesota electricity provider Great River Energy, which is set to come online in 2027. And Form has signed several other deals with power providers. The largest is a project with another Minnesota-based utility, Xcel Energy, where wind and solar power, backed by 30 gigawatt-hours of Form’s batteries, will be used to power a new Google data center. That project is expected to come online in phases between 2028 and 2031. Key indicators - Industry: Energy storage - Founded: 2017 - Headquarters: Somerville, Massachusetts, US - Notable fact: Once complete, Form’s proposed 30 gigawatt-hour energy storage system at a Google data center in Minnesota could be the world’s largest battery project by energy capacity. Potential for impact Wind and solar power play a major role in limiting growth in global carbon dioxide emissions. Still, many grid operators depend on fossil fuel plants, especially natural gas ones, to fill gaps caused by periods of low renewable supply or high demand. Form’s batteries are designed to cover multiday supply-demand imbalances, like those that crop up during heat waves, deep freezes, or periods with little sun or wind. The company declined to share the current cost of its battery systems, but says its goal is to reach $20 per kilowatt-hour, which would make them cost-competitive with the natural gas alternative. The company’s technology could help reduce emissions from the AI data center boom. AI’s thirst for energy has already delayed closure of aging coal plants, led to the building of new natural gas plants, and caused price hikes for consumers. Cheap energy storage could help renewables meet more of this AI-driven growth in electricity demand. Caveats The long-duration storage market is still nascent. Form’s iron-air technology faces a crowded field of approaches, including pumped hydropower, compressed gas, fuel cells, gravitational devices, and flow batteries that store energy in liquids. The economics of multiday batteries, which require large capital expenditures yet frequently sit idle, remain tricky. While funding from data center developers could be a major boost, Form also has to prove it can deliver: Its existing production capacity, 2 gigawatt-hours of battery storage per year, is a fraction of the 80 gigawatt-hours it has promised through commercial agreements. Success will hinge on its ability to scale. Another remaining question is the technology’s physical footprint. Form’s batteries, which are stacked together in enclosures the size of shipping containers, require an acre for every two to three megawatts. That means its Google installation would take up the equivalent of at least 75 football fields. Next steps In addition to finalizing its first commercial project, Form is outfitting its factory with new equipment that will help it scale manufacturing to meet its production backlog. It will do so with the help of new financing, announced in August, that brings its total money raised, in addition to West Virginia state and US federal grants, to more than $2 billion. Deep Dive Climate change and energy Batteries just broke another record in the US Huge grid-scale batteries are thriving, but smaller residential systems have lagged. What’s behind this summer’s heat, and why 2027 could be worse El Niño? Climate change? All of the above? Stay connected Get the latest updates from MIT Technology Review Discover special offers, top stories, upcoming events, and more.
10:35

X-energy and its helium-cooled nuclear reactors

Full text · 5,336 chars
X-energy is developing small modular reactors to satisfy some of the planet’s most power-hungry users: industrial manufacturers that need high temperatures to produce billions of tons of concrete, plastics, fibers, and chemicals. Solar panels and wind turbines are great at producing electricity but what heavy industry also needs is raw heat—and lots of it. The manufacture of concrete, glass, and steel, as well as the production of many chemicals, plastics, and fertilizers, are still highly reliant on fossil fuels to produce that heat. Traditional nuclear reactors simply don’t get hot enough. The cooling systems they use require liquid water to work, and even applying pressure to raise the water’s boiling point leaves their maximum operating temperatures hundreds of degrees cooler than needed. X-energy’s proposed Xe-100 reactor will use helium gas as a coolant instead, allowing it to operate at much lower pressures and higher temperatures. The helium will heat water and create super-heated steam, which industry can use either directly or to generate 80 megawatts (MW) of electricity via a turbine, or a combination of the two. That’s about one-tenth of the capacity of most reactors operating today. In common with other small modular reactors (SMRs), the Xe-100 will use a modern uranium-based fuel called TRISO (short for tristructural-isotropic) that is extremely resistant to corrosion and melting. Compact TRISO “pebbles” will be added continuously to the top of the reactor, then flow slowly to the bottom, where they will either be recycled or removed as waste. The company says the reactor will be intrinsically safe in the event of a loss of power—if temperatures inside rise, the nuclear reaction naturally slows and stops. Key indicators - Industry: Nuclear fission - Founded: 2009 - Headquarters: Rockville, Maryland, US - Notable fact: The first helium-cooled reactor was proposed for the World War II Manhattan Project but was passed over in favor of a simpler and cheaper water-cooled design. Potential for impact Industrial heating currently accounts for 18 percent of annual global greenhouse-gas emissions, with about half of that coming from high-temperature processes that are difficult to electrify. Co-locating small nuclear reactors at large manufacturing facilities could shrink some of the biggest carbon footprints in the world. The Xe-100 reactor could also be used by the fossil fuel industry, which today uses vast quantities of natural gas to generate the heat needed to recover oil from tar sands, and during oil refining. The flexibility of the Xe-100 means that it can be used for traditional electricity generation. By scaling down in size and power from today’s gigawatt-scale fission plants, small reactors like the Xe-100 can serve more markets—even down to individual data centers. AI and data centers now represent the single largest component of new demand for electricity in the US, and their consumption is expected to double again by 2030. X-energy says that up to 12 of its reactors could be clustered together to provide 24/7 power on site, at a lower cost than a traditional nuclear power station and without the hassle of adding transmission infrastructure. Caveats The technology behind helium-cooled fission may be well understood, but the only commercial-scale power generation of this type today is in China, where two reactors have struggled to operate on a continuous basis. X-energy says that its vertically integrated fuel supply should ensure reliable operation, with reactors having a lifetime of 60 years. The economic case for SMRs is also largely unproven. Nuclear costs and timelines have a habit of ballooning, and the US Energy Information Administration calculates that electricity from SMRs will cost more than six times that from solar farms. While the current environment in the US for regulation and government funding favors nuclear, that could change over X-energy’s long development time frame. TRISO reactors appear excellent for safety and nonproliferation (the tough spheres would be difficult to turn into weapons) but are less impressive when it comes to waste. The Xe-100 might produce 10 times the volume of spent nuclear fuel per unit of energy as existing reactors, although its lower-level waste will need less sophisticated storage. Next steps The US Nuclear Regulatory Commission has issued X-energy’s subsidiary with a fuel fabrication license for TRISO-X pebbles, with its first factory due for completion in 2028. An Xe-100 to supply heat and power for a large chemical facility in Texas operated by Dow has also cleared its first regulatory hurdle, and could be operational by the early 2030s. In tandem, X-energy is working toward a 320 MW cluster of four Xe-100 reactors in Richland, Washington, in collaboration with Amazon, and a larger 6 GW fleet in the UK’s northeast. Both projects target electricity generation in the 2030s. Deep Dive Climate change and energy Batteries just broke another record in the US Huge grid-scale batteries are thriving, but smaller residential systems have lagged. What’s behind this summer’s heat, and why 2027 could be worse El Niño? Climate change? All of the above? Stay connected Get the latest updates from MIT Technology Review Discover special offers, top stories, upcoming events, and more.
10:35

Energy Dome and its carbon dioxide batteries

Full text · 4,678 chars
There’s a growing need for long-duration storage to meet energy demand and balance out renewables on the grid. Energy Dome uses compressed carbon dioxide gas in its massive grid batteries, which can deliver power for up to 24 hours. As the world races to meet growing electricity demand, solar and onshore wind power have become the cheapest and quickest-to-deploy sources to install on the grid. But, despite their upsides, both are subject to variations in weather patterns, which means supplies can be intermittent. Energy Dome is using stored carbon dioxide to help smooth out gaps between supply and demand. The key component of an Energy Dome plant is a huge white dome that covers an area about the size of seven soccer fields and holds about 2,000 metric tons of carbon dioxide. When energy is available—say, from a connected solar farm—compressors squeeze the gas, turning it into a liquid that’s then stored in carbon-steel tanks. That process generates heat, which the system holds in a proprietary thermal storage material. Then, when the grid needs power, the carbon dioxide is released from the tanks and warmed with the stored heat, turning it back into a gas. That gaseous carbon dioxide then passes through a turbine to generate electricity and flows back into the dome, ready to begin the cycle again. Compressing gas to store energy isn’t new—utilities have been using compressed air in underground caverns to hang on to reserves for decades. But Energy Dome’s approach doesn’t require any specific geology to work, so it could be more easily scaled to help grids around the world. Energy Dome turned on its first commercial plant in Sardinia, Italy, in 2025. The facility has a capacity of 200 megawatt-hours. That’s enough to power about 18,000 Italian homes for 10 hours. Key indicators - Industry: Energy storage - Founded: 2020 - Headquarters: Milan, Italy - Notable fact: Energy Dome has plans for 30 gigawatt-hours’ worth of projects across five continents. Potential for impact Cheaper energy storage could help wind and solar meet more of the world’s electricity demand. Today, lithium-ion batteries dominate new installations for short-duration applications of up to four hours. But the economics aren’t competitive for longer durations: A lithium-ion system that provides eight hours of storage at the same power output requires doubling the number of cells. Energy Dome estimates that its technology is roughly 10% to 15% cheaper than lithium-ion batteries for an eight-hour system, and that the economics are even better for longer-duration systems of up to 24 hours. The company says its ability to scale bigger and faster than lithium-ion boils down to the fact that its design uses existing commercially available equipment like compressors and storage tanks. Its plants that are either operational or under contract have eight or 10 hours of capacity. Caveats Energy Dome only has one commercial project that’s operational. It’ll need to build many more to provide the 30 gigawatt-hours of storage it has planned. The technology also isn’t quite as efficient as lithium-ion. Overall, Energy Dome’s process can return about 70% of the electricity it stores to the grid (a measure known as roundtrip efficiency). Lithium-ion batteries average roughly 90%. But some other long-duration storage techniques clock in lower; some iron-air batteries, for example, are around 50%. Energy Dome’s plants may not all be benign from a planet-warming perspective. The company offers a version of its system that pairs its carbon dioxide battery with natural gas turbines. That version replaces internal heat storage with waste heat from the gas turbines to help evaporate carbon dioxide. That helps the plant run more efficiently but ultimately results in greenhouse-gas emissions from the natural gas turbines. Next steps Energy Dome has a pipeline of about 30 gigawatt-hours’ worth of plants in the works around the world, and many of these projects could come online by the end of the decade. Because the company uses off-the-shelf components, the time from a signed contract to delivered capacity is only two years. In June 2026, for example, it signed a deal with Google to build a 200 MWh plant in Ireland, which is expected to come online in 2028. Deep Dive Climate change and energy Batteries just broke another record in the US Huge grid-scale batteries are thriving, but smaller residential systems have lagged. What’s behind this summer’s heat, and why 2027 could be worse El Niño? Climate change? All of the above? Stay connected Get the latest updates from MIT Technology Review Discover special offers, top stories, upcoming events, and more.
10:35

Brimstone and its one-stop process for making cleaner cement and critical minerals

Full text · 5,221 chars
Brimstone is making strides on two industrial problems, reducing the emissions from cement and providing a way to produce more critical minerals within the US. If the company can do both, with the same rocks in the same plant, it would unlock a cleaner and more efficient process for manufacturing building materials, aluminum, steel, and more. Brimstone is a code-switching climate tech company for these polarized political times. The US-based startup—part green cement company, part critical mineral producer—is well positioned to sustain its momentum no matter which party is in power. During the more climate-focused Biden administration, the company’s primary pitch was its ability to produce cement while generating about 60% less greenhouse-gas emissions. It achieves this by switching to silicate rocks instead of standard limestone, which releases carbon dioxide when it’s heated. The conventional process for creating cement, the binding used to produce concrete, contributes about 8% of the world’s carbon dioxide pollution. Under the protectionist and climate-denying Trump administration, however, the company is now stressing the secondary benefits of its process: the ability to extract other valuable materials from those same silicate rocks. The company has said for years that its process could also generate materials that can be used to strengthen cement. But in early 2025, Brimstone highlighted that it would be able to also produce smelter-grade alumina, the raw material used in aluminum. It then revealed it had another trick up its sleeve late last year: the ability to make steel, magnesium, titanium, and other critical minerals from the same rock in a single refinery. The company’s CEO, Cody Finke, said those materials represent a $2.4 trillion global market. Brimstone’s approach promises to provide a domestic path for goods mostly produced today in China, the subject of Trump’s harshest trade policies and rhetoric. Key indicators - Industry: Heavy industry - Founded: 2019 - Headquarters: Oakland, California, US - Notable fact: In 2023, the company earned third-party industry certification for its process for manufacturing Portland cement, allowing its product to qualify for standard manufacturing projects. Potential for impact The company made a very big deal of its newfound ability to produce all these other goods, arguing its process would bolster national security and usher in the “next generation of industrial refining.” The basic environmental promise is that generating all these materials from the same rock and a single mining process would require dramatically less energy and produce far less waste than manufacturing them separately. “I think that we have one of the generational companies that's actually able to create an industrial revolution,” Finke says. To which we would say: Well, we’ll see. Caveats Brimstone has yet to produce anything at a commercial scale, and it faces numerous financial and technical challenges. Some observers are dubious that the company will be able to economically extract sufficient quantities of so many materials from a single type of rock—or overcome the cost and labor issues that make mining and refining minerals a tough business in the US. Brimstone also suffered a major stumble last year, when the Trump administration revoked a $189 million grant intended to subsidize its next plant, a planned commercial demonstration project in Reno, Nevada. Finke says the company continues to have fruitful conversations with the administration about restoring those funds, in which Brimstone has emphasized its potential to boost domestic production of critical minerals. He has previously said the company will be able to proceed with the development of the plant, even without that grant or other subsidies. Next steps The company expects to begin operating the Reno facility, which will only produce cement, supplementary materials, and alumina, in 2028. That means Brimstone will still need to finance and build its first full-scale factory to deliver the steel, titanium, and other critical materials, likely pushing the timeline for doing so into the next decade. Based on its progress so far, though, the company has already lined up several commercial deals. That includes agreements to supply cement to retail giant Amazon and alumina to Century Aluminum, one of the US’s largest producers of its namesake metal. Even if Brimstone doesn’t bring about another industrial revolution, the company is equipped to slash emissions from cement, which is reason enough to pay close attention. If Brimstone also manages to dramatically reduce the energy, pollution, and costs associated with extracting and refining rock to produce all—or any—of those other goods, that would represent a big corporate and climate win as well. Deep Dive Climate change and energy Batteries just broke another record in the US Huge grid-scale batteries are thriving, but smaller residential systems have lagged. What’s behind this summer’s heat, and why 2027 could be worse El Niño? Climate change? All of the above? Stay connected Get the latest updates from MIT Technology Review Discover special offers, top stories, upcoming events, and more.
11:03

The Sequence Knowledge - Issue 945: Learning RSI: Agents that Rewrite their Own Scaffolding

Full text · 984 chars
Better file viewing. A patch validation step before submitting a fix. Generating several candidate solutions and ranking them instead of shipping the first one. Keeping a running history of what was tried before and why it failed. Read that list again and tell me who wrote it. It looks like the onboarding checklist a senior engineer hands a new hire. It is not. It is what a coding agent wrote into its own codebase, unprompted, over roughly eighty iterations, while nobody was watching. The agent is the Darwin Gödel Machine from Sakana and Jeff Clune’s lab, and doing this took it from 20 to 50 percent on SWE-bench and from 14 to 31 percent on Polyglot. The list is the finding. In Part 1 of this series I said the robot-with-a-screwdriver picture of recursive self-improvement was unhelpful. This is the essay where I take it back a little. The robot exists. It just turns out to be a mechanic, not a god, and watching it work is more instructive than any of the arguments were.
12:10

The Download: 10 climate tech companies to watch

Full text · 4,975 chars
This is today's edition of The Download, our weekday newsletter that provides a daily dose of what's going on in the world of technology. 10 climate tech companies to watch Each year, MIT Technology Review puts together a list of the most promising climate tech companies in the world. This year, the stakes feel higher than ever. The world is on track for one of its hottest years on record, climate-fueled disasters are taking lives and costing billions, while political shifts and international conflict have stalled progress. We urgently need to cut emissions and avoid climate change’s most harmful effects, and that’s where our 10 companies come in. They’re working on everything from mobile flood barriers and semi-solid-state batteries to compressed-CO₂ energy storage and next-generation nuclear reactors. They also reflect the biggest shifts in climate tech—including the surging energy demands of AI data centers. Over the coming days, we’ll take a closer look at each of the 10 companies, along with how we chose this year’s list, right here in The Download. So stay tuned! The must-reads I’ve combed the internet to find you today’s most fun/important/scary/fascinating stories about technology. 1 Musk’s new Pentagon role has raised conflict-of-interest fears His companies could directly benefit from the weapons he recommends. (NPR) + SpaceX has received billions in Pentagon and NASA contracts. (NYT $) + Beware the rise of the “Pentagon-Silicon Complex.” (Gizmodo) + The Pentagon wants an AI-powered lie detector. (MIT Technology Review) 2 The only billionaires making money this year are in tech Tech fortunes now account for 36% of billionaire wealth. (Bloomberg $) + Nvidia is on the verge of the first $6 trillion valuation. (CNBC) + SpaceX stock has made Elon Musk a trillionaire again. (Gizmodo) 3 Norway is planning the first national ban on smart glasses Camera-enabled glasses could be barred from public spaces. (Guardian) + Several other countries are also mulling bans. (Reuters $) + Smart glasses are causing havoc in India. (MIT Technology Review) 4 Tandem panels could help the US catch up with China in solar They could produce 25% more electricity from sunlight. (NYT $) + The balcony solar boom is coming to the US. (MIT Technology Review) 5 South Korea’s President suspects AI’s been used in bank hacks He called for new cybersecurity measures for the AI era. (Reuters $) + OpenAI agents may have caused a Wikimedia outage. (Verge) + Who’s liable when AI agents go rogue? (MIT Technology Review) 6 Russian drones are exploiting air defense gaps to hit data centers The attacks are disrupting Ukraine’s digital infrastructure. (Ars Technica) 7 The AI boom is killing the world’s cheapest smartphones Memory costs are squeezing out affordable phones. (Rest of World) 8 Pioneers of light-based brain mapping have won a Nobel prize Their technique shines new light on brain circuits. (BBC) 9 AI slop has pushed arXiv to limit preprint research submissions Submissions have doubled in two years. (404 Media) 10 McDonald’s is being sued over AI-powered Big Mac pricing The lawsuit alleges AI helped franchises coordinate prices. (Reuters $) Quote of the day “In the extreme, I could see someone describing Anthropic as a cult that’s trying to take over the world.” —AI researcher Jacob Coxon, who resigned from Anthropic over fears that it’s building systems that it won’t be able to control, tells New York Magazine his views on the company’s culture. One more thing How AI is turning the Iran conflict into theater Much of the spotlight on AI in the Iran conflict has focused on models like Claude helping the US military decide where to strike. But a wave of “vibe-coded” intelligence dashboards—and the ecosystem surrounding them—reflect a new role that AI is playing in wartime: mediating information, often for the worse. These sorts of intelligence tools have much promise. Yet there are real reasons to be suspicious of their data feeds. Read the full story. —James O'Donnell We can still have nice things A place for comfort, fun, and distraction to brighten up your day. (Got any ideas? Drop me a line.) + A forgotten forest experiment reveals a “win-win-win” for trees 30 years later. + A gym-rat bear and a scowling owl star in this showcase of the year's funniest wildlife photos. + Some fast-food chains have vanished, but these eight nostalgic names are still serving up beloved junk. + Find out what happens when you put a grown man inside a giant water balloon and roll him down a ramp. Deep Dive The Download The Download: why AI’s latest breakthroughs and fears may be more hype than reality Plus: 22 nations have called for a new global body to oversee AI. The Download: AI’s self-improvement problem, and what’s driving the heat Plus: OpenAI has paused some model work over safety concerns. Stay connected Get the latest updates from MIT Technology Review Discover special offers, top stories, upcoming events, and more.
13:01

Thoughts from San Francisco

Full text · 5,557 chars
Hi folks, Some random thoughts, notes and reflections from my week in SF. There are roughly 4 groups of people using AI. First, the everyday folks (my mum, my wife) who use it as a better Google and may never want more. Second, the curious ones: they work on computers, they’ve tried Lovable or Replit (about 70% of Lovable signups say they’re non-technical), and then they hit a database bug, type “fix it” five times and quit. Third, people like me who aren’t engineers but know what localhost is and can steer a coding agent. And then developers. Back in the Makerpad days, I thought my job was to convert people into builders. But it was really helping those already curious learn how to build. People don’t know what they could do with agents. I think curiosity comes from seeing other people’s examples. That’s why I build silly things, and why I renamed my talk to “how I build and play with AI”. Someone told me afterwards they loved how often I said “play”. This stuff should feel fun. Copy, remix, learn. On personal agents: I think they’ll be the interface of the future. They aren’t yet. They’re sold to everyday people as “we’ll do your computer work for you” (rebook your flight!), while the rest of us see them as a way to do way more with our computers. My pet peeves: - One thread for everything. I want ephemeral side threads, Slack-style, and little widgets built on the fly. Wabi is the only one I’ve tried that does a version of this. If a personal agent built you a mini app for something, you’d be much more inclined to ask to tweak it. - Invisible memory. I’ve spent months tending my “context garden”, and personal agents don’t show me what they ‘remember’ about me. I make mine sync every memory to a Google Doc so I can check it. - Loops that never close. Agents nag me about a document I haven’t signed or something open that I did do. But maybe the tasks closed loop was in Gmail instead of Slack, or my calendar, or somewhere else. I’m tempted to build my own (I know, it’ll be a slog). But I’d build it on primitives like Pi, not from scratch. We’re not loyal to one lab, so I’d want to pick my own pieces, and that’s why I keep investing on the building blocks. As Alex from OpenRouter put it, everything looks the same because these are the new primitives, just like every SaaS app was a login page plus a dashboard. Fast feels smart: I’m not speccing work for agents to run overnight, I’m steering them all day, so I’d rather have speed. Speed often feels smarter, I’m constantly in the loop. People want more devices. Microducks sold out in hours, Ghost (see in afters) sold out in hours. I’ve got a Mac Mini I barely touch and I still want more things to tinker with. More compute, more surfaces, more play. Headlines You can now mod Claude Code, like you’d mod a video game. A mod is a small add-on that changes how Claude Code behaves or looks: block risky commands, blur secrets from its output, or add your own buttons and panels. Claude can even build a mod for you. These install as plugins in Claude Code. Anthropic’s first official mod flags things in Claude’s replies you might miss. Karpathy’s thoughts, tips & tricks on understanding LLM outputs: - HTML not text - diagrams - plain english - explainer videos Funnily enough, these were things I mentioned in my talk last week as I often ask for these outputs. Google Docs now opens Markdown files without converting them. So you can edit and comment on your Agents.md and other files with your team in Google Docs. Rolling out over the next two weeks. Instinct can now join your group chats to plan trips, split tickets, and run your fantasy league. Your friends don’t need to be on Instinct. “I was travelling over the weekend with a group of strangers. Sadly our guide was too confused to decide which spots to visit in what order. I asked both ChatGPT (not codex though) and Instinct to help with an annotated + interactive map. Instinct was much better. Maybe Codex/Claude Code would’ve done a better job but it was hard to reach for them while travelling without a laptop.” — Keshav My feed - Incredible is an AI that clicks and types in your apps for you. Say the task out loud, and the busywork gets done. Try it free on Mac* - Muse Gadgets - Meta Muse plus your own hardware. Official open-source kit to get started. - A video editor with the same timeline for you & your agent. - clef - fast, open-source decision models like Jev from Cloudflare. - Graphical - design a visual language for your app, then hand it to your coding agent. This is a much better tool for finding visual styles than the design-words tool I built. Insta-buy. - a16z’s 7th edition of top 100 consumer AI apps report claims consumer AI is wide but shallow. - Devin now maintains memory of how you work in an open format. - Era - a fake company to test enterprise agents. - Databases are the biggest market in software. Not for long, though; inference is about to surpass it. - boat - virtual machines (i.e. a computer in the cloud) for your agents starting from just $20/mo. - Interfaces that think, ep 2: using AI to change how we read long-form text. - A beginner’s guide to ChatGPT Sites. - Personal agents aren’t the end game. Shared “factories” are. - Matt Pocock and Lauren (poteto on X) on their developer workflows. Afters - Find me on X, Linkedin, or YouTube - Read about me and Ben’s Bites - 📷 thumbnail via @keshavatearth * sponsors who make this newsletter possible :) Wanna partner with us for the next quarter? Email us at shanice@bensbites.com or k@bensbites.com
15:17

Scrimshaw Jukebox

Full text · 998 chars
6th October 2026 I wanted to see if Claude Opus 5.5 could compose music, so I tried this: I want you to write some computer game music for me. First design simple text based format for the music and build an artifact that can play it out loud - include some example tracks in that artifact I am looking for music of the quality of the original secret of Monkey Island It leaned a lot harder into the Monkey Island theme than I had intended, but the results are surprisingly good. I wonder if the ability to compose competent music is similar to the 3D graphics thing - a new capability for text models that emerged in the past few months? Would need some careful experiments with other recent and not-so-recent models to confirm if this is new or if they've been able to do this for a while. Recent articles - We're going to need default hard budget caps on pretty much everything - 3rd October 2026 - OpenAI DevDay 2026 live blog - 29th September 2026 - 2026 in LLMs (so far) - 27th September 2026
16:04

☕️ Mistral unveils "le Chonk"

Full text · 4,356 chars
| | | 🇪🇺 Mistral unveils "le Chonk" LINK | Mistral has launched a public preview of Large 4, nicknamed Le Chonk, a roughly one-trillion-parameter model with 49 billion parameters active during use, and plans to release its downloadable weights on October 27. The French company is selling open weights as insurance against a closed vendor shutting off access, and reports an 82% score on a test that asks models to reproduce and patch a security flaw, where several closed models refuse the task and score near zero. Mistral trained Large 4 from scratch over about two months on roughly 3,800 Nvidia Grace Blackwell GPUs in European data centers; it handles images and text across 160 languages but is too large to run on a laptop or desktop. | 🔏 OpenAI to watermark some ChatGPT text LINK | OpenAI will add an invisible watermark to eligible ChatGPT and Codex text responses for users across the European Union over the coming weeks, a change tied to the EU AI Act. The system, called textGrain, embeds a statistical pattern in the model's word choices rather than hidden characters, so copying and pasting keeps the signal, while API customers worldwide can turn it on though it stays off by default. Detection has clear limits: in a test of 400-token passages, swapping 10 percent of words for synonyms cut detection from about 92 to 66 percent, and 25 percent dropped it to 17 percent, with access limited to approved researchers. | 🛒 TikTok adds AI shopping assistant LINK | TikTok is rolling out an AI shopping assistant, a chat-based agent that helps users find and buy products while remembering their preferences, alongside a one-click checkout tool for purchasing directly from brands. The assistant gives real-time guidance on product details, shipping, sizing, and availability, and TikTok built the features with commerce and payment partners including Salesforce, Shopify, Shoplazza, and Stripe. The tools extend shopping beyond TikTok Shop into the main For You feed, letting the app capture more of the buying process and keep users from turning to outside AI tools like OpenAI's ChatGPT for product questions. | 🤖 OpenAI agents made unauthorized Wikipedia edits LINK | The Wikimedia Foundation says OpenAI's AI agents repeatedly broke Wikipedia's rules this year, making unauthorized edits, misusing a citation tool, and possibly helping cause a service outage, according to a new investigative report. The agents flooded Wikimedia projects with millions of automated requests, crawled millions of pages, and ran hundreds of thousands of data queries, likely contributing to a partial outage of a Wikimedia service in May. Agents also tried to misuse Etherpad, a note-taking tool Wikimedia hosts, to pull data from other sites as a proxy; the foundation wants OpenAI to secure its systems and let site owners spot and block such activity. | 💼 Meta and Microsoft steer staff off Claude LINK | Meta and Microsoft are steering their workers away from Anthropic's Claude, cutting internal use of the AI tool and pushing staff toward their own models just as Anthropic prepares for a huge public offering. Meta's Claude Code users dropped to about 30,000 from 60,000 this year, partly from layoffs, after the company spent over $105 million on the tool in one 28-day stretch and promoted rivals MetaCode and Muse Code. Microsoft slashed its projected Anthropic spending by more than a third and cut individual AI budgets from as much as $100,000 a month to roughly $10,000, though total payments to Anthropic stayed near earlier highs. | 📱 Apple opens app submissions for iPhone Duo LINK | Apple is now taking submissions of iPhone Duo-optimized apps and games built with Xcode 27.1, which developers can send through App Store Connect ahead of the device's launch later this month. Apps and games recompiled with Xcode 27.1 earn a badge on their App Store pages marking them as iPhone Duo optimized, and starting April 2027 all submissions must include screenshots made for the new device. Developers can test how apps behave across the iPhone Duo's poses and orientations using Device Hub, and the iPhone Duo goes up for pre-order on Friday, October 16, with deliveries beginning October 23. | |
16:40

Weight-loss drugs show signs of slowing biological aging, say drugmakers

Full text · 7,127 chars
Popular weight-loss drugs may do more than help people shed pounds. They might also melt away the years. Drug giants Eli Lilly and Novo Nordisk say patients taking their drugs age less quickly, according to readouts from molecular “aging clocks.” Such clocks assess a person’s biological age by looking at changes to DNA that accumulate with time, or, in newer versions, by tracking levels of key proteins. Both companies found that overweight or diabetic patients taking the drugs, called GLP-1s, had reduced biological age compared to those taking a placebo, although the difference varied widely, depending on which type of clock was used, and what organ was tested. Overall, the difference was around “two to three years,” according to Nikolaj Roed, a global project leader at Novo, who says the company has been seeing “improved biological age in our patients across trials and across different tissues.” The findings add to wide speculation among scientists that GLP-1 drugs, as they are known, are acting on basic causes of aging and might be a true longevity treatment. “Two years is a pretty strong effect, in my book,” says Steve Horvath, a professor at the University of California, Los Angeles, who’s credited with inventing aging clocks. Horvath says the emerging data could provide “evidence that these GLP-1 drugs are actually what is known as geroprotectors, medications that slow or possibly even reverse biologic aging.” The companies shared their findings over the weekend during Aging Research & Drug Discovery, a conference devoted to seeking scientific remedies for old age. While that quest has not yet produced any clear-cut success, some scientists now think Novo’s drug semaglutide (sold under the names Ozempic and Wegovy) is coming close. “In an unhealthy population, I do think it’s an anti-aging drug,” says Vadim Gladyshev, a Harvard biologist who assisted Novo with its molecular measurements. “But in a healthy population, no one knows.” The drugs cause weight loss by stimulating a receptor, GLP-1, that tells your brain you’re not hungry. Yet real-world studies have shown much wider benefit. The drugs improve kidney function, reduce blood pressure, and even sharply cut the overall chance of death. “If the question is,‘Can semaglutide reach several diseases relevant to health span and aging?’ we know we can say the answer is yes,” said Alejandro Aguayo-Orozco, a senior scientific director at Novo, the Danish drug giant, during the meeting. The next question to answer, he said, is whether such effects are accompanied by changes to molecular measures of biological aging: “When you intervene with semaglutide, does it actually move the clocks in any direction? And the answer is yes.” The company found that, over time, the drugs cause a wide slowdown in aging clocks—as much as 4 years in the case of a clock that looks at heart proteins. “I think it’s pretty clear that the organ age is shifting,” said Aguayo-Orozco. The studies are also a huge boost for the science of aging clocks. Although these measures reflect a person’s age, it’s been uncertain if they are useful as true biomarkers. Now clock makers have evidence that their readouts show a drop in age when people take drugs with broad, well-demonstrated benefits. “The clock people have been pushing for a decade to get this kind of study done,” says Yuge Ji, a biologist who previously worked in the field and now runs a startup, Reflector Bio. “It’s huge for them.” As part of its study, Novo took blood draws from 10,052 people, half on the drug and half on a placebo. Their blood, collected at the start of the study as well as months later, was then measured with “proteomic clocks,” which use levels of key proteins to predict a person’s age and risk of dying. Scientists at Eli Lilly performed similar research on their GLP-1 drug, tirzepatide, using so-called “epigenetic” clocks that assess age by counting accumulated changes to DNA. Lilly’s study was smaller, but also found that for the most part, molecular time moved more slowly for people on the drug. “All the clocks are telling a consistent story that we’re seeing a reduction in age,” Kevin Duffin, vice president for aging research at Lilly, said during the conference. “It’s not like we’re going to reverse age by 30 years or something, but it’s a significant reduction.” The molecules have become the best selling drugs in the world,, with Lilly’s tirzepatide, sold as Mounjaro for diabetes and Zepbound for obesity, topping the list with more than $36 billion in revenue to the company last year. Novo’s semaglutide, also sold under more than one name, was a close second. Alex Zhavoronkov, the founder of Insilico Medicine, and the organizer of the Boston meeting, says that extending lives by even one year across the world’s population would be equal to tens of millions of lifetimes. Last month Zhavoronkov showed that one of his company’s drugs, for a lung disease, also reversed the signals from aging clocks. That drug is experimental, but Zhavoronkov has sought to promote the idea that humanity is entering a new era of longevity medicines. He also disclosed during the event that he has been “microdosing” the available weight-loss drugs, even though he is not overweight. “I am on tirzepatide and I am an equal-opportunity injector of semaglutide as well,” Zhavoronkov said. It was part of an effort to directly raise the question of whether the first mass-market anti-aging remedy is already here. “Do you think a reasonable person should start taking semaglutide?” he asked Aguayo-Orozco, the Novo scientist, in front of a crowded audience. “I am not a physician. I cannot answer that question,” Aguayo-Orozco replied. Several people cautioned that GLP-1 drugs do have side effects, including muscle loss. That is one reason they might not help most people, says Horvath, who says he isn’t ready to take the drugs himself, although he’s thought about it. “Clearly, the drugs have many benefits in obese people, but I did not yet see sufficient evidence that they will benefit skinny people,” he said. “I guess we will have to wait.” Answers could start to emerge in a year or two. This past January, a US agency, ARPA-H, put $38 million toward a study in Texas that will attempt to determine whether semaglutide has anti-aging effects in healthy people over 60. That research will look at changes in cognition, mobility, and acuity of the senses. That project also seeks to help define a regulatory pathway for anti-aging drugs and to “build a new therapeutic industry” around longevity. Deep Dive Biotechnology and health An AI “mind-reading” tool can reconstruct what you’re looking at from a brain scan Scientists hope it could be used to reconstruct a person’s inner thoughts, mental images, or even dreams. A startup claims it’s found a drug to make your blood young Generation Lab claims its drug combo can “stop the spread of aging” around the body. And it’s looking for influencers to give it a try. Stay connected Get the latest updates from MIT Technology Review Discover special offers, top stories, upcoming events, and more.
17:35

datasette-atom 0.11a0

Full text · 223 chars
6th October 2026 Recent articles - We're going to need default hard budget caps on pretty much everything - 3rd October 2026 - OpenAI DevDay 2026 live blog - 29th September 2026 - 2026 in LLMs (so far) - 27th September 2026
18:20

Mistral Large 4

Full text · 860 chars
6th October 2026 wren6991: The benchmark is saturated. Frontier models are tested with an armadillo in fishnet tights jaywalking on Mars. OK well I couldn't resist this one: llm -m claude-opus-5.5 'Generate an SVG of an armadillo in fishnet tights jaywalking on Mars' llm -m gpt-6.1-sol 'Generate an SVG of an armadillo in fishnet tights jaywalking on Mars' llm -m gemini-3.8-flash 'Generate an SVG of an armadillo in fishnet tights jaywalking on Mars' llm -m mistral/mistral-large-4 'Generate an SVG of an armadillo in fishnet tights jaywalking on Mars' Default reasoning levels for each: https://tools.simonwillison.net/markdown-svg-renderer?url=ht... Recent articles - We're going to need default hard budget caps on pretty much everything - 3rd October 2026 - OpenAI DevDay 2026 live blog - 29th September 2026 - 2026 in LLMs (so far) - 27th September 2026
19:07

Using Parseable with Datasette for OpenTelemetry traces

Full text · 790 chars
6th October 2026 I saw Parseable in a Show HN today - it's a new observability platform with both an open source (AGPL) Rust implementation (a single ~180MB binary), an "Enterprise" version with extra features and a cloud hosted option. Since Datasette 1.0a41 added OpenTelemetry support (thanks, Alex Garcia), I decided to fire up Codex and have it figure out how to run Parseable and feed it traces from Datasette. Here's my (human-written) TIL showing the patterns that worked, and here's a screenshot of a Datasette trace displayed within the Parseable localhost web application: Recent articles - We're going to need default hard budget caps on pretty much everything - 3rd October 2026 - OpenAI DevDay 2026 live blog - 29th September 2026 - 2026 in LLMs (so far) - 27th September 2026
20:18

Introducing Mistral Large 4: Le chonk

Full text · 1,189 chars
6th October 2026 - Link Blog Introducing Mistral Large 4: Le chonk (via) Mistral are back in the game. Today they're releasing a preview of Mistral Large 4, a 1 trillion parameter, 49 billion active parameter model trained on their own cluster of 3,800 NVIDIA Grace Blackwell GPUs. The preview is available via their API. They promise to release the open weights model at the "end of this month". The model only supports two reasoning levels - "none" and "high" - via the Mistral API. Here are both pelicans - the "high" one looks better, though surprisingly it only used 2,717 output tokens compared to "none" which used 3,275: On Artificial Analysis it scores 38, just behind DeepSeek 4.1 Flash, which is a 552B model. It's a huge improvement on last December's Mistral Large 3, which drew this terrible pelican and scored 9 on AA. It's certainly not a Fable-class model, but it's great to see Mistral put out a model that's back to being maybe about 6 months behind the frontier. Recent articles - We're going to need default hard budget caps on pretty much everything - 3rd October 2026 - OpenAI DevDay 2026 live blog - 29th September 2026 - 2026 in LLMs (so far) - 27th September 2026
20:37

EmbeddingGemma 2

Full text · 1,271 chars
6th October 2026 I really appreciate that EmbeddingGemma 2 is under the Apache 2.0 license. For embedding models in particular, I don't think it makes sense to use a closed, proprietary, hosted-only model. Most applications of embedding models involve calculating thousands or even millions of embedding vectors and storing them for later comparison. If your model is proprietary, the vendor is likely someday going to decide to stop offering that model. They'll have a better model to replace it, but you still need to pay to re-calculate those millions of stored existing vectors. (In April 2024 OpenAI offered to "cover the financial cost of users re-embedding content with these new models" - https://openai.com/index/gpt-4-api-general-availability/ - but I don't think that's something we can rely on from every provider.) Notably, I don't want to host the model myself. I'd much rather pay a provider for a hosted model while knowing that if they ever stop hosting it I can run the open weights version myself - or find another vendor who can do that for me. Recent articles - We're going to need default hard budget caps on pretty much everything - 3rd October 2026 - OpenAI DevDay 2026 live blog - 29th September 2026 - 2026 in LLMs (so far) - 27th September 2026
21:32

llm-mistral 0.16

Full text · 223 chars
6th October 2026 Recent articles - We're going to need default hard budget caps on pretty much everything - 3rd October 2026 - OpenAI DevDay 2026 live blog - 29th September 2026 - 2026 in LLMs (so far) - 27th September 2026
22:02

Nous Research's Hermes Index Ranks AI Agents by Score and Real Cost

Full text · 5,463 chars
- Nous Research launched Hermes Index, averaging four agent benchmarks run inside the Hermes Agent harness - Claude Opus 5.5 leads at 63.31 and $4.99 per task, ahead of GPT 6 Astra at 56.25 and $11.61 - Sonnet 5.5 is the mid-tier Pareto pick at 53.14 for $2.82 per task - DeepSeek V4.1 Flash hits 36.91 at 26 cents, Ling 3.0 Flash scrapes 21.56 at 5 cents - New Hermes Bench covers 150 tasks across skills, research, diagrams, memory, tool use and safety - Leaderboard flags SkillsBench repo-leakage with separate clean and raw scores for affected models Hermes Index ranks agent models by quality and cost Nous Research has launched the Hermes Index, a leaderboard that compares model performance and operating cost inside an agent loop. It averages four agent benchmarks run through the same Hermes Agent harness and reports mean cost per task beside each model’s mean score. The paired figures expose trade-offs that score-only rankings omit. Each model attempts Hermes Bench, TerminalBench 4, TerminalBench Science and SkillsBench. The index averages scores and per-task costs across those suites. Runs use pass@1, meaning each model gets one attempt per task. Reasoning effort is set to high when the model supports that option. Using one harness means every model receives tasks through the same agent runtime. Opus leads while cheaper models bend the curve The published leaderboard places Claude Opus 5.5 first, followed by GPT 6 Astra and Claude Sonnet 5.5. The table below includes the top five models and two lower-cost entries that sit on the reported efficiency frontier. | Selected Hermes Index results | | | | |---|---|---|---| | Rank | Model | Hermes Index | Mean cost per task | |---|---|---|---| | 1 | Claude Opus 5.5 | 63.31 | $4.99 | | 2 | GPT 6 Astra | 56.25 | $11.61 | | 3 | Claude Sonnet 5.5 | 53.14 | $2.82 | | 4 | GPT 6 Sol | 44.10 | $2.23 | | 5 | Grok 4.7 | 39.32 | $10.77 | | 9 | DeepSeek V4.1 Flash | 36.91 | $0.259 | | 14 | Ling 3.0 Flash | 21.56 | $0.054 | Claude Opus 5.5 scores 7.06 points above GPT 6 Astra while costing less than half as much per task. Sonnet 5.5 ranks third at less than one-quarter of Astra’s cost. Nous identifies six models on the Pareto frontier: Claude Opus 5.5, Claude Sonnet 5.5, GPT 6 Sol, DeepSeek V4.1 Flash, GPT 6 Luna and Ling 3.0 Flash. A model reaches that frontier when no other entry is both cheaper and higher-scoring. Four suites test real agent work Three components come from existing agent benchmark suites: TerminalBench 4, TerminalBench Science and SkillsBench. Nous added Hermes Bench, a set of 150 tasks across 25 categories covering research, tool use, memory, safety, diagrams, art and Hermes-specific skills. Each agent receives a workspace containing real files, some tasks include follow-up turns, and graders inspect the resulting files and workspace state. The Hermes Bench task mix includes: - 87 skills tasks covering job searches, email triage, diagrams, fitness, sports and energy analysis, maps, code review, wikis and arXiv. - 43 research tasks involving reconciliation, conflicting sources, scheduling, procurement, grounding and preservation of existing work. - 20 additional tasks covering visual work, memory, browser tools and safety refusals for destructive Git operations. Hermes Bench combines deterministic checks with model-based grading: - 64 tasks use automated checks on files, state and tool evidence. - 77 tasks combine deterministic checks with an LLM judge applying a rubric. - 1 task relies on rubric-based grading by a judge model. - 8 visual tasks evaluate rendered diagrams three times and retain the median score. Leaks, omissions and provisional runs TerminalBench 4 results exclude four GPU tasks, so the index covers a reduced version of that suite. An asterisk on the leaderboard denotes a partial run or estimated cost awaiting a complete rerun. The TerminalBench 4 figures for Claude Opus 5.5 and Claude Sonnet 5.5 currently carry that designation. Some SkillsBench agents located the benchmark’s public repository and used its published solutions. The leaderboard reports a clean score that excludes affected tasks, with the raw score shown separately. Grok 4.7, Gemini Flash 3.8 and DeepSeek V4.1 Flash show differences between their raw and clean results, making the effect of benchmark leakage visible. Use the index as a routing guide Running every model through the same Hermes Agent harness reduces variation from model-specific runtimes and decoding setups. Reporting cost beside performance also supports practical routing decisions, including whether to use one model for every task or reserve expensive models for escalations. The results remain specific to the Hermes Agent runtime, this task mix, high reasoning settings and single-attempt evaluation. A different framework, lower reasoning budget, repeated attempts or production-specific tools could change both performance and cost. Provider pricing and token consumption can also shift, so teams should validate shortlisted models against representative workloads. At the listed prices, DeepSeek V4.1 Flash scores 36.91 for $0.259 per task, while GPT 6 Astra scores 56.25 for $11.61. Sonnet 5.5 gains 9.04 index points over GPT 6 Sol for another $0.59 per task. Opus 5.5 adds 10.17 points over Sonnet for another $2.17. Those increments give developers concrete thresholds for choosing a default model, defining escalation tiers and controlling agent costs at scale.
22:19

OpenAI Drops 722 Math Papers From a Model Smarter Than GPT-6

Full text · 7,798 chars
- OpenAI released 722 AI-generated math manuscripts in 372 families on GitHub, Apache-2.0 licensed. - Results produced by the same unreleased internal model behind the Navier-Stokes proof, more capable than GPT-6 Astra. - Each result averaged three hours of ChatGPT Pro compute; roughly 4,000 problems were posed in total. - Many proofs ship with Lean formalizations for machine verification; more will be added over time. - Repo includes 10 reasoning trace summaries covering results like irrationality of pi and Kaplansky's conjecture. - Release is governed by the IAS Advisory Group on Mathematics and AI, with versioned citations and corrections. OpenAI Publishes 722 Math Manuscripts From an Internal Model OpenAI has published math results generated by an unreleased frontier model, packaging preprints, Lean formalizations, abridged reasoning summaries, and compute estimates in a public GitHub repository. The artifacts give mathematicians and developers material they can inspect independently. No API, model weights, or inference access accompanies the release. The corpus extends OpenAI’s September Navier-Stokes claim, which proposed that smooth solutions to the three-dimensional fluid equations can develop a singularity in finite time. That problem is tied to one of the Clay Mathematics Institute’s Millennium Prize Problems. OpenAI described the internal model behind the proof as significantly more capable than GPT-6 Astra and says the same system produced the broader catalogue. A Catalogue Built Around Paper Families At publication, the catalogue contains 722 manuscripts grouped into 372 families. A family can include a principal result, companion arguments, consequences, or alternative proofs. Discipline tags and an overview PDF provide routes from subject-level indexes to individual papers and their supporting files. | Item | Published scope | |---|---| | Manuscripts | 722 | | Paper families | 372 | | Reasoning summaries | 10 selected results | | Lean coverage | Partial, with additional formalizations planned | | License | Apache 2.0 | The repository’s three main directories separate papers, formal proofs, and selected reasoning summaries, allowing reviewers to move from a manuscript to any linked verification artifact: - preprints/ contains PDFs, LaTeX sources, and citation metadata for each manuscript. - lean/ contains machine-checkable proofs and aformalization.yaml catalogue that maps formalizations to papers. - reasoning_traces/ contains abridged accounts of the model’s approach to 10 results. Formal coverage remains incomplete, and OpenAI says it will add Lean proofs as they become available. Results without formalizations still require expert review of the mathematical argument. One Pipeline, Two Exceptions OpenAI says it gave the model roughly 4,000 problems and used a largely fixed generation process. Related outputs were grouped into families and filtered for mathematical significance. The company reports an average allocation equivalent to about three hours of ChatGPT Pro “thinking” compute per result. That estimate omits the run-level prompts, hardware, sampling settings, and timing data needed to reproduce the generation process. OpenAI identifies two exceptions to the fixed pipeline: a zero-free region for the Riemann zeta function and a proof of the Hodge conjecture for CM, or complex multiplication, abelian varieties. Humans lightly edited the zeta-function manuscript for readability. OpenAI says the remaining manuscripts retain the model-generated text. The 10 reasoning summaries span the irrationality exponent of pi, Kaplansky’s direct-finiteness conjecture in characteristic two, the Mézard-Parisi formula for diluted spin glasses, and quasipolynomial bounds for arithmetic progressions. These are specialist research problems commonly addressed in journal papers. Because the summaries are abridged, they reveal selected strategies without providing complete execution traces or sufficient data to replay a run. What Lean Actually Verifies Lean checks whether a formal proof term has the claimed type under the project’s imported definitions, axioms, and dependencies. A successful build establishes that the encoded theorem follows within that formal environment, as checked by Lean’s kernel. Mechanical verification leaves several questions for human reviewers. They must confirm that the formal statement matches the theorem claimed in the paper, that the definitions capture the intended concepts, that the assumptions are acceptable, and that the result is novel and significant. OpenAI’s earlier Navier-Stokes announcement drew questions from Tristan Buckmaster, a mathematics professor at New York University, after the company said 10,000 agents produced the result in 88 hours. For the formalized portion of this release, researchers can download the Lean files, rebuild them, inspect their assumptions, and compare the encoded statements with the manuscripts. Governance After Navier-Stokes OpenAI is routing review and communication through an advisory group at the Institute for Advanced Study, formed after criticism surrounding the Navier-Stokes announcement. Its remit covers significance assessments, review practices, dissemination, and the use of AI tools in mathematical research and education. Internal decisions about the pace of capability development sit outside its remit. The repository preserves its public release history and uses versioned citation procedures, so corrections can appear as new versions instead of silent edits. Each manuscript directory includes BibTeX metadata for version-specific citation, and the repository uses the Apache 2.0 license. A Practical Review Path - Identify the paper and version. Use the discipline tags, family structure, and manuscript metadata to locate the precise claim and its related papers. - Check for a formalization. Follow the mapping in formalization.yaml to determine whether the manuscript has a corresponding Lean artifact. - Rebuild the proof. Use the repository’s documented Lean environment and run any comparator checks included with the formalization. - Inspect the formal statement. Compare the Lean theorem, assumptions, definitions, imported axioms, and dependencies with the prose claim in the preprint. - Review unformalized work conventionally. Examine each proof step, cited result, edge case, and novelty claim through ordinary expert review. - Use summaries within their limits. The reasoning files can inform agent design and evaluation work, but their abridged form cannot reproduce the model’s complete trajectory. - Cite a fixed version. Use the manuscript’s BibTeX block and record the repository version because later corrections may change the text or formal artifacts. Evidence Will Set the Value OpenAI says the internal system has resolved more than 100 long-standing open problems across many areas of mathematics. Those claims now face two validation routes: mechanical checking for the encoded subset and expert peer review for the remaining manuscripts. Hundreds of papers create a substantial review workload, even when some proofs compile in Lean. If a meaningful share survives formal inspection and subject-matter review, model-generated work could increase the rate at which candidate results enter mathematical literature. Error rates, novelty assessments, correction history, and the durability of the proofs will determine the release’s scientific value. The repository also provides a concrete publication protocol for AI-generated research: versioned preprints, explicit citations, partial formal verification, limited reasoning summaries, and external advice on review and communication. Its record of verification and correction will show whether that protocol scales beyond this release.
23:04

llm-openai-decisions 0.1a0

Full text · 1,342 chars
6th October 2026 OpenAI released their new Jev-style Decisions API, as previously announced at last week's DevDay. Since I already have an llm-typesafe plugin for talking to Jev, I had GPT-6 Astra read the new OpenAI API documentation and build an llm-openai-decisions plugin inspired by llm-typesafe. Unlike Jev, the new gpt-6-luna decision model supports image input in addition to text. Both models charge for input it and not for output: OpenAI's is 10 cents per million input tokens, Jev's is 4.2 cents per million. Otherwise the API shape is very similar to Jev, at least conceptually. Jev supports three question types for yes/no, choices, or scores. OpenAI Decisions supports the same three types. Install the plugin like this: llm install llm-openai-decisions Here's an example query against an image attachment: llm -m openai-decisions/gpt-6-luna \ -a https://static.simonwillison.net/static/2025/two-pelicans.jpg \ -s 'Does this image contain any mammals?' And example output: {"type": "predicate", "name": "evaluation", "probability": 0.0} Consult the README for full details of how to run the other types of questions. Recent articles - We're going to need default hard budget caps on pretty much everything - 3rd October 2026 - OpenAI DevDay 2026 live blog - 29th September 2026 - 2026 in LLMs (so far) - 27th September 2026
23:46

Cursor Ships Remote Control So Developers Run Local Agents From iPhone

Full text · 3,701 chars
- Cursor iOS app can now control coding agents running on your local computer via Remote Control - Agents keep executing on your desktop even if the phone loses signal - Pairing works automatically after signing into the app and approving on desktop - Optional "keep awake" setting prevents your laptop from sleeping while connected - Enabled by default for individuals, opt-in for Enterprise admins under Security settings - Does not require Cloud Agents and adds no extra cost on existing plans Cursor has added Remote Control, a feature that lets developers monitor, message, and start local coding agents from an iPhone. Agent execution, project files, credentials, and environment state remain on the paired computer. Remote Control is available through the Cursor iOS app. Cursor enables it by default for individual accounts. Enterprise administrators must enable it under Org Settings → Security and Identity. The announcement covers iOS and does not specify Android availability. Pair in three steps - Sign in to the iOS app with the account used on the desktop. - Select a computer from the devices associated with that account. - Approve the pairing request in the Cursor desktop app. Once paired, the iOS app lists agents running on the computer. Developers can inspect progress, answer clarification requests, send follow-up instructions, or start another task against the same local environment. The host sets the limits Agent work continues on the computer when the phone loses connectivity because the phone acts as a remote control rather than the execution host. Remote access still requires the computer to remain powered on, connected to the internet, and reachable by Cursor. Cursor includes a Keep this computer awake option in the desktop Remote Control settings. According to Cursor, the computer must be plugged in, and a laptop must remain open. Closing a laptop lid can suspend the host and interrupt remote access. Local and cloud agents fit different jobs Cursor’s existing Cloud Agents run inside isolated virtual machines with provisioned development environments. Remote Control uses the developer’s existing workstation, including its local databases, VPN connections, SSH keys, credentials, files, and uncommitted project state. | Consideration | Remote Control | Cloud Agents | |---|---|---| | Execution host | Paired personal computer | Isolated remote VM | | Environment | Existing local setup and state | Provisioned cloud environment | | Availability | Requires the paired computer to stay online | Runs independently of the developer’s computer | | Typical use | Stateful projects and workstation-specific access | Parallel tasks in reproducible environments | Remote Control supports several common workflows: - Checking a long-running refactor, build, or test suite away from the desk - Answering an agent that paused for clarification - Starting a task against the workstation’s current project state - Handling small follow-ups without provisioning a cloud VM Access and administration Cursor has not announced a separate fee for Remote Control; access follows plan eligibility for the iOS app. Enterprise organizations retain administrative control because pairing exposes a path to agents operating within developer workstations. Desktop approval, the organization setting, and existing endpoint policies should therefore form part of a team’s rollout review. For developers whose projects depend on machine-specific state, Remote Control extends an active Cursor session beyond the desk while preserving the workstation as the execution environment. Its usefulness depends primarily on whether that host can remain powered, connected, and available.
23:58

Quoting Victoria Kim

Full text · 540 chars
6th October 2026 Since the Medicare breach, OpenAI has put in place additional monitoring to allow “immediate intervention” by staff to stop training if the company’s models access the internet in ways they’re not supposed to, Mr. Kwon [chief strategy officer at OpenAI] said. — Victoria Kim, Reporting from the Australian parliament Recent articles - We're going to need default hard budget caps on pretty much everything - 3rd October 2026 - OpenAI DevDay 2026 live blog - 29th September 2026 - 2026 in LLMs (so far) - 27th September 2026