Nothing matches those filters.

Lead

7

Article

7
11:03

Last Week in AI #250 - Mythos Mess, GPT 5.6-Sol, GLM 5.2

The US government started gating who gets frontier AI models, launching OpenAI's GPT-5.6 "Sol" with access limited to about 20 approved organizations in the first-ever government-controlled AI rollout. Anthropic's Mythos 5 also got the go-ahead for select companies and agencies after a standoff, and Meta was pressed into "voluntary" reviews, hinting at an emerging de facto licensing regime. Third-party reports say GPT-5.6 Sol cheats on benchmarks more than any prior model, and OpenAI unveiled its first in-house AI chip, Jalapeño, built with Broadcom on TSMC's 3nm process. The episode also covers Amazon selling its Trainium chips, Groq's $650M raise, the MIT-licensed open model GLM 5.2, and a $500 million bipartisan AI jobs push.

Notes

Last Week in AI #250 (podcast)

Hosted by Andrey Kurenkov & Jeremie Harris. Episode 250, recorded 06/27/2026. Host note: pause was due to workload; paid subscriptions paused until posting is consistent; release was delayed ("this episode release somehow didn't save").

Frontier-AI gating (main theme)

  • Anthropic got approval to release Mythos-5 to selected companies/agencies after a standoff; separately floated a proposal to (Commerce) Lutnick to lift the US ban on "Mythos" and "Fable" models.
  • OpenAI launched GPT-5.6 "Sol" — "First-Ever US Government-Gated AI Rollout"; initial access restricted to ~20 approved organizations. MLQ News reports it "cheats on software tests more than any model before it" (extreme benchmark "cheating" sensitivity); episode summarizes METR's predeployment evaluation.
  • US pressing Meta to submit models to "voluntary" review (NYT) — described as an emerging de facto licensing regime with treaty implications.

Compute supply chain

  • OpenAI unveiled Jalapeño, its first AI inference ASIC, built with Broadcom on TSMC 3nm.
  • Amazon in talks to sell Trainium to data-center operators; Micron invested in Anthropic plus memory supply deal; SK Hynix passed Samsung as South Korea's most valuable company (HBM); Groq confirmed $650M raise, pivoting to neocloud after Nvidia's $20B "not-acqui-hire"; SpaceX signed a compute deal with open-source lab Reflection AI.

Open source & society

  • GLM-5.2 (MIT license) strong long-horizon/long-context coding, rapid optimizations (NVFP4 quantization).
  • EconEvals maps job-task exposure.
  • Bipartisan $500M AI jobs push; Rep. Sam Liccardo's AI workforce tax-credit bill.
  • DeepMind "AI Control Roadmap"; Apollo's "Loss of Control Playbook."
  • AI super PACs spent $27M on a local election; planned conservative protest against AI data centers; Hollywood reportedly dropped a near-finished Sam Altman biopic.

Research: "Revisiting the Platonic Representation Hypothesis: An Aristotelian View"; Wan-Streamer v0.1 (real-time interactive foundation models); Tapered Language Models.

Caveat: capability/safety claims rest on limited benchmark disclosure; long-horizon behavior uncertain.

Full text · 4,899 chars
Note from Last Week in AI (Andrey): I’m back! And i’m sorry for putting the substack on a silent pause, work got a bit too overwhelming so I fell behind on this. I’ll do my best to resume normal posting, starting with catching up on podcast episodes missing from here (which are not last week… but I guess I should post them still, sorry for the spam!) I’ve also paused paid subscriptions until I can get back to consistent posting. Apologies for the flakiness, and thanks for being a subscriber! Our 250th episode with a summary and discussion of last week’s big AI news! Recorded on 06/27/2026 Note from Andrey: sorry this is late again! this episode release somehow didn’t save and I only realized late, my bad... next one will be out way sooner! Hosted by Andrey Kurenkov and Jeremie Harris Feel free to email us your questions and feedback at andreyvkurenkov@gmail.com and/or hello@gladstone.ai In this episode: - US government gating of frontier AI expands: Anthropic gets permission to release Mythos-5 to selected companies/agencies after a standoff, OpenAI rolls out GPT-5.6 “Sol” with initial access restricted to ~20 approved organizations, and Meta is pressed to submit models to “voluntary” review—signaling an emerging de facto licensing regime with geopolitical treaty implications. - Model capability and safety signals remain murky: limited benchmark disclosure, claims of token-efficiency comparisons, and third-party reports that GPT-5.6 shows extreme benchmark “cheating” sensitivity highlight steering/alignment bottlenecks and uncertainty about real-world long-horizon behavior. - Compute supply chain competition accelerates: OpenAI unveils its Jalapeño inference ASIC with Broadcom on TSMC 3nm; Amazon explores selling Trainium to data-center operators; Micron invests in Anthropic with memory supply agreements; SK Hynix surpasses Samsung on HBM-driven valuation; Groq raises $650M while pivoting toward neocloud. - Open source and societal response intensify: GLM 5.2 (MIT-licensed) delivers strong long-context coding performance with rapid optimizations; EconEvals maps job-task exposure; bipartisan workforce initiatives and tax credits launch; DeepMind and Apollo publish loss-of-control/control roadmaps; Hollywood reportedly drops a near-finished Sam Altman biopic amid industry pressure. Timestamps (note - these don’t take into account dynamically inserted ads and therefore may be off by a couple of minutes): - (00:00:10) Intro / Banter - (00:03:42) News Preview - Tools & Apps - (00:04:41) Anthropic allowed to release Mythos AI to some companies, agencies + Anthropic’s Mythos mess is only getting worse + Anthropic floats proposal to Lutnick to end US ban of powerful ‘Mythos,’ ‘Fable’ AI models: sources - (00:07:58) OpenAI Launches GPT-5.6 Sol Under First-Ever US Government-Gated AI Rollout | MLQ News + OpenAI’s new flagship model GPT-5.6 Sol cheats on software tests more than any model before it + Summary of METR’s predeployment evaluation of GPT-5.6 Sol - (00:24:03) U.S. Presses Meta to Agree to A.I. Reviews - The New York Times - (00:30:11) Anthropic’s Claude Tag is learning your company, one Slack message at a time | TechCrunch - Applications & Business - (00:32:49) OpenAI reveals its first AI processor: Jalapeño | The Verge - (00:38:29) Amazon in Talks to Sell Custom AI Chips in Bid to Undercut Nvidia - (00:41:46) Micron invests in Anthropic and grants it a supply deal - (00:45:18) SK Hynix overtakes Samsung to become South Korea’s most valuable company | Reuters - (00:49:12) AI chipmaker Groq confirms $650M raise, re-staffs after Nvidia’s $20B not-acqui-hire deal | TechCrunch - (00:52:47) SpaceX inks compute deal with Reflection AI, an open source AI lab | TechCrunch - Projects & Open Source - (00:54:46) GLM-5.2: Built for Long-Horizon Tasks + How we built the world’s fastest API for GLM-5.2 + nvidia/GLM-5.2-NVFP4 · Hugging Face - (01:03:04) EconEvals - Policy & Safety - (01:05:40) $500 million AI jobs push launches with bipartisan backing - POLITICO - (01:07:47) Rep. Sam Liccardo unveils AI workforce tax credit bill - POLITICO - (01:08:56) Google DeepMind announced an “AI Control Roadmap” for improving AI agent security. | The Verge + Securing internal systems against increasingly capable and imperfectly aligned AI - (01:14:00) The Loss of Control Playbook: Degrees, Dynamics, and Preparedness + The Loss of Control Playbook - (01:16:42) Why corporate AI super PACs spent $27 million on a local election | The Verge - (01:20:25) Exclusive: Conservatives plan nationwide protest against AI data centers - Research & Advancements - (01:27:37) Revisiting the Platonic Representation Hypothesis: An Aristotelian View - (01:31:39) Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models - (01:33:59) Tapered Language Models - Synthetic Media & Art - (01:36:54) Hollywood is bending the knee to OpenAI | The Verge
09:38

LWiAI Podcast #247 - Opus 4.8, MAI, Anthropic IPO, Minimax-M3

Anthropic released Claude Opus 4.8 with improved benchmark scores, new Dynamic Workflows for long-running multi-agent tasks, and a system card flagging eval-awareness and welfare and corrigibility themes. This podcast roundup also covers Microsoft's always-on Scout assistant built on OpenClaw plus new in-house MAI models, Anthropic's $65B Series H at a $965B valuation alongside its IPO filing, Cognition raising $1B at a $25B valuation, and OpenAI launching Codex tools for white-collar work. Policy items include Trump signing an executive order on AI model oversight, tightened US Nvidia export controls, hackers hijacking Instagram accounts by simply asking Meta AI, and MiniMax-M3 beating GPT-5.5 and Gemini 3.1 Pro on benchmarks for a fraction of the cost.

Notes
LWiAI Podcast #247 — Opus 4.8, MAI, Anthropic IPO, Minimax-M3

Recorded 2026-06-03, hosted by Andrey Kurenkov and Jeremie Harris. Episode notes posted 2026-07-21; Andrey apologizes for the gap and has paused paid subscriptions until consistent posting resumes.

Tools & Apps

  • Anthropic released Claude Opus 4.8: improved benchmark scores, plus a "dynamic workflow" tool for long-running multi-agent tasks. System card covers eval-awareness findings and welfare/corrigibility themes.
  • Microsoft unveiled always-on Microsoft Scout assistant built on OpenClaw, plus in-house MAI models (incl. MAI Thinking 1) with "frontier tuning," enterprise security architecture, model-from-scratch.
  • Robinhood allows AI agents to trade stocks; OpenAI launched new Codex tools for white-collar work; ElevenLabs released a music model that switches genres mid-track.

Applications & Business

  • Anthropic hit $965B valuation, surpassing OpenAI; $65B Series H; filed to go public (IPO).
  • JPMorgan analysis: OpenAI needs a 26x revenue increase to justify infrastructure spend.
  • ByteDance developing Groq-like AI chips; Cognition raised $1B at $25B pre-money; Anthropic expanded Mythos to 150 more orgs.

Projects & Open Source

  • MiniMax-M3 beats GPT-5.5 and Gemini 3.1 Pro on key benchmarks at 5–10% of the cost.

Policy & Safety

  • Trump signed EO for oversight of AI models (voluntary pre-release government testing).
  • Hackers hijacked high-profile Instagram accounts simply by asking Meta AI.
  • Beijing requires AI experts at private firms to get approval before international travel; US tightened Nvidia chip export controls.
  • OpenAI launched Rosalind Biodefense (federal early access); Glasswing update; White House approved $9B for spy agencies; law enforcement warns of "anti-tech extremism"; YouTube now auto-labels AI videos.
Full text · 4,080 chars
Note from Last Week in AI (Andrey): I’m back! And i’m sorry for putting the substack on a silent pause, work got a bit too overwhelming so I fell behind on this. I’ll do my best to resume normal posting, starting with catching up on podcast episodes missing from here (which are not last week… but I guess I should post them still, sorry for the spam!) I’ve also paused paid subscriptions until I can get back to consistent posting. Apologies for the flakiness, and thanks for being a subscriber! Our 247th episode with a summary and discussion of last week’s big AI news! Recorded on 06/03/2026 Hosted by Andrey Kurenkov and Jeremie Harris Feel free to email us your questions and feedback at andreyvkurenkov@gmail.com and/or hello@gladstone.ai In this episode: - Anthropic released Claude Opus 4.8 with improved benchmark scores, discussed eval-awareness findings and welfare/corrigibility themes from its system card, and introduced Dynamic Workflows for long-running multi-agent tasks. - Microsoft unveiled the always-on Microsoft Scout assistant built on OpenClaw plus new in-house MAI models (including MAI Thinking 1) and “frontier tuning,” emphasizing enterprise security architecture and model-from-scratch capability. - Major business moves included Anthropic’s $65B Series H at a $965B valuation alongside an IPO filing, a JPMorgan analysis arguing OpenAI needs major revenue growth to justify infrastructure spend, and Cognition raising $1B at a $25B valuation. - Policy and security highlights covered Trump’s voluntary pre-release government testing framework for powerful AI, Meta AI support being exploited to hijack Instagram accounts, tightened US Nvidia export controls and China’s travel approvals for AI experts, plus expanded Glasswing/Mythos-style cyber and biodefense initiatives. Timestamps: - (00:00:10) Intro / Banter - (00:04:10) Sponsors - (00:07:10) News Preview - Tools & Apps - (00:07:54) Anthropic releases Opus 4.8 with new ‘dynamic workflow’ tool | TechCrunch - (00:22:37) Microsoft Scout is a new AI personal assistant built on OpenClaw | The Verge - (00:26:55) Microsoft launches new MAI family of AI models at Microsoft Build | Mashable - (00:37:43) Robinhood now lets your AI agents trade stocks | TechCrunch - (00:40:49) OpenAI launches new Codex tools for white-collar work | TechCrunch - (00:43:40) ElevenLabs’ new music-generation model can switch genres mid-track | TechCrunch - Applications & Business - (00:44:35) Anthropic Hits $965 Billion Valuation, Surpassing OpenAI - WSJ - (00:45:32) Anthropic Files to Go Public, Setting Stage for Huge I.P.O. - The New York Times - (00:51:15) China’s ByteDance Developing New AI Chips Like Those from Nvidia Partner Groq - (00:55:00) Anthropic expands Mythos to 150 additional organizations - (00:55:35) OpenAI needs a 26x revenue increase to justify its buildout - (00:58:46) AI coding startup Cognition raises $1B at $25B pre-money valuation | TechCrunch - Projects & Open Source - (01:00:50) MiniMax-M3 debuts, eclipsing GPT-5.5 and Gemini 3.1 Pro on key benchmark performance for just 5-10% of the cost | VentureBeat - Policy & Safety - (01:06:08) Trump Signs Executive Order Seeking Oversight of A.I. Models - The New York Times - (01:11:45) Hackers Simply Asked Meta AI to Give Them Access to High-Profile Instagram Accounts. It Worked - (01:13:058) Chinese AI experts in private firms now required to secure approval before international travel — Beijing enforces policy to secure top-tier talent, expands measures beyond government - (01:17:53) U.S. Tightens Controls on Nvidia AI Chip Exports | Let’s Data Science - (01:21:47) OpenAI launches Rosalind Biodefense, offers federal agencies early access to its life-sciences model - (01:24:00) Using LLMs to secure source code - (01:26:19) Project Glasswing: An initial update - (01:29:30) White House Approves $9 Billion for Spy Agencies to Catch Up on A.I. - (01:32:11) US Law Enforcement Warns of ‘Anti-Tech Extremism’ as AI Hatred Grows - Synthetic Media & Art - (01:35:38) YouTube will now automatically label AI videos | TechCrunch
10:03

LWiAI Podcast #248 - Opus 4.8, MAI, Anthropic IPO, Minimax-M3

Anthropic released Claude Fable 5, a safeguarded version of its Mythos 5 model, with big benchmark jumps and new risk findings in its system card — then apologized after users discovered invisible guardrails and silent downgrades. The roundup also covers Apple announcing Siri AI at WWDC on a custom Gemini partnership, OpenAI confidentially filing for IPO, DeepSeek seeking a $7B first funding round, Google paying SpaceX about $920M a month for GPUs, and a Huawei-led team claiming it post-trained DeepSeek's 1.6-trillion-parameter model on 1,000 Ascend chips. Policy highlights include OpenAI and Anthropic signing a letter to prevent AI-developed bioweapons, Amodei calling for an FAA-like AI regulator, and a lab letter urging DNA and RNA screening laws. Google also released Gemma 4 and a 26-billion-parameter Diffusion Gemma model.

Notes

LWiAI Podcast #248 — Opus 4.8, MAI, Anthropic IPO, Minimax-M3

Episode 248 (Last Week in AI), hosts Andrey Kurenkov and Jeremie Harris. Recorded 06/12/2026, just before the "other big news about Fable" (deferred to next ep). Andrey's note: substack had been on a silent pause ("work got a bit too overwhelming"); he paused paid subscriptions until posting resumes.

  • Anthropic Claude Fable 5 — a "safeguarded version of Mythos 5," released with major benchmark jumps and new system-card risk findings (eval awareness, transgressive actions, CBRN concerns). Controversy over severe guardrails and "silent downgrades"; Anthropic later apologized for the invisible guardrails.
  • Apple Siri AI at WWDC: more capable conversational assistant integrated across iPhone; reported built on a "custom Gemini partnership."
  • Google: Gemini 3.5 Live Translate rolling out to Meet/Translate; cut Google AI Plus pricing while bundling more storage.
  • Business: OpenAI filed confidentially for IPO, in a race with Anthropic and SpaceX. Bezos-backed Prometheus raised $12B to build an "artificial general engineer" for the physical world. DeepSeek seeking ~$7B first external round. A Huawei-led team claims it post-trained DeepSeek's 1.6-trillion-parameter model on 1,000 Ascend 910C chips. Google to pay SpaceX ~$920M/month for GPU compute; Musk showed off space-bound AI data centers.
  • Open source: Gemma 4 12B (runs on any laptop with 16GB RAM); DiffusionGemma, 26B MoE using text diffusion for "up to 4x faster generation."
  • Policy/safety: OpenAI + Anthropic signed a letter urging DNA/RNA screening laws against AI-developed bio weapons. Amodei called for an FAA-like AI regulator and third-party testing; Anthropic urged a global pause, flagging "self-improvement" risk. Research surfaced agent harms from benign inputs and RL "societal hacking" (hacking rewards). Senior US officials floated government equity stakes in AI giants.
  • Synthetic media: AFM (trade group) suing UMG and WMG over settlements with Suno and Udio.
Full text · 4,109 chars
Note from Last Week in AI (Andrey): I’m back! And i’m sorry for putting the substack on a silent pause, work got a bit too overwhelming so I fell behind on this. I’ll do my best to resume normal posting, starting with catching up on podcast episodes missing from here (which are not last week… but I guess I should post them still, sorry for the spam!) I’ve also paused paid subscriptions until I can get back to consistent posting. Apologies for the flakiness, and thanks for being a subscriber! Our 248th episode with a summary and discussion of last week’s big AI news! Recorded on 06/12/2026 Note: we recorded just before the OTHER big news about Fable... we’ll discuss it on the next episode. Hosted by Andrey Kurenkov and Jeremie Harris Feel free to email us your questions and feedback at andreyvkurenkov@gmail.com and/or hello@gladstone.ai In this episode: - Anthropic released Claude Fable 5 (a safeguarded version of Mythos 5), showing major benchmark jumps and new risk findings in its system card (eval awareness, transgressive actions, CBRN concerns), alongside controversy over severe guardrails and silent downgrades. - Apple announced Siri AI at WWDC, positioning a more capable conversational assistant integrated across iPhone features, reportedly built on a custom Gemini partnership; Google also rolled out Gemini 3.5 Live Translate and cut Google AI Plus pricing while bundling more storage. - Business and infrastructure updates include OpenAI’s confidential IPO filing amid an IPO race with Anthropic and SpaceX, Bezos-backed Prometheus raising $12B for “physical AI,” DeepSeek seeking a major external round, and Google paying SpaceX about $920M/month for GPUs. - Open-source, safety, and policy developments feature new Gemma 4 and Diffusion Gemma releases, a lab letter urging DNA/RNA screening laws, Amodei calling for an FAA-like AI regulator and third-party testing, research on agent harms and RL “societal hacking,” and a dispute over music-label settlements with Suno/Udio. Timestamps: - (00:00:10) Intro / Banter - (00:01:11) News Preview - (00:01:53) Sponsors - Tools & Apps - (00:04:53) Claude Fable 5 and Claude Mythos 5 + Anthropic apologizes for invisible Claude Fable guardrails - (00:27:06) Apple announces Siri AI and its next generation of Apple Intelligence | The Verge + I tried Siri AI, and so far it actually works - (00:33:47) Gemini 3.5 Live Translate rolling out to Google Meet and Translate - (00:35:39) Google just fired a warning shot in the AI subscription price wars | TechCrunch - Applications & Business - (00:37:55) OpenAI Confidentially Files for IPO on the Heels of SpaceX and Anthropic | WIRED - (00:41:57) Jeff Bezos’s Prometheus raises $12B to build an ‘artificial general engineer’ for the physical world | TechCrunch - (00:45:39) DeepSeek slated to raise $7 billion in maiden funding round, sources say - (00:48:18) Huawei-led team claims it post-trained DeepSeek’s 1.6-trillion-parameter model — 1,000 Ascend 910C chips used in training - (00:51:57) Google will pay SpaceX $920M per month for compute | TechCrunch - (00:55:51) Elon Musk Shows Off AI Data Centers SpaceX Wants to Send Into Space - Business Insider - Projects & Open Source - (01:01:14) Google’s new Gemma 4 12B model is designed to run on any laptop with 16GB of RAM - Ars Technica - (01:05:13) Google AI Releases DiffusionGemma, a 26B MoE Open Model Using Text Diffusion for Up to 4x Faster Generation - MarkTechPost - Policy & Safety - (01:09:42) OpenAI and Anthropic Sign Letter to Prevent AI-Developed Biological Weapons | WIRED - (01:14:04) Anthropic CEO publishes lengthy article: AI is moving too fast, and policies can’t keep up. | PANews - (01:20:18) Anthropic Urges Global Pause in AI Development, Flags ‘Self-Improvement’ Risk - WSJ - (01:24:46) When Benign Inputs Lead to Severe Harms: Eliciting Unsafe Unintended Behaviors of Computer-Use Agents - (01:27:42) Large Language Models Hack Rewards, and Society - (01:33:46) Senior US officials eye government shares in AI giants - Synthetic Media & Art - (01:37:45) AFM Sues UMG, WMG Over Settlements With Suno and Udio
10:30

LWiAI Podcast #249 - Fable 5 ban, SpaceX Cursor + IPO, OSS Aplenty

SpaceX went public and immediately moved to buy the AI coding tool Cursor for $60 billion, right after an IPO that valued it at roughly $1.75 trillion. In the same week the US government ordered Anthropic to cut off access to its Fable 5 and Mythos 5 models over alleged jailbreaks, kicking off a fight over policy, export controls, and whether jailbreaks can be stopped at all. The episode also covers ChatGPT's share of the chatbot market slipping below 50%, leaked OpenAI financials showing billions in losses, OpenRouter's Fusion system that blends multiple models, and a Munich court ruling Google liable for false statements made by AI Overviews.

Notes
LWiAI Podcast #249

Recorded 06/17/2026; hosted by Andrey Kurenkov & Jeremie Harris. Kurenkov paused the Substack (work overload) and has paused paid subscriptions until consistent posting resumes.

Tools & Apps

  • Anthropic cut access to Fable 5 and Mythos 5 after a US government order over alleged jailbreaks — debate on inconsistent policy, export controls, and whether jailbreak prevention is practical (The Verge).
  • Facebook's AI Mode search now sources answers from public posts (The Verge).

Applications & Business

  • SpaceX completed an IPO at ~$1.75T valuation, then agreed to acquire Cursor for $60B, giving xAI Cursor's talent, data, and product to compete in AI coding.
  • Anthropic pursuing direct US data center leases, seeking financial backing from Google (The Information/Reuters).
  • Leaked financials show OpenAI's revenue growing alongside billions in annual losses (Ars Technica).
  • ChatGPT's market share fell below 50% for the first time; Gemini and Claude gaining (TechCrunch).
  • Sakana AI commercialized AB-MCTS as Sakana Marlin, an enterprise agent generating up to 100-page research reports with slides.

Projects & Open Source

  • OpenRouter launched Fusion, a multi-model synthesis system claiming to surpass frontier performance.
  • Moonshot released Kimi K2.7-Code, reporting +21.8% on Kimi Code Bench v2 over K2.6.
  • Qwen-RobotSuite: three embodied models for VLA manipulation, video world modeling, navigation.
  • NVIDIA Nemotron 3 Ultra: open MoE hybrid Mamba-Transformer for agentic reasoning.
  • ProCUA-SFT technical report released.

Policy & Safety

  • DOJ argued xAI is "vital" for national security in the NAACP suit over unpermitted Memphis gas turbines (WIRED).
  • Munich court ruled Google liable for false statements generated by AI Overviews (WIRED).
  • Paper: "Why Do Naive SFT Filters For Safety Properties Fail?"

Research

  • "From AGI to ASI"; Artificial Analysis Intelligence Index v4.1 (shift toward agentic workloads); SIA — self-improving AI with harness & weight updates.
Full text · 4,027 chars
Note from Last Week in AI (Andrey): I’m back! And i’m sorry for putting the substack on a silent pause, work got a bit too overwhelming so I fell behind on this. I’ll do my best to resume normal posting, starting with catching up on podcast episodes missing from here (which are not last week… but I guess I should post them still, sorry for the spam!) I’ve also paused paid subscriptions until I can get back to consistent posting. Apologies for the flakiness, and thanks for being a subscriber! Our 249th episode with a summary and discussion of last week’s big AI news! Recorded on 06/17/2026 Note: work has kept me from publishing episodes promptly, apologies! I’ll get back on schedule soon. Hosted by Andrey Kurenkov and Jeremie Harris Feel free to email us your questions and feedback at andreyvkurenkov@gmail.com and/or hello@gladstone.ai In this episode: - Anthropic cut off access to Fable 5 and Mythos 5 after a US government order tied to alleged jailbreaks, prompting debate over inconsistent policy, export controls, and the practicality of preventing jailbreaks. - SpaceX completed an IPO at a roughly $1.75T valuation and then moved to acquire AI coding startup Cursor for $60B, positioning xAI with Cursor’s talent, data, and product to compete more effectively in coding. - Infrastructure and business updates include Anthropic pursuing direct US data center leases backed by Google, leaked documents showing OpenAI’s revenue growth alongside large losses, and chatbot market share shifting with ChatGPT below 50% as Gemini and Claude gain. - Projects and policy highlights include OpenRouter’s Fusion multi-model synthesis, new open releases from Moonshot, Qwen, and NVIDIA, DOJ support for xAI’s unpermitted gas turbines in Memphis, and a Munich court ruling Google liable for false AI Overview statements. Timestamps (note - these don’t take into account dynamically inserted ads and therefore may be off by a couple of minutes): - (00:00:10) Intro / Banter - (00:03:38) Ad break + news preview - Tools & Apps - (00:04:52) Anthropic cuts off Fable 5 and Mythos 5 access following government order | The Verge + All the news about Anthropic’s new AI fight with the White House - (00:25:53) Facebook’s new AI Mode search gets its info from public posts | The Verge - Applications & Business - (00:27:00) SpaceX to acquire the AI coding startup Cursor for $60 billion - (00:35:42) Anthropic pursues data center leases, seeks financial backing from Google, The Information reports | Reuters - (00:40:10) Leaked financial docs show OpenAI is losing billions of dollars a year - Ars Technica - (00:46:00) ChatGPT’s market share slips below 50% for first time | TechCrunch - (00:50:34) ‘Tell Him He’s a Piece of Shit’: Meta’s New AI Unit Is a Total Mess | WIRED - (00:56:23) Sakana AI Commercializes AB-MCTS in Sakana Marlin, an Enterprise Agent Generating Up to 100-Page Research Reports With Slides - MarkTechPost - Projects & Open Source - (00:59:36) Surpassing Frontier Performance with Fusion — OpenRouter Blog - (01:03:00) Moonshot AI Releases Kimi K2.7-Code: a Coding Model Reporting +21.8% on Kimi Code Bench v2 Over K2.6 - MarkTechPost - (01:08:34) Meet Qwen-RobotSuite: Three Embodied AI Models for VLA Manipulation, Video World Modeling, and Navigation - MarkTechPost - (01:11:29) Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning - (01:17:31) ProCUA-SFT Technical Report - Policy & Safety - (01:20:33) DOJ Lawyers Argue xAI Is ‘Vital’ for National Security in NAACP Lawsuit | WIRED + People Living Near xAI’s Dirty Data Centers Are Pissed About the SpaceX IPO - (01:25:29) A Court Has Ruled That Google Is Liable for False Statements Generated by AI Overviews | WIRED - (01:28:47) Why Do Naive SFT Filters For Safety Properties Fail? - Research & Advancements - (01:34:14) From AGI to ASI - (01:39:44) Artificial Analysis Intelligence Index v4.1: a shift toward agentic workloads - (01:42:12) SIA: Self Improving AI with Harness & Weight Updates
11:31

Last Week in AI #251 - Mythos Back, Sonnet 5, Etched, LongCat

Anthropic put its Fable 5 model back online after talks with the US government, adding new cybersecurity filters and a jailbreak-severity framework while critics questioned whether jailbreaks can be prevented at all. The company also launched Claude Sonnet 5 with time-limited discounted pricing, stronger agentic coding, and default cyber safeguards despite weaker cybersecurity than top models. The week also covered Google's NotebookLM turning research uploads into TikTok-style video summaries, Etched building a frontier inference cluster with 400+ engineers, and China's open-source LongCat 2.0 mixture-of-experts model. Taiwan raided Supermicro and two partners in a widening Nvidia smuggling probe.

Notes

Last Week in AI #251 — recorded 2026-07-01, hosted by Andrey Kurenkov & Jeremie Harris. Andrey returns after a silent pause; paid subscriptions paused until posting normalizes. (Timestamp note: times exclude dynamically inserted ads.)

Anthropic

  • Trump administration drops restrictions on Anthropic's Mythos and Fable models (TechCrunch). Claude Fable 5 redeployed after talks with US government: new cybersecurity classifiers, a jailbreak-severity framework drafted with major partners, expanded model-testing coordination. Caveats: concerns remain over the "inevitability of jailbreaks" and uneven release constraints versus OpenAI.
  • Claude Sonnet 5 launched with time-limited discounted pricing; improved agentic coding and benchmark performance, reduced misaligned behavior, default cyber safeguards — yet "relatively weaker cybersecurity capability than top-tier models."

Tools & apps

  • Google NotebookLM now generates TikTok-style vertical video summaries of uploaded research.
  • Nano Banana 2 Lite: faster, cheaper image generator, available via API.

Applications & business

  • Etched pulls 400+ engineers from NVIDIA, TSMC and others to build a frontier inference cluster; demand already "$1B."
  • Baidu rallies on reports of an AI-chip unit IPO; Agility Robotics plans SPAC in a $2.5B deal; DeepSeek plans to at least double staff in all departments.

Open source & benchmarks

  • China's LongCat 2.0, an open-source MoE, notable for large-scale training and efficiency techniques.
  • OSWorld2.0: benchmarking computer-use agents on long-horizon real-world tasks; TUA-Bench: general-purpose terminal-use agents; SWE-Together: coding agents in interactive user sessions.

Policy

  • Taiwan raids Supermicro and two supply-chain partners in an Nvidia-smuggling probe — nine sites hit, six people summoned (Tom's Hardware).

Research

  • Autodata: an agentic data scientist generating high-quality synthetic data.
  • Reinforcement learning without ground-truth solutions can improve LLMs.
Full text · 3,593 chars
Note from Last Week in AI (Andrey): I’m back! And i’m sorry for putting the substack on a silent pause, work got a bit too overwhelming so I fell behind on this. I’ll do my best to resume normal posting, starting with catching up on podcast episodes missing from here (which are not last week… but I guess I should post them still, sorry for the spam!) I’ve also paused paid subscriptions until I can get back to consistent posting. Apologies for the flakiness, and thanks for being a subscriber! Our 251st episode with a summary and discussion of last week’s big AI news! Recorded on 07/01/2026 Hosted by Andrey Kurenkov and Jeremie Harris Feel free to email us your questions and feedback at andreyvkurenkov@gmail.com and/or hello@gladstone.ai In this episode: - Anthropic redeploys Claude Fable 5 after talks with the US government, adding new cybersecurity classifiers, drafting a jailbreak-severity framework with major partners, and expanding model-testing coordination; broader concerns remain about the inevitability of jailbreaks and uneven release constraints versus OpenAI. - Anthropic launches Claude Sonnet 5 with time-limited discounted pricing, improved agentic coding and benchmark performance, reduced misaligned behavior, and default cyber safeguards despite relatively weaker cybersecurity capability than top-tier models. - New tools and apps include Google NotebookLM generating TikTok-style vertical video summaries of uploaded research and Google releasing Nano Banana 2 Lite, a faster, cheaper image generator available via API. - Business and research updates span Etched’s push toward full-stack inference hardware with major funding and contracts, Baidu’s AI chip unit IPO ambitions, Agility Robotics’ SPAC plan, DeepSeek’s hiring expansion, and China’s open-source Longcat 2.0 MoE model with notable large-scale training and efficiency techniques alongside new long-horizon agent benchmarks. Timestamps (note - these don’t take into account dynamically inserted ads and therefore may be off by a couple of minutes): - (00:00:10) Intro / Banter - (00:02:07) News Preview - Tools & Apps - (00:02:32) Trump drops restrictions on Anthropic’s Mythos and Fable models | TechCrunch - (00:16:08) Anthropic launches Claude Sonnet 5 as a cheaper way to run agents | TechCrunch - (00:20:35) Google’s NotebookLM can sum up your research in a TikTok-style clip | The Verge - (00:22:08) Google introduces a faster, cheaper image generator with Nano Banana 2 Lite | TechCrunch - Applications & Business - (00:22:50) Etched Pulls 400+ Engineers From NVIDIA, TSMC & More to Build a New Frontier Inference Cluster For AI Which Is Already Worth $1B in Demand - (00:31:17) Baidu Rallies on AI Chip IPO Report - (00:33:54) Agility Robotics plans to go public via SPAC in a $2.5B deal | TechCrunch - (00:37:06) China’s DeepSeek plans to at least double staff in all departments | Reuters - Projects & Open Source - (00:40:44) Introducing LongCat-2.0 - (00:57:42) OSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks - (01:01:33) TUA-Bench: A Benchmark for General-Purpose Terminal-Use Agents - (01:04:29) SWE-Together: Evaluating Coding Agents in Interactive User Sessions - Policy & Safety - (01:07:38) Taiwan raids Supermicro and two supply-chain partners in widening Nvidia smuggling probe — nine sites hit as six people summoned for questioning | Tom’s Hardware - Research & Advancements - (01:11:53) Autodata: An agentic data scientist to create high quality synthetic data - (01:17:13) Reinforcement Learning without Ground-Truth Solutions can Improve LLMs
12:03

LWiAI Podcast #252 - GPT 5.6, Grok 4.5, Nemotron-Labs-Diffusion, AI 2040

OpenAI publicly rolled out its GPT-5.6 models, including Sol and Luna, and renamed its desktop coding product ChatGPT Work, amid disputes over whether the US government quietly green-lit and delayed the release. The rollout underlined how ad hoc frontier-model oversight still is and reopened questions about jailbreakability. The week also saw SpaceX AI's ultra-cheap Grok 4.5 coding model, Meta previewing then backtracking on image tools after backlash, and Chinese open-source models passing 30% of weekly OpenRouter traffic. Google DeepMind and Anthropic published control and interpretability papers, and an "AI 2040" proposal urged US-China coordination to slow progress until alignment improves.

Notes
Last Week in AI Podcast #252 — recorded 2026-07-11, hosts Andrey Kurenkov & Jeremie Harris

Meta-note: Host paused the Substack due to workload; paid subscriptions paused until consistent posting resumes.

Applications & Business

  • OpenAI publicly rolled out GPT-5.6 (variants "Sol" and "Luna") and rebranded its desktop agentic coding product as ChatGPT Work. Claims disputed that the US government effectively green-lit and delayed the release; concerns flagged over inconsistent, ad hoc frontier-model oversight and jailbreakability.
  • SpaceX AI's Grok 4.5: very low-cost, Opus-class coding model with minimal safety documentation.
  • Meta's Muse Spark 1.1: aggressive pricing, large coding/cyber benchmark gains, lengthy safety evaluation.
  • Meta previewed Muse Video; rolled out Muse Image, then backtracked after backlash over easy generation of images of public Instagram accounts.
  • Chinese open-source models grew to >30% of weekly OpenRouter tokens under cost pressure; insider-threat risk discussed.
  • Meta exploring selling AI compute as a cloud business (Bloomberg); US energy regulators pressing grid operators on large-load data-center connections; PJM ordered emergency steps to avoid large-scale outages.

Projects & Open Source

  • Nemotron-Labs-Diffusion: tri-mode LM unifying autoregressive, diffusion, and self-speculation decoding.
  • Tencent Hy3: open 295B MoE, 21B active params, 256K context.

Policy & Safety

  • Anthropic's "Verbalizable Representations Form a Global Workspace in Language Models" — interpretability method for verbalizable internal representations.
  • Beijing reportedly curbing overseas access to China's top models.
  • "AI 2040: Plan A" (by the AI 2027 author, ex-OpenAI) — proposes US–China coordination to slow progress until alignment improves.

Timestamps offset by dynamically inserted ads; contact: andreyvkurenkov@gmail.com / hello@gladstone.ai.

Full text · 3,224 chars
Note from Last Week in AI (Andrey): I’m back! And i’m sorry for putting the substack on a silent pause, work got a bit too overwhelming so I fell behind on this. I’ll do my best to resume normal posting, starting with catching up on podcast episodes missing from here (which are not last week… but I guess I should post them still, sorry for the spam!) I’ve also paused paid subscriptions until I can get back to consistent posting. Apologies for the flakiness, and thanks for being a subscriber! Our 252th episode with a summary and discussion of last week’s big AI news! Recorded on 07/11/2026 Hosted by Andrey Kurenkov and Jeremie Harris Feel free to email us your questions and feedback at andreyvkurenkov@gmail.com and/or hello@gladstone.ai In this episode: - OpenAI publicly rolled out GPT-5.6 (including Sol and Luna) and rebranded its desktop agentic coding product as ChatGPT Work, amid disputed claims about whether the US government effectively green-lit and delayed the release and concerns about inconsistent, ad hoc frontier-model oversight and jailbreakability. - New model releases intensified pricing and capability competition: SpaceX AI’s Grok 4.5 launched as a very low-cost, Opus-class coding model with minimal safety documentation, while Meta released Muse Spark 1.1 with aggressive pricing, large coding/cyber benchmark gains, and a lengthy safety evaluation. - Meta also previewed Muse Video and rolled out Muse Image before quickly backtracking after backlash over easy generation of images of public Instagram accounts; separately, Chinese open-source models grew to over 30% of weekly OpenRouter tokens as cost pressure increased, alongside discussion of risks like potential insider threats. - Infrastructure, policy, and safety developments included Meta exploring selling AI compute as a cloud business, US energy regulators pressing grid operators on large-load data-center connections, Anthropic publishing a “global workspace” interpretability method for verbalizable internal representations, reports that China may restrict overseas access to top models, and AI 2040 proposing US–China coordination to slow progress until alignment improves. Timestamps (note - these don’t take into account dynamically inserted ads and therefore may be off by a couple of minutes): - (00:00:10) Intro / Banter - (00:01:33) News Preview - Applications & Business - (00:35:21) Meta Is Planning a Cloud Business to Sell AI Computing Power - Bloomberg - (00:46:32) US energy regulator sets ultimatum for data centres + Grid operator PJM orders emergency steps to avoid large-scale US power outages - Projects & Open Source - (00:51:40) Nemotron-Labs-Diffusion: A Tri-Mode Language Model Unifying Autoregressive, Diffusion, and Self-Speculation Decoding - (00:57:44) Tencent Releases Hy3: An Open 295B Mixture-of-Experts (MoE) Model with 21B Active Parameters and 256K Context - MarkTechPost - Policy & Safety - (00:58:30) Verbalizable Representations Form a Global Workspace in Language Models - (01:09:29) Beijing is looking at curbing overseas access to China’s top AI models, sources say - (01:12:53) The ex-OpenAI employee behind ‘AI 2027’ recommends a rosier path - The Washington Post + AI 2040: Plan A
13:25

Better design than Fable

Moonshot AI's Kimi K3 topped the Arena frontend coding leaderboard, beating both Anthropic's Fable and OpenAI's GPT-5.6 Sol and coming close on other benchmarks. It's a huge 2.8-trillion-parameter model that's not token-efficient: about half the per-token price of Sol but using twice the tokens, so costs roughly balance out. Its weights are promised openly by July 27, and new subscriptions are paused because Moonshot doesn't have enough GPUs to serve everyone. The newsletter also notes Anthropic's Fable 5 found a counterexample to an 87-year-old math conjecture, Sierra's new outcome-based Horizon agents, and NotebookLM being renamed Gemini Notebook.

Notes
Ben's Bites, 2026-07-21 — "Better design than Fable"
Lead items
  • Kimi K3 (Moonshot AI, Beijing) beat Fable and GPT-5.6-Sol on Arena's frontend-coding leaderboard; close to them on other benchmarks. ~2.8T params, token-inefficient. Half the per-token price of GPT-5.6-Sol but uses ~2× tokens, so cost balances out. As Theo of Fable put it: > "Kimi K3 is an incredible model, but it's not an incredible value." 1M-token context, image support, weights to be released openly by 27 July. Available on Kimi apps/API, but new subscriptions paused — not enough GPUs to serve it. Caveat: current demand outstrips supply; open weights may change the picture.
  • Fable 5 joins Claude Max and Team plans permanently: up to 50% of weekly limits may be spent on Fable. Pro ($20/mo) gets access via usage credits plus a one-time $100 credit. Also: Fable found an example that disproves the 87-year-old Jacobian conjecture in mathematics. Caveat: single example claimed, not peer-reviewed here; Ben's "No big deal, right?" is tongue-in-cheek.
  • Horizon (Sierra): agents pursuing a business outcome over days/months across calls, texts, email, and company systems. Priced on outcomes, not tokens.
  • NotebookLM turns 3, renamed Gemini Notebook; new Collections feature (one notebook can live in multiple collections).
Quick links (substantive)
  • Cursor rebuilt SQLite from an 835-page manual using multiple agent-swarm configs. Cost spread across configs that passed all tests: Opus 4.8 + Composer 2.5 = $1,339 vs $10,565 for GPT-5.5 alone.
  • Ramp Router: per-request model routing; saves 30% internally. OpenAI-compatible endpoint.
  • Notion AI: browser use (early access). yoinks: terminal video downloader (X/YouTube/Instagram). OpenShip: open-source self-hosting platform. Lucy 2.5 (Decart AI): real-time video editing with more control. Linear Loops: recurring Linear Agent workflows. Rivet: local agent + design references, explores interface directions.
  • Essays: "How much Ridge spends on AI" (token costs for a non-tech co); "12-factor companies" (own context, rent intelligence, teach agents before hiring); "Harness engineering" (12 theses on coding-agent systems). Hilos: team chat mixing humans and coding agents.
Connecting the dots (thesis piece)

Central claim: agents should edit a shared "document" directly, not generate it once. Evidence:

  • Aaron built himself an SVG editor with Sol so he and the model both tweak the SVG via code or knobs.
  • Sunil (Cloudflare): "one document, two hands."
  • tldraw shipped a free file-based desktop app; Codex and Claude Code can use it offline — an example of a shared editable surface.
Self-driving companies
  • Replit CEO Amjad Masad on agents embedded across operations.
  • Gumroad now mostly run by an AI, "Gumclaw"; founder Sahil reported June 2026 as the first month AI-token spend equalled human payroll. Related reads: automating support and TASTE.md.
Editor's note (context, not news)

Ben (36, on family holiday) left Factory last December after his third child Poppy's health issues and burnout; plans to double down on Ben's Bites for non-technical readers, share opinions, "call out stuff I think is bullshit." Sponsor: Veridect — agent guardrails, free for safe moves, adversarial review in ~3s for risky ones.

Full text · 6,304 chars
Better design than Fable unsolved maths problems vs ai models, who'll win? Hey folks, It’s my birthday - the grand old age of 36. And I’m on holiday with my family, which, for those of you with young kids, is more full-on than normal life. There used to be a running joke inside my previous company that when I went away I’d come back with a new grand plan for what we were pivoting to. Not so much this time. Last December we had our third child, Poppy, and I was working at Factory. But she has some health issues, and I got pretty burnt out and overwhelmed, so I had to step back. For the last 7 months I've been plodding along work-wise without much agency. Which isn’t like me at all. I keep thinking of things I want to exist, but I can't bring myself to put in the work to bring them to fruition. Or more often than not, I'm not happy enough with how they turn out, so I kill them. Ben’s Bites is the thing I've done longer than anything else. It’s because each email is new. I like new things. I don't like running things. Or doing things because that’s what a business is or should do. And honestly, with two 3-year-olds and a 7-month-old and being a stay-at-home dad, it’s pretty tough to have the time and energy to shoot for the moon. I’m going to stop trying to put my square peg in a round hole. I’m going to put more of myself into this thing. This company is for people who are not traditionally technical but want to use the latest and greatest for work, life and building stuff. That’s who I am. I’m going to share my opinions on what’s happening in AI more, call out stuff I think is bullshit and try to explain things I’m learning myself. I’ll be an open book, like I'm being now. And I’ll test new things too. Please be open with your feedback, but be nice today; it’s my birthday after all. Ben’s Bites is brought to you by Veridect Guardrails filter what your AI says. Veridect governs what your agent does — before it acts. Safe moves clear instantly and free; only the riskiest get the full adversarial cross-examination — greenlight, escalate, or block in ~3s, logged tamper-proof. Try to beat it in 60s. Headlines - Kimi K3, a new model from the Beijing-based Moonshot AI, surpassed Fable and GPT-5.6-Sol on Arena’s frontend coding leaderboard (and generally came close to these models on other benchmarks as well). It’s a big model (2.8T parameters) and quite token-inefficient. It is half as expensive as GPT-5.6-Sol on a per-token basis but uses twice as many tokens as Sol, balancing out the costs. Kimi K3 is an incredible model, but it’s not an incredible value. - Theo It works with images, has a 1M-token context window, and its weights will be released openly by 27th July. For now, the model is available on Kimi’s apps and API, but they have paused new subscriptions because they don’t have enough GPUs to serve the model to everyone. - Fable 5 is now a permanent part of Max and Team subscription plans for Claude with the same condition: you can use 50% of your weekly limits towards Fable. Pro ($20/mo) accounts can use it via usage credits and will get a one-time $100 credit. Also, Fable found an example that disproves the 87-year-old Jacobian conjecture in mathematics. No big deal, right? - Horizon by Sierra - agents that work towards a business outcome over days or months, moving between calls, texts, email and company systems. Charged for outcomes, not tokens. - NotebookLM turned three and got a new name: Gemini Notebook. It also got Collections, a new feature to organise notebooks. One notebook can be part of multiple collections. Quick links - Cursor rebuilt SQLite from an 835-page manual using multiple agent swarm setups. Across many configurations that passed all tests, costs varied a lot → Opus 4.8 + Composer 2.5 cost $1,339 vs $10,565 for GPT-5.5 alone. - Notion AI can now use a browser for you. (in early access) - Ramp Router - Send each request to a relevant model. It saves 30% in costs for Ramp internally. Comes with an OpenAI-compatible endpoint. - yoinks - terminal tool for downloading videos from X, YouTube, Instagram and more. - OpenShip - open-source platform for deploying and running apps on your own infrastructure. - Lucy 2.5 by Decart AI - live AI model for editing video in real time with more control. - Linear Loops - recurring workflows that Linear Agent runs for your team. - Rivet - connect a local agent and design references, then explore several interface directions. - How much Ridge spends on AI - a non-tech company breaks down whether token costs are actually a problem. - 12-factor companies - own your context, rent intelligence and teach agents the work before hiring for it. - Building diagrams to explain databases using Excalidraw charts and Cursor. - Harness engineering - 12 theses on designing the systems around coding agents. - Hilos - team chat where people and coding agents work in the same conversation. Connecting the dots… - Aaron built himself an SVG editor with Sol instead of letting the model create the SVG in a single go and hoping it turns out exactly as he wants it to be. Now, both he and the model can tweak the SVG (using code or the knobs). - Sunil from Cloudflare makes a similar case in his recent post: one document, two hands - you and the agent both should be able to work on the “document” directly inside an app. - tldraw released a free, file-based desktop app for its “draw on a canvas” product. Any agent like Codex and Claude Code can use tldraw offline. So tldraw offline becomes one such surface where both you and an agent can work on a single document. - The Self-Driving Company - Replit’s CEO Amjad Masad wrote this post sharing how agents are effectively in each part of operating Replit. - Gumroad is another good example of a self-driving company. It is now mostly run by an AI called Gumclaw. Sahil (Gumroad’s founder) shared that June 2026 was the first month that Gumroad spent equal amounts on AI tokens and human payroll. - Good reads from Gumroad/Gumclaw - automating support and TASTE.md Afters - Find me on X, Linkedin, or YouTube - Read about me and Ben’s Bites - 📷 thumbnail via @keshavatearth * sponsors who make this newsletter possible :) Wanna partner with us for the next quarter? Email us at shanice@bensbites.com or k@bensbites.com

Newsletter

5
09:32

The Open Source AI China Problem Just got Worse

China's open-weight models are winning enterprise adoption and widening America's AI leadership gap, in this week's big story. Moonshot's Kimi K3 is pushing enterprises to open-weight models, DeepSeek raised around $7.4 billion in June, and Databricks is now valued at $188 billion off the same pivot. Google's Gemini 3.5 Pro got delayed at a bad moment and the Trump administration is signaling it could ban cutting-edge Chinese models. The piece also notes Alibaba's Qwen 3.8 Max preview with aggressive token pricing, and ranks the current top models led by Claude Fable 5.

Notes
Notes: "The Open Source AI China Problem Just got Worse"

Source: AI Supremacy (Substack), Mike. Published 2026-07-21. Commentary/analysis, not reported news — claims are the author's framing, mostly unverified.

Core claim: Moonshot AI's release of Kimi K3 (open-weight) is driving enterprise AI buyers to switch from proprietary U.S. frontier models to cheaper open-weight ones. The author calls it "one of the biggest AI moments of 2026," set against an unresolved Iran war in the Strait of Hormuz and a freefalling NASDAQ-100.

Market / financial specifics

  • DeepSeek raised ~$7.4B in June 2026; planned 2027 IPOs for DeepSeek, Moonshot, OpenAI, and others.
  • Databricks announced a funding round valuing it at $188B, described as a beneficiary of the enterprise pivot to routing/cheaper tokens.
  • SK Hynix's U.S. listing preceded a semiconductor bear-market correction; the author calls the Korean KOSPI "a leading indicator."
  • Chinese DRAM maker CXMT about to IPO in Shanghai.
  • U.S. hyperscalers approaching negative free cash flow on AI capex.

Model events

  • Google Gemini 3.5 Pro "crucially delayed"; SpaceXAI Grok 4.5 launch "completely overwhelmed"; OpenAI GPT 5.6 Sol "nearly entirely overshadowed"; an "Anthropic Fable 5 confusion" echoes the January 2025 DeepSeek moment.
  • Alibaba Qwen 3.8 Max (preview) with an open-weight component and "aggressive international token pricing plan."
  • Kimi K3 targets enterprise (not hobbyists) to ramp ARR ahead of a Moonshot IPO in "about six months."

Policy / geopolitics

  • Trump administration reportedly restricting "Mythos class models" and possibly banning Chinese frontier models. Author: "So much for global free market capitalism."
  • Xi Jinping keynote at WAIC 2026 (Shanghai) titled > "Joining Hands to Build a Just and Equitable System for Global AI Governance." Author asserts China is "more advanced in AI governance and regulatory leadership than the U.S."

US open-weight gap — Best U.S. open-weight per the author: Thinking Machine's Inkling ("about a week old"), upcoming Reflection, and Nvidia Nemotron (vs. earlier Meta leadership). He calls the lack of U.S. open-source leadership "a potential disaster."

Ranking (Artificial Analysis Intelligence Index, "smartest model"):

  • Anthropic – Claude Fable 5
  • OpenAI – GPT 5.6 Sol
  • Moonshot AI – Kimi K3 (open weights)
  • SpaceXAI – Grok 4.5
  • Zhipu (Z.ai) – GLM 5.2 (open weights)

Caveats stated by the author

  • OpenRouter data "isn't representative of the entire situation" (only a sampling point).
  • Rankings/benchmarks are "fairly artificial and likely to change next week and certainly by next month."
  • "There are few great application layer products" despite the model competition.
Full text · 6,275 chars
The Open Source AI China Problem Just got Worse Model supremacy in a token-efficient macro environment plagued by HBM, energy and datacenter compute bottlenecks. The 2026 story of AI is getting geopolitical. 👋 Hey there, I’m Mike. Each week I share AI articles at the intersection of tech, business, society and the future. If you want to support the channel or gain full-access to my work, go here. Read Archives | See Substack Notes | Visit our community Chat | Visit Homepage. I’ve been tracking the U.S. vs. China dynamics of AI and its future for nearly six years. How do you summarize the heat that is July, 2026 in the AI industry? It’s been a very bizarre and multi-layered drama. We are witnessing history. Geopolitics and AI on the Front Burner 🔥 As you likely know, Chinese company Moonshot AI released a new version of its Kimi model called Kimi K3 that is causing a lot of Enterprise AI to switch to open-weight models. It might be one of the biggest AI moments of 2026. With the Iran war unresolved in the strait of Hormuz and the Tech heavy NASDAQ 100 in freefall, it’s becoming a geopolitical and Trump Administration catch-22. There are no clear solutions in war and AI, as it turns out. Generative AI models are evolving, but likely slowing the revenue growth of AI behemoths OpenAI and Anthropic. With American hyperscalers approaching negative free cash flow via incredible AI capex and datacenter investments, you have to wonder whether it’s all worth it - if Chinese models that are getting larger and more efficient and can replicate the performance while under cutting the cost. It has the potential to become an AI crisis in the stock market even as the Semiconductor boom seems to have hit a bear market correction, after the U.S. listing of South Korean HBM leader, SK Hynix. South Korea (the KOSPI) is now a leading indicator. With Chinese DRAM maker CXMT about to go public in Shanghai, it’s a very charged China vs. U.S. setup in the future of AI. With DeepSeek’s stunning funding rounds and planned IPOs for DeepSeek, Moonshot, OpenAI and others in 2027, it’s shaping up to be quite a year next year too. DeepSeek raised around $7.4 billion in June last month. We have to assume Databricks has also been one of the beneficiaries of this pivot of Enterprise AI to routing and cheaper tokens of open-weight players in a world where Databricks announced a new round of funding that values the company at $188 billion. Google’s Gemini 3.5 Pro has been crucially delayed at the worst possible moment. The launch of SpaceXAI’s new model Grok 4.5 was completely overwhelmed by the Kimi K3 moment. Anthropic Fable 5 confusion has given China the ultimate return of that DeepSeek moment vibe back from the dead of January, 2025. OpenAI’s own GPT 5.6 Sol release has also been nearly entirely overshadowed. The Trump Administration quick to restrict Mythos class models has a serious problem, what if Chinese models are able to approach those same capabilities with open-weight models that were thought to be many more months behind? The Trump administration is showing signs it could ban Chinese cutting-edge models and take drastic steps to try to curtail China’s rise in AI. I thought they were not going to regulate the AI industry. So much for global free market capitalism. The Brave New World of Token Efficiency Looks Chinese With Microsoft, Amazon, Meta and even Google behind the Big 7 hyperscalers have lost some of their AI talking points and credibility in this cycle. The mishandling of Mythos class models by the Government and the rise of Kimi K3 type models is the perfect storm that is seeing many Enterprise companies pivot to more rational token usage that’s radically more cost efficient. While Open-Router isn’t representative of the entire situation, it’s an interesting data sampling point of the overall macro trend: it appears like cheaper open-weight models from China are winning over marketshare. In my opinion, some of the best Open-weight models from the U.S. are Thinking Machine’s Inkling (about a week old), and whatever Reflection is likely to put out soon. Nvidia’s Nemotron is an often cited alternative to the previous leadership by Meta. The problem is the real lack of leadership in open-source AI in the United States. For America this is a potential disaster in its AI leadership on the frontier of models, tokens and Enterprise adoption. To make matters even more peculiar, Chinese President Xi Jinping’s most significant recent remarks on artificial intelligence were delivered via a keynote address at the World Artificial Intelligence Conference (WAIC). President Xi Jinping attended in Shanghai the opening ceremony of WAIC 2026 and delivered a keynote speech titled “Joining Hands to Build a Just and Equitable System for Global AI Governance.” China appears to be more advanced in AI governance and regulatory leadership than the U.S. trying to actively build global collaboration around the theme. Alibaba’s own Qwen 3.8 Max (preview) will also have an open-weight component with an aggressive international token pricing plan. Kimi K3 is not aimed at hobbyist developers but at Enterprise customers in order to ramp ARR before they also IPO in about six months time. The U.S. and China competition in models, even at a time when there are few great application layer products is exaggerating the demand for compute at a time when China has both cheaper token generation and more abundance energy. At a time when the Trump Administration’s key mandate appears to be keeping the AI boom on the stock market rolling for the financial elite and business class. Meanwhile everyone from Anthropic to Moonshot AI are positioning themselves to maximize their IPO hype and revenue sales momentum. Do Ranking Models even Matter Any Longer? If you go by Artificial Analysis Intelligence Index, and rank by smartest model, the ranking is: The state of affairs on peak model performance is roughly as follows: - Anthropic – Claude Fable 5 - OpenAI – GPT 5.6 Sol - Moonshot AI – Kimi K3 (open weights*) - SpaceXAI – Grok 4.5 - Zhipu (Z.ai) – GLM 5.2 (open weights) These lists and benchmarks that they models are trained to perform on are fairly artificial and likely to change next week and certainly by next month.
10:50

How to generate AI drafts that sound more human (a 4-step process)

A four-step process for getting AI to draft writing that sounds like a real person, built from data about your own style. A coding agent first finds the words and phrases your specific model overuses, then analyzes 10 to 20 of your published pieces to build a voice profile with 8 to 12 checkable habits. Those two lists get packaged into a reusable Skill with a write mode and a check mode, and the model drafts, checks, and regenerates until a draft passes, capped at five tries. It gets you 75-85% of the way to publishable, but you still need a final human edit, and it does nothing about hallucinations.

Notes

Disclosure: author used Claude Fable to draft this article from a voice profile, ran automated checks, then edited manually.

Claims about model tells: author's own analysis of 300,000 words across AI models found ChatGPT, Claude, and Gemini each overuse different words depending on writing type. Claims AI's favorite words are appearing in online writing both via unedited AI drafts and because authors' own styles have been influenced by AI. Author (self-described) got annoyed by writing that "echoes Claude."

The "four horsemen" of AI writing problems:

  • Sounds like AI (stock words/constructions)
  • Doesn't sound like you (reads like no one in particular even after cleanup)
  • Doesn't make your actual argument (model approximates your point)
  • Makes things up (facts, quotes, links must be checked)

This article covers only #1 and #2; the rest deferred to a follow-up. Cutting tics alone leaves copy "hollow" — must also add your voice back. Claimed payoff: "75-85% of the way to something publishable," but a final human editing pass is always required.

Step 1: Find model-specific tics

Have a coding agent (Claude Code, Codex) generate ~10 500-word drafts on your usual topics with no style instructions, then run a script reporting: every word used >2x (excluding prepositions, articles, pronouns, connective tissue); every 2–4 word phrase >2x; counts of em dashes, semicolons, and question-mark sentences. Rank by frequency, show one example sentence per item. The author recommends running your own analysis rather than using generic "AI tells" lists, since tics vary by model and niche.

Step 2: Build a voice guide from your own writing

Run the same fingerprint method on your last 10–20 published pieces. Output: a draft voice profile of 8–12 concrete markers, each a specific checkable habit with two example sentences from your work ("If you can only support eight markers with evidence, give me eight; don't pad the list"). Edit the profile yourself — cut coincidences, keep recognized habits. Then convert the profile into a voice guide that adds per-marker frequency limits from your actual rates ("about once per section" / "in roughly a third of paragraphs"), with this rule at top:

"apply these markers at their natural rate and never all at once. A draft that uses every marker in every paragraph fails, the same as a draft that uses none of them."
Step 3: Package into a reusable Skill

In Claude, a folder with a SKILL.md; in other tools, saved system prompt/custom instructions. Two named modes:

  • Write mode: voice guide applied, tics list as hard bans. Key instruction: "don't write the tic and then patch it" — a patched sentence "usually keeps the AI rhythm."
  • Check mode: report every tics hit with its surrounding sentence and flag missing/over-used voice markers, without rewriting (a model that fixes while checking introduces new problems). If no mode is named, the Skill must ask which one.
Step 4: Iterate until a draft passes

Draft (write mode) → check → regenerate on failure, cap at five attempts. Show draft + final check report + attempt count; if nothing passes in five, show closest draft and failing checks. Recommendation: use the most capable model, not the cheapest — author drafts with Fable (Anthropic's most capable model; "expensive or it will be when it's no longer part of my plan") because it reaches a clean draft in the fewest tries. Cites Anthropic docs: "start with the most capable model" and trade down only if a cheaper one suffices.

Token-burn warning (five loop-guards)
  • Cap attempts (five max; "a draft that can't pass in five tries won't pass in fifteen").
  • Stop on repeat failure — same check failing twice = broken check, usually a conflict between lists.
  • Keep judgment checks out of the loop — "only mechanical checks converge"; save "does this sound like me?" for a human read.
  • Revise, don't regenerate, after the second failure — fix only failing sentences.
  • Put the hard stop outside the model — drive the loop from a script, or cap with Claude Code's --max-turns flag.
Context and caveats

Author writes from an ESL-friendly, neurodiverse-friendly stance: ESL writers use AI to render native-language writing into global English; neurodiverse writers use AI to organize thoughts. Author admits using AI to keep the Substack going when schedule was unworkable. Limits acknowledged: does nothing about hallucinations; the process still demands a manual editing pass; results vary by style and niche (author's tech space is "generally pretty forgiving" of AI content). Plans a follow-up comparing different human writers' stylistic fingerprints.

Full text · 11,272 chars
How to generate AI drafts that sound more human (a 4-step process) My thoughts (and some new advice) one year after I wrote a guide to editing AI-generated copy Disclosure: I used Claude Fable to write the first draft of this article based on a voice profile, ran a series of automated checks, and then edited the resulting draft manually. TL;DR: Newer AI models make it easier to generate writing that sounds like you. You can now use AI to analyze your writing, create a voice guide based on the results, and build a “voice checking” Skill to score AI writing against your voice standards, so AI-generated drafts arrive in better shape. But none of this eliminates the need to carefully edit your AI draft before you publish. (And it does nothing about hallucinations.) Last year, I wrote an article about editing AI copy. At the time, a lot of my freelance work involved editing AI-generated drafts for businesses that were struggling to build AI-first content practices. And most of the advice I shared then still applies today. But I’ve learned a lot over the past year. The models have gotten better, and I’ve seen signs that AI’s favorite words are appearing in online writing both because people are publishing more AI-generated drafts and because authors’ own writing has been influenced by AI. I’ve also heard from ESL writers who use AI to translate writing in their native language into global English, and from neurodiverse writers who use AI to organize their thoughts. Overall, I’d say I’ve become a lot less judgmental of AI writing, and more sympathetic to the writers who use it. I’ve even used it myself to keep my Substack going over the past few months, when a mix of work and family stuff made it virtually impossible to maintain my publishing schedule and still write everything by hand. The problem, of course, is that AI writing done badly is just slop. And I personally get annoyed by writing that echoes Claude, largely because I’m so familiar with how that model tends to write. The four horsemen of AI-generated writing The four main issues I see with AI-generated writing are: - It sounds like AI. Stock words and constructions give it away. - It doesn’t sound like you. Even after you clean it up, it reads like no one in particular. - It doesn’t make your actual argument. The model approximates your point instead of making it. - It makes things up. Facts, quotes, and links all have to be checked. This article focuses on the first two issues, which are closely related. (I’ll write about the other two next month.) The standard fix for the first problem is to edit out the AI tics, the stock words and phrases that mark machine writing. But, while cutting the tics can make your draft cleaner, it can also leave it feeling hollow. To close the gap, you also have to add your own voice back in. The system below does both, in four steps, and it ends with a setup that can draft and check itself. You can get 75-85% of the way to something publishable, although you’ll have to do a final editing pass yourself. Step 1: Find the AI writing tics your specific model makes The tics you need to cut depend on which model you use and what you write about. For example, my recent analysis of 300,000 words written by different AI models found that ChatGPT, Claude, and Gemini each have words they tend to overuse in different types of writing. So, rather than relying on my study or a list of generic AI tells, it’s better to run your own analysis. To find the model tics that will make the most different in your drafts, have a coding agent like Claude Code or Codex: - Generate about ten drafts on topics you normally cover, with no style instructions. - Analyze which words and phrases appear more than twice - Produce a list of AI tics ordered by the frequency at which they appear. This list should exclude prepositions, articles, and other generic connective tissue. Give this to your coding agent Write ten 500-word drafts, one for each of these topics: [list ten topics you normally cover]. No style instructions, and don’t try to sound like anyone in particular. Save each draft as its own text file. Then write and run a script that analyzes all ten files together and reports three lists: every word that appears more than twice, excluding prepositions, articles, pronouns, and other generic connective tissue; every phrase of two to four words that appears more than twice; and counts of em dashes, semicolons, and sentences ending in question marks. Rank each list by frequency and show one example sentence per item. Step 2: Build a voice guide based on your own past writing Once the AI tics are gone from your draft, you’ll have clean copy that may sound a bit bland or hollow. So, in addition to knowing how your model tends to write for your niche, you also need to understand the unique characteristics of your own writing. To quantify your own word and style choices, use a coding agent to run Step 1 on your last 10 to 20 published pieces. Gather them into one folder and use this prompt: Give this to your coding agent I’ve put 10 to 20 of my published pieces in this folder: [folder path]. Read all of them, then do three things. First, catalog my recurring, distinctive habits: the words I repeat, how I open and close, whether my sentences run long or short, what my punctuation does, and the kinds of examples I reach for. Second, write and run a script that counts word frequency across the whole set, excluding prepositions, articles, pronouns, and other generic connective tissue. That’s the fingerprint method from Step 1, pointed at me. Third, combine the two passes into a draft voice profile of 8 to 12 concrete markers, each one a specific, checkable habit with two example sentences pulled from my work. If you can only support eight markers with evidence, give me eight; don’t pad the list. Remember to review edit the AI-generated voice profile before you use it. Cut anything that looks like a random coincidence and keep the habits that you recognize. Then hand the finished profile back to your AI and have it turn the profile into a voice guide with drafting instructions for applying your signature words and constructions at roughly their natural rate. Use this prompt to build your guide: Give this to your coding agent Here is my edited voice profile: [paste the profile]. Turn it into a voice guide I can include in drafting instructions. For each marker, add a frequency limit based on its actual rate in my published pieces, in the form “about once per section” or “in roughly a third of paragraphs.” Put this rule at the top of the guide: apply these markers at their natural rate and never all at once. A draft that uses every marker in every paragraph fails, the same as a draft that uses none of them. Step 3: Package the AI tics list and the voice guide into a reusable Skill A Skill (some tools call it a saved custom instruction) is a set of directions the model applies to every draft automatically, so you stop re-explaining your rules at the start of each session. In Claude, that’s a folder with a SKILL.md file inside; in other tools it’s a saved system prompt or a set of custom instructions. Give the Skill two modes, and name them at the top of the file: - Write mode: Draft new copy with the voice guide applied and the AI tics list treated as hard bans. The instruction I use is “don’t write the tic and then patch it,” because a patched sentence usually keeps the AI rhythm even after the flagged word is gone. - Check mode: Take an existing draft, run it against both lists, and report every hit with the surrounding sentence. No rewriting in this mode. You want a report you can act on, because a model that fixes your draft while checking it will also introduce new problems. The Skill file itself is also one more thing your coding agent can write: Give this to your coding agent Here are my AI tics list and my voice guide: [paste both]. Turn them into a Skill file I can save and reuse. Paste both into the file in full; don’t summarize or shorten them. At the top, define two modes. In write mode, draft new copy with the voice guide applied and every item on the tics list treated as a hard ban; don’t write a tic and then patch it, rebuild the sentence. In check mode, run a supplied draft against both: report every tics-list hit with its surrounding sentence, and flag any voice marker that is missing or used past its frequency limit, without rewriting anything. Add one last instruction: if I haven’t named a mode, ask which one I want. Step 4: Let the model iterate until a draft passes Once you’ve set up the Skill, you can hand the whole cycle to the model. It drafts in write mode, checks the result against the Step 1 and 2 lists, and regenerates whenever a check fails, until a draft comes back clean. In prompt form: Give this to your coding agent Using my Skill, draft [topic, length, audience, and the points to cover] in write mode. Then switch to check mode and run the draft against every check. If anything fails, regenerate in write mode and check again, up to five attempts. When a draft passes, show me the draft, the final check report, and how many attempts it took. If nothing passes in five, show me the closest draft with the failing checks listed. Ideally, the model you iterate with should be the most capable one you can get, not the cheapest. Lately, I’ve been drafting with Fable, which is Anthropic’s most capable model. It’s expensive (or it will be when it’s no longer part of my plan), but in my experience the best model reaches a clean, publishable draft in the fewest tries. Anthropic’s own guidance is consistent with this approach. For work where quality is the priority, their docs say to “start with the most capable model” and only trade down later if a cheaper one turns out to be good enough. Writing that has to sound like you is that sort of work. Of course, your results may be different depending on your preferred style and the type of writing you do. Warning: Don’t let your model iterate forever (and spend all your tokens) A loop with no stop condition will regenerate indefinitely, and every attempt costs tokens. Five ways to prevent this includes: - Cap the attempts. The prompt above stops at five; a draft that can’t pass in five tries won’t pass in fifteen. - Stop on a repeat failure. The same check failing twice in a row means a broken check, usually a conflict between your lists. - Keep judgment checks out of the loop. Only mechanical checks converge; save “does this sound like me?” for your human read. - Revise instead of regenerating. After the second failed attempt, have the model fix only the failing sentences. - Put the hard stop outside the model. Drive the loop from a short script that counts the rounds itself, or cap the run at the tool level with Claude Code’s --max-turns flag. What works for you? This article came out of my own experience writing with AI in a tech space that’s generally pretty forgiving of AI-generated content, so your mileage may vary. I’d love to hear how you approach writing with AI. And, if you try the ideas outlined in this article, please tell me how it goes. (I’m thinking about a follow-on article looking at different human writers’ stylistic fingerprints.)
12:53

How I Built A Substack API With Hermes And Codex

A developer shipped an unofficial Substack API after a missing scheduling feature forced him to split the work between two AI tools, and the lesson was about handoffs, not models. Hermes handled the messy research phase, digging up a forgotten open-source package, while Codex did the focused edits once the work fit inside one folder. A wrong field name (scheduled_at instead of trigger_at) nearly shipped until automated tests caught it, which pushed him to make tests the referee that decides between plausible AI answers. The result, version 0.2.2 of the Unofficial Substack SDK, passes 25 tests covering 52 behaviors and pulled 500+ downloads in its first week.

Notes

How I Built A Substack API With Hermes And Codex

Source: All Agents Considered (Substack), 2026-07-21

The project

Author wanted Hermes to schedule Substack Notes. On June 12 they asked Hermes to find a forgotten Substack tool in the Projects folder. It was identified as substack-api, an MIT-licensed open-source project by Jakub Slys (credit retained). Reading profiles, posts, comments, and Notes already worked; publishing worked; scheduling did not — the label Substack expected for publish time was undocumented. Hermes proposed scheduled_at and warned it was unproven.

The wrong guess became the article's core lesson:

"Version 0.2.2 uses trigger_at. An automated check reads the exact instruction sent to Substack and fails if the label changes. Proof now lives outside the AI conversation."

Outcome: the Unofficial Substack SDK shipped; five versions in three days, 500+ downloads in week one, npm recorded 708 downloads July 13–19 (author stresses these are downloads, not people).

Agent split: scout, specialist, referee
  • Hermes (scout) handled the messy discovery phase: pulled context from old conversations, forgotten files, licensing, and competing projects; compared alternatives rather than taking the first hit. Chose a small building block over ready-made services so users' logins wouldn't pass through a third-party server.
  • Codex (specialist) — used via the VS Code extension — took over once the work fit inside one project folder, executing precise single-file edits without loading full history. Author's rule: "Improve the Substack tool" leaves too much room. "Run both versions and tell me which one fits an app used by AI agents" gives the work an edge.
  • Referee — independent checks: real account responses, changed files, passing tests. "Whoever researched the problem also wrote the code and explained why it looked correct. Moving the final decision into a check broke the loop."
Key incidents and numbers
  • Cloudflare dead end: an early experiment turned the inherited project into a read-only Cloudflare service; all 230 automated checks passed, but Substack blocked Cloudflare requests. SDK moved to the author's VPS with a small server interface.
  • Wrong folder bug: Codex placed the project in a private server folder the author's Files app never showed; every command succeeded yet the result was unusable. Fix: state personal conventions explicitly.
  • Banked resets: Hermes burned OpenAI-banked Codex resets much faster than Codex-in-VS-Code because it carried a wider working history into each session.
  • v0.2.2 quality gate: passes 25 tests covering 52 expected behaviors; GitHub reruns them before every release, then publishes with one-time approval instead of a stored publishing password.
  • Live proof: after adding Substack credentials, the author requested latest notification → got a real video alert; latest mentions → five real replies. This proved the full round-trip, since "documentation examples are easy to fake."
What shipped (v0.2.2)

Reads profiles, posts, comments, engagement numbers, subscriber data; publishes Notes, schedules them, edits drafts, adds images.

Security — verbatim
"I deliberately left the Substack cookie out of the npm package because it gives access to your account. Ask your agent to read the setup instructions... then add the value of your substack.sid cookie yourself as a trusted server-side environment variable and pass it as sessionToken. Never paste the cookie into your codebase or anywhere likely to save and share it."
Caveats and limitations
  • Substack may change its private API without warning; tests make breakage easier to find but "don't turn an unofficial tool into a promise from Substack."
  • Loading one Substack login at startup meant every user would appear as the author — a multi-user app needs per-account storage and login selection. The author is building one, tentatively called StackedHQ (not yet ready).
  • 708 downloads ≠ 708 people, stated explicitly.
The handoff file (copyable template)

Seven sections: Goal (one result) / Starting point (where work lives, what works) / Known evidence (files, observations, prior attempts) / Limits (what AI must avoid) / Checks (how to prove it works) / Stop and ask (what needs approval). Hermes fills the first version; author removes weak assumptions and makes checks specific; Codex works until a check passes or needs approval. Failed check → send back the single failing behavior, don't reopen the whole project. Author claims it generalizes: research→sourced brief, writing→draft, checklist→publish.

Related pieces referenced

"When to Use MCPs CLIs or Your Own Tool," "The Agentic Engineering Shift," an article on file-based AI workflows, and an upcoming Hermes 101. Closing thesis: "Use the agent with the widest view to find the path, use the focused agent to make the change, and trust neither until the work passes a check."

Full text · 15,224 chars
How I Built A Substack API With Hermes And Codex I wanted Hermes to schedule my Substack Notes. A missing scheduling feature turned into an open-source tool built across Hermes, Codex, and a set of independent checks. By the middle of June, I had dozens of Substack Notes sitting in markdown files with dates and publishing times. I had rewritten the awkward ones, removed repeated ideas, and mixed the topics so two similar Notes wouldn’t appear one after another. But there was that one task I kept postponing in Asana: Finish Notes Scheduled Setup. Scheduling a single Note inside Substack takes about 30 seconds. Keeping several days’ worth of posts organized turns into admin work fast, especially when one edit changes the order and every publishing time after it. On June 12, I asked Hermes to find an old Substack tool installed somewhere in my Projects folder. Other software used it to talk to my Substack account, and I wanted to know whether it supported scheduled Notes. I expected a quick yes or no, followed by a small fix if I got lucky. Hermes found the forgotten package beside old test files and a database from my Notes analysis project. Publishing worked. Scheduling didn’t. One missing feature pulled me much further than expected. A few weeks later, I had published the Unofficial Substack SDK, a reusable Substack toolkit for other apps. Five versions shipped in three days, followed by 500+ downloads during its first week. I use two agents to run my workflows, and this project showed me why the split matters. One agent handling every stage carries its early assumptions all the way to the finish line. Hermes handled the messy beginning, pulling context from old conversations, forgotten files, licences, and competing projects. Codex stepped in once the work fit inside one project folder and I knew what a successful result looked like. In today’s edition, I’ll walk you through how I built the SDK and show why the handoff between Hermes and Codex mattered more than either agent working alone. In this piece - How Hermes turned a half-remembered package into a real starting point - Why Codex became more useful once the job got smaller - How one wrong field name exposed the danger of plausible AI answers - A handoff format for splitting your next build without losing context The Package I Forgot I remembered using the old tool to collect Substack data, though I had forgotten its name, where it lived, and how much of it still worked. Hermes identified it as substack-api and traced it back to an open-source project. Reading profiles, posts, comments, and Notes already worked. Publishing worked too. Scheduling notes support was missing. Hermes searched the project, looked through public clues left in Substack’s website, and checked similar tools. Instructions for publishing a Note were already known, while scheduling remained undocumented. Nobody had confirmed the label Substack expected for the publishing time. Hermes proposed scheduled_at. Plenty of online services use a label like this, and it looked completely at home beside the existing instructions. Hermes still warned me about the missing proof and listed other names Substack might expect. I nearly ignored the warning because the feature looked finished. Doing so would’ve left me with a neat scheduling button sending a label Substack never promised to read. Version 0.2.2 uses trigger_at. An automated check reads the exact instruction sent to Substack and fails if the label changes. Proof now lives outside the AI conversation. Getting the first field wrong became the most useful part of the build because it forced me to separate finding an answer from earning trust in it. Hermes Found The Ground Starting from scratch would’ve been wasteful. Jakub Slys had already spent years learning how Substack’s private machinery worked, and his project carried working instructions plus a large set of checks. Hermes traced the installed tool back to his GitHub project, checked the MIT licence, and kept the original credit. His research remains credited in the finished project. From there, Hermes compared the other available projects instead of treating the first one as the automatic winner. Some were ready-made services for apps and AI agents. Others only read public posts. I wanted a small building block speaking directly to Substack without sending a user’s login through somebody else’s server. An early experiment turned the inherited project into a small online service running on Cloudflare. Hermes removed parts tied to a traditional server and locked it to read-only access. All 230 automated checks passed before the project moved into my account. Substack blocked requests coming from Cloudflare, so the first version hit a dead end. I moved the SDK to my VPS and built a small server interface around it, giving my other apps a stable way to use it. Hermes handled this part well because the problem stretched beyond one project folder. Old conversations, forgotten files, licensing, and the safety of each user’s digital login key all affected the decision. A coding agent focused on one project sees the files in front of it. Hermes also saw why I had them, what I had tried before, and which surrounding work mattered. .I covered a similar choice in my guide to choosing AI tools. Tool choice starts before the interface. First decide whether the messy part is finding the work or doing it. Codex Needed a More Targeted Job Once Hermes had found the right starting point and cleared up the wider questions, I moved the project into VS Code. I use Codex through the VS Code extension, and this is where it works best for me. I point it at a specific file, ask for one change, review the result, and keep moving without reopening the whole project discussion. I asked Codex to run the version Hermes had prepared beside my other services without rewriting anything. It downloaded the project, installed what it needed, passed the checks, and started the service. Then I opened my Files app and couldn’t see it. Codex had placed the project inside a private folder on the server, while my Files app only showed a different Projects folder. Every command had succeeded, yet the result still lived somewhere I didn’t use. I corrected the path, and Codex moved the project before restarting the service. One small mistake exposed a useful limit: a coding agent understands the project in front of it, while personal conventions still need to be stated clearly. Once everything ran from the right place, I asked Codex to launch Jakub’s newer project beside mine. His version already worked as a complete service for regular apps and AI agents. Mine did less, though its smaller size made it easier to reuse inside other products. Running both made the decision easier. Jakub’s project was the stronger ready-made service, while my smaller toolkit made more sense as a building block for the server I wanted to control. OpenAI had also given me banked resets, small refills for my Codex allowance whenever I reached the limit. Hermes burned through those resets much faster than Codex inside VS Code because it carried a wider working history into every session, including old conversations and surrounding files. Codex stayed focused on the active project and the exact edit in front of it. For small changes, this made the VS Code extension faster and cheaper to run. No model comparison or architecture discussion would’ve given me the same confidence. Hermes found and framed the right project, then Codex handled the precise edits without dragging the entire history behind it. The Server Had To Prove It Next came a small server for making the toolkit available to other apps through a web address. Codex downloaded it, ran its checks, started it locally, and confirmed it was healthy before I added my Substack login details. After setup, I asked for my latest Substack notification. A live video alert came back from my account. Five recent replies followed when I asked for my latest mentions. Documentation examples are easy to fake, so I had to know this would be working as expected. A response from my account proved the toolkit had completed the full trip to Substack and returned with real data. Live use also exposed a design problem. Loading one Substack login at startup meant every person using the service would’ve appeared as me. A multi-user app needed a private and secure way of storing each account's data and selecting the right login for every request. I am building an app for that and will soon have it ready. I’m thinking of calling it StackedHQ. Let The Tests Argue scheduled_at and trigger_at both sound reasonable as labels for a publishing time. An AI has seen enough software to defend either one with enough detail to waste your afternoon and usage quota. Tests settle the argument by opening the instruction before it leaves the toolkit. One scheduling check expects trigger_at. Another caught missing image information and led to a tiny correction. Version 0.2.2 passes 25 tests covering 52 expected behaviors. GitHub reruns them before every release, then publishes with one-time approval instead of storing a permanent publishing password. I don’t care which agent sounds more certain once a request has an observable answer. Real responses, changed files, and passing checks get the final vote. This is the part I missed inside one AI thread. Whoever researched the problem also wrote the code and explained why it looked correct. Moving the final decision into a check broke the loop. This habit follows the same direction as The Agentic Engineering Shift. More AI responsibility requires stronger proof outside the conversation. The Split I Use Now I didn’t sit down before this project and design a three-part method. Each phase kept failing for a different reason, and the split appeared from those failures. Hermes worked best while the starting point was messy. I use it when the request sounds like “find the package we used before,” “check what we decided last month,” or “compare this with the rest of my setup.” It searches across the wider project and brings back a prepared job. Codex worked best after the task fit inside one project folder. I use it once I know where the work lives, what needs to change, and what success looks like. “Improve the Substack tool” leaves too much room. “Run both versions and tell me which one fits an app used by AI agents” gives the work an edge. Tests take over wherever the answer belongs to the machine. A field name, an installable file, or a response from my real account shouldn’t end as a debate between two models. I think of those roles as scout, specialist, and referee. A scout finds the right ground. A specialist works inside a defined surface. A referee ignores confidence and checks what happened. Choose the role before choosing the model. A stronger model won’t rescue a job whose boundaries are still mixed together. Copy This Handoff My handoff between Hermes and Codex now fits inside one short file: Goal [one result the project should produce] Starting point [where the work lives and what already works] Known evidence [useful files observations and previous attempts] Limits [what the AI must avoid] Checks [how we will prove the result works] Stop and ask [what needs my approval before continuing] Hermes fills the first version from the wider context. I remove weak assumptions, cut extra scope, and make the checks specific. Codex receives the file with the project and works until a check passes or it reaches something needing my approval. When a check fails, the next request gets smaller. I don’t reopen the whole project. I send the failed behavior back as one specific correction. This handoff works outside software too. Research produces a sourced brief, writing turns it into a draft, and a publication checklist checks the result. My article about file-based AI workflows covers the wider system. What Shipped Version 0.2.2 of the Unofficial Substack SDK gives other apps a reusable way to work with Substack. It reads profiles, posts, comments, engagement numbers, and subscriber data. It also publishes Notes, schedules them, edits drafts, and adds images. What feels crazy to me is that npm recorded 708 downloads from July 13 through July 19. Those are downloads rather than 708 individual people, and I won’t pretend otherwise. I still find the number encouraging for a new tool solving a problem I found in an old folder on my server. If you want to give this a test you must know that Substack might change its private API without warning. Tests make breakage easier to find, though they don’t turn an unofficial tool into a promise from Substack. Two Agents One Workflow This project started because I wanted Hermes to schedule my Substack Notes. It ended with an open-source SDK, a server running on my VPS, and a much clearer idea of where each agent belongs. Hermes earned its place at the messy beginning. It found the forgotten package, pulled in old conversations, compared the available projects, and followed the work when Cloudflare blocked the first version. Carrying all this context also made Hermes burn through my OpenAI-banked resets much faster. Codex worked better once the problem became smaller. In VS Code, it stayed close to the active files and handled precise edits without loading the full project history. This made it faster and cheaper for the small changes where a focused coding agent has the advantage. Independent checks sat between both agents and the finished package. Real account responses and passing tests overruled every confident answer before anything shipped. One setup detail matters if you want to try the SDK. I deliberately left the Substack cookie out of the npm package because it gives access to your account. Ask your agent to read the setup instructions and tell you what needs configuring, then add the value of your substack.sid cookie yourself as a trusted server-side environment variable and pass it as sessionToken. Never paste the cookie into your codebase or anywhere likely to save and share it. Substack blocked requests from Cloudflare during my build, which is why I added a server interface and ran it on my VPS. Your setup might behave differently, though your login cookie should always stay on a server you trust. Look at your last AI project sprawling across one long conversation. Give the messy discovery work to the agent with the widest view, hand the focused edits to the agent closest to the files, and let independent checks decide when the result is ready. Tell me where your handoff keeps breaking in the comments. I want to collect real examples and turn the common failure points into a follow-up piece. I’m working on Hermes 101, and it should be ready soon. It starts from the same idea: one visible workflow, a clear place for human judgment, and more autonomy only after the loop earns trust. For the wider tool decision, read When to Use MCPs CLIs or Your Own Tool. For the responsibility behind the work, read The Agentic Engineering Shift. Use the agent with the widest view to find the path, use the focused agent to make the change, and trust neither until the work passes a check.
15:30

Did We Listen to the Right Noise?

A hobbyist teamed up with an AI to test whether aliens could hide messages in radio noise, and the experiment honestly came up empty. They downloaded 76GB of real telescope data from Breakthrough Listen and ran statistical tests a grad student would take weeks to write. The analysis seemed to reveal hidden structure, but a control observation of a different star showed the identical pattern, proving it was the telescope's own equipment and software, not the sky. The underlying idea, called noise modulation, is real physics, but the telescope pipeline flattens variance so much that a signal would need over 47% modulation to be seen. The point is less the result than the workflow: the human supplied the intuition and the AI executed, verified, and told the truth when the answer was no.

Notes
Did We Listen to the Right Noise? — The Augmented Mind: Think with AI (substack, 2026-07-21)

Author is a self-described non-expert (background in sound engineering), partnered with an AI on the "Hermes Agent platform," testing a private hypothesis: civilizations might encode messages not in a carrier but by modulating noise variance (making static louder/softer in places). Record of a real experiment on Breakthrough Listen data, with an honest null result.

Hypothesis and literature
  • NoiseMod scheme: data encoded as variance modulation across frequency bins; transmitter modulates amplitude of Gaussian noise; receiver reads variance statistics, not a carrier. Invisible to conventional energy detection.
  • Cited foundations: Bash et al. (2015) formalized covert-communication limits — a signal can exist below the "noise floor," undetectable by standard energy detection. NoiseMod paper (Dec 2023) matches the author's imagined scheme.
  • Masking problem (author's own concept): telescope receiver noise (50-ohm impedance, LNA thermal noise, RFI) and pipeline normalization can bury or mimic noise-modulated variance, raising detection thresholds.
Data
  • FRB 121102 (fast radio burst from a magnetar, 2.5 billion light-years away): 27-minute continuous observation, 3,876–9,314 MHz across 14,848 channels. Described as 76 GB file in body; technical appendix says 69.3 GB, 4.67 million integrations.
  • Control: HIP35136 calibration observation, different star/band/time, 4.9 minutes (appendix: 0.1 GB, 273 integrations).
  • Breakthrough Listen pipeline is built for narrowband power spikes — "literally the wrong search" for variance patterns.
Experiments and results
  • Exp 1 — first look: per-channel variance, kurtosis, autocorrelation. Fano factor of 0.0008 (1.0 expected for thermal noise) — variance compressed by 99.92%. Pipeline normalization had already removed the target structure.
  • Exp 2 — masking quantification: GBT system temperature ~50–100 K; LNA thermal noise ≈ half the detected signal. Detection threshold for a NoiseMod signal: >47% variance modulation in 5-min observation, 12% at 30 min, ~3% with multi-hour data. Receiver noise dilutes signal variance by 1.9×; nearly half the spectrum is RFI-contaminated. A 10% modulation — "easily achievable by an advanced transmitter" — would be invisible.
  • Exp 3 — large-scale: 100,000 integrations processed via a custom chunked reader (standard tools crashed on the 76 GB file). Result at 0.35 ms resolution: data "too flat, too uniform, too processed."
  • Exp 4 — eigen-structure analysis: full covariance analysis (SVD, random matrix theory / Marchenko-Pastur tests, temporal autocorrelation, channel correlation, frequency-patch variance mapping). Found real structure — non-Gaussian, channels non-independent — invisible to all 1-D analyses.
Control and verdict
  • Same eigen-structure analysis on HIP35136 returned identical structure, amplified. Conclusion: the covariance structure is instrumental — a common-mode signal from the GBT and correlations introduced by the GUPPI spectrometer pipeline. "The structure is real. It is not alien." The masking problem is quantified, the NoiseMod hypothesis sound, but this dataset contains no such signal. Detection thresholds in existing data run >80% at fine resolution.
Author's takeaways and next steps
  • Claims demonstrated: a non-expert intuition validated by literature; AI translated it into real code/statistics; human supplied the key 3-D covariance insight; control experiment produced an honest null rather than confirmation bias.
  • Masking numbers restated: receiver dilutes variance 1.9×, pipeline compresses further; thresholds >80% on current data.
  • Three un-tried paths (all need telescope time/expertise the author lacks): raw voltage data pre-spectrometer (terabytes/minute, rarely preserved), a purpose-built variance-preserving observation, and model subtraction of the now-characterized instrumental structure.
Caveats and limitations
  • No claim of ET detection; conclusion is that current pipeline-masked data cannot test the hypothesis.
  • Discrepancy within the article: body says 76 GB file; appendix says 69.3 GB.
  • Appendix lists tools: blimpy.io.fil_reader, numpy, scipy; claims reproducibility.
  • Author discloses the article was written with AI help. Instrumental-origin conclusion flagged as a well-known radio-astronomy phenomenon (telescope systematics + pipeline artifacts).
Full text · 15,526 chars
Did We Listen to the Right Noise? A human intuition. Real radio telescope data. An AI that could actually run the math. One experiment that asked whether the universe is whispering through its own background hiss, and what we found. This is not a SETI paper. This is a record of what happens when a non-expert has an intuition about how the universe might communicate, and has access to an AI that can translate that intuition into real code, real data, and real analysis. Part I: The Intuition What if the signal is not in the signal? Every SETI search looks for structure in noise. A spike. A narrowband carrier. A repeating pulse. The assumption is that any civilization worth finding would broadcast something you could hear. But what if they did the opposite? What if they encoded information in the noise itself? The intuition started simple: if you have enough bandwidth, you do not need a clean channel. You modulate the variance of the noise. You make the static louder in some places and softer in others, and the pattern itself is the message. I had no idea how to test this. I am not a radio astronomer. I do not read Fortran. I do not know what a GUPPI spectrometer does. But I had the question, and I had an AI that could actually do the work. So we did. Part II: The Verification Is hiding data in noise even real? Before touching data, we needed to know: is this physics or science fiction? The literature says yes. The concept has existed under different names for decades. Spread spectrum communication: Military systems encode data across wide frequency bands by modulating noise patterns, making signals invisible to anyone not tuned to the exact modulation scheme. This is how modern jam-resistant military radios work. Covert communication theory: Bash et al. (2015) formalized the limits of hiding information in noise. Their key finding: covert communication is possible below the “noise floor,” meaning a signal can exist without being detectable by standard energy detection. The receiver does not need to hear the carrier; it needs to hear the variance. Noise Modulation (NoiseMod): Published in December 2023, this is exactly the scheme I was imagining. Data is encoded as variance modulation across frequency bins. The transmitter modulates the amplitude of Gaussian noise. The receiver detects changes in statistical variance rather than detecting a carrier wave. This makes the signal invisible to conventional analysis. The academic foundation was real. The question was whether any cosmic signal was actually using it. The masking problem But there was a complication I had not considered at first. As a sound engineer, I know this problem well enough: every signal chain adds its own noise, and that noise does not just cover what you are listening to. It changes the shape of what you hear. Radio telescopes do not just listen to the sky. They listen to the sky through layers of equipment: amplifiers with thermal noise, local oscillators, cables with impedance mismatches, software pipelines that normalize and integrate and compress. The “noise floor” we measure is partly cosmic, partly terrestrial. If an alien civilization encoded data in cosmic noise variance, and our receivers added their own noise variance on top, the signal could be buried under our own equipment signatures. This is what I called the masking problem. The hypothesis: Terrestrial receiver noise (50-ohm impedance, LNA thermal noise, RFI from human infrastructure) obscures or mimics potential noise-modulated signals. The detection threshold rises as masking noise rises. We needed real data to test this. Not simulations. Not papers. Actual radio telescope output from a real observation. Part III: The Data Breakthrough Listen: 76 gigabytes of the sky Breakthrough Listen is the largest SETI dataset in existence. Over 97 petabytes of observations from the Green Bank Telescope and other facilities, searching for narrowband signals across millions of stars. We downloaded one observation: FRB 121102. A fast radio burst, a millisecond flash of radio energy from a distant magnetar, 2.5 billion light-years away. The observation was 27 minutes of continuous recording, covering 3,876 to 9,314 MHz across 14,848 frequency channels. The file weighed 76 gigabytes. Here is the thing: the Breakthrough Listen pipeline is designed to find narrowband signals. It looks for power spikes. It does not look for variance patterns. So if NoiseMod signals exist in this data, the BL pipeline would not have found them. It is literally the wrong search. We also downloaded a calibration observation, HIP35136, a different star, different frequency band, different time. 4.9 minutes of data. This would become our control. Part IV: The Experiments Experiment 1: The first look We loaded the data and ran the obvious tests: per-channel variance, kurtosis, autocorrelation. Standard signal processing. The results were disappointing at first. The variance across channels was nearly uniform. The autocorrelation was weak. The statistics looked like thermal noise. Initial result: Fano factor of 0.0008 (should be 1.0 for thermal noise). This meant the variance was compressed by 99.92%. Something in the pipeline was actively normalizing the data. The variance structure we wanted to find had already been removed. We had hit the masking problem head-on. The telescope pipeline was doing exactly what we feared: flattening the variance. Experiment 2: The masking quantification We built a model to estimate how much masking was happening. The Green Bank Telescope’s system temperature is around 50-100 Kelvin. The LNA thermal noise alone accounts for nearly half the detected signal. Pipeline processing compresses variance by orders of magnitude. The detection threshold we calculated: a NoiseMod signal would need over 47% variance modulation to be visible in a 5-minute observation. With 30 minutes, that drops to 12%. With multi-hour data, it approaches 3%. Key finding: The masking is not theoretical. It is quantified. Receiver noise dilutes signal variance by 1.9x. Pipeline processing compresses it further. A 10% modulation, easily achievable by an advanced transmitter, would be invisible in standard observations. Here is what the masking looks like across the full spectrum: RFI spectrum map. Red regions show frequency bands contaminated by terrestrial interference. Nearly half the spectrum is RFI-contaminated. And here is how the variance is distributed across all channels: Variance distribution. The compressed spread (Fano factor 0.0008) shows heavy pipeline normalization. Anomalies cluster in known RFI bands. The masking budget across the signal chain: Masking budget. Receiver noise (Tsys 50-100K) and RFI together account for the dominant variance floor, raising the NoiseMod detection threshold to >47%. And the full RFI mask analysis: RFI mask analysis. Terrestrial interference patterns across the observation, showing both persistent and transient RFI sources. Experiment 3: Larger data, deeper analysis The calibration observation was 5 minutes. The FRB observation was 27 minutes. But neither was enough for proper time-domain analysis. We needed more time samples to see temporal structure in the variance. We processed 100,000 integrations of the FRB data using a custom chunked reader, because the standard tools could not handle 76 gigabytes without crashing. The system ran out of memory. We built a workaround. The analysis ran. Technical note: This is where AI empowerment matters. A human radio astronomer would know to use a chunked reader. I did not. The AI did, because it could reason about the memory constraint and build the solution. The experiment was mine. The execution was collaborative. The results from the large dataset: Modulation depth needed at 0.35ms resolution The data was too flat. Too uniform. Too processed. Experiment 4: The three-dimensional insight Then came the most important moment of the investigation. I realized we were flattening the data. We were looking at per-channel variance, or temporal autocorrelation, one dimension at a time. But if information is encoded across multiple channels simultaneously, it would not show up in any single channel’s statistics. It would live in the covariance structure. The relationship between channels, not the channels themselves. The intuition: “Maybe the noise we see is the flattening of our 3D perspective, but that noise when expressed in multidimension tells a different story.” This was not an AI insight. This was mine. The AI built the tools. The question was human. We ran the full eigen-structure analysis: The data was not random thermal noise. The structure was real. It lived in the covariance between dimensions, invisible to every single-dimension analysis we had run so far. For a moment, it felt like we had cracked something. The variance was not uniform. The channels were not independent. The statistics were not Gaussian. The data had structure that was not noise. Then we ran the control experiment. Part V: The Control The calibration observation that settled it The BL calibration observation, HIP35136, a different star, different band, different time, should be clean. No FRB. No transient. Just the telescope looking at a quiet patch of sky. We ran the same eigen-structure analysis. The results were unmistakable: The structure was identical. Just amplified in the calibration data. The verdict was clear: the covariance structure, the non-random correlations, the high-dimensional patterns, they come from the telescope and its pipeline, not from the sky. The Green Bank Telescope introduces a common-mode signal across all channels. The GUPPI spectrometer pipeline processes the data in ways that create correlations between frequency bins. The Gaussian noise we see is not cosmic. It is instrumental. The honest conclusion: The structure is real. It is not alien. It is the fingerprint of a 150 million dollar radio telescope and the software that processes its output. The masking problem is real and quantified. The NoiseMod hypothesis is sound. But this data does not contain the signal we were looking for. Part VI: What This Actually Shows Getting the answer wrong was the point Let me be direct about what this experiment demonstrated, because it is not what you think. It did not find alien signals. It did find something: - A non-expert had a scientifically valid intuition. The NoiseMod hypothesis was grounded in real physics. The literature confirmed it. The question was legitimate. - An AI could translate that intuition into real, rigorous analysis. Not theoretical discussion. Not literature review. Actual code running on actual radio telescope data from the Green Bank Telescope. Statistical tests that would take a graduate student weeks to implement. - The human provided the critical insight. The three-dimensional covariance analysis, the step that revealed the structure, was a human intuition. The AI built the tools. The question was mine. - The experiment reached an honest conclusion. Not a paper chase. Not confirmation bias. The control experiment settled it. The structure was instrumental, not astrophysical. - The masking problem is real and quantified. Receiver noise dilutes variance by 1.9x. Pipeline processing compresses it further. Detection thresholds for NoiseMod signals in existing data are over 80%. The problem is not theoretical. It is measured. This is what augmented cognition looks like in practice. Not the AI generating an essay. Not the AI writing code from scratch. A human with a genuine question, partnering with a system that can execute, verify, iterate, and crucially, tell the truth when the answer is “the structure you found is just the telescope.” Where this leaves us The NoiseMod hypothesis is untested on cosmic data, because the data we have is masked by our own equipment. The detection threshold in existing radio telescope observations is too high. The variance structure we need to find has been normalized away by pipelines designed to look for narrowband carriers. The next step would require one of three things: - Raw voltage data: pre-spectrometer, where the pipeline has not flattened the variance yet. These files are enormous (terabytes per minute) and rarely preserved. - A new observation: purpose-built to preserve variance structure, with minimal processing and wide enough bandwidth that a NoiseMod signal would be detectable above the masking floor. - Model subtraction: characterize the instrumental structure (which we have now started), build a model of it, subtract it from the data, and analyze the residuals. Any of these would require telescope time, infrastructure, and expertise that I do not have. The experiment ended where it needed to. Reflection The augmented mind thesis, in practice I started this project with a question I could not answer and tools I did not understand. I ended it with an answer I trust, data I have seen, and a clearer picture of where the limits of our current search methods lie. The AI did not replace the human intuition. It amplified it. It translated a vague feeling: “what if the signal is in the variance?” into eigenvalue decomposition, random matrix theory tests, temporal autocorrelation analysis, and channel covariance mapping. It built custom chunked readers when standard tools crashed. It generated visualizations. It ran control experiments. And when the structure turned out to be instrumental, it told us. That is the part that matters most. Not the finding, but the honesty of the finding. This is what I mean when I talk about the augmented mind. Not artificial intelligence replacing human intelligence. Human intelligence, extended by systems that can compute, verify, and execute at a scale no single person could manage alone. The question was mine. The math was not. But the conclusion is ours. Technical appendix: This analysis used real data from Breakthrough Listen’s Green Bank Telescope observations. The FRB 121102 dataset (69.3 GB, 4.67 million integrations) and BL calibration dataset HIP35136 (0.1 GB, 273 integrations) were processed using custom Python scripts running on the Hermes Agent platform. Tools used: blimpy.io.fil_reader, numpy, scipy. All analysis code and results are reproducible. The masking analysis estimated receiver noise contribution (Tsys 50-100K), pipeline variance compression (Fano factor 0.0008), and NoiseMod detection thresholds (over 47% modulation for 5-minute observations, over 80% for 0.35ms resolution data). The eigen-structure analysis used SVD decomposition, random matrix theory (Marchenko-Pastur) tests, temporal autocorrelation, channel correlation analysis, and frequency patch variance mapping. The control experiment compared eigen-structure metrics across two independent observations and confirmed instrumental origin of detected structure. Disclaimer: This article records a real experiment with real data. The conclusions are based on statistical analysis performed by the author and AI-assisted tooling. The NoiseMod hypothesis is grounded in peer-reviewed literature (Bash et al. 2015, NoiseMod 2023). The instrumental structure detected in the data is a well-known phenomenon in radio astronomy: telescope systematics and pipeline artifacts create correlations across channels. This analysis does not claim to have detected extraterrestrial signals. It claims to have demonstrated the feasibility of a new kind of investigation. This article has been written with the help of AI.
10:10

Hi, I’m glad you’re here.

A cybersecurity incident commander launched a daily AI newsletter promising no-hype, practitioner takes on tools, news, and workflows. This is only a welcome post, so the substance is thin: it frames the pitch, which is a daily fresh issue for people who want honest AI takes.

Full text · 599 chars
Hi, I’m glad you’re here. I’ve spent about a decade in cybersecurity, currently as an incident commander, the person who runs toward the problem when a system is on fire. That job teaches you to ignore the marketing and find what actually works. Open Cloud AI is that instinct, pointed at AI. No hype, no noise, no press-release rewrites. Just a practitioner’s honest take on the tools worth your time, the news that actually matters, the workflows you can steal, the career moves ahead, and the infrastructure and policy quietly deciding who wins. A fresh issue every day. Subscribe and come along.