Nothing matches those filters.

Lead

1

Article

2
13:10

Opus 5 >> Fable 5

Claude Opus 5 is out, and Anthropic says it lands close to its flagship Fable 5 at half the price, with more than 80% of Claude Code's system prompt stripped out and no measured coding loss. In a blind ranking a tester still put Opus 5 first even though she disliked using it, and reviewers note it argues, stops early, and fights skills built for older models. Also in this roundup: ChatGPT Voice can now drive the ChatGPT desktop app, Kimi K3's weights and technical report shipped, and NVIDIA's Jensen Huang posted on X for the first time backing open-source AI, a statement OpenAI, Google, and Microsoft signed but Anthropic did not, drawing accusations of hypocrisy.

Notes
Opus 5, ChatGPT Voice, and tldraw canvas building — Ben's Bites, 2026-07-28

Personal build (Keshav/Ben): Using tldraw's offline canvas app so agents produce visuals (mascot "bites" animating from the logo, timezone widget, weather timeline). Motivation: text-only agent output is boring and ignored; wants a one-stop canvas for todos/emails/widgets, then a canvas-based personal site. Plans to wire the mascot to an agent he can drag tasks onto.

Sponsor (Ability AI): "17 AI agents" running real ops (marketing, client health, competitive intel, finance), git-tracked actions, human approval for external actions. Apache 2.0: github.com/Abilityai/trinity.

Claude Opus 5: Anthropic claims it approaches Fable 5 at half the price. Claire "hated using" it but ranked it first in a blind model-output test. Every's review: it argues, stops early, and "fights the skills people built for older models"; possibly being overprompted. Anthropic cut more than 80% of Claude Code's system prompt for Opus 5 and Fable 5 with no measurable coding-eval loss.

ChatGPT Voice: Now runs the ChatGPT desktop app, works in ChatGPT Work and Codex. A "Voice" session can spawn new sessions for individual tasks and report results back in the main thread. Keshav: "feels a lot like OpenClaw with voice," but delegation "doesn't feel as smart as doing it in a chat with 5.6 Sol in Codex" — no control over which model/thinking level each spawned session uses. Verdict: great for productivity (email checking, draft replies, dashboard chart pulls); for serious work he prefers his own speech-to-text app (optionafk.com) plus full model/thinking control.

Claude voice mode: previously Haiku-only, now supports Sonnet and Opus and can call Gmail, Calendar, Slack mid-conversation.

Other news: Kimi K3 weights + technical report released; in Droid at 50% off until 10 August. Jensen Huang made his first-ever X post backing open-source AI amid rumored US ban on Chinese open-weight models; OpenAI, Google, Microsoft and 150+ signed his statement — Anthropic did not. Backlash followed; Anthropic published an essay claiming it never pushed for an open-weights ban. Critics call hypocrisy: fine with "cute little 8b toys" but not "big, dangerous models" that threaten its business.

Quick links (notable): FLUX 3 — single model for image/video/audio/action prediction, video in early access; Health in ChatGPT (US, Apple Health + medical records, context-aware questions); BackSearch (web search as of a past date, news-only for 2026 for now); open-git (agent code host with autopilot that fixes PRs when CI is red); OpenWorker (open-source Mac agent, asks before consequential actions); Pilot Protocol (250,000+ agent network for app distribution — Ben is an investor); Notion-as-code (TypeScript-defined workspace, beta); The Stack v3 (5T filtered tokens, 770 languages); box (VM sandboxes at $0.00001/sec with snapshots/SSH/desktop); Mintlify (resolved support tickets → doc fixes); Cognition acquired Interaction (Poke team joins Devin makers).

Skills: Interface Craft (4 animation lessons + updated DialKit); /prototype (multiple UI versions, shareable URL); Transitions.dev (33+ CSS/React transitions or agent skill); sharechat (share Claude Code/Codex sessions as unlisted, no-signup links).

Full text · 6,041 chars
Opus 5 >> Fable 5 Codex plus Voice is the new openclaw Hey folks, I’m back from holiday and straight back to messing around building. I’ve been kinda obsessed and enamoured with the demos from tldraw. It’s a free canvas/whiteboard app but their offline app (recently updated) is cool because agents can do all sorts with it; create visualisations to explain stuff, make games, interactive blog posts and a ton of other stuff. Like turning our logo into a little animating mascot. He’s called ‘bites’ btw. or little widgets like a timezone checker or a daily weather timeline I’m not good using agents when everything is just text in a file. It’s boring, I don’t read it and I sure as shit don’t do enough (anything?) with it. So if it’s visual and all on one canvas, maybe I will? My idea here is to create a little ‘one-stop’ canvas for some widgets I use regularly as well as showing me my todos/emails to sort and things like that. I’m then planning on making my personal site an interactive canvas too, then some more explainer-like stuff too. And maybe I’ll get the mascot hooked up to an agent so i can drag him to my todos and ask him to do stuff… I’ll share more soon 😊 Ben’s Bites is brought to you by Ability AI This startup runs on 17 AI agents - and open-sourced the platform. Marketing, client health, competitive intel, finance - real ops, not demos, with every agent action git-tracked and a human approving anything external. Agents you own, not rent. Apache 2.0: github.com/Abilityai/trinity Headlines Claude Opus 5 is out - Anthropic says it gets close to Fable 5 at half the price. Claire hated using Opus 5, but when asked to rank model outputs in a blind test, she ended up rating it first. Every’s review lists some behaviours that might explain this: it argues, stops early and fights the skills people built for older models. And maybe we are overprompting it. Anthropic says it removed more than 80% of Claude Code’s system prompt for Opus 5 and Fable 5 with no measurable coding-eval loss. ChatGPT Voice can now run the ChatGPT desktop app. You can use it in ChatGPT Work and Codex both; just start a “Voice” session and speak to it. It can create new sessions for individual tasks and then report the results to you in the main thread. It feels a lot like OpenClaw with voice. I didn’t have a great time using ChatGPT Voice for “serious” work. The delegation of tasks to new sessions via ChatGPT Voice doesn’t feel as smart as doing it in a chat with 5.6 Sol in Codex. Maybe it’s the lack of control on what models/thinking level is used for each new session. I imagine Voice to work great for productivity type tasks, checking emails, creating draft replies, pull some info from a dashboard and create a chart etc. For serious work, I still prefer using a speech to text app (currently using one I built — optionafk.com) and sending that to the agent of my choice with full control over model/thinking effort. — Keshav Also, I wouldn’t blame you if you missed it, but Claude also has a voice mode. Till now, it could only use Haiku, but now it can use Sonnet and Opus as well as call multiple tools like Gmail, Calendar and Slack mid-conversation. I don’t remember when I last used Claude’s chat option. It’s been Claude Code for a long time now. Kimi K3’s weights and technical report are out. It’s already available in Droid at 50% off until 10 August. Following rumours that the US might ban companies from using Chinese open-weight models, NVIDIA’s CEO Jensen Huang made his first-ever post on X - a long statement backing open-source AI. Over the weekend, everyone signed that statement (OpenAI, Google, Microsoft and 150+ more) - except Anthropic. Of course, that led to some backlash, and Anthropic did what it does: published an essay to save face, saying it never pushed for an open-weights ban. The reaction to it is people calling Anthropic out on their hypocrisy — they are okay with open models as long as they are cute little 8b toys, but not if they are “big, dangerous models” (to Anthropic’s business). Quick links - FLUX 3 - one model for image, video, audio and action prediction; video is in early access. - Health in ChatGPT - US users can optionally connect Apple Health and supported medical records, then ask about changes in context. - BackSearch - search the web as it was on a particular date in the past. Right now, it only works for news in 2026. - here.now Workspaces - private team home for the sites, docs and dashboards your agents publish. - Why software factories fail - part 1, part 2, part 3. - open-git - a code host built for agents, with an autopilot that fixes your PR when CI goes red. - OpenWorker - open-source Mac agent that finishes everyday work and asks before consequential actions. - Pilot Protocol - a network of over 250,000 agents that you can use to gain distribution for your app. (I’m an investor) - Notion as code - define your whole workspace in TypeScript and deploy it through the API. (beta) - The Stack v3 - open code dataset with about 5T filtered tokens across 770 languages. - box - full VM sandboxes for your agents at $0.00001 a second, with snapshots, SSH and a desktop. - Mintlify - turns resolved support tickets into proposed documentation fixes. - Cognition acquired Interaction - the team behind Poke joins the makers of Devin. Skills section - Interface Craft - four new lessons on advanced animation, plus an updated DialKit skill for timelines. (tweet) - /prototype - asks your agent for several UI versions and saves the chosen one in a shareable URL. (example run) - Transitions.dev - 33+ polished UI transitions as copyable CSS or React, or an installable agent skill. (tweet) - sharechat - share a Claude Code or Codex session as an unlisted link; no signup, no login to read it. Afters - Find me on X, Linkedin, or YouTube - Read about me and Ben’s Bites - 📷 thumbnail via @keshavatearth * sponsors who make this newsletter possible :) Wanna partner with us for the next quarter? Email us at shanice@bensbites.com or k@bensbites.com
11:03

The Sequence Knowledge #902: Learning About Distillation: When the Dataset Becomes the Teacher

A cheap model can learn a big model's skill without the big model ever running again, because the big model just produces the training data offline first. Called 'synthetic data as distillation,' this turns data from something you dig up into something a model manufactures on demand. The teacher's answers become a dataset, the smaller student learns from it, and the teacher disappears at runtime while some of its behavior stays embedded in the student.

Full text · 1,177 chars
The Sequence Knowledge #902: Learning About Distillation: When the Dataset Becomes the Teacher Once models can generate their own curricula, data stops being a static resource and becomes a transmission medium for intelligence. For most of machine learning history, data was treated as geology. It already existed somewhere in the world—in books, websites, code repositories, conversations, photographs, and databases. The researcher’s job was to excavate it, clean it, tokenize it, and feed it into a model. Large language models changed this relationship. A capable model is not only a consumer of data. It can produce questions, answers, explanations, critiques, preference labels, tool traces, textbooks, code exercises, and entire miniature curricula. This creates a new training primitive. Instead of asking an expensive model to answer every production query forever, we ask it to manufacture the experience from which a smaller model learns. The teacher runs offline. Its outputs become a dataset. The dataset trains the student. The teacher disappears at inference time, but some of its behavior remains embedded in the student. That is synthetic data as distillation.

Newsletter

4
12:37

How Pangram Decides What Looks Like AI

Substack's new 'Scan for AI' button uses a trained detector called Pangram that judges writing against its own training data, not by hunting em dashes or repetitive rhythm like older detectors. To avoid cheating, it pairs each human document with a matching AI-written 'twin' about the same subject, and retrains on the human passages it wrongly flags. Independent University of Chicago testing found near-zero misclassification on matched passages, beating GPTZero and OriginalityAI, but the percentage it shows is the share of text segments classified as AI, not the odds a human used AI. The open weights are only a smaller research model; the full commercial version 3.3 stays closed.

Notes

How Pangram Decides What Looks Like AI

All Agents Considered (Substack), 2026-07-28

Substack added a "Scan for AI" feature (Pangram) to every post, Note, comment, and reply. Pangram estimates how much of a text is human, AI-assisted, or AI-written. It reads only the finished text — it does not observe typing, check browsing history, or detect model signatures. The author's framing: Pangram is a serious pattern detector, not a yes/no authorship test.

Why older detectors earned distrust

Early detectors relied on perplexity (how surprised an LM is by the next word) and burstiness (variation in sentence length/structure). Both are flawed heuristics. 2023: Ars Technica reported a detector labeling part of the US Constitution as likely AI; a 2023 academic review found serious reliability problems after editing/paraphrasing. Pangram argues these two measures alone fail.

How Pangram works
  • A trained classifier (transformer adapted for text classification), not a rules/checklist detector.
  • It does NOT call OpenAI, recover prompts, or search a database of model outputs.
  • Visible clues (em dashes, lists, headings, stock phrases, Markdown) are separate from the flagship detector — shown to readers but not inputs to the main model. Analogy: removing a prisoner's striped shirt doesn't defeat a facial-recognition scan.
  • Synthetic mirroring: for each human document it creates an AI counterpart on the same subject/tone/style, so the detector must find differences between "twins" rather than comparing a shopping list to a sales page. Human training material: essays, reviews, books, creative writing, news, scientific papers, Wikipedia, general web text; AI side generated in-house.
  • Hard-negative mining: it scans large collections of known human writing, finds passages wrongly called AI, feeds those mistakes (plus AI mirrors) back into training. Invokes Abraham Wald's WWII survivor-bias story (reinforce engines, not wings).
Open model vs. commercial product
  • Public EditLens repo (code, data, model weights) accompanies a paper accepted at ICLR 2026 on measuring degrees of AI editing.
  • Public release is two smaller research models and explicitly "a research baseline [that] should not be used to enforce AI policies in schools or workplaces."
  • Commercial model is version 3.3. Model card documents a larger production system. Substack's help page does not name the exact version or threshold used.
What the percentage means

EditLens: take human text, ask models to edit at different intensities, learn to estimate editing degree from final text. Key limitation: the detector only sees the final text, never the original draft ("like a restorer guessing how much a painting was retouched").

  • Long documents are broken into overlapping windows (~50-word resolution per 3.2 card), with a finer pass near uncertain boundaries.
  • Output categories: Human, Lightly AI-assisted, Moderately AI-assisted, or AI. "30% AI" = share of classified segments in that category.
  • It does NOT mean: 30% chance AI was used, machine typed 30% of words, prompt recovered, or writer contributed 70% of thinking.
Evidence it works
  • **2025 University of Chicago working paper, *Artificial Writing and Automated Detection***: 1,992 verified human passages (blogs, novels, résumés, essays, business writing) matched with AI versions from GPT-4.1, Claude Opus 4, Claude Sonnet 4, Gemini 2.0 Flash. On medium/long passages: essentially zero false positives and false negatives; outperformed GPTZero, OriginalityAI, and an open RoBERTa detector.
  • Caveats: working paper; controlled matched set; tested the May 2025 Pangram service; not every voice/mixed workflow/humanizer/format.
  • Pangram's 3.3 card (company-reported): 0.01% FP long-form creative, 0.02% academic, 0.04% multilingual how-to, 0.49% poetry; 1.5% false-negative rate on a random Chatbot Arena set.
  • Feb 2026 GPTZero benchmark: GPTZero 4.3b beat Pangram 3.2 on its datasets. Author: this proves the winner changes with test set/model/threshold, not that GPTZero is better.
Known failure modes
  • Best on long-form prose in complete sentences; 3.3 warns about bullet lists, technical instructions, ToCs, references, templates, equations, headers, footers.
  • Short text is unreliable.
  • **May 2026 preprint *Base Models Look Human To AI Detectors***: base LM output looked human to both Pangram and GPTZero; instruction-tuned chat models were easier to identify; repeated paraphrasing made AI text look human. Suggests detectors recognize chat-model training habits, not a universal "AI writing."
  • The Atlantic: ChatGPT/Claude output run through the Walter Writes humanizer was labeled human (Pangram says 3.3 improved humanizer detection). Arms race asymmetry: "AI" can be a false accusation; "Human" can be disguised AI.

Pangram CEO Max Spero told The Atlantic the detector should never be the final arbiter.

Author's rules for reading a result
  • Check the material — detector intended for long-form prose; lists/instructions/templates/equations more FP-prone.
  • Read the label as resemblance, not observation of who typed.
  • Don't mistake the score for reconstruction of the writing process (high ≠ model's ideas, low ≠ no AI).
  • Ask what question you want answered — accuracy/originality/value must be judged by the reader.

Verdict: distinct from crude detectors (synthetic mirrors, hard-negative mining, calibration favoring avoiding false accusations, EditLens modeling mixed editing); independently performed well on tested passages — but sees only finished text, research is partial, accuracy varies by domain, new models/paraphrasing evade it, and the percentage is not a probability. Treat as evidence, not verdict.

Full text · 20,717 chars
How Pangram Decides What Looks Like AI From its training data to the percentage on your screen, this is how Substack’s new AI detector works. I have no doubt you already know that last week Substack just attached an AI detector to every article on the platform. The new “Scan for AI” feature lets readers analyze a post, Note, comment, or reply with Pangram, which estimates how much of the writing is human, AI-assisted or written by AI. And the percentage it returns might looks precise, like a speed camera telling you exactly how fast you were going. But Pangram did not watch the person type, check their browsing history, or find a hidden signature left by ChatGPT or Claude. It looked at the finished text and decided what the writing resembled. That distinction sits at the center of the argument happening across the Substack, professor, and AI research subreddits. Some people think these detectors generate little more than a random number. Others treat the score like a confession. Both sides are missing the point. Pangram is pretty good at identifying obvious AI writing, and it has more behind it than the early tools that made AI detection a laughing stock. In independent testing from the University of Chicago, it identified writing from widely available AI models with low error rates under the conditions studied. But it is not perfect. Some model output and repeated paraphrasing can evade it, and human writing can still be classified incorrectly. A low AI score does not prove that no AI was used, just as a high score does not prove an algo wrote the text. Pangram is a serious pattern detector, not a yes-or-no authorship test. It cannot magically settle the question of authorship. So today I am going a bit out of the way with the main topics of AAC to show you the truth behind Pangram and AI detectors. In this edition I’ll cover how Pangram is trained, what its score measures, why it differs from older detectors, and where it fails, without pretending the tool can prove with 100% certainty who wrote a piece a text. In this article: - Why older AI detectors earned so much distrust - How Pangram learns to identify AI patterns - What its percentage does and does not mean - The strongest evidence that it works - The known ways it can fail - How to use a result without accusing the wrong person Why Everyone Hates AI Detectors Early AI detectors often leaned heavily on two ideas called perplexity (not the tool) and burstiness. Perplexity asks how surprised a language model is by the next word. AIs are trained to predict the next word, that’s literally written in the DNA of all LLMs. On the other hand, our human writing patterns are less predictable just because our writing tends to be messier. Burstiness asks whether sentence lengths and structures vary. Humans may write one clipped sentence, then an overgrown paragraph, followed by a shorter sentences. AI models usually settle into a steady rhythm (like the Rule of 3’s). You don’t need me to point out that these “AI detection” techniques are fundamentally flawed. Imagine judging whether bread came from a factory by checking whether every slice has the same thickness. That clue may help in some cases, but it doesn’t tell you exactly who baked the loaf. A safety manual is predictable because it needs to be. A legal form repeats itself because consistency matters. A non-native writer may use simpler sentence patterns. A poet may produce language so unusual that the detector has little familiar ground beneath it. This is why old detectors produced absurd headlines. In 2023, Ars Technica reported that one detector labeled part of the US Constitution as likely AI-generated. A 2023 academic review of detection tools also found serious reliability problems, especially after text was edited or paraphrased. Pangram argues that those two measures alone fail as a basis for detection. Thus an online myth was born: AI detectors don’t work. Before we go into why Pangram is truly different, we really have to get this out of the way. Pangram is not the only trained detector in 2026. GPTZero describes a multiclass model for human, mixed, polished, generated, and paraphrased text. Originality says its current detector uses a trained transformer model built from human and generated examples. Pangram Does Not Search for an Em Dash Pangram is a trained classifier. Researchers train it on many documents labeled by source: human-written or AI-generated. From those examples, the detector learns which combinations of patterns tend to distinguish one group from the other. Think of it like teaching someone to recognize counterfeit passports rather than giving them a checklist of what a document should contain. Side note: this reminded me of the game Papers, Please. Pangram says its flagship detector uses a transformer model adapted to classify sequences of text. Its technical explanation and original research paper describe a system that learns from human writing and model-generated writing. Basically they created an AI that can detect AIs. It does not call OpenAI to ask whether ChatGPT wrote the paragraph. It does not recover the prompt. It does not compare the post with a giant database of every answer an AI company has produced. There’s literally no need for that. Pangram learns what different kinds of writing tend to look like, then it analyzes the text you feed it against that map. This should also clear up the em dash frenzy. Pangram can show human-readable supporting clues such as em dashes, lists, headings, stock phrases, Markdown, and groups of. But Pangram says those visible clues are separate from the flagship detector. They help a reader inspect the text and show that they are not the inputs driving the main model. To better put this into perspective, imagine a prisoner just escaped from incarceration. Deleting every em dash is like removing their striped shirt before a facial-recognition scan. The prisoner changed something visible without necessarily changing what the system recognizes. Pangram Trains on Twins The most interesting difference is how Pangram constructs its training set. Suppose we trained a detector using personal diaries as the human examples and software landing pages as the AI examples. It might appear accurate while learning the wrong lesson. Instead of detecting AI, it could learn that feelings are human and product features are synthetic. Pangram tries to prevent that shortcut with what it calls synthetic mirroring. For each human document, it creates an AI counterpart about the same subject, in a similar tone and style, carrying similar information. A human book review is compared with an AI book review of the same material. A human essay is paired with its AI twin. The detector must find differences between twins rather than compare a shopping list with a sales page. Pangram’s model card says its human training material covers essays, reviews, books, creative writing, news, scientific papers, Wikipedia, and general web text. Its AI side is generated in-house to create those matched comparisons. Then it makes the exam even harder. Pangram scans large collections of known human writing and looks for passages the detector wrongly calls AI. Researchers add those mistakes back into the training data, create AI mirrors for them, and train again. This technique is called hard-negative mining. There’s a famous story from World War II based on the statistician Abraham Wald. The military engineers were asked to examine aircraft that returned from battle and map where they had been hit. In most cases, the wings and fuselage were covered in bullet holes, so the obvious response was to reinforce those areas. Wald argued that the military was looking only at surviving aircraft, but if those planes could return despite being hit in the wings and fuselage, then those areas were not the most vulnerable. That’s how he figured out the places with few bullet holes (like the engines) needed more protection. Hard-negative mining follows a similar principle. The most useful evidence often comes from the cases a system mishandles. So instead of repeatedly training a detector on obvious AI writing, researchers study the difficult human passages Pangram (initially) falsely identified as AI and feed those mistakes back into the training process. Pangram’s Open Model Is Not the Substack Model Pangram has released public code, data links, and model weights through its EditLens repository. The project accompanies a paper accepted at ICLR 2026 on measuring degrees of AI editing. That is meaningful transparency. It lets researchers inspect and reproduce part of the method without trusting a marketing page. What it doesn’t reveal is the complete system running inside Pangram’s commercial product. The public repository contains two smaller research models and instructions for training them. Pangram’s Open Pangram announcement explicitly says the release is a research baseline and should not be used to enforce AI policies in schools or workplaces. The current commercial model described by Pangram is version 3.3. Its model card documents a larger production system with later training and scanning changes. Substack publicly confirms the Pangram integration, but its help page does not name the exact model version or threshold it uses. What the Percentage Means We can all agree that teal writing has made binary detection obsolete. A writer may create the first draft, ask Claude to tighten one section, run the result through Grammarly, dictate a new ending, and then rewrite every sentence by hand. Which words belong to whom? Pangram’s EditLens research tries to measure that grey area. Researchers begin with human text, ask models to edit it at different intensities, and compare each edited version with the original. A classifier then learns to estimate the degree of editing from the final text. The full EditLens paper states the important limitation of its method. When the detector checks your article, it doesn’t receive your original draft. It only sees the final text. The estimate is like a restorer examining a painting and guessing how much has been retouched. They may recognize signs of later work, but they can’t count the extra brushstrokes. For long documents, Pangram breaks the text into overlapping windows, checks each window, then makes a finer pass near uncertain boundaries. Its 3.2 model card describes an approximate resolution of 50 words. Pangram can then describe shares of a document as Human, Lightly AI-assisted, Moderately AI-assisted, or AI. A result such as 30% AI refers to the share of classified text segments assigned to that category. It does not mean: - There is a 30% chance the writer used AI - A machine typed exactly 30% of the words - Pangram found the original prompt - The writer contributed only 70% of the thinking Ideas, wording, structure, research, and final judgment are different contributions. No detector can reduce all of them to one authorship percentage by reading the final copy. This is where Substack’s interface creates a social problem. A percentage looks like a measurement, but underneath it’s just a classification. The Evidence That Pangram Works The strongest independent evidence I found comes from the 2025 University of Chicago working paper Artificial Writing and Automated Detection. The researchers gathered 1,992 verified human passages across blogs, novels, résumés, essays, business writing, and other genres. They produced matched AI versions using GPT-4.1, Claude Opus 4, Claude Sonnet 4, and Gemini 2.0 Flash. On medium and long passages in that test, Pangram produced essentially zero false positives and false negatives. It outperformed GPTZero, OriginalityAI, and an open RoBERTa detector in the same experiment. That should clear up the myth that “AI detectors never work”. But this should also be taken with a grain of salt. The paper was, after all, a working paper. The test used a controlled set of matched human and synthetic documents. It evaluated the Pangram service available in May 2025 and it did not test every Substack voice, mixed editing workflow, new model, humanizer, or unusual format. Pangram’s current numbers come from Pangram. Its 3.3 model card reports a 0.01% false-positive rate for long-form creative writing, 0.02% for academic writing, 0.04% for multilingual how-to articles, and 0.49% for poetry. It also reports a 1.5% false-negative rate on a random Chatbot Arena set. Those are impressive company-reported results, but they shouldn’t be treated as promises about the next text you scan. The clearest warning comes from a competing test. In February 2026, GPTZero published its own benchmark and reported that GPTZero 4.3b beat Pangram 3.2 on its datasets. That does not prove GPTZero is better. It proves the winner changes with the test set, model version, threshold, and person running the test. I made the same argument in my breakdown of why AI benchmarks fail real agent workflows. A leaderboard compresses the conditions into one number, but conditions change and can be picked arbitrarily. Where Pangram Messes Up Pangram’s own documentation says it works best on long-form prose written in complete sentences, like a Substack post. Its 3.3 limitations warn about bullet lists, technical instructions, tables of contents, references, templates, equations, headers, and footers. Poetry has a much higher reported false-positive rate than long creative prose. Short text creates another problem. Just as a smoke detector cannot identify much from a few molecules in the air, Pangram can not correctly classify a short text. Then there are false negatives. A May 2026 preprint titled Base Models Look Human To AI Detectors found that output from base language models often looked human to both Pangram and GPTZero. Instruction-tuned chat models were easier to identify than base models. Researchers also used repeated paraphrasing to make AI text appear more human to the detectors. That suggests detectors may be especially good at recognizing the habits produced by chat-model training, rather than detecting one universal substance called “AI writing.” The Atlantic found another crack. A reporter sent ChatGPT and Claude output through the Walter Writes humanizer, then watched Pangram label the result human. Pangram has since described better humanizer detection in version 3.3, so one test does not freeze the system forever, but it does show the shape of the fight. The detector learns the disguise, and the disguise keeps changing. It’s just like with police figuring out a new way to transport illicit substances, then traffickers come up with something new and the wheel keeps spinning. This arms race creates an important asymmetry: - A result of “AI” can be a false accusation - A result of “Human” can be successfully disguised AI Neither side provides proof. Pangram CEO Max Spero told The Atlantic the detector should never be the final arbiter. The company setting limits on its own tool carries more weight than a critic attacking it from outside. How to Read a Pangram Result A Pangram result is closer to weather radar than a fingerprint. Radar detects conditions associated with a storm. It can tell you when the signal is strong, where to look, and whether you should carry an umbrella. A fingerprint connects a person to an object through physical evidence, and Pangram has no such access to the writing process. I also don’t use Pangram as a moral scorecard. I write about technical subjects, and I use AI heavily to structure my thoughts and explain difficult concepts to less technical readers. My first question when I read something is much simpler: Did I get anything valuable from it? If an article teaches me something new, gives me a useful idea, or makes a complicated subject easier to understand, I’m satisfied with the time I spent reading it; whether AI helped produce it comes second. I take the same approach to Duolingo, where I’m currently learning German. Duolingo has openly described using generative AI to create and validate course content. I also know that more than 1,000 volunteers helped create dozens of its original language courses. I have serious ethical problems with that history. I personally believe that Duolingo owes an enormous debt to the community that helped build it, and embracing automation does not erase that debt. But I cannot reverse the company’s decisions, and I still want to learn German, so I keep on using the app and embrace the technology for what it can do. That is roughly how I approach AI-assisted writing. I care about the quality of the thinking, the usefulness of the result, and whether the article rewards my attention. A detector score may tell me something about how the prose was produced, but it cannot decide whether the prose was worth reading. With that in mind, I use four rules when reading a Pangram result: - Check the material. Is it long-form prose, or a short Note, poem, list, technical manual, reference section, or template? Pangram’s own model card says the detector is intended for long-form writing in complete sentences and warns that lists, instructions, reference sections, templates, and dense equations are more susceptible to false positives. - Read the label as resemblance. Pangram is a trained classifier. It studies patterns in human and AI writing and estimates which group a new passage most closely resembles. Pangram describes the process as turning the text into a numerical representation and passing it through a classifier. It does not observe who typed the words or watch the document being created. - Do not mistake the score for a reconstruction of the writing process. A high result does not tell me which ideas came from a model, how much work the writer contributed, or whether AI only helped with editing. A low result does not prove that no AI was involved. The detector sees the finished prose, not the conversation, research, prompting, or revision behind it. This is also why my best agent workflow is mostly files. Files preserve the trail between research, drafting, and editing, even when AI is part of the process. - Ask what question I actually want answered. If I want to know whether the prose resembles AI-generated writing, Pangram provides a useful signal. If I want to know whether the article is accurate, original, insightful, or worth reading, I have to judge those things myself. A strong idea does not become worthless because AI helped express it, just as a completely human-written article does not become valuable merely because a person typed every word. I apply the same discipline when AI gives me a technical answer because accepting AI’s first answer without checking it replaces judgment with convenience. I do not reject an answer because a model produced it, and I do not trust it merely because it sounds polished. I ask whether it is accurate, useful, and worth my attention. A Pangram result can describe patterns in the writing, but it cannot tell me whether I learned something. That would be dystopian. My Verdict On Pangram Pangram is different from the crude detectors people remember. Its classifier learns from paired human and AI documents. - Its synthetic mirrors reduce topic shortcuts. - Its hard-negative mining feeds human false positives back into training. - Its calibration favors avoiding false accusations. - EditLens attempts to model mixed editing rather than forcing every document into a human-or-machine box. Independent researchers found that this approach performed extremely well on their tested medium and long passages. Those are facts. Pangram still sees only the finished text. It cannot observe the actual writing process. Its published research does not reveal every part of the production system, and its accuracy varies by domain. Short and formulaic formats remain particularly difficult. New model types and paraphrasing can also evade detection. Not to mention independent benchmarks do not always agree with Pangram’s own results. And even the percentage can be misleading. Readers may mistake it for a probability or a record of the writing process, but it is neither. Those limitations are facts too. The reasonable conclusion lies somewhere between “all detectors are useless” and “the algorithm knows the absolute truth”. Pangram has earned the right to be treated as evidence, but it hasn’t earned the right to be treated as a verdict. Has Pangram scanned something you wrote from scratch? Tell me what kind of text it was, how long it was, and whether any part of your workflow involved AI. I want to compare the writing history with the score, not collect percentages without receipts.
13:10

The 20-Year-Old Who Took an App to 1M Downloads and $500K/Month — Then Handed the CEO Seat to Someone Else

A 20-year-old founder built QUITTR, an app that helps people quit watching porn, to a million downloads and about $500,000 a month fully bootstrapped, then handed the CEO job to someone else. The playbook: a 12-step onboarding that makes users self-diagnose before a hard paywall with no free trial, a 25% download-to-paid conversion, and influencer deals with view guarantees and ad repurposing rights, one of which earned $100,000 in 65 hours. Early traction leaned on Reddit guerrilla posts that got the team banned and a $3,000 Instagram influencer bet that returned $37,000. He has since launched an AI flight-booking app called Soar that booked $250,000 in its first 30 days.

Notes
Alex Slater / QUITTR — 1M downloads, $500K/mo, CEO handoff (Substack, Bootstrapped AI Builders, 2026-07-28)

Subject: Alex Slater (20, London → SF), founder of QUITTR, a porn-quitting app. Author first covered him Feb 2025 at 19, at ~$250K/month revenue five months after launch (referenced Medium post).

The numbers now

  • ~1M downloads within ~1 year of launch; ~$500K/month revenue; author says "almost certainly higher now" — likely past $1M/month per Alex's X posts.
  • Zero venture funding, fully bootstrapped.
  • Sept 2025: appointed a new CEO, stepping back ("I'm taking my creativity to the next level").
  • New app Soar (AI flight booking): $250K booking volume in first 30 days.

Founding story

  • Alex taught himself to code at 17, dropped out of UK university at 18, moved to SF; met Connor McLaren via Y Combinator's co-founder matching.
  • Team: Alex code (19), Connor marketing (22), brother Chris design (17). Learned TikTok/influencer marketing from Zach Yadegari (Cal AI, sold under 2 years) met at a Miami hotel; Zach learned from Blake Anderson — "Blake's fingerprints are on QUITTR's success."
  • Initial version built in ~10 days.

Early launch (gritty, not influencer-first)

  • Product Hunt launch "completely flopped." Moved to guerrilla posts in r/NoFap ("this app called QUITTR changed my life"). Author flags this as astroturfing; Reddit bans stealth marketing and they got banned — after generating ~$2,500 in revenue.
  • Put $3,000 into a Christian influencer on Instagram → turned into $37,000. Then an "infinite reinvestment loop": ~90% margins, ~100% of profit reinvested, founder comp kept to minimum.

Onboarding (the famous part)

  • 25% download-to-paid conversion; 99% completion on the 12-step onboarding.
  • Flow: 12 pages of questions ("When did you start watching?" "Has it been escalating?"), symptom selection (low motivation, diminished interest in real-life partners), user testimonials, negative effects/benefits of quitting, App Store review prompt, personalized plan → paywall. Mechanism: users self-diagnose, resolve to fix it, then hit the paywall.
  • Hard paywall, no free trial. They tried a free trial once; "usage rates collapsed" and they pulled it immediately. Claim: without the hard paywall, revenue would be "1/25th" of what it is.
  • Author caveat: his own Japanese-learning app runs a hard paywall too, but 1-star reviews ("can't use it for free") are pushing him to test a free trial; gut says pure hard paywall is stronger, but a free trial may be smarter until a healthy review base exists. Notes ~80% of other apps he's covered offer free trials.
  • Onboarding was reverse-engineered: they screenshotted top apps doing $5–6M/month, laid them out in Figma, extracted common patterns. QUITTR's onboarding is now widely copied.

Influencer deals (real numbers)

  • Targets: Christian, fitness, self-improvement influencers. Angle by niche: betrayal of faith / lowered testosterone / performance gains.
  • First contact: paid for a $50 consultation with a Christian creator to get a live call → that call generated $1,000/day; repeated with a $30 fitness consultation, similar results. (Cold DMs get ignored.)
  • Deal terms: flat fee $2,000–$10,000/month; performance bonus $500 per 500K views; CPM $2–3 per 1,000 views; 20% upfront, remainder on delivery; view guarantees (e.g. "$1,500 for four videos with a guaranteed 500K combined views"); proration/extra posts if short; app must appear on screen; contract grants rights to repurpose the video as a paid ad.
  • Signature case: Jeremiah Jones (2M followers), $3 CPM + bonus deal, 9.9M views → $40K in one day, $100K in 65 hours.
  • Systematized: $40K/month into Meta ads at 3–4x return ("discover winning creative via influencers → scale with paid ads"). Prospects tracked in a Google Sheet; VAs handle DM outreach.

Brand phase

  • Past $250K/month they pivoted to brand, modeling Liquid Death (punk-branded water → billion-dollar). Eyeing research collaborations with YouTubers and neuroscientists. Goal: "quit porn = QUITTR" top-of-mind recall.

The handoff & what came after

  • Aug 2025 X post: "I wish I could run multiple storylines at once. I want to be a footballer, an artist, an entrepreneur, and a politician simultaneously." Sept 2025: handed QUITTR to a new CEO. Author sees a pattern: Zach sold Cal AI; young founders convert winners into assets and move on.
  • Reset (habit app, "rebuild your good habits in 60 days", near-identical onboarding, conceptually close to Life Reset): appears to have failed — abandoned/unmaintained, 3.5 rating, nothing new, no niche. Author's guess: overconfidence, over-belief in onboarding, corners cut on core functionality behind the paywall. Lesson: onboarding can't compensate for product quality.
  • Soar: one-tap booking, claims cheaper fares than Google Flights/Skyscanner, flight updates pushed to iMessage, auto AI check-in with boarding pass via iMessage, auto fare-drop refund/credit. Author tested it: "UI/UX is excellent."

Author's takeaways: target deep-pain blank markets; ship in 10 days and polish onboarding/marketing first; don't neglect the product; study winning apps rather than reinventing; hard paywall for commitment products; structure influencer deals as view guarantee + ad repurposing rights; when something hits, systematize it, hand it off, climb the next mountain.

Full text · 19,269 chars
The 20-Year-Old Who Took an App to 1M Downloads and $500K/Month — Then Handed the CEO Seat to Someone Else A 25% download-to-paid conversion rate with no free trial, and one video that made $100,000 in 65 hours. Here's the entire QUITTR playbook — including the app Alex launched This Substack breaks down real-world cases of people making serious money with apps in the AI era. Today’s subject: Alex Slater. https://x.com/alexsllater I actually covered him once before, back in February 2025. He was 19 at the time. He’d launched an app with an unusual premise — an app that helps people quit porn, called QUITTR — and had just hit roughly $250,000 a month in revenue five months after launch. - https://medium.com/@yumaueno/alex-a-19-year-old-genius-who-reached-250k-in-monthly-revenue-in-just-5-months-after-launch-2e2160e8f8bb That was about a year and a half ago. So what happened to the app since? Short version: QUITTR crossed one million downloads within roughly a year of launch, and monthly revenue reached about $500,000. (It’s almost certainly higher now — based on his recent posts on X, he appears to be well past $1 million a month.) Zero venture funding, ever. Fully bootstrapped the entire way. Then, in September 2025, he made a surprising announcement: I’ve appointed a new CEO for QUITTR. I’m taking my creativity to the next level. At 20 years old, he casually handed over the top job at an app he built himself — an app that was firing on all cylinders. Here’s his post: Walking away from running a $500K-a-month app at 20 is a genuinely remarkable decision. And more recently, he launched a new AI flight booking app called Soar, which processed $250,000 in booking volume in its first 30 days. In this piece, I’m updating everything from my earlier coverage and rebuilding the full picture — the behind-the-scenes of his success and his specific growth tactics, with everything that’s come to light since. Stick with me to the end. There’s a lot here. ⚡️ A Refresher: What Exactly Is QUITTR? Let’s start with the basics. QUITTR is an app for people who want to stop watching porn. The core feature is a streak counter tracking how many days you’ve gone without watching. On top of that: a panic button you hit when the urge strikes, an anonymous community where users encourage each other, an AI therapist named Melus available 24/7, and a virtual plant that grows as your streak grows. Honestly, there’s nothing technically remarkable happening in this app. But it landed. Hard. We live in an era where OnlyFans and social platforms have exploded, and porn is flowing to young people constantly. An enormous number of young men are struggling with it — and yet the only apps addressing the problem were relics that hadn’t been meaningfully updated in over a decade. QUITTR was a Gen Z developer walking straight into a wide-open Gen Z market that nobody else would touch. 🤝 The Founding Story So how did QUITTR come together? Let’s walk through it. In my earlier piece, I covered how Alex met Zach Yadegari at a hotel in Miami and learned his TikTok marketing playbook. Zach, of course, is the founder who built Cal AI at 17 and sold it in under two years. And Zach, in turn, learned from Blake Anderson, who I covered last week — so it’s fair to say Blake’s fingerprints are on QUITTR’s success too. The chain of young app founders spreading out from Blake is honestly ridiculous. But there’s another key figure behind QUITTR who shouldn’t be overlooked: Connor McLaren. Alex is from London. He taught himself to code at 17, dropped out of university in the UK at 18, and moved to San Francisco. There, he met Connor through Y Combinator’s co-founder matching platform. The two hit it off, and the concept for QUITTR was born. The third key player was Alex’s younger brother, Chris, who was 17 at the time. He handled design. So the founding team was: Alex on code (19), Connor on marketing (22), and Chris on design (17). Inject Zach’s and Blake’s influencer marketing and TikTok virality know-how into that team, and They get an app that scaled to $500,000 a month in a single year. One more detail worth noting: Alex built the initial version of QUITTR in roughly 10 days. In this era, if you have the idea and the willingness to move, a product can take shape in ten days. 🔥 A Product Hunt Flop → Reddit Guerrilla Posts → $3,000 Turning Into $37,000 If you assume QUITTR grew on influencer marketing from day one, you’d be wrong. What’s surprisingly little known is how gritty and unglamorous the early launch phase actually was. QUITTR first launched on Product Hunt — and completely flopped. So they moved to guerrilla posting on Reddit. They embedded themselves in r/NoFap, the main community for people quitting porn, and seeded posts and replies along the lines of “this app called QUITTR changed my life.” To be clear, astroturfing like this isn’t something to celebrate. Reddit explicitly bans stealth marketing, and they were in fact banned — after the tactic had generated about $2,500 in revenue. That said, it’s also true that plenty of founders do some version of this early on. Companies that are publicly traded today were, in their first months, growing on exactly this kind of thing. Among the app success stories I’ve covered in this Substack, Some people used Reddit marketing to get early traction. Now, they got banned. But the reaction they got through Reddit — plus $2,500 in revenue — gave them conviction that the app had legs. So they made an unexpected bet. With very little revenue in the bank, they put $3,000 into a promotion with a Christian influencer on Instagram. I couldn’t track down the specific account or video, but the result speaks for itself: that $3,000 turned into $37,000. From there, it became an infinite reinvestment loop. They funneled every dollar of revenue back into influencer marketing, and the whole thing snowballed. Margins ran around 90%, and essentially 100% of that profit was reinvested. Founder compensation was kept to the bare minimum, with everything else poured into growth. That level of discipline is what powered the bootstrapped rocket ride. Now let’s get into the real substance. 🎯 Best-in-Class: The Full Breakdown of a 25% Conversion Onboarding Let’s start with QUITTR’s famously long onboarding, which I touched on briefly in the earlier piece. The download-to-paid conversion rate is 25%. One in every four people who downloads the app pays. What’s even more absurd: the completion rate on that 12-step onboarding is 99%. Here’s the flow, laid out: - 12 pages of questions — “When did you start watching?” “Has it been escalating?” — questions that force you to confront yourself - Select the symptoms that apply to you — low motivation, diminished interest in real-life partners, and so on - Testimonials from existing users whose lives changed - The negative effects of porn, and the benefits of quitting - An App Store review prompt - A personalized plan, leading straight into the paywall The mechanism is elegant. By answering the questions, users effectively diagnose themselves as having a serious problem. By the time the paywall appears, they’ve already resolved to fix it — and then they’re asked to pay. And here’s the critical piece: they run a hard paywall with no free trial. They tried a free trial once. Usage rates collapsed, and they pulled it immediately. Because users pay upfront, they commit — “I am absolutely going to quit” — and both retention and ratings go up as a result. They’ve gone as far as to say that without the hard paywall, revenue would have been 1/25th of what it is. A 25%+ conversion rate with a hard paywall and no free trial is, frankly, insane. For what it’s worth: the Japanese-learning app I run also grew on a no-free-trial hard paywall, but I started getting a steady trickle of one-star reviews complaining that you can’t use it for free — so I’m currently testing a free trial. My gut says revenue is stronger with the pure hard paywall. But until you’ve accumulated a healthy base of reviews, adding a free trial may be the smarter move. Most of the other apps I’ve covered do offer free trials, by the way. My sense is that close to 80% of them do. On Onbo Hub — the product I’ve been building and running — you can inspect each app’s onboarding and check which ones use free trials. Worth a look: One more thing about QUITTR’s onboarding: they didn’t invent it from scratch. They screenshotted the onboarding flows of top apps doing $5–6 million a month, laid them all out side by side in Figma, extracted the common patterns, and adapted them for their own app. Onboarding is one of those areas where simply studying what already sells will completely change your results. I’d strongly recommend analyzing a wide range of app onboardings yourself. These days, QUITTR’s onboarding has become such an industry reference point that it gets copied constantly. 💃 The Influencer Deals, With the Actual Numbers Now the acquisition side. I mentioned last time that influencer marketing is QUITTR’s primary growth channel — this time, let’s get into the real numbers. Who they approach: Christian, fitness, and self-improvement influencers — the audiences most aligned with quitting porn. And they change the angle by niche. For religious audiences, it’s framed as a betrayal of faith. For fitness audiences, it’s lowered testosterone. For self-improvement audiences, it’s the performance gains that come from abstinence. Here’s the clever part — how they make first contact. They paid for a $50 consultation with a Christian creator to get a direct conversation. That single call ended up generating $1,000 a day in revenue. They repeated it with a $30 fitness consultation, with similar results. Cold-DMing influencers gets you ignored. So instead, they put their own money down on whatever consulting or advisory offer the creator was already selling. And then they pitched, live, on the call. It’s a phenomenal approach. Here are the deal terms they used: - Compensation is proposed as a mix: flat fee ($2,000–$10,000/month), performance bonuses ($500 per 500K views), and CPM ($2–3 per 1,000 views) - Payment is 20% upfront, remainder on delivery of results - View guarantees built in — e.g. “$1,500 for four videos with a guaranteed 500K combined views” - If views fall short, the fee is prorated or additional posts are required - The app must be shown on screen in the video - The contract includes rights to repurpose the created video as a paid ad That combination — view guarantees plus ad repurposing rights — is genuinely brilliant. Influencer campaigns are wildly hit-or-miss. View guarantees cap the downside, while the right to scale winning creative through paid ads maximizes the upside. The signature example: Jeremiah Jones, a creator with 2 million followers. On a $3 CPM plus bonus deal, his post hit 9.9 million views — a mega-viral moment. It drove $40,000 in a single day and $100,000 in 65 hours for QUITTR. Here’s the video: Today, QUITTR has systematized this to the point where they put $40,000 a month into Meta ads and pull a 3–4x return. Discover winning creative through influencers → scale it with paid ads. At this point, that’s the standard winning formula for consumer apps. Their operations are equally systematized: influencer prospects are tracked in a Google Sheet, and virtual assistants handle DM outreach. Even a team of two or three can run campaigns worth hundreds of thousands of dollars a month if you build the machine properly. 🧨 The Goal Is “Liquid Death”: Brand Strategy Beyond the App Now for what’s emerged since my earlier piece — the next phase. Once they passed $250,000 a month, their focus shifted to brand awareness. Their model? A water brand, of all things: Liquid Death. Liquid Death is the American monster brand that took plain water and, with punk skull-can branding, built it into a billion-dollar company. They’re also reportedly eyeing research collaborations with YouTubers and neuroscientists to build authority. From competing on app features to competing on brand and credibility. And given QUITTR’s market position, that makes sense — features can be copied endlessly, so the only durable moat left is the brand. They’re going after top-of-mind recall: quit porn = QUITTR. At this level, it’s completely outgrown anything resembling indie development. 👑 At 20, He Hands Over the Crown and Moves On Back to where we started. In August 2025, he posted this on X: I wish I could run multiple storylines at once. I want to be a footballer, an artist, an entrepreneur, and a politician simultaneously. Unfortunately each requires decades of focus. Not impossible — but definitely not short-term. He’d scaled QUITTR at breakneck speed and made serious money at 19. But his ambition ran much further than that. And then, in September 2025, he made the call to hand QUITTR to someone else: Handing off management of a thriving $500K-a-month app to return to a phase where he can exercise his own creativity. This move is quietly becoming a shared pattern among young consumer app founders. Zach Yadegari sold Cal AI. Many of these successful founders aren’t clinging to the app that hit — they’re converting the winner into an asset and moving on to the next big swing. For them, the app isn’t the goal. It’s the vehicle for a bigger ambition. Broke at 19, hit app at 19, handed over the company at 20, already onto the next concept. The pace is unreal. 🌃 The Next Habit App — and Why It Failed After handing off QUITTR, the next thing Alex launched was an app called Reset: The concept: rebuild your good habits in 60 days. Thematically close to QUITTR, with a nearly identical onboarding flow. It also looks conceptually quite close to Life Reset, the very successful app I’ve covered before. So how did it do? Alex hasn’t explicitly commented, so it’s hard to say definitively — but it appears to have failed. The app seems to have been abandoned and left unmaintained. Let’s dig into why. When Alex launched a new app, it briefly generated buzz on X. But it never had QUITTR’s explosive momentum, and the app’s reputation is poor — a 3.5 rating. The app itself offered nothing new. It was essentially an inferior version of Life Reset. Having built an explosive hit like QUITTR, there may have been a degree of overconfidence at work. My best guess is that he became a true believer in onboarding and cut corners on the app’s core functionality behind the paywall. Onboarding matters enormously for sales — but this was a stark demonstration that great onboarding is meaningless if the product quality doesn’t back it up. On top of that, habit apps are everywhere, and there was nothing novel in the feature set. Nor was it narrowed to a niche target the way QUITTR was. If you had to name the differentiator, it’s “change your habits in 60 days” — but as noted, that’s a second helping of Life Reset, and it reads as unremarkable to the exact users it needs to reach. Which means it would have been extremely hard to make it go viral in short-form video. The likely result: a temporary bump of traffic from Alex’s own X following, and a struggle to acquire users after that. The lesson from Alex’s case is clear — no matter how strong your onboarding and TikTok skills are, if you miss on the core app concept or cut corners on the features, it doesn’t work. Enormously instructive. ✈️ The Next Move: Launching Soar, an AI Flight Booking App But Alex isn’t the type to be derailed by that. The new venture he co-founded is Soar, an AI flight booking app. Soar handles everything from flight search to booking in one shot. Here’s how he describes it: - Book flights in one tap without re-entering your personal details every time - Find cheaper fares than Google Flights or Skyscanner - Get flight updates — delays, gate changes — pushed to iMessage - An AI agent checks you in automatically and sends your boarding pass via iMessage - If the fare drops, it automatically secures the airline refund or credit for you His claim: use it once and you’ll never go back. I tried it myself, and the UI/UX is excellent. And the early traction is already wild. He’s publicly reported processing $250,000 in booking volume within 30 days of launch: With Reset, I couldn’t shake the feeling that it was derivative — a second helping of an app that was already selling. Soar, though, I find genuinely exciting. Travel booking is a massive market, occupied by giants like Google Flights and Skyscanner. But none of them have meaningfully evolved their UX in years, and no one is yet delivering an AI-native experience. I travel constantly myself, and I feel the friction and awkwardness of travel tools every single day. There’s every reason for an app with a genuinely AI-native, optimized UX to exist. Maybe Alex is the one who replaces the legacy incumbents and cracks open that enormous market. Taking on huge markets with huge problems, armed with Gen Z instincts — that’s how Alex and the founders like him fight. Watching this makes me want to go after a giant market myself. Let’s not let them have all the fun. 📝 Wrapping Up So that’s the follow-up on Alex Slater and QUITTR, about a year and a half after my first piece. Pulling together what his success teaches: - Target the blank markets where the pain is deep but nobody has solved it. The themes that look strange are exactly where the opportunity lives - Build the product in 10 days. What you polish first is onboarding and marketing - But don’t neglect the product itself, or you’ll fail. Sharpen the core features - Study successful apps relentlessly and adopt what works. Don’t reinvent the wheel - Apps that require commitment don’t need a free trial. Use a hard paywall to make people buy in - Structure influencer deals as “view guarantee + ad repurposing rights” to cap downside and maximize upside - When something hits, systematize it, hand it to someone else, and go climb the next mountain At the time of my last piece, he was “a 19-year-old who hacked TikTok.” Eighteen months later, he’d evolved into “a 20-year-old who built a brand, built an organization, handed over management, and launched his next product.” Success in indie development isn’t the finish line. It might be where things actually begin. Personally, this was a fresh reminder of how much the post-hit phase matters — systematizing and letting go. Plenty of us can grind our way to building and growing something. Very few indie developers have a plan for what comes after. Time to get to work. So — that’s this week’s deep dive on Alex Slater. Thanks so much for reading! If anything caught your attention or you have questions, just hit reply to this email — I read everything. And if you mention my X account with your thoughts, it genuinely makes my day. I always respond. See you again. References https://flysoar.ai/ https://www.linkedin.com/in/alexsllater/ https://www.linkedin.com/company/flysoar/ https://boringcashcow.com/interview/interview-with-the-founder-of-quittr https://startupspells.com/p/porn-addiction-app-quittr-250k-mrr-4-months https://www.laweekly.com/from-broke-to-bold-how-alex-slater-built-quittr-into-a-1m-digital-wellness-powerhouse-at-19/ https://www.starterstory.com/quittr-breakdown
18:32

Everything You Need to Know to Master Claude's Fable 5

Claude Fable 5 is Anthropic's new flagship tier built to run alone for days on a million-token context, priced at $10 per million input tokens and $50 per million output. The guide's core advice is to spend Fable only on planning and verification, grind the middle with cheaper models, and use /goal with a measurable finish condition plus a turn cap so unattended runs don't loop forever. An over-asking model takes the cheapest path to 'done,' as its demo showed when it brute-forced Pokémon without reading the game. The catch: a mandatory 30-day data retention window with no opt-out, safety-classified requests rerouted to Opus 4.8, and a two-week US export-control suspension right after launch.

Notes

Claude Fable 5 — Master Guide (The AI Corner, 2026-07-28)

"An autonomous model gives you exactly what you specified, at whatever cost the specification allowed."

What it is. Fable 5 = first public model in Anthropic's "Mythos-class" tier, sits above Opus (not a replacement). Shares its underlying model with Claude Mythos 5, the restricted sibling for vetted cybersecurity/infrastructure partners — same brain, different guardrails. Specs: 1M-token context, up to 128,000 output tokens/request, built for multi-day autonomous work (planning in stages, delegating to sub-agents, self-checking).

Safeguards. Safety classifiers screen cybersecurity, biology, chemistry, model-distillation requests. Flagged requests are not refused — they're answered by Claude Opus 4.8 instead (in <5% of sessions) and not billed at Fable rates.

Pricing. Fable 5: $10/M input, $50/M output (~1 token ≈ ¾ English word). Opus 4.8 = exactly half. Sonnet 5: $2/$10 intro pricing through end of August (~5x cheaper than frontier). Stripe early-access ceiling: 50M-line codebase migration compressed from ~2 team-months → 1 day. Routing rule: expensive model at the edges (plan/verify), cheap models in the long middle — routine chat belongs on Sonnet.

Loops (Claude Code).

  • /goal — completion condition, works turn after turn until met. Use for bounded work (migration, refactor, backlog to zero).
  • /loop — reruns a prompt on schedule until cancelled. Use for watching/polling, not project completion.
  • Auto mode approves tool calls without clicking — required for unattended runs.
  • After each turn a small fast model reads the transcript and returns yes/no on your condition; it runs no commands, opens no files. So conditions must be machine-checkable: "npm test exits 0 and git status is clean," not "properly refactored."
  • A durable condition carries 4 parts: measurable end state, stated check that proves it, constraints (what must not change), cap on turns/time. No built-in token budget exists — skip the cap and the model loops pointlessly. Trick: ask Claude to interview you and draft its own condition.

Skills. A folder of instructions + reference material teaching a repeatable procedure. Three starts: extract patterns from a past chat where output was right; build from scratch from your repeatable list; build from examples (feed admired work, codify its structure/tone). Refine by correcting output and folding the fix back into the file. Anthropic says Fable updates its own skills and builds its own checks during work. Skills are plain files, vendor-portable ("Models come and go; your playbook only grows").

Vision. Extracts precise values from dense charts (graph → dataset), reads diagrams/tables nested in PDFs, rebuilds a working web app's source from interface screenshots. Self-checking trick: Fable checks its code output against the original mockup/design — image-as-acceptance-test, screenshot-as-evidence for the evaluator. Slay the Spire with persistent file-based notes: reached final act 3x more often than Opus 4.8.

Fine print (three constraints).

  • Data retention: mandatory 30-day retention on all traffic, no zero-retention option even for enterprises that previously had it. Won't be used for training — but don't feed it medical records/client files.
  • Refusal handling: classifiers may decline; you need a fallback (formal integration retrying on another model, or workflow that doesn't assume an answer arrives).
  • Access: June 12, 2026 — US export controls suspended the model for all users; access only returned July 1. Two weeks of downtime = argument for keeping skills/context in portable files.

Note: content includes a paid-style promotion for a "Claude-athon" (Outskill) offering a free 16-hour live curriculum, Sat–Sun 10 AM–7 PM EST.

Compressed habit: define done so a machine can check it; demand proof in the transcript; route each job to the cheapest model that clears the bar; read the terms before the first token streams.

Full text · 12,591 chars
Everything You Need to Know to Master Claude's Fable 5 Anthropic's most capable public model can work on its own for days. Here is how to point it at real work without burning your budget, your data, or your patience. Every Claude Fable 5 guide follows the same script: learn the loop commands, build some skills, admire the vision demos, multiply your output by 10. It is right about the features and wrong about where the difficulty lives. The proof sits inside the launch demo everyone shared and almost nobody read. On June 9, Anthropic posted a timelapse of Fable 5 beating Pokémon FireRed on raw screenshots alone. No maps, no helper tools, no reading the game’s internal state. The winning strategy is the funny part. The model pushed a single Charmander all the way into a level 78 Charizard, ignored type matchups entirely, and at one point wasted a Revive on a level 3 Pikachu mid-battle. None of that was an accident. The goal said “beat the game,” so the model found the ugliest path that satisfied it and committed for hours without asking anyone. That is the one lesson to carry through everything below. An autonomous model gives you exactly what you specified, at whatever cost the specification allowed. together with Outskill: Fable 5 can run for days on its own. The edge goes to whoever knows how to point it. This weekend, the world’s first Claude-athon condenses 800+ hours of research into a 16-hour live curriculum: ▫️ Master all three modes: Chat, Cowork, and Code ▫️ Set up Skills, Connectors, and plug-ins to automate your files, Notion, and desktop ▫️ Vibe-code apps plus 10+ tools that pair with Claude Free, Saturday & Sunday, 10 AM to 7 PM EST. This guide covers what the model actually is, what it costs, how to run work that finishes itself, how to teach it your procedures, what its vision is really for, and the fine print that most write-ups skip. Table of Contents 1. What Fable 5 Is, in Plain Language 2. The Price Sheet, and the Routing It Demands 3. Loops: How to Make Work Finish Itself 4. Skills: Teach the Procedure Once 5. Vision: The Most Underused Feature 6. The Fine Print, and the Habit That Ties It All Together 1. What Fable 5 Is, in Plain Language Before the tactics, thirty seconds of orientation, because the naming alone has confused half the internet. A new tier, not a new version Fable 5 is the first public model in a capability tier Anthropic calls Mythos-class, which sits above the Opus line rather than replacing it. It shares its underlying model with Claude Mythos 5, a restricted sibling reserved for vetted cybersecurity and infrastructure partners. The two names describe safeguard configurations, not different brains. Same model, different guardrails. The headline ability is endurance. Anthropic built the model to sustain multi-day autonomous work: planning in stages, delegating pieces to sub-agents, and checking its own output along the way. It ships with a 1 million token context window and up to 128,000 output tokens per request. What the safeguards mean for you Because the capabilities cut both ways, Fable ships with safety classifiers that screen requests touching cybersecurity, biology, chemistry, and model distillation. A flagged request is not simply refused. It gets answered by Claude Opus 4.8, the previous flagship, instead. Anthropic reports this happens in under 5% of sessions, and rerouted requests are not billed at Fable prices. For most people, most of the time, the safeguards stay invisible. What is very visible is the invoice, which is where this guide goes next. 2. The Price Sheet, and the Routing It Demands A guide that never mentions cost is an advertisement. The economics here are unusual enough to shape everything else you do with the model. Three price tags, one decision Fable 5 costs $10 per million input tokens and $50 per million output tokens. For orientation, a token is roughly three quarters of an English word, so a million output tokens is on the order of a long novel. Claude Opus 4.8 runs at exactly half those rates, and Sonnet 5 sits at $2 and $10 under introductory pricing through the end of August. That makes the frontier model roughly five times the price of the mid-tier. It earns the premium in one specific place. The price pays off on long-horizon autonomous work, and that the longer and more complex the task, the wider Fable’s lead grows. Stripe’s early-access result shows the ceiling, where a migration across a 50-million-line codebase, compressed from an estimated two team-months into a single day. A routing rule anyone can apply Here’s the rule to follow: expensive model at the edges, cheap models in the middle. Fable plans the work at the start and verifies the result at the end, the two moments where judgment is worth paying for. Cheaper models grind through the long middle. Everyday questions, quick drafts, and routine chat belong on Sonnet at a fifth of the cost, because putting them on Fable is paying limousine rates for a grocery run. Popular guides dress this split up as a barbell strategy and sell it as an insider trick. It is plain arithmetic wearing a costume, and now you know the arithmetic. The largest bills come from unattended runs, though, and controlling those starts with how the run gets instructed. 3. Loops: How to Make Work Finish Itself The endurance only becomes useful once you stop feeding the model one prompt at a time. Inside Claude Code, two commands handle that, and they solve different problems. Two commands, two different jobs /goal sets a completion condition and keeps the model working, turn after turn, until the condition holds. Use it for bounded work with a clear finish line such as a migration, a refactor, a research question answered in full, a backlog worked down to zero. /loop reruns a prompt on a schedule, every half hour or every morning, until you cancel it. Use it for watching and polling such as flagging the emails that genuinely need you, checking whether a deploy stayed healthy. It is the wrong tool for finishing a project. A third piece completes the setup. Auto mode approves the model’s tool calls without pausing for your click, and pairing it with /goal is what makes a truly unattended run possible. Without it, your overnight job stalls at 2 a.m. waiting for permission. Write conditions a machine can check Here is the mechanism most guides skip. After every turn, a small fast model reads the conversation and returns a yes or no on your condition. That evaluator sees only what has surfaced in the transcript. It runs no commands and opens no files of its own. So “the module is properly refactored” hands the judge a vibe, while “npm test exits 0 and git status is clean” hands it evidence, because the test output lands in the conversation where the evaluator can read it. A condition that survives a long run carries four parts: - A measurable end state - A stated check that proves it - Constraints naming what must not change along the way - A cap on turns or time. The cap is important because there is no built-in token budget. If you skip it, you’ll end up with a model looping for hours without progress, or an evaluator declaring victory over nothing concrete. One practical trick before launching is to ask Claude to interview you and draft the condition itself, with follow-up questions until “done” is specific and measurable. It writes better conditions for itself than most people manage on the first try. A condition governs a single run. The asset that compounds across every future run lives in a different file. 4. Skills: Teach the Procedure Once A skill is a folder of instructions and reference material that teaches Claude a repeatable procedure. Think of it as a recipe you write once and reuse forever. What goes in one, and how to start Anything you do more than twice is a candidate. Your weekly report format, a grant application structure, a trip-planning checklist, the house style your team writes in, the exact way you like meeting notes summarized. Three starting points work well: - Turn a past chat into a skill by asking Claude to extract your preferences and patterns from a conversation where the output came out right. - Build one from scratch by listing what you do repeatedly and describing the steps. - Or build one from examples by feeding in a pile of work you admire and asking the model to codify its structure and tone. Then refine it the way you would train a new colleague. That means by correcting the output, and folding the correction back into the file so the fix persists. Why skills matter more on this model Anthropic describes Fable as updating its own skills based on what it learns during work, and building its own checks and evaluations along the way. That turns skill-building from a documentation chore into a flywheel, because every completed run can sharpen the procedure the next run inherits. The quieter advantage is ownership. Skills are plain files in a folder you control, welded to no vendor, and the same library can brief whichever model wins next year’s benchmark race. Models come and go on a release cycle. Your accumulated playbook only grows. Procedures cover work that has a checklist. The next section covers verification for work that has to be seen. 5. Vision: The Most Underused Feature The launch coverage treated vision as demo material. In daily use it works more like a second pair of hands. What it can read for you Fable 5 extracts precise values from dense charts and scientific figures, which effectively turns a published graph back into a dataset. It understands diagrams and tables nested inside PDFs, the daily terrain of finance, legal, analytics, and research work. And it can rebuild a working web app’s source code from screenshots of the interface alone. So here’s the advice: Stop retyping numbers from a report chart and photograph it instead. Screenshot the confusing dashboard, the dense form, the diagram in the appendix, and ask what it means. Hand it a picture of your slide, your layout, or your spreadsheet and ask what a careful reviewer would flag. The self-checking trick The key is what the model does with vision during its own work. Anthropic says Fable checks its coding output against the original design or goal, using its eyes as a critic rather than a camera. That plugs directly into the loop mechanics we mentioned earlier. A mockup image becomes an acceptance test, and a screenshot of the finished result becomes evidence the evaluator can judge inside the transcript. Memory follows the same pattern. Given persistent file-based notes while playing the card game Slay the Spire, Fable reached the final act three times as often as Opus 4.8, which suggests the model genuinely uses the context you hand it rather than letting it pile up. All of this is impressive, and none of it matters if the deployment never clears the fine print. 6. The Fine Print, and the Habit That Ties It All Together Three constraints decide who can actually rely on this model, and none of them appear in the productivity threads. 1. Data retention Fable and Mythos carry mandatory 30-day retention on all traffic, with no zero-retention option, even for enterprises that held such agreements before. Anthropic frames the window as a misuse-detection measure and says the data will not train models, but the practical advice holds for everyone: think twice before feeding it medical records, client files, or anything you would not want held for a month. 2. Refusal handling Because the safety classifiers can decline a request, anything built on top of Fable needs a fallback plan for the moment it says no, whether that means a formal integration retrying on another model or simply your own workflow not assuming an answer always arrives. 3. Access On June 12, days after launch, US export controls forced Anthropic to suspend the model for all users, and access only returned on July 1. Two weeks of a flagship tool vanishing overnight is the strongest argument in this guide for keeping your skills and context in portable files rather than inside any one vendor’s walls. Now if you put all the pieces together, the whole guide compresses into one working habit. Define done so precisely that a machine can check it. Demand proof in the transcript, not confidence in the summary. Route each job to the cheapest model that clears the bar. Read the terms before the first token streams. The model supplies the endurance. You supply the definition of done. Every result you get from Fable 5, good or bad, will trace back to which of you did their job better.
10:30

The Great AI Disappointment

A newsletter author lays out a "Great AI Disappointment" thesis, arguing Americans feel let down by AI and the technology isn't living up to its promises. The post itself is mostly a request to share the essay rather than a summary of its arguments, so this entry is drawn from the title and framing. The author frames the current AI path as the wrong future for the next generation.

Full text · 588 chars
The Great AI Disappointment Americans’ feelings about AI are clear and the tech isn't living up to the promises or the hype. While the U.S. is damaging its reputation with the world and its own consumers... If you think this is a story worth sharing, please spread the word. We need to do better. This is not the right future for our descendents. If you can do me the favor of sharing this article with a friend, family member, coworker or acquaintance, it would mean the world to me. Or restacking it or sharing it elsewhere. In this document I provide my Great AI Disappointment thesis.