Nothing matches those filters.

Lead

1

Article

1
13:01

How to use GPT-5.6

OpenAI's GPT-5.6 models are now available to everyone, and the new ChatGPT desktop app merges ChatGPT and Codex into one tool with a Work mode and a Sites plugin that hosts web pages with optional login. The model family comes in three tiers — Luna, Terra, and Sol — each with five thinking levels plus an Ultra mode that spawns subagents, and heavier settings burn through usage caps fast enough that even the writer nearly ran out. Elsewhere: Anthropic again extended Claude Fable 5's limits, Claude Code gained an in-app browser, Meta shipped its Muse Spark 1.1 multimodal coding model, Apple is suing OpenAI over alleged trade-secret theft, and Notion launched Ship OS to run product development from feedback to merged PR.

Notes
How to use GPT-5.6 (Ben's Bites, 2026-07-14)

OpenAI released GPT-5.6 models to all users with several product changes.

Product changes

  • ChatGPT's macOS app and the Codex app merged into one combined app; updating the Codex app delivers it. New mode called ChatGPT Work.
  • Codex and Work look similar; details fine-tuned for coding-related vs non-coding work.
  • New plugin "ChatGPT Sites": build hosted websites with optional "Login with ChatGPT" feature. Author found it useful for sharing but disabled it after it captured everything he asked on ChatGPT Sites.

GPT-5.6 model series — three models, five thinking levels, one new mode

  • Luna, Terra, Sol.
  • Each ships with 5 thinking levels: light, medium, high, xhigh, max, plus Ultra mode — "basically allows these models to go haywire with subagents."

Usage/caveats

  • Higher thinking levels burn usage limits much faster. Author: "don't rip ultra; first time I've ever been nearly out of usage in Codex."
  • Author's defaults: Sol medium for most building/creativity, background agents for harder tasks, Luna xhigh for day-to-day productivity.
  • Weekend bug: OpenAI reset usage 4–5 times during the app-merge fixes and temporarily removed the 5h limit — so you can exhaust weekly limits in one go.

Per-model observations

  • Sol: good at UI, better with references; at Max thinking has strong writing and is fun to chat with.
  • Terra: "feels like a replacement for 5.5 with minor improvements in UI and writing skills"; more steerable, so skills would be useful with it.
  • Luna: "a bit of a mini model smell" — sometimes misses intent in ambiguous prompts, but doesn't fail clearly defined tasks.

Computer Use

  • Codex-app models are strong at Computer Use (self-driving cursor: opening apps, clicking buttons, using them by looking at the screen). Recommended trial: Sol medium/high on a small task.

Headlines

  • Anthropic extended Fable 5 on paid plans and Claude Code's 50%-higher weekly limits through July 19.
  • Claude Code gained an in-app browser (open/click through docs, designs, apps); Artifacts now support public links and multiplayer editing, incl. pages made via Claude Tag in Slack.
  • Meta released Muse Spark 1.1: multimodal for coding, computer use, agent tasks, 1M token context; now offered via API at a comparatively cheaper price point.
  • Apple suing OpenAI and two former employees over alleged trade-secret theft for its upcoming AI hardware; Apple claims candidates were told to share confidential info and bring physical parts to interviews. OpenAI says it has no interest in others' secrets, is reviewing the case.
  • Notion Ship OS: runs product development from customer feedback to merged PR in one workspace; agents triage/route/summarise while people make judgment calls.

Feed highlights

  • Tend: open-source ChatGPT Work loop for inboxes/hiring/support that learns from approvals.
  • v0 Design Systems 2.0: import components, Figma, Storybook so generated apps use your real system.
  • Inference AutoTune: claims to distil a frontier model into your own 1–30B specialist in ~2 hours for under $250.
  • Cloud Run sandboxes: isolated computers for agents (Google Cloud).
  • Control ideas, not code: agents good at local code; humans own design, testing, direction.
  • Satya Nadella's "reverse information paradox": firms should own the corrections/feedback that make AI useful.
Full text · 5,347 chars
How to use GPT-5.6 new desktop app and hosted sites in ChatGPT Hey folks, OpenAI finally released the GPT-5.6 models to everyone, with many product changes. ChatGPT’s macOS app and the Codex app are now merged, with a new mode called ChatGPT Work. You can update your Codex app and get the new combined app. Codex and Work look similar, with the details fine-tuned for coding-related vs non-coding-related work. There’s a new plugin called “ChatGPT Sites” that lets you build hosted websites with an optional “Login with ChatGPT” feature. It’s nice for sharing but kept annoying me by making everything that I ask on ChatGPT Sites, so I turned it off. GPT-5.6 series has three models: Luna, Terra and Sol. Each of them comes with 5 thinking levels - light, medium, high, xhigh, max, and a new mode: Ultra mode. Ultra mode basically allows these models to go haywire with subagents. Now… these models eat up your usage limits much faster when used at higher thinking levels. So don’t rip ultra; first time I’ve ever been nearly out of usage in Codex. I’m defaulting to sol medium for most building/creativity, background agents for harder tasks and day-to-day productivity using luna xhigh. Some general patterns I’ve noticed: - Sol is pretty good at UI, but it’s even better once you give it some references. Sol at Max thinking has really good writing, and it’s fun to chat with. - Terra feels like a replacement for 5.5 with minor improvements in UI and writing skills. It also feels more steerable, so skills would be useful with it. - Luna has a bit of a mini model smell, like sometimes it doesn’t “get” what you meant in ambiguous prompts. But it doesn’t fail at tasks you clearly define. Over the weekend, OpenAI reset usage 4-5 times while they were fixing bugs introduced during the app merge, and they have temporarily removed the 5h limit on your usage. Take note of that because you might end up using up your weekly limits in one go. These models in the Codex app are really good at Computer Use, i.e., self-driving your cursor so it can open apps, click buttons, and use them by looking at the screen. You should definitely try this with Sol medium/high for some small task to see it in action. Further reading: Ben’s Bites is brought to you by Voices Building or training voice models? Stop scraping. Voices has professionally directed, fully consented voice data ready now—part of 100,000+ hours of custom, premium data. Character performances, 43 emotional states, 8 languages. Request a sample today. Headlines - Anthropic again extended Fable 5 on paid plans and Claude Code’s 50% higher weekly limits through July 19. This feels like a joke at this point. - Claude Code got an in-app browser - let Claude open and click through docs, designs or your app. Claude Code Artifacts now support public links and multiplayer editing, including pages made through Claude Tag in Slack. - Meta released Muse Spark 1.1 - a multimodal model for coding, computer use and agent tasks with a 1M token context window. Meta is also starting to offer its models in the API now, and Muse Spark 1.1 is available at a relatively cheaper price point vs its peers. - Apple is suing OpenAI and two former employees over alleged trade-secret theft for its upcoming AI hardware. Apple says candidates were told to share confidential information and bring physical parts to interviews; OpenAI says it has no interest in others’ secrets and is reviewing the case. - Notion Ship OS runs product development from customer feedback to a merged PR in one workspace. Agents triage, route, and summarise work while people make the judgment calls, using Notion’s existing docs and databases. My feed - The future worth building is human - New mission statement from Thinking Machines making the case for many AIs shaped by local knowledge. - Tend - open-source ChatGPT Work loop for inboxes, hiring or support that learns from your approvals. - v0 Design Systems 2.0 - import components, Figma and Storybook so generated apps use your real system. - NameThat - visual dictionary for UI patterns you can see but do not know how to describe to an agent. - OptionAFK - Private and fast dictation + transcription toolkit that your agents can use. - shadcn/typeset - one editable CSS file for styling Markdown in blogs, docs and streaming chat. - Inference AutoTune - claims to distil a frontier model into your own 1-30B specialist in ~2 hours for under $250. - Cloud Run sandboxes - isolated computers for agents from Google Cloud. - Hot takes on AI memory - why retrieval and longer context will not solve company-wide memory. - A software factory that works - scheduled agents create Linear issues; a label sends Claude to fix them. (What is a software factory?) - Control ideas, not code - agents are good at local code; humans should own the design, testing and direction. - AI’s biggest winners may have the lowest margins - agents can cut coordination costs in physical businesses. - The reverse information paradox - Satya Nadella thinks firms should own the corrections and feedback that make AI useful. Afters - Find me on X, Linkedin, or YouTube - Read about me and Ben’s Bites - 📷 thumbnail via @keshavatearth * sponsors who make this newsletter possible :) Wanna partner with us for the next quarter? Email us at shanice@bensbites.com or k@bensbites.com

Newsletter

1
12:49

OpenAI Is Coming for Hermes One Codex Update at a Time

An agent enthusiast finds himself spending more time in OpenAI's Codex than in his own self-hosted Hermes agent, because one $20 ChatGPT subscription now bundles the model, remote access, browser control, and plugins. The merged ChatGPT desktop app puts Chat, Work, and Codex under one roof, Codex Remote went generally available in late June, and OpenAI temporarily dropped the five-hour usage cap on July 12 as Codex passed six million active users. He still keeps Hermes for scheduled, unattended, provider-agnostic work like his morning briefing, and sorts every workflow into three buckets: bounded interactive tasks for Codex, autonomous persistent tasks for Hermes, and rare unstable work left manual.

Notes

OpenAI Is Coming for Hermes One Codex Update at a Time — research notes

Author: a self-described "All Agents Considered" substack writer who runs the open-source agent Hermes (on a self-hosted VPS) alongside OpenAI's Codex inside a $20/month ChatGPT Plus subscription. Published 2026-07-14. Thesis: OpenAI is folding the self-run agent layer into one subscription, one Codex update at a time.

The July sprint (concrete releases)
  • GPT-5.6 released July 9, 2026 across ChatGPT, Codex, and the API, in three variants: Sol (frontier), Terra (balanced), Luna (efficient).
  • GPT-5.6 landed in Codex alongside faster computer use; author argues the pairing matters more than benchmarks.
  • ChatGPT desktop app merges Chat, Work, Codex in one window: Markdown/code editing with inline annotations and selected-text revision; GitHub PRs in the sidebar; related repos share one project; plugins managed in Settings.
  • Developer Mode: with explicit approval, Codex gets controlled access to Chrome devtools — console, network activity, page structure, styles, performance data.
  • Codex Remote reached GA on June 25; tasks start/continue from the ChatGPT mobile app while running on a paired Mac/Windows machine.
  • DigitalOcean Droplet Workspace plugin: provisions a remote machine, configures SSH, connects it as a Codex workspace.
  • Plugin Directory replaced the App Directory; plugins can package skills, apps, templates; improved plugin loading and remote plugin catalogs.
Usage-limit changes (with the stated caveat)
  • June 11: eligible Plus/Pro users got reset banking incl. one free launch reset.
  • June 11–24: referral promo awarded resets; earned resets expire after 30 days.
  • July 12: Codex lead Tibo Sottiaux announced temporary removal of the five-hour restriction for Plus/Business/Pro; weekly limits remained; the pricing page still documents the five-hour structure; a usage reset followed Codex hitting 6M active users.
"Temporary is the word doing the work here. OpenAI can bring the restriction back or change the allowance again."
Hermes ↔ Codex via OAuth
  • Hermes officially supports an OpenAI Codex provider via ChatGPT OAuth (hermes model → select OpenAI Codex → device-code login); no separate OpenAI API key; can import existing Codex CLI credentials.
  • Limited to subscription-exposed Codex models, not the full OpenAI API.
  • Author also runs GLM-5.2 in Hermes through a separate provider route; rule: the workflow should survive the model swap.
  • Contrasts with the author's earlier $30 Hermes stack breakdown; notes Nous Portal now offers a bundled route, making several API keys optional.
The decision framework

Three buckets for every workflow:

  • Codex — bounded, interactive work done in the author's presence: one-off research, repository work, browser testing, documents, short analyses.
  • Hermes — unattended/scheduled runs, external triggers, private services, cross-session persistence, provider-switch survival.
  • Manual — rare/unstable work, automated only after three real runs without changing inputs or judgment points.

Six pre-build checks:

  • Needs to run without you?
  • Needs a schedule/external trigger?
  • Touches files/services under your control?
  • Would losing one vendor break it?
  • Need to switch models/providers?
  • Does control repay the maintenance cost?

Test results: morning research passes 5 of 6 (→ Hermes); single-article research passes almost none (→ Codex). Author's method: write three recurring tasks, mark each C/H/M, pause one duplicated layer, see if the workflow still finishes.

What moved vs. what stayed
  • Moved to Codex: one-off article research (researches against the local brief in Obsidian, writes back to the vault — this article is given as the concrete example), repo/browser work, one-off office files/assets. Halfway: recurring research stays in Hermes, but Codex takes over turning observations into finished assets.
  • Stays in Hermes: the morning briefing workflow (scheduled, filters chosen sources, saves briefing, delivers via a gateway the author controls — on the VPS, "even when my laptop is closed"); Telegram access; persistent file chains with review notes; provider choice (GLM-5.2 ↔ Codex models via OAuth).
Stated limitations / position
  • Author explicitly does not allege copying: "an expanding product overlap rather than alleging that OpenAI copied a specific feature or set out to kill an open-source agent."
  • The overlap risk cuts both ways: "July's friendlier terms prove both sides of that bargain because the company can remove friction or restore it quickly"; points to the author's prior vendor-lock-in article.
  • One-directional counterpoint: "Ownership has to earn its maintenance now."
Full text · 15,120 chars
OpenAI Is Coming for Hermes One Codex Update at a Time OpenAI keeps folding more of my agent stack into a $20 subscription. I still run Hermes as I want complete control over my workflows and the freedom to choose any model. Last week I caught myself spending more time inside Codex than using Hermes, and I couldn’t pinpoint when the shift happened. GPT-5.6 had just landed, and what used to be a coding tool inside my ChatGPT subscription had become something closer to a full agent workspace. What made it strange was that the same $20 bill could also feed a bunch of models into my Hermes agent. One subscription covering two competing stacks, with OpenAI shipping features every week that made one of them feel redundant. Browser control, remote access, plugins, banked resets, and looser usage limits all arrived over a single month, and every one of those used to be a separate purchase or, in some cases, an implementation I had to build myself. Now I still use Hermes every day, but I just started spending more of my week inside Codex because OpenAI keeps adding features that I use all the time. Today I’ll share why I believe Codex is catching up with Hermes and why I believe the $20 CHatGPT subscription is really a great deal. In this article: - The recent Codex releases that shifted how I compare the two - How my ChatGPT subscription supplies Codex models inside Hermes through OAuth - Which workflows moved to Codex and which ones I refuse to move - A three-way test I run before building another agent workflow The July Sprint That Reshaped My Stack GPT-5.6’s July 9 release pushed this piece to the front of my queue. OpenAI added the GPT-5.6 family across ChatGPT, Codex, and its API, with three versions called Sol, Terra, and Luna. OpenAI positions Sol as the frontier model, Terra as the balanced option, and Luna as the efficient one (OpenAI’s GPT-5.6 announcement). Placement matters more than early benchmarks. GPT-5.6 arrived inside Codex alongside faster computer use. OpenAI improved the model and its interface for acting on a computer at the same time. That pairing matters to me more than another leaderboard. An agent becomes useful when its brain and working environment stop feeling like separate purchases. OpenAI expanded the desktop surface too. Its new ChatGPT desktop app puts Chat, Work, and Codex under one roof. Editing happens directly in Markdown and code, with inline annotations and selected-text revision. GitHub pull requests sit in the sidebar, related repositories share one project, and plugins are managed in Settings (Codex changelog). Those sound like small interface updates when you read them one at a time. Together, they remove handoffs. Review now stays inside the task. I can inspect the diff where the work happened and return feedback without moving changes through a separate editor. Codex can hold several related repos in one project instead of treating every codebase like a separate room. Browser work became more serious too. With explicit approval, Developer Mode gives Codex controlled access to Chrome’s developer tools, including the console, network activity, page structure, styles, and performance data (Codex browser documentation). That turns the browser from a page the agent can click into an environment it can inspect. For anyone building or testing a site, this removes another reason to wire up a separate browser setup for bounded work. Codex Reaching Beyond Desktop Codex also gained more reach. Codex Remote reached general availability on June 25. A task can start or continue from the ChatGPT mobile app while the work runs on a paired Mac or Windows computer. OpenAI also released a DigitalOcean Droplet Workspace plugin that provisions a remote machine, configures SSH, and connects it as a Codex workspace (ChatGPT release notes). That closes part of the gap I used to describe as simple. Hermes lived on my VPS and stayed available from Telegram. Codex lived on the computer in front of me. Remote access and remote workspaces make that boundary less clean. Plugins are moving the same direction. OpenAI replaced the old App Directory with a Plugin Directory, and its plugins can package skills, apps, and templates. Codex also improved plugin loading and made remote plugin catalogs easier to use. A workflow that once forced me to choose between an MCP, CLI, or custom tool can increasingly arrive as one installable bundle. Then OpenAI softened the usage wall. On June 11, eligible Plus and Pro users received reset banking, including one free launch reset. A separate referral promotion ran from June 11 through June 24 and awarded resets after invited users sent their first Codex message. Earned resets expire after 30 days, so I wouldn’t treat them as permanent monthly allowance (Codex pricing). On July 12, Codex lead Tibo Sottiaux said OpenAI was temporarily removing the five-hour restriction for Plus, Business, and Pro users. Weekly limits remained, and OpenAI’s standing pricing page still documented the five-hour structure. His announcement also included a usage reset after Codex reached six million active users (Sottiaux’s announcement, Codex pricing). Temporary is the word doing the work here. OpenAI can bring the restriction back or change the allowance again. Still, the immediate price calculation changed. My $20 subscription stretches further during heavy weeks, with banked resets and fewer interruptions while the temporary change lasts. This is bigger than GPT-5.6 or a reset button. OpenAI shipped the model and workspace upgrades alongside remote control and friendlier limits. That’s what turns Codex into an agent workspace rather than a coding interface. Work that previously started with choosing four services now starts with opening one app. One Subscription Feeding Two Agents Hermes officially supports an OpenAI Codex provider authenticated through ChatGPT OAuth. I can run hermes model, choose OpenAI Codex, complete the device-code login, and use the Codex models available through my ChatGPT subscription. Hermes can also import existing Codex CLI credentials when present (Hermes provider documentation). I don’t need a separate OpenAI API key for that route. It’s limited to Codex models exposed through the subscription rather than every model sold through the OpenAI API. That boundary still changes the economics. My ChatGPT subscription now pays for two different layers. It pays for the serviced Codex workspace, and it supplies one model route inside the Hermes runtime I control. Codex and Hermes are competing for my workflows while sharing part of the same model bill. I also use GLM-5.2 inside Hermes, and this is one of the main reasons Hermes is staying online. My workflow remains in place while I change the model serving it. Hermes’s model catalog includes GLM-5.2 through supported provider routes, while its CLI lets me switch among models I’ve configured (Hermes CLI documentation). I’ve tested model choice inside a real Hermes workday, and the result kept pointing back to the same rule: the workflow should survive the model swap. GLM-5.2 uses a separate provider route, leaving the ChatGPT subscription as another option beside it. That distinction was missing from my earlier $30 Hermes stack breakdown. Hermes still has visible hosting and provider costs. Nous Portal now offers a more bundled route, so several separate API keys are optional. Maintenance time remains part of the bill either way. Codex hides more of those decisions inside a single price. I covered the wider plan economics in my comparison of six AI subscriptions, but the practical difference is simple. Codex supplies a serviced workshop. Hermes gives me the keys to one I own. Serviced workshops keep adding equipment. My own workshop lets me decide which engine runs it. The Workflows Codex Took From Hermes Research moved first. My Hermes research setup used a custom search route and a saved-file workflow that sorted everything against my brand filter. I still use it for recurring research, but one-off article research is usually faster in Codex now. Codex researches against my local brief and writes the result back into the same Obsidian workspace. Research and writing live in one task, giving most bounded questions one search route. This article is a concrete example. Its outline, research brief, old posts, and brand files all live in my vault. Codex can research the release claims against those files and write the article into the correct folder without me carrying context between tools. Repository and browser work followed. Codex already had an advantage here because code is its home territory. Inline review and Browser Developer Mode widen that advantage. I can inspect a site and its console, edit the code, and review the diff in one working session. One-off files became obvious too. When I need an office file or visual asset once, building a permanent Hermes workflow around it makes little sense. Codex has the file tools and task context ready. I verify the output and leave without creating another piece of agent infrastructure. My biggest change is how rarely I prepare infrastructure before starting. Codex removes the provider and output-routing decisions for bounded work. Some work moved only halfway. Hermes still collects recurring research and saves the briefing on schedule. Codex often takes over when I turn one of those observations into a finished asset. That handoff keeps both systems useful without maintaining the same production setup twice. The Workflows Hermes Still Defends My morning workflow stays. It runs before I sit down, applies my filters to the sources I chose, saves a briefing, and delivers it through the gateway I control. I documented the full version in How My Hermes Agent Plans My Morning Before I Have My Coffee. Codex now supports scheduled automations and remote work of its own. My reason for keeping this workflow in Hermes survives those additions because the complete runtime already lives on my VPS. Moving it would trade a working system I control for a product surface whose limits and behavior OpenAI controls. I keep that runtime dependable with my Hermes maintenance routine, which checks the layers a bundled product manages for me. Telegram stays too. Hermes remains available where I already communicate, even when my laptop is closed. It can call my own scripts and services, work with the files on my server, and keep the result in a location another workflow already knows how to find. Provider choice matters most here. I use GLM-5.2 in Hermes and can switch the same workflow to a Codex model through ChatGPT OAuth when I want to. If one provider changes its terms or performs poorly on a task, I’ve got another route. That freedom has a maintenance cost, but here it buys continuity instead of technical decoration. Persistent files finish the case. My workflows leave briefs and outputs in folders I own, with review notes beside them. Codex can work inside those folders, but Hermes is the runtime connecting them over time. One task hands a file to the next without depending on a single product account to remember the whole chain. Sorting Every Workflow Into Three Buckets I stopped choosing one agent for everything. I sort each workflow into one of three groups. Bounded, interactive work where I’m present to start it and approve the result goes to Codex. Research for one article, repository work, browser testing, a document, or a short analysis usually lands there. Work that needs to run without me, start from a schedule or outside trigger, reach private services, persist across sessions, or survive a provider switch stays in Hermes. Morning research, Telegram access, and multi-step file workflows land here. Rare or unstable work stays manual for three runs. I automate it only after the inputs, judgment points, and output stop changing. Six Checks Before You Build I use six questions before deciding where a workflow belongs. Does the workflow need to run without me? Does it need a schedule or external trigger? Does it touch files or services I want under my control? Would losing one vendor break the workflow? Do I need to switch models or providers? Does the control repay the maintenance cost? Results map to Codex for bounded work, Hermes when continuity or control shows up several times, and three manual runs when the pattern remains unclear. Two Workloads Through the Same Test My morning research passes five of the six ownership checks. It runs unattended on a schedule and feeds later workflows from my files, so provider switching changes the result. Hermes earns its place there. Researching a single article passes almost none. I’m present, the task is bounded, and the output goes into a draft I’ll review. Codex wins. This test takes less time than configuring one API, and it has stopped me from maintaining the same capability twice. Run Your Own Audit Write down three recurring AI tasks and mark each one C, H, or M. Pause one duplicated layer this week, then check whether the workflow still finishes cleanly. The Expanding Overlap Open Source Needs to Answer When I say OpenAI is coming for Hermes, I’m describing an expanding product overlap rather than alleging that OpenAI copied a specific feature or set out to kill an open-source agent. Codex is swallowing the layer of self-run agent work where convenience was the main payoff. Every new piece OpenAI bundles into the working environment makes the ownership case work harder. Open source needs to protect a result I’d lose inside the rented product, because more switches alone no longer win. For me, those results are provider choice, persistent files, and an always-on runtime I control. Codex carries the opposite risk. OpenAI controls the product and its limits. July’s friendlier terms prove both sides of that bargain because the company can remove friction or restore it quickly. I wrote the longer version of that risk in my vendor-lock-in article. Codex’s current sprint has made me more selective about where independence pays while leaving the underlying risk intact. Codex is shrinking the part of my stack worth maintaining. Hermes protects the workflows I refuse to rent. Ownership has to earn its maintenance now. Where the Line Sits Today Codex handles bounded research, article production, repository work, browser inspection, and one-off files. These jobs start with me, end with a reviewed output, and benefit from the serviced bundle. Hermes handles scheduled research, Telegram, persistent file chains, private services, and workflows where I want GLM-5.2, a Codex model, or another provider without rebuilding the system. Manual covers everything else until it survives three real runs. Which workflow has a subscription recently pulled out of your self-run stack? Codex is the serviced workshop, and OpenAI keeps delivering new equipment. Hermes is the workshop I own, where I control the keys, files, and engines. I build fewer tools myself and reserve ownership for the workflows where it protects the result. But that line keeps moving one Codex update at a time.