Dots got a computer. Argon still waits behind Fairwind.
OpenAI’s personal agents keep working when your laptop sleeps, Claude Code can rewrite itself, and Google’s million-token flagship is still not for sale.
Evening collect closed the day at 149 items. The morning story still holds — Gemini 4 Argon can aim at a million output tokens and you still cannot buy it — and the second wave made the agent layer louder. OpenAI’s DevDay centered on Dots. Anthropic let Claude Code load TypeScript mods that intercept the agent. Ai2 open-sourced a trillion-parameter MoE training stack. If you build or teach, today is about jobs that keep running, hooks that rewrite the harness, and research that says your system prompt is still mostly a costume.
Personal agents that own a job, not a chat
Ben’s Bites from DevDay puts Dots at the center. A Dot is a never-ending GPT-6 Astra chat that lives in the ChatGPT sidebar, spins Codex and Work threads on its own cloud computer, and can message you first. Pro plans get one Dot that does not burn the base usage pool; the chats it spawns still do. The same day brought GPT-6.1 Sol, a $500 Pro tier, Sign in with ChatGPT for third-party apps, Plugins, ChatGPT Space, and Ultrafast mode that runs Astra eight times faster for six times the usage.
That is the shift Opinion AI’s Dots guide names cleanly: stop asking for better answers and give the agent one recurring responsibility. The Meta Muse version of the same advice is Cloud AI’s job brief — a Friday Weekly Decision Brief with draft-only permissions, not “manage my life.”
Claude Code can rewrite the agent from inside
Anthropic’s Claude Code Mods are TypeScript middleware. A mod can observe, rewrite, or short-circuit tool calls, turns, and UI renders, and it ships through the normal /plugin flow. Built-ins like /diff and AGENTS.md are now mods you can disable or replace. Reference samples cover a token-usage meter, a destructive-command guard, and a diff replay theater. Mods are unsandboxed with full machine access. Team and Enterprise plans load a sec-default mod first so user extensions cannot override permission denies.
Pair that with the morning folder-not-chat Claude Code note. The product is no longer only a coding chat. It is a programmable harness.
Google’s flagship is still a gated demo
Latent Space’s Argon recap remains the cleanest map. Gemini 4 Argon claims first place on 13 of 19 published benches and a 1M-token output ceiling, up from 64K. Intro price is $2/$10 per million tokens (standard $4/$20), with 95% off cached input. Access starts with government users and Fairwind cyber defenders. There is no general-availability date. Artificial Analysis only hit 1M through Long Decode Continuation — pause and resume across calls. AI Supremacy still names the seven-month gap to a consumer answer for Muse, Instinct, Grok Bot, and Dots.
Open training, watermarks, and exact image layouts
Ai2’s Olmo-core 3 redesigns MoE training so experts stay resident under DDP. Growing the expert pool from 8 to 128 (still about 3.2B active) cut throughput by less than 5% while capacity rose from 4.6B to 47B, and the stack has been benchmarked past a trillion total parameters. AlphaSignal’s write-up adds the 2.7× gain on eight B300s and the failure modes Ai2 bothered to publish.
DeepMind’s SynthID Bio stamps AI-designed proteins so synthesis houses get a provenance clue without wrecking binders on VEGF-A, SARS-CoV-2 RBD, and PD-L1. Black Forest Labs’ FLUX 3 Image takes bounding boxes on a 0–1000 grid, keeps untouched pixels bit-identical across edits, and renders up to about 16.8 megapixels.
Your preamble is still a costume. Quiz the holes.
If you treat the system prompt as governance, read The System Prompt Illusion. Across 17 models, persona and formatting restructure middle layers. Safety barely moves them. Restrictive safety and “you have no restrictions” used nearly the same pathways (mean CKA 0.997). Matthew Green, quoted by Simon Willison, adds that boxed agents can still pass a worm through anything they share.
For teaching, Why Try AI’s gap-card prompt still beats a chatbot that writes a deck for the whole syllabus. Ten oral questions, a score-out-of-5 table, then cards only for the holes.
I am watching whether Dots become connector-dependent companions or just proactive chat, whether Claude Code mods stay a trusted-publisher world, and whether Argon’s coherence holds after a few hundred thousand tokens once Fairwind opens the door.
Also worth a click
- 😸 Trump renamed AI "Super Intelligence" — The NeuronFederal rename plus a voluntary lab pledge; FTC probe confirmed the next day. Meta’s experimental data-center tax credit cut $3.9B from its 2025 bill.
- Barclays scales Claude to upgrade operations and improve client experience — AnthropicClaude Code toward half of developers by end-2026; Colleague Knowledge Assistant already past a million searches for 16,000+ UK colleagues.
- An AI “mind-reading” tool can reconstruct what you’re looking at based on a brain scan — MIT Technology ReviewWeizmann decoder rebuilds viewed images from high-res fMRI; ~1 hour of calibration on a new person. Consent risk if it moves to EEG.
- Boston Dynamics Gives Atlas a 13-Motion Hand Built for AI Training — AlphaSignalFour fingers, 13 DoF, direct-drive and built for sim-to-real. No comparative success rates published.
- Upstage's Solar Mini 4 Triples Its Benchmark Score With 3B Active Parameters — AlphaSignal35B/3B MoE scores 24 on AA Intelligence Index; weak on Terminal-Bench; huge reasoning traces inflate task cost.
- The Missing Talent Tree — The Augmented MindEveryone has the same AI superpower, so it feels like none. Commit to a path.
- What AI Agents See When They Look at Your Pricing — The AI Corner7,675 Cloud 100 pricing prompts: own pricing page in 46% of answers, first in only 12%.
- ☕️ Pentagon taps Musk to shape future warfare — TechpressoProject Meridian with Luckey and Gingrich; report due 2027-01-28. Same cup: California No Robo Bosses Act, Reddit RSS off Nov 13.
- How smaller, distributed batteries could help the grid — MIT Technology ReviewConsumer-scale packs when utility batteries hit NIMBY. Not an AI product story.
- Fermion Research's Phonon-2 Beats Whisper at 5.21% WER in Just 164 MB — AlphaSignalOpen English ASR, CC-BY-4.0, within 0.25 pp of a 2.5 GB teacher.
New on arXiv
- TomasuLLM: Out-of-Order Speculative Execution for LLM Agents — arXivDraft tool calls in copy-on-write sandboxes; 1.31× on 100 SWE-bench Verified; 4,010 validations, zero false accepts.
- The System Prompt Illusion — arXivSafety preambles barely move internals; restrictive vs unrestricted CKA 0.997.
- EVOKE: Eliciting World Knowledge in Agents for Transferable Decision-Making — arXivRank the same actions under different goals so pretrained world knowledge has to show up.
- Evaluating Language Model Safety Across Long Adversarial Conversations — arXivFirst-turn safe 85–100%; by turn 101, 15–44%.
- Framing the Narrative: Ideological Mimicry in Large Language Models — arXivChanging terminology alone flipped stance in 16.9% of matched comparisons.
- ContextAdapt — arXiv95.6% mean appropriateness; correct professional justification 25.6–76.9%.
- Large Language Models are Approximate Survival Estimators — arXivZero-shot Survprompt cMAE near specialist forests; c-index still weak.
- Conformal Factuality Control for Multi-Hop RAG — arXiv95% target: 95.8–97.2% supported claims, but most answers go empty.
- TutlAit v1 — arXivMoroccan Tamazight speech, ~20.9 hours, accent labels, Arabic text.
- Halluscoring 2026 — arXivArabic hallucination shared task; winning detection AUC-ROC 0.772 / 0.767.
- Which Models Work Well Together? — arXivTeam selection by error decorrelation, not just individual quality.
- NinaXander — arXivBolt Pythia’s first 5 layers to RWKV; 84.4% less KV cache, quality drops off-domain.
- Doc2LoRA — arXivEach paper becomes a talkable LoRA; mixtures stay decodable.
- Evaluating Whether LLMs Can Reliably Connect the DOTs? — arXiv~9.2K narrative-infill bench; Gemma-2-2B beat models 10× larger.
- Automatic estimation of verbal fluency index in people with Motor Neuron Disease — arXivWhisperX + Silero; P-words R² 0.9.