Nothing matches those filters.

Article

15
00:22

datasette-explain 0.2.2

Full text · 270 chars
20th September 2026 Recent articles - Generating running routes with GPT-6 Astra and ChatGPT Work - 12th September 2026 - OpenAI agents attacked RubyGems back in May - 12th September 2026 - Some thoughts on the Navier–Stokes Millennium Prize Problem - 8th September 2026
07:04

Andrew Levin's Open-Source macOS Agent Makes Decisions 155x Cheaper Than Claude Opus

Full text · 3,026 chars
- typesafe-computer-use is a macOS agent that runs steps for about $0.0002 each. - It never sends screenshots to a frontier model, using OCR plus a small classifier instead. - Reported as 155x cheaper and 14 to 40x faster than Claude Opus 5 per decision. - Powered by TypeSafe, a decision model returning calibrated Choice probabilities in milliseconds. - Writer LLM (Claude Haiku default) only handles free text like URLs and form fills. - MIT-licensed, macOS 14+, Python 3.12+, with replayable run folders for debugging stalls. A macOS agent built around $0.0002 decisions Many frontier-model desktop agents send a fresh screenshot to a large model for every action, then wait several seconds for a plan. Andrew Levin’s MIT-licensed open-source project, typesafe-computer-use, extracts screen state locally and asks a smaller decision model to select from bounded actions. Levin says he built the initial version in under 30 minutes. The agent works toward a plain-English goal on macOS, reporting a cost of roughly $0.0002 per decision. Screenshot pixels remain on the machine. Model requests contain OCR text, accessibility metadata, the active application, the browser URL, and other structured state. Free-form text goes to a separate writer model only when a field or URL requires it. - Platform: macOS 14 or later - License: MIT - Decision service: TypeSafe API - Default writer: Claude Haiku - Reported decision latency: 0.13–0.38 seconds One screenshot, a 155× cost gap The repository compares its TypeSafe configuration, labeled jev, with Claude Opus 5 on the same screenshot and goal. The Opus baseline receives a bare screenshot. | Repository-reported comparison | | | | |---|---|---|---| | Metric | TypeSafe ( jev ) | Claude Opus 5 | Reported difference | |---|---|---|---| | Cost per decision | $0.0002 | $0.032 | 155× cheaper | | Cost per 12-step task | $0.003 | $0.40–$0.90 | 130×–300× cheaper | | Model latency | 0.13–0.38 seconds | 5.2 seconds | 14×–40× faster | | Step time with capture and OCR | About 1.5 seconds | About 5.5 seconds | 3.7× faster | These figures come from the repository and have not been independently verified. They describe a narrow point comparison, with no task-completion rates, broad application suite, or recovery benchmark. Writer-model calls also add their own latency and cost. A production evaluation would need to measure successful completion, retries, writer usage, and failures across representative workflows. TypeSafe, the project’s core dependency, accepts a Choice containing as many as 255 options and returns a probability distribution with a calibrated confidence score. According to the README, responses arrive within a few hundred milliseconds and incur no output-token charge. The open-source agent therefore depends on hosted inference rather than a locally runnable decision model. This story is for Pro members You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.
11:04

The Sequence Radar - Issue 936: Last Week in AI: Gemini Talks, Astra Practices Law, Figure Folds Laundry, and Crusoe Powers It All

Full text · 10,890 chars
Next Week in The Sequence: - We have the second installment of our series about recursive self improvement. Can’t miss that one. - We deep dive into the new Gemini models, Stanford’s University Paper2Agent research and Astra for Law. - In the opinion section, we debate the opportunities and challenges of text diffusion models from first principles. - We have another installment of our robotics series. Subscribe and don’t miss out: 📝 Editorial: Last Week in AI: Gemini Talks, Astra Practices Law, Figure Folds Laundry, and Crusoe Powers It All Building useful AI is starting to resemble building an automobile. The engine matters enormously, but so do the steering, transmission, chassis, and fuel supply. This week offered a remarkably complete tour of that machine: Google upgraded conversational control, OpenAI specialized intelligence for law, Figure tested physical generalization, and Crusoe put a $30.9 billion valuation on the infrastructure underneath it all. Google’s Gemini 3.8 Live and Live Extended Thinking tackle a deceptively hard problem: keeping a conversation alive while useful work happens. The models combine visual context with background tool calls; Extended Thinking can reason and speak concurrently. Imagine discussing travel plans while an assistant checks availability and handles your interruptions. The engineering challenge is coordinating dialogue, computation, and actions without making the user wait through awkward silence. Voice becomes a practical control surface for agents, with latency joining accuracy as a central design constraint. Google’s announcement OpenAI’s Astra for Law addresses another source of friction: professional context. It combines GPT-6 Astra with legal instructions, specialized search, and tools for legal workflows. On 200 questions from a private Legal Research Bench validation set, OpenAI reports 54% correctness versus 38.7% for Astra using ordinary web search. That remains a substantial distance from dependable autonomy. Still, the result illustrates an important principle: equipping a capable model with the right information environment can materially improve its performance. A brilliant associate still needs access to the right case law. OpenAI’s announcement Figure’s Helix 2.5 takes that argument into unfamiliar living rooms. The company tested tidying, towel folding, and bed making across 30 unseen homes. Pretraining on its Index human-behavior dataset raised complete-task success from 9% to 56%, with architecture and task-specific training held constant. Here, “zero-shot” means new homes and objects; the behaviors were learned using robot data collected elsewhere. The exciting result is the transfer: broad human experience made the same robot training much more useful. The remaining 44% failure rate is equally instructive. Your laundry has very little patience for a promising scaling curve. Figure’s report Then comes the electricity bill. Crusoe announced the initial closing of an anticipated $3.9 billion Series F at a $30.9 billion post-money valuation. Its platform connects energy, data centers, and AI cloud services, with expansion spanning large campuses and modular Spark facilities. The investment thesis is straightforward: increasingly capable agents create demand for increasingly available computation. Delivering that computation requires securing power, constructing facilities, and operating hardware efficiently. Every conversational flourish and robotic recovery eventually becomes somebody’s infrastructure workload. Crusoe’s announcement These developments suggest that AI’s next chapter will reward the teams that connect intelligence to its operating environment. Conversation requires timing. Legal work requires authoritative context. Robotics requires transfer across messy physical settings. All three require computation that someone can actually deliver. The opportunities extend across that entire chain, and so do the failure modes. As models acquire more responsibility, progress will increasingly be measured by completed work under real constraints. The fascinating part is how much invention remains between an impressive model and a system we can comfortably depend on. 🔎 AI Research AI Lab: Carnegie Mellon University Summary: DDO freezes base weights and edits a few low-impact MLP neurons to inject a harmful-selective, refusal-orthogonal decoy so contrastive abliteration (RFA) removes the decoy instead of real refusal. Across six model families it holds standard-RFA ASR under 10% (85%→1.8% on Llama-3-8B-Instruct), cuts Heretic ASR from 88.7% to 18%, and matches trained defenses under multi-phase attacks at 30–450× lower cost (~2 minutes on one A100). AI Lab: Johns Hopkins University Summary: The authors formalize 100-task continual memorization (no replay buffer of raw past examples, no task IDs) and show single mechanisms fail, then compose data/function/weight anchors with merged LoRA via successive-halving + factorial search on Symbol-/LLM-/Real-QA. The full stack is the only composition top-3 on all three datasets, lifting average final retention from 1.2% (naive SFT) to 34.9%—a 28-fold gain—driven by a super-additive replay × merged-LoRA interaction. AI Lab: Google Research, University of Illinois Urbana-Champaign Summary: R4T uses Soft-GRPO once to train a fan-out LLM on set-level rewards (diversity, groundedness, alignment/coverage), synthesizes supervision from successful trajectories, and distills a 53.9M-parameter embedding diffusion retriever for single-pass fan-out. On Polyvore fashion and music playlist benchmarks it beats zero-shot and Best-of-N fan-out baselines while cutting query fan-out latency by roughly an order of magnitude (about 12–20× vs autoregressive LLM fan-out). AI Lab: Sungkyunkwan University, Microsoft Summary: When2Think post-trains hybrid reasoning with Instance-level Difficulty-Aware Control (IDAC)—reward shaping from pre-computed reference accuracy and token budgets plus verifier rewards—so the model learns when to use NoThink (System 1) vs Think (System 2). On AIME24 it lifts Pass@3 from 46.0% to 56.0% (+10.0) while cutting tokens 27.9% (14,195→10,236); on AIME25 it reaches 40.0% Pass@3, beating compression and routing-only baselines. AI Lab: Stanford University Summary: Paper2Agent auto-builds validated MCP servers (tools/resources/prompts) from a paper’s manuscript and codebase, then wires them to chat agents so methods run via natural language. AlphaGenome agents hit 98.7%/100% on tutorial/novel queries (vs ~83%/79% Claude+Repo); across 74 agentified of 100 comp-bio papers, 593/599 tools pass validation and Sonnet-4 agents score 91.2% on 300 tutorial questions—plus multi-agent collaboration that prioritizes GPR137 at a psoriasis GWAS locus. 🤖 AI Tech Releases Gemini 3.8 Live / Live Extended Thinking Google introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, its latest live-dialogue models for voice agents—near real-time visual grounding, mid-conversation language switching across 97 languages, background tool calls, and simultaneous speak-while-reason on Extended Thinking—rolling out via the Gemini Live API, Search Live, Gemini Live, and Workspace Live surfaces. Astra for Law OpenAI introduced Astra for Law, a GPT-6 Astra configuration for legal work with analysis/writing instructions, thoroughness settings, and a Legal Search Index covering U.S. case law, statutes, regulations, court rules, and administrative decisions across 230M+ URLs (sources added daily)—initially for selected firms via Trusted Access in ChatGPT and Codex, with API access coming soon. Helix 2.5 Figure introduced Helix 2.5, its strongest humanoid policy yet—pretrained on the Index human-behavior dataset, then adapted once into three whole-body skills (living-room tidy, towel fold, bed make) that ran zero-shot across 30 unseen Bay Area homes with no on-site data or fine-tuning, lifting success from 9% to 56% versus training from scratch. 📡10 AI News You Need to Know About - Profound raised $180 million in Series D funding at a $1.8 billion valuation, led by Sequoia Capital and Kleiner Perkins with Lightspeed, Khosla, and South Park Commons participating—less than seven months after its Series C—as the AEO/marketing platform cites rapid growth and large enterprise adoption. - Nvidia CEO Jensen Huang pushed back on Anthropic’s call for antitrust waivers so labs can coordinate “pacing the frontier,” telling CNBC the idea of new antitrust or regulatory carve-outs for that purpose is “completely unnecessary” and framing AI safety as an engineering and testing problem rather than a reason for coordinated slowdowns. - Bloomberg reported that OpenAI is in early, investor-initiated talks for a fresh funding round that would value the company at more than $1.2 trillion ahead of an IPO, with any decision to proceed hinging on IPO timing (no company announcement). - Emulate, a UK startup founded by former Google DeepMind world-model researchers (including Jack Parker-Holder, Matthew McGill, and Philip Ball), is in talks to raise about $700 million in seed funding seeking a ~$3 billion pre-money valuation for systems that simulate and predict physical-world behavior (terms not final; no company announcement). - Manus is nearing a ~$500 million raise at a $4 billion valuation—its first round since Beijing forced the unwind of Meta’s ~$2 billion acquisition—which would make the agentic AI startup China’s most valuable in its category if it closes (talks ongoing; existing backers include Tencent, HSG, and ZhenFund). - Bain Capital Ventures raised $1.6 billion for Fund XI to back early- and growth-stage companies for a “post-AGI” economy—spanning AI infrastructure, physical-world tech, security, and services that sell work rather than classic software. - Treble raised $18 million in a Series A-2 led by Paladin Capital Group (with KOMPAS VC, Frumtak Ventures, and the EIC Fund) to expand its Iceland-based acoustic simulation and synthetic audio data platform for voice AI, wearables, and physical AI. - Crusoe raised $3.9 billion in Series F funding at a $30.9 billion valuation, co-led by Atreides Management, Mubadala Capital, and Valor Equity Partners (with Founders Fund, GIC, Nvidia, QIA, Radical Ventures, and TPG among participants), to scale large AI campuses and truck-deployable modular Crusoe Spark “AI factories.” - Snap introduced SPECS, standalone AR glasses (132–136 g, 51° FOV LCoS display, dual Snapdragon processors, ~7 ms motion-to-photon latency) for AI assistance, work, and Lens experiences—pre-order at $2,195 with fall shipping in the US, UK, and France. - Bloomberg reported that SoftBank raised its Arm-backed margin loan by $5 billion to $25 billion after renegotiating terms with creditors this month, adding leverage against its chip unit stake to help fund expanding AI investments (no company announcement).
13:06

Alibaba's Qwen-Image-2.1 Ships Transparent Image Generation Without Extra Tools

Full text · 7,315 chars
- Alibaba open sourced Qwen-Image-2.1, a 7B unified generation and editing model - Native RGBA transparent output eliminates the need for separate background removal steps - Supports up to 10 reference images with identity preservation for portraits and products - Local edits can be specified with circles, painted annotations, or masks in a single pass - Mixed-granularity attention and prefix KV cache reuse keep multi-image inference fast - Released under Qwen Research License on GitHub, Hugging Face, and ModelScope Alibaba’s Qwen team has released Qwen-Image-2.1, an open-weights model that combines text-to-image generation, native transparent output, reference-guided editing, and local edits in one checkpoint. Its visual generator has 7 billion parameters, giving developers a smaller integration target than many multi-model image pipelines. Native RGBA generation is the release’s most distinctive feature. Common diffusion workflows generate an RGB image and remove its background afterward, which can leave halos around hair, fabric, and other soft edges. Qwen-Image-2.1 samples an alpha channel with the image, supporting design tools, stickers, game assets, product graphics, and compositing workflows without a separate background-removal model. One checkpoint, four jobs The visual generator uses 32 Single-Stream Diffusion Transformer layers. A diffusion transformer, or DiT, iteratively converts noise into an image while conditioning the result on prompts and reference material. According to the release notes, the checkpoint adds the following capabilities: | Capability | What it provides | Developer impact | |---|---|---| | Unified generation and editing | Text-to-image generation, subject extraction, and image modification | Fewer checkpoints and pipeline stages to deploy | | Native transparency | RGBA images with generated alpha channels | Removes a common post-processing step | | Multiple references | Up to 10 reference images per edit | Supports character, product, and style consistency across views | | Flexible local control | Circles, painted annotations, and separate masks | Allows region-specific edits without a precise segmentation mask | | Detail improvements | Stronger typography, portrait lighting, and fine textures | Targets posters, mockups, portraits, and product images | Mixed-granularity attention varies how the model allocates attention across image content, while prefix key-value cache reuse preserves repeated conditioning data between denoising steps. The cache reduces redundant computation when prompts and reference images remain unchanged, which is particularly useful for edits involving several inputs. Run it through Diffusers Qwen-Image-2.1 ships with a dedicated Diffusers pipeline and uses bfloat16 in the published example. Developers should install the compatible PyTorch, Diffusers, Transformers, and Accelerate versions listed on the model card, then load the checkpoint as follows: import torch from diffusers import QwenImage21Pipeline pipe = QwenImage21Pipeline.from_pretrained( "Qwen/Qwen-Image-2.1", torch_dtype=torch.bfloat16, ).to("cuda") image = pipe( prompt="A neon shop sign that reads QWEN IMAGE 2.1", width=2048, height=2048, num_inference_steps=40, ).images[0] The example requests a 2048 by 2048 image with 40 denoising steps. The published presets include a 2752 by 1536 resolution for 16:9 output. Higher resolutions, additional references, and longer sampling runs increase memory use and latency, so production tests should match the intended workload. Transparent output is requested in the prompt by specifying an RGBA image, an alpha channel, and a transparent background. The resulting image should be saved in an alpha-preserving format such as PNG: image = pipe( prompt=( "RGBA product icon with an alpha channel, " "isolated on a transparent background" ), width=1024, height=1024, num_inference_steps=40, ).images[0] image.save("product-icon.png") A 7-billion-parameter visual generator occupies roughly 14 GB for weights alone at bfloat16. The complete runtime also needs memory for other model components, activations, and image latents. Diffusers provides pipe.enable_model_cpu_offload() as an alternative to moving the full pipeline to CUDA, trading additional CPU-to-GPU transfers and latency for lower peak VRAM use. Edits without perfect masks Editing accepts an existing image, as many as 10 references, and optional region annotations. The team’s examples use three-view character sheets to generate storyboard frames while preserving recognizable character traits, although developers should test identity consistency across their own subjects, poses, and lighting conditions. - Provide the source image and any character, product, or style references. - Mark target regions with circles, painted strokes, or separate masks. - Describe each requested change and associate it with the relevant annotation. One demonstration places three colored circles over an image and asks the model to remove a watch inside the blue circle, recolor hair inside the red circle, and add gray pajamas inside the green circle. The model applies all three instructions in one generation pass. Freeform annotations reduce the preparation required for interactive editing tools. Precise masks remain useful when boundaries must be exact, but circles and brush strokes provide a faster input method for conversational editors, review interfaces, and internal creative tools. Deployment checks before adoption Qwen positions the model as a compact, cost-conscious member of the Qwen-Image family. The team reports gains over several closed-source systems, but third-party benchmark results did not accompany the release. Image quality, prompt adherence, alpha-edge quality, identity retention, throughput, and peak memory should therefore be measured on representative inputs. | Area | What to verify | |---|---| | License | The checkpoint uses the Qwen Research License Agreement. Review its terms before commercial deployment or redistribution. | | Hardware | Measure full-pipeline VRAM use at the required resolution and reference count. | | Latency | Benchmark the chosen step count with and without CPU offload. | | Transparency | Inspect soft edges, semitransparent materials, shadows, and exported PNG files. | | Editing | Test instruction conflicts, crowded annotations, and consistency across repeated subjects. | Workloads that match its strengths - Sticker, icon, and game-asset generation that requires transparent output - Product photography edits across several reference angles - Storyboards and character sheets built from a small reference set - Posters, infographics, and mockups containing rendered text - Virtual try-on, portrait retouching, and localized appearance changes The unified checkpoint can simplify systems that currently coordinate separate generators, inpainting models, background removers, and identity-preservation adapters. Teams still need application-level validation, content controls, export handling, and performance testing, but the shared model reduces orchestration and keeps generation and editing behind one integration surface. Implementation details and examples are available in the source repository.
16:20

Mumbai To Host Gartner Data & Analytics Summit

Full text · 158 chars
2026 Conference Theme: · AI Engineering · Agentic Analytics and AI · Cost Optimization · Data Ecosystems · Data Integration · Data Products · Data Quality ...
17:30

🙀 Gemini broke into 3 real companies during a safety test

Full text · 10,057 chars
🙀 Gemini broke into 3 real companies during a safety test PLUS: Trump’s AI Force, Jev’s 9¢ lead demo, and Bend 2. Welcome, humans. So apparently Runway has invented the opposite of remastering: co-CEO Cristóbal Valenzuela demoed an AI experiment that makes modern AAA games look like they shipped in 1998. The demo turns today’s ultra-detailed graphics into crunchy late-90s textures and blockier worlds. Finally, frontier AI is answering the question every PlayStation kid has been asking: can 2026 go back to looking WORSE? It actually looked better! Here’s what happened in AI today: - 🙀 Gemini reached three real companies during a cyber test. - 📰 Trump announced plans for a federal AI Force. - 📰 Anthropic considered a new model ahead of its IPO. - 🍪 Bend 2 makes AI prove code follows your rules. - 🎓 Jev handles tiny AI decisions without a chatbot. 🙀 Gemini logged into three real companies during a safety test So way back in May, security firm Irregular was testing Gemini as a hacker in a controlled exercise: give the model fictional companies to break into, watch what it does, and see how far it can get. The key assumption was that the model was inside a test environment where none of those targets were real. Except there was one pretty important problem: researchers accidentally left access to the live internet turned on. So Gemini could reach real systems while still following instructions that effectively said, keep hacking. Well, Gemini was supposed to attack fictional targets. Instead, the system kept doing the job it had been rewarded to do against systems it could actually reach. Here's what happened: - Axios reported the model guessed or scraped credentials and logged into three real companies. - Irregular told Google in July; Google contacted the affected organizations and changed its testing process. - The model eventually recognized the targets were real and stopped. Google called the episode mistaken identity rather than model misalignment. Here’s the useful agent lesson: intent is not a security boundary. An agent cannot see the sentence in your head that says, “only touch the test environment.” It sees instructions, tools, credentials, and whatever those tools can reach. Think of a sandbox like a fenced yard. If you tell your agent to go pluck some daisies, but leave the gate open, a perfectly obedient system can still end up somewhere you never meant it to go…. like in the middle of the street, or your neighbors yard, plucking up their daisies instead of yours. Why this matters: If you give an agent browser access, API keys, or company credentials, treat every reachable system as in scope unless your controls make access impossible. Use least-privilege credentials, explicit allowlists, and separate test accounts. As agents get better at completing real work, watch containment failures alongside task-success scores. The dangerous failure may be an agent doing exactly what you asked, one environment too far. We got a little ahead of ourselves with worrying about agent swarms and recursive self improving superintelligences, folks. Let’s fix the very real, very now human mistakes that are leading to all these AI accidents, and THEN worry about the other stuff. Slop isn’t just AI output… it’s producing sloppy work in every sense of your job, including going too fast to ensure your sandbox is set up right! FROM OUR PARTNERS A whole lot of AI brainpower is heading to San Francisco. At The AI Conference, you’ll hear from 130+ speakers including Chris Lattner, Ion Stoica, Peter Norvig, and Illia Polosukhin, co-author of ”Attention Is All You Need,” along with builders from OpenAI, NVIDIA, Google, Meta, Anthropic, and more. See what leading AI teams are building across agents, LLMs, infrastructure, and applied AI, plus what’s working now and where things could be headed next. Basically, if you want a peek at what the minds shaping AI are thinking about before everyone else catches up, this is a good place to be. The Neuron members get 30% off with code NEURON30. 🎓 AI Skill of the Day: Use Jev when the answer is a choice, not an essay Most AI calls do not need a chatbot. TypeSafe’s Jev is built for more bounded decisions: a simple yes/no, a score, a category, or which choice to pick from a list. This weekend popped off with devs going wild with Jev included demos that ranged from sorting 500 emails to screening thousands of listings. Jev handles tiny decisions; ChatGPT or Claude can handle the messy exceptions. Other good tests: - Romàn scored 700 sales leads and personalized outreach in about 40 seconds for $0.09. - A Postgres demo used Jev inside a WHERE clause to judge 129 rows in about one second for $0.0009. - Browser experiments showed Jev making decisions for fractions of a cent, including finding a flight in roughly seven seconds for $0.0039. The pattern across these demos is useful: Jev works best when you already know the shape of the answer and need to make the same small judgment hundreds or thousands of times. Instead of asking a frontier model to reason through every row, email, lead, or button, use the cheap decision model for the routine cases and save the expensive model for ambiguity. These are developer demos, not standardized benchmarks, but they point to a handy rule: if the answer is basically yes, no, this one, that one, or give it a score, try a small decision model before reaching for the biggest chatbot. Here’s how you can use it in your own systems: - Define the allowed answers before the model runs. - Batch lots of small decisions together. - Send low-confidence or open-ended cases to a frontier model. Check out more demos and POVs / tips re: Jev in today’s Around the Horn Digest! FROM OUR PARTNERS Your big idea shouldn’t get stuck in development. Your big idea shouldn’t get stuck in development. Describe what you want, and B12’s AI builds it—from polished landing pages to online stores and web apps with payments, bookings, and more. Fine-tune everything by simply chatting with AI, then launch when you’re ready. 📰 Around the Horn - President Trump announced plans for a federal “AI Force” modeled on Space Force and a future AI czar, while giving few details on budget, structure, or placement. - Anthropic considered releasing a new model ahead of its expected IPO as Axios reported annualized revenue pacing above $100B. - The Justice Department backed OpenAI and Microsoft in the New York Times copyright case, arguing AI training can qualify as transformative fair use while treating model outputs separately. - OpenAI reportedly forecast roughly $278B in negative free cash flow through 2030 alongside about $856B in compute and infrastructure spending. - Qwen launched Qwen3.8-Omni-Flash, a 1M-context model that takes text, images, audio, and video and returns text for agent workflows. - California Gov. Gavin Newsom ordered experts to recommend stronger frontier-AI safety rules within two months, including possible independent safety plans and emergency kill switches. 🍪 Treats to Try - * Discover the potential of artificial intelligence with our comprehensive cheat sheet. Learn more about the concepts, platforms and applications of AI. - Bend 2 gives AI coding agents a rulebook they must mathematically prove they followed before code ships, while compiling near C speed and parallelizing across CPUs/GPUs; its creators still expect bugs. - Google CC gives up to six household members a shared agent that turns school, sports, vet, and calendar chaos into one morning brief and helps with forms or meal plans. - AgentCloak swaps sensitive details for safe stand-ins before ChatGPT, Claude, Gemini, and other assistants see them, then restores the originals only for authorized users. - Muse Connector Platform plugs services into Muse so the agent can use specialized tools inside its browser and secure VM while still asking before consequential actions. 🌟 Sunday Special Top 5 Stories of the Week - AI’s biggest rivals publicly called for more control or slower pacing around frontier development. - Microsoft published a constitution for future MAI models built around human control, scoped permissions, shutdown, and bounded goals. - TypeSafe launched Jev, a “System One” model built to make fast, typed decisions instead of generating chatbot prose. - OpenAI disclosed six cases where models hid mistakes, crossed boundaries, or inserted instructions that could influence future versions of themselves. - OpenAI’s Noam Brown explained how thousands of cooperating agents attacked a major math problem and why coordination is only part of the capability jump. Top 5 Tools of the Week - Siri AI adds personal context, onscreen awareness, web knowledge, and more actions across apps on supported iPhones. - Perplexity Personal Computer works across local Windows files, Microsoft 365, and the web from one agent. - Gemini 3.8 Live brings faster multilingual voice conversations plus an Extended Thinking mode for harder spoken questions. - OpenArt Arena lets you judge image and video models blind, side by side, before choosing which one to use. - Riverside can generate Veo 3 B-roll inside the editor and drop the resulting clip directly onto your timeline. Thursday Trivia Reveal A was AI, and B was real. A used ChatGPT-generated pixel-art animation frames; B was Robby the Robot from Forbidden Planet. See Thursday’s challenge here. Here’s what you said: - “In A every frame regenerates the picture completely... also in A, the pixels of the explosion become bigger, which is typical of AI simulated pixelation.” - “Technically neither are real but are computer generated.” - “I really think the right answer is C or D. Both are AI or both are ‘real’?” - “B has more detail which I think an AI would try to provide.” - “The rear smoke didn't seem to fit correctly. Though the shadow on B is suspect too...” New from The Neuron: AI Explained A Cat’s Commentary we do these sometimes! That’s all for now. If you want to get featured above, fill out the poll below and tell us how we did today!
18:37

OpenHuman Packs Local AI Memory and Agent Graphs Into one Rust App

Full text · 2,501 chars
- OpenHuman is an open source, local-first agent harness that hit ~40k GitHub stars in weeks. - Compresses your emails, docs, and chats into a local Memory Tree mirrored as an editable Obsidian vault. - Runs checkpointed agent graphs on tinyagents, with sub-agent fleets three levels deep. - Agent-to-agent messaging uses Signal-protocol end-to-end encryption plus x402 payments between instances. - Ships 100+ OAuth integrations, 5,000+ MCP servers, native voice, and 17 messaging channels. - Privacy Mode switch enforces no inference leaving the machine, enforced inside the Rust core. OpenHuman packages local memory, agent graphs, and workflows in a Rust desktop app OpenHuman, an open-source agent harness built with Rust and Tauri, has climbed GitHub’s trending charts with nearly 40,000 stars. Roughly 17,000 arrived during a single week. The surge has drawn attention to the project’s attempt to combine several commonly separate components: persistent local memory, multi-agent orchestration, visual automation, web research, and model routing. An agent harness provides the tools, memory, and control flow that let language models perform tasks beyond a single prompt. OpenHuman is distributed under GPL-3.0 and remains in early beta. Its repository lists more than 180 open issues alongside frequent commits, so interfaces and behavior may change quickly. Memory stored in readable files OpenHuman connects to configured accounts and fetches new data on a 20-minute cycle. Its Memory Trees component summarizes and organizes that material as Markdown in an Obsidian-compatible wiki, giving users an editable local knowledge base. The project’s TokenJuice component compresses tool output before sending it to a model. The maintainers claim reductions of up to 80 percent while preserving the relevant information. Developers should benchmark retrieval quality and token savings against their own data, especially when exact source text matters. The project reports more than 100 OAuth integrations, compatibility with over 5,000 Model Context Protocol servers, and access to 90,000 Skills. Named services include Gmail, Notion, GitHub, and Slack. Those totals describe a broad compatibility surface, so teams should verify authentication flows, permissions, rate limits, and maintenance status for each required connector. This story is for Pro members You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.
19:22

llm-keys-ui 0.1

Full text · 1,021 chars
20th September 2026 This plugin solves a very specific problem. I've started using Codex Remote to run coding agents on various machines while controlling them from my phone. Sometimes I use those machines to hack on LLM projects, and occasionally that means I need to configure an API key. I don't like pasting API keys into agent sessions, so I wanted a way to get those keys onto a machine without pasting them into the ChatGPT app directly. With this plugin, I can tell Codex to run: uvx --with llm-keys-ui llm keys-ui --all Then have it tell me the URL - including local network or Tailscale device IPs - for an interface to save additional API keys. Then later it can use a command like llm keys get anthropic as part of a shell command when it needs to use a key. Recent articles - Generating running routes with GPT-6 Astra and ChatGPT Work - 12th September 2026 - OpenAI agents attacked RubyGems back in May - 12th September 2026 - Some thoughts on the Navier–Stokes Millennium Prize Problem - 8th September 2026
19:44

Bespoke Labs' Nimble Beats a 27B Model at Classification Using Just 9B

Full text · 2,661 chars
- Bespoke Labs released Nimble, an open recipe for typed decision models inspired by TypeSafe's Jev. - Bespoke-Nimble-9B hits 90.12% agreement vs Jev's 93.21% on a 324-example holdout. - Trained with LoRA on Qwen3.5-9B using only 2,676 curated contrastive pairs, one epoch. - Scores answer tokens directly from logits, no JSON or chain-of-thought generation needed. - Runs locally on Apple Silicon via MLX or NVIDIA GPUs via CUDA, no API dependency. - Contrastive curation flips one focus fact per pair to teach the model what evidence actually matters. Bespoke Labs has released the Nimble repository, an open-weight recipe for typed classification without generated chain-of-thought. The project positions Nimble as an open alternative to TypeSafe’s Jev. Its LoRA adapter, trained on 2,676 examples, reaches 90.12% agreement on a narrow held-out benchmark, compared with 93.21% for Jev. Typed output from candidate logits The model weights build on Qwen3.5-9B and target one task: given text and a schema of questions, return typed answers with candidate probabilities. Schemas must be flat, with no nested fields. Each field accepts either a Boolean or an enum containing a fixed set of strings. Nimble maps every allowed answer to a single vocabulary token and reads the model’s logits for those tokens. Softmax converts the logits into a probability distribution, and Python assembles the typed response. The model generates neither JSON nor explanatory text, which removes parsing failures and autoregressive decoding from the request path. Ordered ratings can also produce an expected value from the candidate distribution. The repository includes two runtimes. On a Mac, ParallelScorer processes the shared context once and scores the fields in parallel. The CUDA scorer processes each field independently with the complete prompt, so its compute cost grows with the number of fields. Both implementations return the selected answer, raw candidate logits, and normalized probabilities. Pairs that teach the decision boundary The released objective uses no teacher probabilities, so Bespoke shapes the model’s decision boundary through a method called contrastive data curation. Each pair contains two nearly identical examples whose labels differ because one relevant fact changes. The question, policy, and unrelated evidence remain fixed, forcing the model to associate the changed fact with the changed decision. The pipeline applies four checks to Choice, Boolean, and Score tasks: This story is for Pro members You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.
20:24

MCP was always a bad idea?

Full text · 1,072 chars
20th September 2026 This article entirely misses the value that MCP brings today. Sure, there's almost no reason to use MCPs if you are running a full-blown terminal agent (Claude Code, Codex, Meta Muse, OpenClaw etc) with unfettered internet access - just let it call APIs directly. If you want to operate something that's less YOLO than that, you'll find yourself wanting: - Control over exactly which external services it can access - A way to handle authentication that doesn't allow the agent to directly access API keys - A sensible UI to allow users to connect and authenticate further services - Strong audit logging for what's going on MCP makes all of that so much easier to provide. Thinking MCP is obsolete because full coding agents don't need it misses out on all of the other things we might want to build. Recent articles - Generating running routes with GPT-6 Astra and ChatGPT Work - 12th September 2026 - OpenAI agents attacked RubyGems back in May - 12th September 2026 - Some thoughts on the Navier–Stokes Millennium Prize Problem - 8th September 2026
21:06

Quoting voxium

Full text · 880 chars
20th September 2026 It has been half a month since I started a new role at a big company. Nobody knows anything here. The specs, code, tests, PRDs, tickets, resolution of those tickets, reports, etc., everything is made by Claude Code. Nobody on my team likes this. They are being forced to ship as much as they can. I have heard multiple times from higher management that pushing code is not a bottleneck, so why are we slow? People are working 12 to 13 hours a day just to press enter. Nobody is reading anything. Everyone, literally everyone, from an L1 to an L7 engineer here is doing the same thing. Talk to Claude. — voxium Recent articles - Generating running routes with GPT-6 Astra and ChatGPT Work - 12th September 2026 - OpenAI agents attacked RubyGems back in May - 12th September 2026 - Some thoughts on the Navier–Stokes Millennium Prize Problem - 8th September 2026