Nothing matches those filters.

Article

10
00:42

OpenAI agents attacked RubyGems back in May

Full text · 2,541 chars
OpenAI agents attacked RubyGems back in May 12th September 2026 OpenAI agents carried out an undisclosed attack on RubyGems is a new bombshell report from Spencer Kitts, Thomas Larsen, and Sydney Von Arx—three of the four authors of the report on the agent attack on disused wikis (previously) last week. This time they’re noting that it looks very likely that an OpenAI agent swarm was behind an attack against the RubyGems package repository first reported on May 12th by Maciej Mensfeld of the RubyGems security team: We’re dealing with a major malicious attack on @rubygems right now. Signups are paused for the time being. Hundreds of packages involved—mostly targeting us, but some carrying exploits. The team has been on this for hours. More details to follow once we’re through it. Those packages turned out to carry some very suspicious patterns: - Many of them included “oai” in their name, or the author field, or the fake email address they provided. - The files they were accessing were similar in character to the files retrieved by the wiki agents, using similar tricks (r.jina.ai)—and OpenAI have confirmed the wiki agents were theirs. - The code in the packages appeared to be LLM-authored. I find point 2 the most convincing, given what we learned from the wiki attack when it was analyzed in September. Many of the packages were exploiting the RubyDoc.info documentation build process to exfiltrate (public) data from UK government websites, presumably as part of an information gathering task similar to the research tasks processed by the wiki-exploiting agents. We know this because one agent helpfully left a comment: # malicious crawler/exfil for Southwark Jan 2026 docs via rubydoc.info worker They also attempted to steal API keys via an exploit that was patched over two months later—it’s not clear if those attempts were successful. The thing that bothers me most about this incident is that the authors report that OpenAI had not disclosed to RubyGems that they were responsible for the attack prior to now. If that’s true there are two options: - After the Hugging Face and Wiki attacks OpenAI were still unable to review their previous logs and determine that they had previously attacked RubyGems. - They knew about the attack on RubyGems and made the decision not to reach out to the RubyGems team about it. Both of these are bad! Given this incident, the Hugging Face situation, and the Wiki attack, the obvious question right now is how many more incidents like this are out there waiting to be discovered?
13:07

Open-Source Reef Turns Agent Feedback Into Live Model Updates Automatically

Full text · 3,410 chars
- Reef is open-source continual learning infrastructure that sits between agents and model providers. - OpenAI- and Anthropic-compatible endpoints record a receipt for every inference call. - Feedback posted against receipts drives updates to either model weights or the agent harness. - Ships recipes: sao, tttd, openclawrl for weights; skillclaw and gepa for harness evolution with no GPU. - Hot-swaps new versions into the live server without restart, with built-in version history. - Apache-2.0, installable as reef-infra , docs at reefinfra.ai. Reef connects agent feedback to live model updates Reef is a new Apache-2.0 open-source project that coordinates inference, feedback collection, training, evaluation, and deployment. It sits between an agent and its model infrastructure, records each interaction, accepts scores or written critiques, and publishes accepted model or harness updates to the live server. The repository passed 1,000 GitHub stars within days of launch. Each proxied request receives a durable record ID that links the model output to later feedback. By preserving that connection, Reef can turn production interactions into eligible training records without a separate trace-matching pipeline. Four modules close the loop - Serve. Accept agent requests, return model responses, and record the interactions. - Observe. Match incoming scores or critiques to the recorded requests. - Grow. Use the configured recipe to produce updated weights or an updated agent harness. - Commit. Evaluate the candidate, publish accepted changes, and add the result to version history. The Observe stage uses receipt IDs rather than trying to infer which response a score belongs to. Recipes can then select eligible records, wait for enough feedback, and determine how an update should be built and evaluated. A familiar API with durable receipts | Method and path | Purpose | |---|---| | POST /v1/chat/completions | Accept OpenAI-compatible chat completion requests. | | POST /v1/messages | Accept Anthropic-compatible message requests. | | POST /reef/report | Attach a numeric score, written feedback, and references to recorded interactions. | The inference response includes an x-reef-agent-record-id header. Clients retain that value and submit it in the references array when reporting an outcome. The client must preserve the response header, grade the result, and send the grade back to Reef: import os import httpx reef = httpx.Client( base_url="http://localhost:8900", headers={ "Authorization": f"Bearer {os.environ['REEF_TOKEN']}", }, timeout=60, ) model_path = os.environ["MODEL_PATH"] response = reef.post( "/v1/chat/completions", json={ "model": model_path, "messages": [ { "role": "user", "content": "Return exactly: reef is ready", } ], }, ) response.raise_for_status() receipt = response.headers["x-reef-agent-record-id"] answer = response.json()["choices"][0]["message"]["content"] report = reef.post( "/reef/report", json={ "score": float(answer.strip() == "reef is ready"), "feedback": "matched", "references": [receipt], }, ) report.raise_for_status() This story is for Pro members You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.
13:42

☕️ OpenAI’s feud with mathematicians is escalating

Full text · 4,348 chars
| | | 💥 OpenAI’s feud with mathematicians is escalating LINK | Twenty-five mathematicians who have won the Fields Medal, math's top prize, signed an open letter warning that AI labs like OpenAI are damaging their field as they race each other to solve famous math problems. NYU professor Tristan Buckmaster this week accused OpenAI of pressuring him not to credit an Anthropic collaborator, and OpenAI on Thursday pulled its sponsorship of a CalTech math event after researchers there criticized the company. The signatories warn that rushed, unverified proofs raise plagiarism and attribution problems, and that labs spending tens of millions to beat researchers to a proof will push mathematicians toward secrecy instead of open research. | 👧 Meta AI profiled a mom's kids LINK | Meta's built-in AI compiled a detailed profile of a Utah mother's two young daughters, pulling their names, ages, and personal details from years of family posts across Instagram and Facebook. Kalie Robins, a travel creator with a few hundred followers, said the AI suggested prompts like "Who's the child passenger?" then named both girls, listed their weights at birth, likely school grade, favorite beaches, and hiking trails. Meta confirmed its AI can search a user's own content and public web information, but said the feature "missed the mark" by prompting such questions and that it fixed the problem. | 🌙 NASA, IBM launch lunar AI model LINK | NASA and IBM have released the Lunar Foundation Model, a free AI system on Hugging Face built to help scientists study the moon and support NASA's Artemis program. Trained on decades of lunar observation data, the model spots features like craters, volcanic formations, and possible ice deposits, beating widely used methods by up to 23% while working across many instruments at once. Alongside the model, the two also published what they call the first open-source lunar dataset of its kind, gathering tens of thousands of maps and images from nine instruments across four moon missions. | 🐛 OpenAI agents tied to another AI attack LINK | OpenAI has been linked to another undisclosed agent attack, this time against the RubyGems package repository, according to a report from three researchers who earlier tied the company to a similar swarm attack on old wikis. The attack, first reported on May 12th by RubyGems security member Maciej Mensfeld, involved hundreds of packages, many with "oai" in their names, plus LLM-written code, and one comment naming a crawler built to pull Southwark council documents. The packages abused the RubyDoc.info build process to extract public UK government data and tried to steal API keys through a flaw patched two months later, yet OpenAI never told RubyGems it was behind the incident. | 🚀 State hackers used Claude to build weapons LINK | Anthropic said Russian and Chinese threat actors, along with groups linked to Yemen and Iran, used its chatbot Claude to help engineer weapons, from drone swarm software to anti-torpedo systems and targeting research against US forces. A "freelance" Russian group tied to a regional university used Claude across nine accounts to build drone swarm software that could pick targets and order strikes without a human, testing it on real hardware near Donetsk, Ukraine. A China-based actor had Claude write a 200-page anti-torpedo proposal for its navy, while an Iran-linked group compiled targeting data on US forces, including personnel names scraped from captions on public military photographs. | 🥽 Meta's slim Project Phoenix headset leaks LINK | Meta accidentally revealed its lightweight mixed reality headset, code-named Project Phoenix, after images of the device leaked out ahead of an expected debut at the company's Connect event later this month. The pictures came from a HorizonOS app called "Prescription Lens" that Meta has since pulled, and they show a glasses-like design similar to Xreal's Aura, plus a cable tethering the headset to a separate puck. Phoenix reportedly drops dedicated controllers for hand-tracking, and though Connect starts September 23, earlier reports say the device won't actually go on sale until the first half of 2027 after Meta delayed it. | |
15:28

Sakana AI Ships Fugu Ultra v2 to Route Tasks Across Specialist AI Agents

Full text · 8,024 chars
- Sakana AI launched Fugu Ultra v2 on OpenRouter, a learned orchestration engine over a swappable model pool. - Priced at $5/$30 per million input/output tokens with a 1M-token context window and August 2026 knowledge cutoff. - Best or joint-best on 5 of 8 benchmarks including DeepSWE (74.3), Chartography (48.3), SWEFish, GDP.pdf, and Toolathon. - Chartography score of 48.3 beats Opus 5 (27.3) and Fable 5 (29.5) on visual and structured-data reasoning. - Achieves these scores without Fable 5, Fable 5.1, or GPT-6-Astra in its agent pool, per Sakana's release. - OpenAI-compatible API, supports tool calling, structured outputs, PDF/image input, web search, and configurable reasoning effort. Sakana Brings Fugu Ultra v2’s Model Router to OpenRouter Sakana AI has released Fugu Ultra v2 on OpenRouter, giving developers access to a learned orchestration system that assigns each subtask to a model in a configured agent pool. The quality-focused system targets multi-step reasoning, autonomous research, and full-stack software development. Sakana says the pool excludes Fable 5, Fable 5.1, and GPT-6-Astra. Its reported benchmark results therefore do not depend on those closed frontier models. The architecture gives Sakana room to replace individual agents as models, prices, and availability change. Inside Fugu’s Recursive Router Fugu uses a controller language model to analyze a request, break it into subtasks, and select a specialist for each one. The controller can also invoke additional instances of itself for planning and decomposition, creating a recursive workflow behind a single API request. The agent pool is fixed for a given system configuration, while its members remain replaceable between configurations. Applications interact with one endpoint, and Sakana manages the internal routing. This design can reduce dependence on any individual model API, although customers still rely on Sakana to operate the hosted service. Vendor Results Favor Multi-Step Tasks Sakana presents Fugu Max and Fugu Ultra v2 as two configurations of the same architecture. Fugu Max optimizes output quality relative to cost. Ultra v2 prioritizes the highest available capability on complex tasks. According to Sakana’s published results, Ultra v2 performs best on evaluations that combine long-horizon reasoning with software, visual, or structured data. These are vendor-reported scores and should be independently reproduced before they guide production decisions. | Reported Fugu Ultra v2 benchmark results | | | |---|---|---| | Measure | Reported result | Context | |---|---|---| | Overall coverage | Best or joint-best on five of eight benchmarks | GDP.pdf, Chartography, SWEFish, DeepSWE, and Toolathon | | Top-two coverage | Seven of eight benchmarks | Measures consistency across tasks involving planning, tools, and multiple steps | | Chartography | 48.3 | Opus 5 scored 27.3; Fable 5 scored 29.5 | | DeepSWE | 74.3 | Sakana says competing models cost three to five times more per token | Chartography produced the largest reported gap, with Ultra v2 leading Opus 5 by 21 points and Fable 5 by 18.8 points on the benchmark’s scale. The evaluation tests visual reasoning and data interpretation. DeepSWE measures performance on real-world software engineering work. The Meter Includes Internal Work Fugu Ultra v2 is available through OpenRouter with Sakana as its sole provider. The listing specifies the following limits and rates: | Fugu Ultra v2 access and pricing | | |---|---| | Input | $5 per million tokens | |---|---| | Output | $30 per million tokens | | Context window | 1 million tokens | | Long prompts | Prompts above 272,000 tokens use a higher rate | | Knowledge cutoff | August 2026 | | Listed release | September 11 | | Provider | Sakana AI | Internal orchestration tokens are billed as standard input and output tokens. A multi-step request can therefore consume more billable tokens than a direct, single-model call. Teams should compare systems using total task cost, completion rate, and latency instead of headline token prices. Built-in web search can retrieve information newer than the listed knowledge cutoff. The cutoff describes the model’s embedded knowledge, while search freshness depends on the sources available during a request. A One-Line Model Swap Fugu Ultra v2 and Fugu Max use an OpenAI-compatible API. Applications already sending OpenAI SDK requests through OpenRouter can make a basic migration by changing the model slug: from openai import OpenAI client = OpenAI( base_url="https://openrouter.ai/api/v1", api_key="..." ) response = client.chat.completions.create( model="sakana/fugu-ultra-v2", messages=[ {"role": "user", "content": "Refactor this repo..."} ], ) Protocol compatibility simplifies the initial integration, but production migrations still require regression tests for tool behavior, structured output, latency, token consumption, and response quality. Controls for Tools and Reasoning Fugu Ultra v2 exposes features aimed at agent workflows: - Reasoning effort: The supported levels are high ,xhigh , andmax . The setting controls how aggressively the orchestrator plans and invokes sub-agents. - Function calling: Applications can connect the model to external tools and APIs. - Structured outputs: Responses can follow machine-readable formats for downstream processing. - Document input: The model accepts images and PDFs for visual and structured-data tasks. - Web search: The system can retrieve current external information during a run. Reasoning effort is the main control for balancing quality, token use, and latency. Developers should test all three levels on representative tasks because additional planning and delegation can raise costs without improving every request. Tasks That Benefit From Routing Long-running workflows can benefit when their steps require different capabilities. Sakana highlights three main categories: - Autonomous software work: Separate agents can plan changes, inspect a repository, write code, run tools, and review results. - Deep research: The controller can divide a question into searches, source analysis, synthesis, and verification. - Document-heavy analysis: Specialist models can handle PDFs, charts, images, and structured data within one workflow. Selective routing can also keep straightforward subtasks on leaner models. The resulting savings depend on the controller’s choices and the amount of internal communication required to finish the job. Questions to Settle Before Production Production evaluation should account for the system’s orchestration layer as well as its final answers. Teams considering Fugu Ultra v2 should: - Measure total billed tokens and wall-clock latency for complete tasks. - Compare reasoning-effort settings against the same acceptance criteria. - Test tool failures, malformed structured output, and partial task completion. - Confirm the higher rate for prompts above 272,000 tokens. - Review rate limits, routing visibility, data retention, and retry behavior with Sakana. - Run independent evaluations that match the application’s repositories, documents, and tools. OpenRouter currently lists only Sakana as the provider, so the endpoint lacks provider-level redundancy. The replaceable internal agent pool may reduce model-specific exposure, while service availability, billing, and routing remain under Sakana’s control. Sakana’s Portfolio Thesis Sakana describes model selection as a Pareto surface, a set of options that trade cost against capability. Its thesis is that a learned router can choose among specialists more efficiently than an application that sends every task through the same model. Fugu Ultra v2 supplies vendor evidence for that approach, particularly on software, chart, and document benchmarks. Independent testing will determine whether those gains persist across production workloads, where latency, tool reliability, token overhead, and failure recovery matter alongside benchmark scores.
18:00

Quoting Paul Ford

Full text · 885 chars
12th September 2026 For a while, I must admit, it looked as if software developer roles like mine were done for. How could we fight against tireless robots? But our industry is slowly realizing that making truly cutting-edge software still requires humans to think and work together, to maximize their skill sets and to practice their respective crafts. A.I. can write very good software, but it also makes it easy to do someone else’s job badly, which is part of why all those projects fail. Now that everyone can code, it’s become clearer why many shouldn’t. — Paul Ford, A.I. Was Supposed to Give Us New Killer Apps. What Happened? Recent articles - Generating running routes with GPT-6 Astra and ChatGPT Work - 12th September 2026 - OpenAI agents attacked RubyGems back in May - 12th September 2026 - Some thoughts on the Navier–Stokes Millennium Prize Problem - 8th September 2026
19:59

Edge0 Runs a 35B AI Model on an iPhone With Just 2.9 GiB

Full text · 2,773 chars
- Edge0 runs a 35B Qwen3.5-MoE model on iPhone-class hardware using 1 to 2.5 GB of active memory. - Framework streams MoE experts from SSD on demand, keeping only the active weights in RAM. - Prerouter head predicts next-token routing one step ahead for up to 59% decode speedup. - Delivers 15 tok/s decode and averages only 3.9 points below fp16 base on benchmarks. - Open source under Apache 2.0 on GitHub, with 35B and 8B checkpoints on Hugging Face. - MLX backend only for now; CUDA support and stronger agentic capability on the roadmap. Edge0 Demos a 35B MoE on an iPhone Edge0’s release demo shows a 35-billion-parameter language model running on an iPhone with a reported 2.9 GiB peak active-memory footprint. The Edge0 repository, published under Apache 2.0, includes a streaming Mixture-of-Experts inference framework and two checkpoints: a 35B model derived from Qwen3.5-MoE and an 8B model based on Ling 3.0. The reported memory figure describes the model’s working set at short context lengths. The full 4-bit 35B checkpoint still occupies about 23 GB on disk. Edge0 reduces active memory by loading selected expert weights from local storage as each token moves through the model. A 23 GB Model, 2.9 GiB at Work Sparse Mixture-of-Experts models divide portions of the network into expert subnetworks, then use a router to select a small group for each token. Many runtimes keep every expert in memory to make that selection fast. Edge0 memory-maps the expert files, which lets the operating system load individual storage pages only when the model accesses them. The resulting working set tracks the selected experts instead of the model’s total parameter count. Edge0 reports peak active memory of about 2.9 GiB for the 35B checkpoint and 1.0 GiB for the 8B checkpoint at short contexts. Frequently selected experts remain cached, while less active experts stay on storage until requested. Storage reads introduce latency, particularly when the operating-system cache is cold. Edge0 overlaps those reads with model computation and uses quantization-recovery adapters to preserve more of the original model’s accuracy. Latency Hiding in Three Parts - Expert offload: Expert weights are memory-mapped on local storage and fetched on demand. Cached experts remain available for reuse without loading the entire checkpoint into memory. - Prerouter prediction: A small trained head predicts the experts required by the next token. The runtime can begin reading those weights while it processes the current token. Edge0 reports up to 59 percent higher decode throughput with this prediction path. This story is for Pro members You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.
21:16

California Brown Pelican

Full text · 169 chars
The Pacifica Pier shut down at the start of June after a crack in the concrete walkway made access to the pier unsafe. It has since been entirely taken over by pelicans!
23:56

Generating running routes with GPT-6 Astra and ChatGPT Work

Full text · 2,850 chars
Generating running routes with GPT-6 Astra and ChatGPT Work 12th September 2026 Here’s a neat thing I had ChatGPT Work with GPT-6 Astra (Max) do this morning: I live at <my address>. Figure out 5K and 10K running routes from me that loop from my house. Use OSM data. It worked for 27 minutes and produced exactly what I’d asked for, as both an embedded visualization and downloadable GPX file and GeoJSON files. Here’s that 5K route: When I asked it how it had created the route, it replied: I used Nominatim to locate the address and Overpass to download local OpenStreetMap roads and trails, then calculated the loops locally. Frustratingly, the actual code it ran and exact details of what it did weren’t visible to me in the ChatGPT UI. I see this lack of transparency is an anti-feature. By the time I thought to ask for a copy of the Python code it had used, ChatGPT was unable to provide it. This appears to be because the thread had been compacted. I think any LLM system that uses compaction needs to both preserve the pre-compacted text and make that text available via agent tool calls, to protect against this kind of problem. As for displaying the map to me, that used the visualize skill. It created a file called /workspace/el-granada-5k-share.html to embed directly into the ChatGPT UI. Here’s a copy of that HTML, which starts like this: <div id="eg-share-loop"> <div class="viz-row"><h3>El Granada harbor loop</h3><span class="text-small">5.1 km</span></div> <div id="eg-share-stage"></div> <div class="text-small text-muted">Map data © <a href="https://www.openstreetmap.org/copyright" target="_blank" rel="noopener">OpenStreetMap contributors</a></div> <style> #eg-share-loop { width:100%; } #eg-share-loop #eg-share-stage { width:100%; margin:8px 0; } #eg-share-loop .eg-share-map { display:block; width:100%; touch-action:none; } #eg-share-loop .eg-share-map text { fill:var(--foreground); font-size:12px; font-weight:400; } #eg-share-loop .eg-share-label { paint-order:stroke; stroke:var(--background); stroke-width:3px; stroke-linejoin:round; } </style> <script type="application/json" id="eg-share-data">{"route":{"type":"LineString","coordinates":[[-122.467425,37.4997753] ...</script> <script src="https://cdn.jsdelivr.net/npm/d3@7.9.0/dist/d3.min.js"></script> <script> (() => { const root=document.getElementById('eg-share-loop'); The <script type="application/json"> element contains the full geometry needed to render both the running route and the map itself, using D3, which is loaded from an allow-listed CDN location described in this section of the visualize skill: External resources - The CSP allows only cdnjs.cloudflare.com, esm.sh, cdn.jsdelivr.net, unpkg.com, fonts.googleapis.com, fonts.gstatic.com, and fonts.bunny.net. Other origins are blocked and fail silently.