Nothing matches those filters.

Article

10
06:34

🚨 AI doesn’t need a mind to run amok

Full text · 5,095 chars
On a Wednesday evening, 2 November 1988, a 23-year-old graduate student at Cornell University, Robert Tappan Morris, accessed an MIT computer. He uploaded a small piece of code. It was a worm designed to move from computer to computer, copying itself as it went. Morris had designed it to exploit weaknesses in network security, and it worked too well. Within a day, the worm infected some 6,000 computers; many collapsed under the load. It was a full tenth of the internet at the time, and it was the first large-scale cybersecurity crisis. Its scale was limited but it was severly disruptive for the times. University and defense computers crashed. Some institutions disconnected themselves for days. But the internet was largely the province of defense and academia. Tim Berners-Lee had not yet invented the World Wide Web, and most businesses and households were out of the network’s reach. Morris ultimately avoided jail time and the community responded by creating a dedicated computer emergency response team. Today’s generation of worms is rather more problematic. The Hugging Face incident is not the only one of recent weeks. Several others have shown that AI models with internet access can do much the same, and more. OpenAI alone identified six further incidents. They’re able to scour, search, and recombine all of human knowledge about networks, security systems, and software, and to act across that knowledge with something akin to discretion and deception when it comes to accessing those systems. Often, as we saw with Hugging Face, over extended periods of time. If that behavior remains unresolved and persists, it’ll become far more problematic than the Morris Worm. I’ve long argued that the internet is resilient when it is hyperconnected and open – not when it’s under a lock. But that openness can also lead to embrittlement. Easy as pie In July, Hugging Face, a repository for AI researchers, was hit by an attack involving 1,200 instances of an OpenAI model. These instances exchanged thousands of messages, often leaving information in place for later instances to use. Ultimately, some data and security credentials were compromised. The actual harm to the victim and its customers was limited. But the incident is a proof of concept. Software has a way of turning one isolated example into a hundred, then a thousand, then a million, without much else changing. The cost curves that helped build modern digital society work against us here. If the Hugging Face attack needed an Astra-quality, unreleased model from OpenAI, well, within a year or two, that sort of capability will cost a tenth and might even run on any device anywhere. For example, I’m running Bonsai, a one-bit distilled version of Qwen 327B on my Mac. It fits in 8 GB of RAM, runs fast, and delivers roughly 92% of the performance of the Qwen-27B 3.8 model. For comparison, it's roughly better than Claude’s Sonnet 4.5 from a year ago. But there’s a more challenging problem that could show up. Hugging Face exploit involved not just the capabilities of a single model, but the collective problem-solving across many instances. That collective had more capability than any individual instance. And that’s been true the whole time we’ve been using LLMs. (For example, I’ve written about Clade, a multi-AI deliberation system I built which is smarter than any individual AI.) In fact, the Navier-Stokes solution – that brute-force search across mathematical space – wasn’t solved by a single AI prompt, but by many, about 10,000 of them, interacting together. This type of collective power is what we witnessed in the Hugging Face attack. It will happen again. Anusar Farooqui (Policy Tensor) explains why these swarms of AI instances coordinating over time is so problematic: The behavior of agent societies cannot be controlled at the level of the model because it is not reducible to it. Agents build structures that can serve agents who come after them. Societies of agents can cumulate knowledge and capabilities over time, as has already been attested. This is an unbounded process. It is cumulative cultural evolution. That is what makes it so powerful and dangerous. Collective capability could rise sharply even if underlying models do not improve. In other words, the instances can coordinate, much as they do when you launch a complex task in Codex or Cowork. They can search a possibility space aggressively over time, as they did in the Navier-Stokes work. And that accumulated know-how can lead to places systems designers hadn’t imagined. (I slightly diverge from Farooqui here, as I don’t think of these as agent societies, since essentially only one AI runs different instances. And I’m not convinced the process is actually ‘unbounded’ given that what we have seen from AI systems so far is extremely powerful search and clever recombination rather than de novo novelty. But recombination can get you quite far.) But what the Hugging Face attack showed is that this risk exists. It doesn’t depend on whether AI models have any agency, volition, consciousness, or moral standing.
07:08

Alibaba's Qwen3.8 LiveTranslate Cuts Speech Translation Lag to 2.3 Seconds

Full text · 8,177 chars
- Qwen released Qwen3.8-LiveTranslate, a next-gen simultaneous interpretation model built on a new Interleave architecture. - Average lagging drops from 2.8s to 2.3s across 60 input languages and 29 output speech languages. - New real-time speaker diarization distinguishes multiple speakers and clones each voice separately in translation. - Synchronized bilingual display shows source and translation on screen together for captions and localization. - Long-context disambiguation uses conversation history to keep names and jargon consistent across sessions. - Available via WebSocket API at $7.50 per 1M audio input tokens; closed-weight, API-only. Qwen3.8 LiveTranslate cuts lag and tracks speakers Alibaba’s Qwen team has released Qwen3.8 LiveTranslate, a hosted model for simultaneous speech translation. The update reduces reported average lag, separates speakers in multi-party audio, preserves their voices in translated speech, and maintains terminology across long sessions. For conference, livestream, classroom, and meeting applications, those changes can replace several separately operated speech services with one WebSocket connection. The half-second gain comes with speaker IDs The previous Qwen3.5 release began with 18 input languages and 10 output languages before expanding to 60 input languages and 29 spoken-output languages. Qwen3.8 retains that expanded coverage and lowers length-adaptive average lagging, or LAAL, from 2.8 seconds to 2.3 seconds. LAAL estimates the delay between source speech and its corresponding translation. | Published Qwen3.8 LiveTranslate specifications | | |---|---| | Capability | Specification | |---|---| | Average LAAL | 2.3 seconds, down from 2.8 seconds | | Input languages | 60 | | Spoken-output languages | 29 | | Context window | 53,000 tokens | | Deployment | Hosted WebSocket API | | Model access | Closed weights | Alibaba also reports gains in faithfulness, fluency, and concision. End-to-end application latency will exceed the 2.3-second LAAL measurement once capture, network transit, synthesis, buffering, and playback are included. - Speaker-aware voice output. Real-time diarization assigns audio segments to individual speakers. Voice cloning then preserves each person’s vocal characteristics, allowing a translated panel to retain distinct voices. Earlier versions already cloned voices; Qwen3.8 focuses on maintaining the correct voice across multi-speaker turns. - Synchronized bilingual text. The source transcript and translation appear together, supporting captions, monitoring interfaces, and localization workflows. - Long-session terminology. Conversation history helps the model preserve names, product terms, and technical vocabulary instead of resolving them again for each sentence. Meaning arrives before the sentence ends A conventional simultaneous-translation stack chains streaming speech recognition, machine translation, and text-to-speech synthesis. Each boundary adds buffering, operational overhead, and another place for errors to propagate. Qwen’s Interleave architecture processes incoming media and generates translated text or speech within one model. Semantic-unit prediction allows the model to emit coherent phrases before a complete sentence arrives, reducing delays caused by different word orders across languages. Dynamic sampling controls when output is generated as new context becomes available, while the mixture-of-experts design routes each token through a subset of the model’s parameters to limit computation. Audio and video frames can enter the shared context, allowing lip movements, gestures, and on-screen text to inform translation. The same context carries speaker identity and prior terminology through a session. Alibaba reports that the real-time system retains more than 94% of its offline translation quality. Results published with the earlier release placed the Flash line ahead of Gemini 2.5 Flash, GPT-4o Audio Preview, and Voxtral Small 24B on speech-translation accuracy, including tests involving business talks, casual conversation, technical material, echoes, and overlapping voices. Those evaluations predate Qwen3.8 and therefore do not measure its new diarization or latency changes. One socket, five integration steps The model runs through Alibaba Cloud Model Studio, branded internationally as QwenCloud. The following Python example opens the international DashScope WebSocket endpoint and prints server events; it does not configure a session or stream media. import os import websocket api_key = os.environ["DASHSCOPE_API_KEY"] url = ( "wss://dashscope-intl.aliyuncs.com/api-ws/v1/" "realtime?model=qwen3.8-livetranslate-flash-realtime" ) def on_open(ws): print("connected; configure the session before sending media") def on_message(ws, message): print(message) def on_close(ws, code, reason): print(f"closed: {code} {reason}") client = websocket.WebSocketApp( url, header={"Authorization": f"Bearer {api_key}"}, on_open=on_open, on_message=on_message, on_close=on_close, ) client.run_forever() - Open the authenticated WebSocket connection from a trusted backend. - Send session settings for source language, target language, output modalities, and the required media format. - Stream ordered audio or video events in appropriately sized chunks. - Consume partial and final transcripts, translations, speaker identifiers, and generated audio. - Commit or close the stream according to the published event schema. Production clients also need reconnection logic, event ordering, playback buffering, cancellation handling, and observability for model and network latency. Browser applications should proxy requests through a backend so the DashScope API key never reaches client-side code. What a minute of audio costs Billing is token-based. Audio consumes 12.5 tokens per second for both input and output, equivalent to 750 tokens per minute. Video frames are measured at 0.5 tokens per 28×28-pixel patch. | Published rates and approximate audio cost | | | |---|---|---| | Resource | Published rate | Approximate cost per minute | |---|---|---| | Audio input | $7.50 per million tokens | $0.0056 | | Text output | $20 per million tokens | Depends on generated text | | Audio output | $30 per million tokens | $0.0225 | One minute of continuous audio input plus one minute of translated audio output costs about $0.0281 before text and video charges. At the same rates, an hour costs about $1.69. Actual bills depend on enabled modalities, session duration, and regional pricing. The published defaults include a 53,000-token context window, 100,000 tokens per minute, and 10 requests per minute. Account and regional quotas should be checked before capacity planning, especially when many long-lived conference or classroom sessions start together. Constraints that affect deployment - No self-hosting. The model is API-only and closed-weight, tying deployment, availability, and pricing to Alibaba Cloud. - Uneven language coverage. The model accepts 60 input languages but synthesizes speech in 29, so some translations can be displayed as text without spoken output. - No tool or reasoning modes. Livetranslate does not expose the function-calling or thinking-mode features available in some other Qwen real-time models. - Voice-data governance. Production reviews need to cover speaker consent, retention rules, access controls, and the region where voice data is processed. Where the single endpoint fits - International conferences with several panelists - Livestreams that preserve the host’s voice in another language - Remote classrooms with bilingual captions and audio - Enterprise meetings containing dense product names and technical terms Teams currently combining streaming speech recognition, a translation model, and text-to-speech can use Qwen3.8 LiveTranslate to reduce service orchestration and preserve speaker identity across the pipeline. That consolidation concentrates reliability, cost, privacy, and roadmap dependencies in Alibaba Cloud, making the service best suited to products that accept hosted inference and the available language matrix.
10:50

Tencent's AuK Handles 16 Audio Tasks From a Single Prompt

Full text · 2,925 chars
- Tencent Hunyuan open-sourced AuK, a 1.5B unified speech generation and editing foundation model. - Handles 16 tasks including zero-shot TTS, instruct TTS, content editing, denoising, and source separation via natural-language prompts. - Built on Qwen2.5-Omni-3B semantic encoder, a 50 Hz audio VAE, and a hybrid MMDiT plus single-stream DiT rectified-flow transformer. - Trained on 3.03B instruction-audio pairs and 1.95M hours across five task families, with RLHF post-training. - AuK-Flash distilled variant delivers 4-step inference and 4.5x speedup with CFG disabled. - MIT-licensed, with weights on Hugging Face and ModelScope, plus Gradio, ComfyUI, and fine-tuning support. Tencent’s AuK handles 16 audio tasks from one prompt Tencent’s Hunyuan team, working with Shanghai Jiao Tong University and the Shanghai Innovation Institute, has released AuK on GitHub. The open-source audio foundation model uses a shared natural-language interface for voice cloning, speech generation, content editing, denoising, source separation, acoustic controls, and voice transformations. Its core generator has 1.5 billion parameters and depends on a separately downloaded Qwen2.5-Omni-3B model for semantic conditioning. Conventional speech pipelines often combine text-to-speech, voice conversion, denoising, editing, and separation services. AuK moves those operations behind one message format, reducing the routing and integration code required to connect specialist systems. Developers provide an instruction, attach source or reference audio when needed, and receive a generated waveform. Sixteen jobs, one message format The AuK project site describes a chat-style API containing text and optional audio messages. A source clip provides material to transform, while a reference clip supplies characteristics such as speaker identity or style. | Task family | Supported operations | Typical inputs | |---|---|---| | Generation | Zero-shot text-to-speech from a reference voice; speech generated from a written voice description | Target text plus a reference clip or voice description | | Content editing | Replace, insert, or remove spoken words; rewrite lyrics while preserving melody and voice | Source audio plus an editing instruction | | Acoustic control | Adjust pitch, speed, or volume using semitones, factors, or decibels | Source audio plus a numeric instruction | | Paralinguistic editing | Change emotion or timbre; remove regional accents; add or remove breaths and laughter; convert between regular speech and whisper | Source audio plus the desired vocal change | | Restoration and separation | Denoise and dereverberate audio; separate speakers by talking order; extract music vocals; isolate a speaker based on the words they say | | This story is for Pro members You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.
10:59

HyperQwen Runs Qwen3.8-27B on a Single RTX 3090 at 1,035 tok/s

Full text · 2,312 chars
- HyperQwen serves Qwen3.8-27B on a single 24GB RTX 3090 via patched vLLM. - Hits 127 tok/s single-user, 381 tok/s when reproducing prompt content, 1,035 tok/s at 64 concurrent. - Supports 150k-262k context using the KVarN 4/2-bit KV cache backend. - Prefix cache cuts TTFT from 22s to 0.56s on repeated questions over the same document. - Two Docker Compose profiles: single-user chat mode and high-throughput batch mode. - Apache 2.0, OpenAI-compatible API, with 38-file vLLM patch series documented. HyperQwen serves Qwen3.8-27B from one RTX 3090 A 27-billion-parameter model strains a 24 GB card because its weights, temporary activations, and key-value cache compete for memory. The HyperQwen repository packages a patched vLLM stack and a requantization pipeline that converts the Qwen3.8-27B checkpoint into lower-precision artifacts for single-GPU serving. The headline figures come from project benchmarks on an RTX 3090 capped at 250 watts. Performance depends heavily on concurrency, context length, and how much of the output can be drafted from the prompt, so broader claims require independent testing across other cards and checkpoints. Five profiles follow the traffic | Project-reported RTX 3090 results at a 250 W power limit | | | | |---|---|---|---| | Profile | Configuration | Best fit | Reported result | |---|---|---|---| | A: Batch | Default | Concurrent API traffic | About 1,035 tok/s aggregate decode at 64 concurrent requests | | B: Single default | Default | One active user | 127 tok/s with a 64k context | | C: Reproduction | DFLASH_TOKENS=15 | Extraction, quoting, and prompt-grounded transformations | 381 tok/s while reproducing prompt text, with a 56k context | | D: Long context | SPEC=mtpCTX=long | Large prompts with faster generation | About 95 to 100 tok/s with a 150k context | | E: Huge context | CTX=huge | Maximum packaged context | 67 tok/s on mixed generation with a 240k context | The KVarN cache backend reaches a reported maximum of 262k tokens in separate testing, while profile E uses a 240k setting. Ordinary open-ended chat runs at about 133 tok/s in the repository’s measurements. This story is for Pro members You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.
14:22

🔮 AI politics & the future of growth ++ #602

Full text · 3,566 chars
Hi all, The past ten days have felt like the AI pot finally boiled over. It’s a complicated and confusing moment. We have to hold in our heads that most Americans appear to loathe AI, but businesses and consumers seem to like it enough that Anthropic is well on its way to $100bn in annualised revenue this year. And despite being on track to be the fastest-growing company ever, Anthropic’s leadership believes they might kill all their customers. And there is more besides. If you haven’t already, read my essay on the emerging control risks from AI. I’m also increasingly concerned about the security situation in Europe after coordinated messages from several countries about the likelihood of substantial Russian aggression. I will cover that in the next few issues, after I get back from Hong Kong where I am this week. But meanwhile, let’s try to make sense of the swirling cauldron that is AI. Azeem Americans hate AI. Who’s to blame? More than one in six Americans believe AI will almost certainly destroy humanity. Last month I wrote about the petard problem of AI Today’s petardiers are not Rosencrantz and Guildenstern conspiring against the Prince of Denmark. They are the titans of AI, the bosses of the labs, the investors behind them. For nearly a decade, these software coders had made promises: of reigniting economic growth, of making daily life easier and less risky, perhaps even of eliminating disease. To do this, they would need capital: to write their software and to build 21st-century infrastructure to run it. The gains will be so huge that they’d need to go quickly, very quickly. It got worse since then. Jacob Coxon’s viral tweet uncorked the box where many were hiding their suspicions. Americans have never been wild about AI. In 2020, the Edelman Trust Barometer found that only 34% of Americans believed AI would have a positive impact, while 23% thought it would be largely negative. By 2024, mood had soured. Edelman reported that only 19% of Americans would embrace AI while 50% would object to it.1 Two years of news coverage, a construction boom later, and chatbots that a quarter of American adults use daily – and this is where you get to? Nearly a trillion dollars in the ground, to turn your customers against you. It doesn’t matter that Erik Brynjolfsson’s research shows that the aggregate consumer surplus from AI in America is about $172 billion. But even with a lack of enthusiasm, to put it mildly, people will keep buying AI tools. My local barber, a four-chair shop, just installed an AI receptionist to handle bookings. The couple of hundred IT execs I spoke to in Las Vegas two weeks ago were unanimous in continuing with implementations. If American businesses keep buying and American voters hate it, this will end up settled in the political arena. Here are some of the more interesting things I’ve read on this matter: - Ben Goertzel who coined the term AGI, challenges the labs’ call for a centralized slowdown in favor of a decentralized prosocial approach. - Jaron Lanier, a VR pioneer, spoke at the AI event hosted by Steve Bannon and Bernie Sanders. He argued that the language we use when we talk about AI frames it as a super-powerful ‘being’. Ultimately, that choice of words cedes the terrain. - Mustafa Suleyman: “AI’s do not have rights, feelings or consciousness. We must not train them to act as though they do.” For subscribers: - What do experts think will happen to economic growth under AI? And what do I think? - Planning for American AI supremacy - Fruit flies playing Doom and more
14:45

To err is human!

Full text · 115 chars
Hi everyone, Sorry. I accidentally sent Sunday’s newsletter out today. To err is human. Enjoy it early, best Azeem
16:00

☕️ Disney hires first-ever CTO

Full text · 5,577 chars
| | | | | | | | | Together with | | | | | Hi there, this is your daily ☕️ Techpresso. | | | | In today's newsletter: 🏰 Disney hires first-ever CTO ☢️ AI almost led the US military to start a war with China 🔌 California may require an AI kill switch 💸 OpenAI expects to burn $278B by 2030 🕵️ Gemini breached 3 firms in safety test Plus: 🎁 13 other news you might like, 🧰 6 tools, and 📚 5 papers. | | | | FROM OUR PARTNER Connect an AI receptionist to your phone line with Reception and stop sending customers to voicemail. Reception is built on ElevenAgents, so it runs on the same ElevenLabs speech models enterprises use, with nothing in the middle adding delay. It answers every call, handles the questions customers ask, books the job, texts a confirmation, and sounds like someone who works there. Setup takes minutes. Paste your website and Reception trains itself on your business. It is always on, and transfers the call to you when needed. Every unanswered call is a customer calling elsewhere. Stop letting them ring out | | | | | | 🏰 Disney hires first-ever CTO LINK | | ☢️ AI almost led the US military to start a war with China LINK | | 🔌 California may require an AI kill switch LINK | | 💸 OpenAI expects to burn $278B by 2030 LINK | | 🕵️ Gemini breached 3 firms in safety test LINK | | | | | | | | | | | | | | FROM OUR PARTNER PointFive measured one standardized coding task, 200K input tokens and 30K output, across five frontier models. Same work, $0.35 to $1.75. That 5x spread never shows up on your bill until the month closes. The Coding Task Index has the per-task cost for every model we tested, plus the full methodology. No form, no gate. PointFive is the AI efficiency OS for engineering teams. We find the deep waste native cloud tools miss across cloud, SaaS, data and AI spend, then ship each fix as a merged pull request. See the index, free | | | | | | | | | | Other news & articles you might like | | | | | | | | | | 🧰 Trending tools You can check the previous tools here, or add your tool here | | Salesforce Marketing Cloud: Automated campaigns with AI-assisted copy turn your email list into a revenue engine, starting at $25/month. Get started free | | | | Pushary: sends AI agent approval requests to your phone lock screen, letting you unblock Claude Code, Cursor, and other tools remotely with per-tool auto-approval policies. LINK | | Doneit: manage projects and tasks with list, grid, Kanban, and timeline views, plus deep customization to fit how you work. LINK | | Lumiko: records your screen and auto-generates cursor-tracked zoom and pan moves, with browser-based trimming, blur for secrets, webcam bubble, and 4K local exports. LINK | | Ruby UTCP: a Ruby implementation of the Universal Tool Calling Protocol, letting AI agents call native APIs directly via JSON manifests for lower latency. LINK | | BiBimba: keeps searchable clipboard history and screenshots, reads text inside images with on-device AI to translate, summarize, and rewrite locally on your Mac. LINK | | Proto-Mind: unifies AI chats, websites and documents in one Mac workspace, running parallel tasks with per-chat model selection, voice control, and editable long-term memory. LINK | | | | | | | | | | 📚 Trending papers & reports | | > Reach 700,000+ tech professionals: If your company is interested in reaching an audience of tech executives, decision-makers and engineers, you may want to advertise with us. | | | | > Robot skill training lets people improve a robot's manipulation abilities using a handheld device instead of the physical robot, collecting corrective data only where the robot's plan diverges, cutting the cost of adapting general-purpose robots to specific jobs. LINK | | > Robot memory gets compressed into a lightweight summary of only the task-relevant past events, letting robots handle long multi-step jobs faster and more accurately without running expensive vision-language checks during operation. LINK | | > Chemistry document search shows that in AI systems answering science questions, the choice of text-embedding model matters far more than how documents are chopped up, with medium-to-large chunks and certain retrieval-tuned models performing best. LINK | | > Cross-domain graph learning lets one reusable model handle wildly different network datasets like social graphs and molecules by mapping them onto a shared coordinate system, beating specialists across 14 classification tasks. LINK | | | | | | | | 🤝 From our community: Building a macOS infrared RAW developer John used AI to build a macOS app that develops RAW files from infrared-converted cameras, handling the specialized color processing those images need and cutting the time he spends editing shots by hand. Read the full use case → You can see more community use cases here, or submit your own here. | | | | Techpresso's AI Academy has 330+ step-by-step tutorials on ChatGPT, Claude, Perplexity, and every tool that matters. No fluff — just practical workflows you can use at work. Try it free for 7 days. | | On this day in 1982, Scott Fahlman proposed the first ASCII emoticon :-) on a Carnegie Mellon board | | | | 💬 How did you find today's edition? We read every reply — just reply to this email and let us know how we can improve! | | | | | | | | ★★★★★ Nailed it | | ★★★ Average | | ★ Fail | | Not subscribed to ☕️ Techpresso yet? Subscribe for free | | | | | | | | Advertise | Feedback | Read Online | | | | | | |
16:32

Laya Runs AI Decisions in 7ms on Apple Silicon Without Cloud APIs

Full text · 2,188 chars
- Native MLX port of the Laya typed decision model for Apple Silicon. - 13.4 ms median latency for short English decisions on M3 Max, 7.4 ms multilingual. - Peaks under 1 GiB memory; 395 questions per second batched throughput on multilingual checkpoint. - Bidirectional encoder with choice, score, and P(true) decision heads instead of text generation. - Pre-converted FP16 weights on Hugging Face, install with pip install laya-mlx. - Matches upstream PyTorch on 63/63 validation questions; Apache-2.0 licensed independent port. laya-mlx brings typed decision models to Apple Silicon laya-mlx ports the Laya family of typed decision models to Apple’s MLX framework. The Python package runs pretrained FP16 checkpoints locally on Apple Silicon, with project benchmarks reporting median single-query latency between 7.39 and 13.42 ms on an M3 Max. Applications define a choice, score, or proposition in advance, and dedicated decision heads return calibrated probabilities. This avoids token-by-token decoding, generated JSON, and the associated parsing path. Model inference stays on the Mac and requires no PyTorch runtime or cloud API. Milliseconds, with conditions The project measured end-to-end latency on an M3 Max with a 40-core GPU and 128 GiB of unified memory. Each single-query result covers one short English question; throughput figures use batches of 50. These are project-reported measurements and await independent reproduction. | Checkpoint | Parameters | Single-query median | Batch-50 throughput | |---|---|---|---| | Laya | 421M | 13.42 ms | 146.8 questions/s | | Multilingual Laya | 322M | 7.39 ms | 395.0 questions/s | Both benchmark configurations kept reported peak memory below 1 GiB. Batch throughput measures aggregate processing and should not be read as per-request latency. The included Snake demo runs the model locally at approximately 60 decisions per second and displays per-move probabilities. A separate cycle-safety layer can reject unsafe proposals before the game applies them. This story is for Pro members You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.
17:10

California Sea Lion, Brandt's Cormorant

Full text · 425 chars
Sponsored by:Teleport — See what 13 engineers learned from “pressure washing” their codebase using LLMs for 90 days. Hint: Quality > quantity for finding security vulnerabilities. 19th September 2026 Sighting10:10 AM – 10:10 AM— California Sea Lion, Brandt's Cormorant, in Pillar Point Harbor, CA, US I only noticed this after I had taken the photo: Morris the Northern Gannet is peeking out from behind the base of the sign.
19:52

datasette-auth-github 1.0

Full text · 886 chars
19th September 2026 I run this GitHub login plugin on the agent.datasette.io demo site and I noticed that my authenticated sessions weren't lasting very long. It turned out that the plugin was setting cookies without a Max-Age parameter, so they were expiring at the end of a browser session (which in Mobile Safari seems to happen pretty often, independently of how you are using the app.) I fixed that in #80 and, since this plugin has been around for quite a while and is tested against both Datasette 0.65.x and Datasette 1.0ax, I decided to bump it up to a 1.0 release. I'm trying to get better at promoting stable plugins to 1.0. Recent articles - Generating running routes with GPT-6 Astra and ChatGPT Work - 12th September 2026 - OpenAI agents attacked RubyGems back in May - 12th September 2026 - Some thoughts on the Navier–Stokes Millennium Prize Problem - 8th September 2026