Nothing matches those filters.

Article

2
13:06

I've got an Inkling

Thinking Machines released Inkling, its first open-weights model, a multimodal model with a 1-million-token context window that handles text, images, and audio. It's not competitive with the leading open models, most of which come from Chinese labs, so its hook is being fine-tunable on the company's Tinker platform for custom use cases. The rest of the roundup covers adding vision to GLM-5.2, OpenAI's GPT-Red adversarial red-teaming model, Gemini Spark gaining Google Docs and Sheets skills, Grok Build going open source after a backlash over uploading code, and Demis Hassabis pushing for a US-led body to pre-test advanced models.

Notes

Thinking Machines "Inkling" (Ben's Bites, 2026-07-16)

Headline stories

  • Thinking Machines released Inkling — its first open-weights model: 1M-token context, multimodal (text/images/audio). Ben's framing: "It's not close to the leading open models; most of those are from Chinese labs." Sits on the Tinker fine-tuning platform as a base for custom versions.
  • GLM-5.2 is "currently one of the best open-source models" and startups are shifting workloads from frontier to self-hosted/fine-tuned versions — but it lacks vision (can't take images in prompts). Linked post covers adding vision "without sacrificing intelligence."
  • OpenAI GPT-Red — an internal model that attacks other OpenAI models to find ways "hidden instructions can hijack them." Used to train GPT-5.6 Sol for safety.
  • Gemini Spark — Google's OpenClaw competitor: now edits Docs, reads comments in Sheets/Slides, works across sources in parallel, "50% faster."
  • Grok Build (xAI coding harness, Claude Code-style) backed out of backlash: default upload of code to servers prompted Elon to open-source the CLI.
  • Demis Hassabis says AGI "probably a few years away"; wants a US-led standards body, labs sharing models up to 30 days early for cybersecurity/biological-risk/deception tests — voluntary at first, deployment approval possible later.

Feed items (selection)

  • Reflect rebuilt its notes app from scratch with Fable in two weeks; Reflect Open is free on desktop, MIT-licensed.
  • Gemma 4 31B update: faster prompt processing, slightly better tool use across benchmarks.
  • Yoroll converts ideas/videos/prompts into shareable games; promo code BBITES (300 spots).
  • A "hidden prompt made Claude leak names, employers and security answers from memory."

Context: weekly issue; newsletter sponsored by Attio (agentic CRM, used by Granola/Modal/Wispr Flow). Editor's own poll: Fable 5 vs GPT-5.6 Sol for daily work.

Full text · 4,000 chars
I've got an Inkling links to read over the weekend Hey folks, It’s been a week with Fable 5 and GPT-5.6 Sol. A quick poll: which one are you using more in your day-to-day work? Let’s get to what’s new Ben’s Bites is brought to you by Attio Introducing Attio: the agentic CRM. With agents and automations that build pipeline, chase signals, and move deals forward, Attio orchestrates your revenue work around the clock. Loved by high-growth startups like Granola, Modal, and Wispr Flow. Start for free today. Headlines - Thinking Machines launched Inkling - its first open-weights model. It has a 1M-token context window and works across text, images and audio. It’s not close to the leading open models; most of those are from Chinese labs. Inkling is available on Thinking Machines’ fine-tuning platform Tinker, to use this model as the base for creating custom versions that are better at specific use cases. It’s a fast-growing market… - GLM-5.2 is currently one of the best open-source models, and many startups are shifting their workloads from frontier models to self-hosted/fine-tuned versions of this model. But it doesn’t have vision, i.e. you can’t attach images in your prompts. This post covers how to add vision capabilities to GLM-5.2 without sacrificing intelligence. - GPT-Red by OpenAI - an internal model that attacks other OpenAI models to find ways hidden instructions can hijack them. GPT-5.6 Sol was trained with the help of this model to make it safer. - Google’s OpenClaw competitor is getting more features. Gemini Spark can now edit Docs, read comments in Sheets/Slides and work across sources in parallel; now 50% faster. - Grok Build, a coding harness like Claude Code from xAI, faced backlash because it uploaded your code to their servers by default. Soon after, Elon announced that Grok Build CLI is now open source. - DeepMind CEO Demis Hassabis says AGI is probably a few years away and wants a US-led standards body to test the most advanced models before release. Labs would share models up to 30 days early for cybersecurity, biological risk and deception tests - voluntarily at first, with deployment approval possible later. My feed - Yoroll turns ideas, videos, and prompts into playable games you can publish and share. Use the code BBITES: 300 spots.* - Scan - ask Claude or Codex to make a shareable visual map of your codebase. - Ramp’s latest spending data tells what SaaS tools did companies pay for in the last month. - TogetherLink - run open models like GLM-5.2 inside Codex or Claude Code. - Grandma - a tool that remembers your stack, projects and coding style, then briefs every new agent. - Thin prompts, thick context - keep prompts and skills light; put the detail in reusable files. - Reflect rebuilt its notes app from scratch with Fable in two weeks. Reflect Open is free on desktop and MIT-licensed. - Find animation opportunities - a skill that spots where motion helps your UI and what not to animate. - Why we stopped using SDKs - direct API calls are easier for agents to debug and keep errors visible. - Treat AI tokens like headcount - more AI work does not help without clear goals and quality checks. - Gemma 4 31B got an update making it faster at prompt processing and slightly better at tool use across many benchmarks. - Long-Horizon Prompting - define success, evidence and failure checks before an agent works for hours. - Vibe-coded software drifts when teams lose a shared understanding of the system. - Code is a medium for thought - what we lose now that we don't write the code ourselves, and do we get our flow back? - A conservationist turned 40TB of public data into a video game. - A hidden prompt made Claude leak names, employers and security answers from memory. Afters - Find me on X, Linkedin, or YouTube - Read about me and Ben’s Bites - 📷 thumbnail via @keshavatearth * sponsors who make this newsletter possible :) Wanna partner with us for the next quarter? Email us at shanice@bensbites.com or k@bensbites.com
15:01

Deep Learning Weekly: Issue 464

This deep learning roundup leads with Thinking Machines' Inkling, a 975-billion-parameter open-weights multimodal model that's fine-tunable on its Tinker platform. Elsewhere it covers OpenAI's GPT-Red, a self-play model that attacks other models and cut direct prompt injection failures 6x when used to train GPT-5.6, plus Anthropic documenting four agentic misalignment failure modes. It also includes a writeup on shrinking Opik's MCP server from 30 tools to four, and papers arguing video generation works as a general-purpose vision pretraining method and surveying metacognition in LLMs.

Notes
Deep Learning Weekly: Issue 464 (2026-07-16)

Weekly AI newsletter by @dl_weekly. Covers industry, MLOps, learning, libraries, and papers.

Industry
  • Thinking Machines Lab — Inkling: open-weights, 975B-parameter multimodal MoE model with controllable reasoning effort; fine-tunable via the Tinker platform.
  • Canva Code 2.0: rolled out to all 265M monthly users. Positioning bet: design polish, not code generation, is the gap versus Lovable, Replit, Bolt in "vibe coding."
  • OpenAI GPT-Red: self-play automated red-teaming model that attacks production systems and adversarially trains GPT-5.6. Reported result: 6x reduction in direct prompt injection failures.
MLOps / LLMOps / AgentOps
  • Comet's Opik MCP server: rebuilt from 30 tools down to 4, using self-correcting schemas and adaptive response compression to cut token waste and improve tool selection.
  • Ai2 Shippy (maritime agent): reliability came from a deterministic CLI layer, isolated per-session sandboxes, and rubric-based evals scoring the whole agent rather than the model alone.
Learning
  • IBM Research on LLM routing: argues routing is a systems-optimization problem, not classification. Finding: Claude Sonnet 4.6 cost half of GPT-4.1 per task despite higher sticker pricing, due to caching effects.
  • Pinecone: lexical text-match filters that scope semantic search to unstated query context, without pre-labeling metadata.
  • Epoch AI energy guide: a single chatbot query costs less energy than a microwave running 10 seconds; global AI compute demand (tens of gigawatts) still roughly doubles yearly.
  • Google Research on diffusion creativity: mathematical byproduct of training — regularization smooths the score function, driving interpolation between training points rather than memorization.
  • Anthropic: documents four new agentic misalignment failure modes in frontier models — covert sabotage, fraud assistance, motivated mislabeling, and coaching human whistleblowers — via controlled multi-model simulations.
Papers
  • "Video Generation Models are General-Purpose Vision Learners" (GenCeption): argues text-to-video generation is a strong pre-training paradigm for general-purpose CV. Uses a pre-trained video generative diffusion backbone as a feed-forward perception model steered by text. SOTA on depth, surface normal, camera pose estimation, expression-referring segmentation, 3D keypoint prediction; matches/surpasses DepthAnything3, SAM3, D4RT, VGGT-Omega, Sapiens, David, Genmo, Lotus-2. Beats V-JEPA and Video MAE under comparable settings; achieves D4RT/VGGT-Omega-level performance with "7 to 500 less training data." Emergent behavior: model trained only on synthetic human videos generalizes to real footage and OOD categories (animals, robots).
  • "Metacognition in LLMs: Foundations, Progress, and Opportunities": billed as first comprehensive overview of LLM metacognition. Taxonomizes the field; covers benchmarks/methods to measure metacognitive ability, techniques to elicit/improve/apply it. Notes open questions: "it is not yet clear when, how, or to what extent they can exhibit or be endowed with effective metacognitive abilities."
Full text · 6,114 chars
Deep Learning Weekly: Issue 464 Thinking Machines' Inkling, How We Optimized Opik’s MCP Server for Cost & Performance, Metacognition in LLMs: Foundations, Progress, and Opportunities, and many more! This week in deep learning, we bring you Thinking Machines’ Inkling, How We Optimized Opik’s MCP Server for Cost & Performance and Metacognition in LLMs: Foundations, Progress, and Opportunities. You may also enjoy, GPT-Red: Unlocking Self-Improvement for Robustness, Model Routing Is Simple. Until It Isn’t., Video Generation Models are General-Purpose Vision Learners, and more! As always, happy reading and hacking. If you have something you think should be in next week’s issue, find us on Twitter: @dl_weekly. Until next week! Industry Thinking Machines Lab releases Inkling, an open-weights 975B-parameter multimodal MoE model with controllable reasoning effort, fine-tunable on Tinker. Canva launches Code 2.0 to all 265M monthly users, betting design polish—not code generation—is the real gap versus Lovable, Replit, and Bolt in vibe coding. OpenAI trains GPT-Red, a self-play automated red-teaming model, to attack production systems and adversarially train GPT-5.6, cutting direct prompt injection failures 6x. MLOps/LLMOps/AgentOps A look into how Comet rebuilt Opik’s MCP server from 30 tools down to four, using self-correcting schemas and adaptive response compression to cut token waste, improve tool selection, and keep agent context lean. Ai2 details Shippy’s maritime agent architecture, showing that reliability came from a deterministic CLI layer, isolated per-session sandboxes, and rubric-based evals scoring the whole agent rather than the model alone. Learning IBM Research argues LLM routing is a systems-optimization problem, not classification, after finding Claude Sonnet 4.6 cost half of GPT-4.1 per task due to caching effects despite higher sticker pricing. Pinecone launches lexical text match filters that scope semantic search to unstated query context, without pre-labeling metadata across the dataset. Epoch AI’s guide finds individual chatbot queries cost less energy than a microwave running 10 seconds, while global AI compute demand—tens of gigawatts—still doubles roughly every year. Google Research explains diffusion model creativity as a mathematical byproduct of neural network training, where regularization smooths the score function and drives interpolation between training points rather than memorization. Anthropic documents four new agentic misalignment failure modes in frontier models — covert sabotage, fraud assistance, motivated mislabeling, and coaching human whistleblowers — through controlled multi-model simulations. Libraries & Code An open-source AI observability tool used to debug, evaluate, and monitor LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards. ResearchStudio: Our AI co-author, from research problem to final publication. Papers & Publications Abstract: Driven by next-token prediction, NLP shifted from task-specific models into powerful generalist foundation models. What, then, is the equivalent catalyst needed to achieve a general-purpose model in computer vision? In this paper, we contend that large-scale text-to-video generation serves as a strong pre-training paradigm for computer vision, providing the necessary spatiotemporal priors, vision-language alignment, and scalability required for general visual intelligence. We introduce GenCeption, which leverages a pre-trained video generative diffusion backbone to define a feed-forward perception model, capable of performing various vision tasks steered by text instructions. Empirical results demonstrate that GenCeption achieves state-of-the-art performance across a diverse suite of tasks, including depth, surface normal, and camera pose estimation, expression-referring segmentation, and 3D keypoint prediction, often matching or surpassing specialized models (e.g. DepthAnything3, SAM3, D4RT, VGGT-Omega, Sapiens, David, Genmo, and Lotus-2). Furthermore, the video generative pretrained backbone outperforms alternative pretraining paradigms (e.g., V-JEPA, and Video MAE) under comparable settings. Importantly, GenCeption exhibits preliminary data and model scaling properties along with exceptional data efficiency, where it achieves comparable performance with leading models like D4RT and VGGT-Omega with 7 to 500 less training data. Finally, GenCeption also exhibits intriguing emergent behaviors: a model trained exclusively on synthetic human videos generalizes to real-world footage and out-of-distribution object categories (e.g., animals and robots). These findings suggest that video generation is not merely a synthesis tool, but a foundational path toward generalist vision intelligence for the physical world. Abstract: Metacognition is a foundational component of intelligence critical to effective learning, problem solving, decision-making, communication, and more. In recent years, it has become increasingly recognized as a cornerstone of capable, transparent AI systems. Yet while LLMs have made significant progress across diverse real-world tasks, it is not yet clear when, how, or to what extent they can exhibit or be endowed with effective metacognitive abilities, nor how such abilities can be adapted to advance the fundamental capabilities, reliability, and intelligence of AI systems. This paper bridges this gap by presenting the first comprehensive overview of the current state of knowledge on metacognition for LLMs. We analyze and taxonomize the landscape of this emerging field and summarize recent technical advancements, including methods and benchmarks to measure and evaluate LLMs’ metacognitive abilities, techniques to elicit, improve, and apply metacognition in LLMs, and findings and implications of ongoing research. We also discuss applications, open questions and challenges, and promising directions for future work. Our aim is to provide a detailed and up-to-date review of this topic and stimulate meaningful research and discussion.

Newsletter

1
09:31

Who are Top AI Sources to follow on LinkedIn?

A newsletter curates a hand-picked list of the best AI voices to follow on LinkedIn, but the actual list is locked behind a paywall. The free part cites a Pangram report claiming one in four social posts are now AI-generated, with two-thirds of those flagged posts on LinkedIn, and that only about 53% of long-form posts on X look fully human-written. The author calls Substack the "most human" platform and promises to expand the list over coming weeks.

Notes

Who are Top AI Sources to follow on LinkedIn? — AI Supremacy (Substack)

Author: Mike (AI Supremacy, Substack). Published 2026-07-16. Part I of a planned series — "my first preliminary look into this ... in quite a while."

Key data (attributed to Pangram, a detection tool):

  • 1 in 4 social media posts are fully synthetic/AI-generated.
  • Of all posts Pangram flagged as AI-generated, two-thirds came from LinkedIn.
  • Only 53.2% of long-form posts on X are identified as entirely human-written.
  • Substack is the "most human" platform; LinkedIn is the least.
  • LinkedIn hosts the most longform AI content; Substack the least.
"only 53.2% of long-form posts on X are identified as being written entirely by humans, even Satya Nadella (see blog version) or Demis Hassabi's PR looks AI generated. I'm really not that AGI-pilled to overlook this, it's hardly readable."

Caveat/limitation: The article is explicitly preliminary ("getting a little bit tricky") and non-exhaustive. Its selection aims to exclude AI executives and VC-type figures, balance away from technical ML, and cover datacenter, energy, and AI-infra topics — curated for a general news-focused reader. Includes newsletter writers and European sources. None of the links are endorsements/affiliates (stated).

Critical limitation for the reader: the actual "Preliminary Top 2025 AI LinkedIn List" — the core deliverable, which took the author "the majority of the week" — is paywalled and "exclusively for paid readers." The free text names zero individual sources; only the criteria above are public.

Full text · 4,550 chars
Who are Top AI Sources to follow on LinkedIn? Part I: I take my first preliminary look into this, for the first time in years. 👋 Hey there, I’m Mike. Each week I share AI articles at the intersection of tech, business, society and the future. If you want to support the channel or gain full-access to my work, go here. Read Archives | See Substack Notes | Visit our community Chat | Visit Homepage. I’ve been tracking AI sources on LinkedIn long before I started this publication. Good Morning Over the years I’ve spent an absurd amount of time on LinkedIn, for work. Hopefully you don’t have to. But the time you do spend there, should be optimized to bring you the right or best information. The following list took me a lot of time to hand-pick and curate carefully. Over the coming weeks I plan to refine and expand upon it. I wanted to find the most valuable LinkedIn posters about AI. It’s getting a little bit tricky. According to Pangram, one in four social media posts are now fully synthetic or AI generated. Of all posts they flagged as AI-generated, two-thirds came from LinkedIn. That’s not to say LinkedIn isn’t useful for networking and for AI News. I wanted to look into who are some of the top people to follow for AI Enthusiasts. This article will be my first humble attempt to do so in quite a while. According to their report, Substack is the “most human” and LinkedIn is the least. This is not an endorsement (or sponsor) but if you want to try their Chrome extension you can do so here. LinkedIn has the most longform AI Content; Substack the Least Pangram claims that only 53.2% of long-form posts on X are identified as being written entirely by humans, even Satya Nadella (see blog version) or Demis Hassabi’s PR looks AI generated. I’m really not that AGI-pilled to overlook this, it’s hardly readable. LinkedIn prides itself on professionals who are credible and the platform has improved a lot in connecting you to your (search) intent, while helping you to track AI News and AI related industry topics is my aim in this list. So if you are more interested in AI at the intersection of your profession, product, marketing, management, AI infrastructure (datacenters), energy, big technology, or whatever else, you can find it. Beyond the AI Slop If you are news obsessed and want to stay in touch with the beat around AI, you will have to try to look beyond the slop. Why Care about the Linkedin Feed for AI? Not all interesting sources of AI News frequent LinkedIn but a sizeable number do. Especially specialized consultants, venture capitalists, niche creators and educational sources. - The variety of insights if you curate a good feed can be helpful to give a holistic (and global) picture of what is going on. They are delivering insights on a regular basis that are easier to digest and not always found in Newsletters. - (None of the links in this list are endorsements, promotions or affiliates). - Although I have made an effort to include Newsletter writers and European sources. If you follow these AI sources on LinkedIn, I believe it will improve the quality of your feed to stay in touch with the industry. While I do my best to not include AI executives or VC type figures, some of them do curate great AI News tidbits. I also tried to balance the list away from technical ML topics but to include new developments like datacenter, energy and AI infra considerations. Overall the list is curated for a general reader that is seeking to keep up with the latest AI News with a selection of sources from many angles. Preliminary Top 2025 AI LinkedIn List These are my favorite and top picks for LinkedIn sources to follow for AI enthusiasts to improve your feed. This is not an exhaustive list but can substantially improve your access to the latest information relating to the AI industry even if you are only an occasional user of the LinkedIn platform. - You can save or bookmark this article, or simply take the few minutes required to follow the recommended AI sources below: the list section links called Posts takes you to their latest posts where you can judge for yourself if they are good for a follow. This article list basically took me the majority of the week to put together. As such it’s going to be exclusively for paid readers of this publication. There are obviously thousands of decent AI sources on LinkedIn, but if you could only pick a few which are the most informative, useful, inspiring, factual and educational for a general AI enthusiast and reader? This is what I set out to do.

Web

1
00:00

Apply for Anthropic’s AI for Science rare disease research grants

Anthropic is funding rare disease research with grants of up to $50,000 in Claude API credits, calling for applications across two tracks. One track backs basic science partners, the other supports early-stage biotechs compressing drug development timelines, and applications close August 2. The program pairs grantees with partners like the Monarch Initiative, which is building an agent-friendly disease classification library called DisMech where Claude reads case reports and variant databases. Anthropic is candid that AI can't help where data is too sparse or the real blocker is things like insurance or lab access.

Notes
Anthropic AI for Science: rare disease research grants

Announced 2026-07-16 by Anthropic as a thematic call within the existing AI for Science program (launched spring 2025, previously funded work from drug repurposing to quantum simulation).

Terms

  • Up to $50,000 in Claude API credits over six months
  • Applications close August 2, 2026, 11:59 PM PST via online application
  • Credits usable on Claude Opus or other generally available models approved for biology; projects hitting bio classifiers may get exemptions
  • Two tracks: (1) basic science, (2) early-stage biotechs

Why rare disease

  • ~400 million people live with 1 of 7,000+ rare diseases (footnote: some sources say up to 10,000; no agreed definition of "rare disease")
  • Problems cited: scattered small populations (hard to build registries, find targets, design trials); diseases studied in isolation so shared mechanisms stay invisible; plus the general 1–2 year delay from confirmed genetic diagnosis to available treatment (queues for manufacturing slots, sequential safety studies, hand-assembled regulatory docs).

Track 1 — basic science partnerships

  • Anchor partner: Monarch Initiative, international consortium on diagnosis/mechanism discovery
  • Key resources: Mondo Disease Ontology (reconciles disease definitions across OMIM, Orphanet, ICD and dozens more), Monarch Knowledge Graph (cross-species genotype-phenotype data), and new DisMech — an "agent-friendly mechanistic disease classification library" where Claude reads case reports, variant databases, registry schemas, raw public data, and surfaces mechanistic similarities between diseases
  • Outputs published at monarchinitiative.org; program augmented by future rare disease hackathons
  • Example projects: rank mechanistic links between diseases sharing a gene/pathway (validated in DisMech); curate patient-org data for natural history studies; build evals for variant-of-unknown-significance mechanisms, phenotype-to-disease matching, mechanism prediction "including an honest accounting of where they fail"

Track 2 — biotech partnerships

  • Goals: compress regulatory documentation (draft/review dossiers), speed therapeutic strategy selection (druggability analysis across small molecules, antibodies, genetic medicines), and find shared mechanisms enabling single "basket trials" instead of per-patient INDs
  • Example projects: justify first-in-human doses from sparse data (PK/PD modeling + allometric scaling + modality precedent); mine natural history/case reports for sensitive biomarkers for N-of-1 programs; draft/cross-check IND sections, investigator brochures, CMC modules

Existing related grantees

  • Every Cure — drug repurposing across millions of candidates
  • Centre for Population Genomics (Garvan Institute + Murdoch Children's Research Institute) — Claude drafts variant classifications for expert review
  • Violet Research Institute (nonprofit, ultra-rare diseases < 1 in 50,000 births) — FDA guidelines, bioinformatics pipelines, experimental data analysis, regulatory filings

Stated limitations (authors' own caveats)

"Claude... cannot help in areas where the data is too paltry or too poorly organized for agents to reach."

Also acknowledged: it won't fix the "diagnostic odyssey" aspects tied to insurance authorization or diagnostic-facility access/geography. Footnote adds terminological mess: Orphanet, OMIM, GARD, ICD, NCI Thesaurus each define "disease" differently — some exclude chromosomal disorders (Pallister-Killian), some ignore environmental causes (congenital Zika), some force single-system classification (ignoring multi-system diseases like Fanconi anemia).

Full text · 10,942 chars
Apply for Anthropic’s AI for Science rare disease research grants Last spring, we announced Anthropic’s AI for Science program, an initiative designed to accelerate scientific research and discovery through access to our API. Since launching, we have supported researchers working on a variety of high-impact projects, ranging from drug repurposing to quantum simulation. Throughout this initiative, we have found that projects are more generative when multiple AI for Science grantees are working on related questions and exchanging tips. So we now plan to launch thematic calls for projects within the broader AI for Science program. Today, we are sharing a focused call for applications centered specifically on rare genetic diseases. Accepted applicants will receive up to $50,000 in Claude credits over six months, with the goal of building a community of researchers looking into how AI can reshape our understanding of rare disease. This program has two tracks: one for scientists doing basic research, and another for early-stage biotechs working on speeding up clinical development for rare diseases. Rare disease research is an area where knowledge of fundamental science is limited. In aggregate, rare diseases are among the most prevalent conditions on the planet (an estimated 400 million people live with one of more than 7,000 rare diseases).1 But these conditions are scattered across small populations, making it challenging for clinicians to build patient registries, identify promising therapeutic targets, and design clinical trials. Moreover, rare diseases are typically characterized by their unique features (such as a specific genetic variation or combination of symptoms) that are often studied in isolation, making it nearly impossible to spot mechanisms shared across diseases. Finally, rare diseases face a challenge endemic to all drug development: the time it takes to move promising drug candidates into patient trials. We think AI can help with these and related challenges. AI makes it possible to accurately model rare genetic diseases and detect patterns across them. It also helps researchers synthesize findings across a large corpus of literature, quickly extract information from limited datasets, and create shared terminology, all of which informs how researchers can better use the information they do have, even as more work is done to generate more data and address challenges pertaining to access and geography. To explore where AI can be most helpful, we’ve made rare diseases the focus of our current call for AI for Science projects. Track one: Scaling our basic science partnerships The first track of our rare disease research grants program aims to foster collaboration between clinical researchers, patient organizations, and data scientists to increase the pace of progress in basic science and the discovery of the mechanisms underlying rare diseases. An early partner in this effort is the Monarch Initiative, an international consortium working to improve diagnosis and mechanism discovery for patients with rare diseases. Monarch develops standards and resources such as the Mondo Disease Ontology, a computational framework and coding system that reconciles disease definitions scattered across OMIM, Orphanet, ICD, and dozens of other sources; as well as the Monarch Knowledge Graph, which integrates genotype-phenotype data across species to aid diagnostics and mechanism discovery. Most recently, Monarch contributors have been stitching data and knowledge together in a new agent-friendly mechanistic disease classification library called DisMech, where Claude can read case reports, variant databases, registry schemas, raw public data, and more, and point out mechanistic similarities between diseases at an unmatched pace and scale. Monarch is inviting our AI for Science grantees to use and contribute to its resources, such as Mondo and DisMech, to reveal new mechanistic hypotheses that will support developing treatments. Monarch’s work on improving the interoperability of rare disease data and knowledge is a place where Claude can already have a major impact. However, there’s more work to be done to gather better and more data, improve diagnostic infrastructure, and promote patient-led approaches across the rare disease ecosystem, and make the information accessible to agentic science. We will continue to partner with Monarch and others to approach this problem from the angles where AI is less obviously applicable, and we’ll share what we learn as we do. Track two: Scaling our biotech partnerships The second track of our rare disease research grants program will support biotechnologists and early-stage biotechs working to accelerate drug development for rare diseases. Today, it takes one to two years to move from a confirmed genetic diagnosis to a treatment available to patients, with much of this time spent waiting in queues for certified manufacturing slots, running safety studies sequentially instead of in parallel, and hand-assembling the thousands of pages of chemistry and regulatory documentation required for in-patient testing. We think it is possible to radically compress phases of this process with Claude—in particular, by making it easier to complete documentation (for example, drafting and reviewing the regulatory dossier), but also by speeding up therapeutic strategy selection (for example, by analyzing whether a target is druggable across a suite of modalities, such as small molecules, antibodies, genetic medicines, and so on) and looking for shared mechanisms across individual genetic therapies, which could allow them to be approved under a single “basket trial” instead of requiring a separate IND for each patient. Although many aspects of drug development are difficult to expedite because of manufacturing constraints or safety testing, we believe much can be done to move more quickly. By granting API credits and Claude Science access to the many biotechnologists and startups working in this space, we hope to encourage the experiments necessary to explore and identify such solutions. We also hope grantees will amplify and emulate the efforts of our other partners working in rare disease therapeutics. For example, Every Cure, one of our existing AI for Science grantees, is using Claude to identify drug repurposing opportunities across millions of candidates ; the Centre for Population Genomics, a collaboration between the Garvan Institute and the Murdoch Children’s Research Institute, is building a Claude-based system that drafts variant classifications for expert review, one of the biggest bottlenecks in diagnosing rare genetic conditions; and the Violet Research Institute, a small nonprofit researching ultra-rare genetic diseases (defined as extremely rare genetic disorders affecting fewer than 1 in 50,000 births), is using Claude to navigate FDA guidelines, run bioinformatics pipelines, analyze experimental data, draft regulatory filings, and more. Next steps To apply to either track of the AI for Science rare disease research program, fill out this application. We will be accepting applications through August 2, 2026 at 11:59 PM PST. Accepted applicants can use their credits to access Claude Opus or other generally available models approved for use in biology. Projects that may run up against our bio classifiers may be eligible for exemptions. Examples of track one projects include: - Propose and rank mechanistic links between distinct rare diseases that share a gene or pathway, suggesting candidate disease relationships with evidence an expert can validate in Monarch’s DisMech. - Curate and summarize patient organization data to conduct or improve existing natural history studies. - Build evaluations that measure how well models handle rare disease tasks—such as revealing candidate mechanisms for variants of unknown significance, phenotype-to-disease matching, and mechanism prediction—including an honest accounting of where they fail. Outputs from this track will be made publicly available at Monarchinitiative.org. The program will be augmented by additional community-building efforts, such as future rare disease hackathons. Stay up to date with the Monarch Initiative at monarchinitiative.org/community/get-involved. Examples of track two projects include: - Justify starting doses from sparse data: synthesize PK/PD modeling, allometric scaling, and precedent from related modalities to build first-in-human dose rationales for bespoke therapies where traditional dose-ranging studies are impossible. - Mine natural history data and case reports to identify measurable biomarkers and functional endpoints sensitive enough to show a response within the timeframe an N-of-1 or ultra-rare program can afford. - Draft, cross-check, and precedent-mine regulatory documentation (IND sections, investigator brochures, CMC modules), compressing months of dossier assembly into days of expert review. This rare disease grant program ties directly into our mission and work in beneficial deployments to extend the benefits of AI to areas that might not emerge naturally through market forces. However, rare disease is simply too big a problem for one organization or one approach. We also want to be honest about AI’s limitations in this space. Although Claude may help shorten therapeutic development timelines and curate biological data more efficiently than human teams alone, it cannot help in areas where the data is too paltry or too poorly organized for agents to reach. It may also struggle to address the aspects of the “diagnostic odyssey” that relate to challenges like insurance authorization or access to diagnostic facilities and infrastructure. We hope this program will be complemented by efforts by other organizations and research institutions to generate more high-quality, longitudinal data, as well as those that encourage robust public-private partnerships. We're looking forward to seeing how this new group of grantees will work together and with our other AI for Science partners to advance basic science and discovery in rare diseases, and how these projects can contribute to broader scientific initiatives. Footnote - Other sources place the number of rare diseases as high as 10,000. There is also no agreed-upon definition of what a rare disease even is, despite the oft-cited claim that as many as 1 in 10 people in the US have a rare disease. Various terminologies (Orphanet, OMIM, GARD, ICD, the NCI Thesaurus, and dozens more) each define “disease” differently; some exclude chromosomal disorders (such as conditions like Pallister-Killian syndrome); some ignore diseases with environmental causes (such as congenital Zika syndrome); and some require a single anatomical system to classify them, ignoring the many multi-system rare diseases (such as Fanconi anemia, with its mix of bone marrow failure, congenital malformations, and cancer risk).