Elman

Mission

Our thesis

Predictable biology will be won by whoever encodes the scientific judgement that no model can learn from data. Everywhere else in AI the opposite has held, with more data beating hand-built expertise across vision, language, code, game play and even protein structure, so the claim needs explaining. The pattern that won those domains rests on conditions therapeutic development cannot supply: it needs a cheap way to generate near-unlimited labelled experience, a simulator or a game an agent can play against itself billions of times, it needs a reset so that a failure costs nothing and the next attempt is free, and above both it needs that experience to be drawn from the same space as the question rather than from a cheaper proxy for it. Therapeutic development denies the first two outright, because no simulator is faithful to a living human and building one would mean solving the biology first, and because when a prediction fails in the clinic a person dies rather than resetting, and it starves the third for every question that matters, because representative human outcomes are far too sparse to train on and what is left beneath them is mostly adjacent, a mouse that is not a human, a cell measured outside its niche, an organoid that is neither.

The data that would let scale win was therefore never collected at the breadth the problem needs and for almost all questions it can never be. A representative sample would mean measuring every biological layer, DNA, RNA, protein, cell, tissue, organ and whole system, inside living humans at once and without perturbing them, which is impossible. What exists instead is sparse evidence scattered across those layers, each source unfaithful in its own way, a cell measured outside its niche, a mouse that is not a human, an organoid that is neither. Roughly 90% of drugs that enter clinical trials never reach approval[1]. The causes are many, from efficacy to safety to commercial choice, but our thesis is that evidence this sparse and this unfaithful is a first-order reason the problem has resisted scale.

The winning move is a system that reasons across the unfaithful layers and judges how far each piece of evidence can be trusted on the way to the human answer. That system is the Judgement Engine, a portable, model-agnostic layer that encodes how an expert scientist reasons and prunes evidence, the judgement that has never been written down in a form a model can train on. It extracts and structures evidence before it reasons, separating fact from the bias baked into a model and the bias carried in the prompt and input. It lays that evidence onto a provenance-tracked causal graph across the biological layers. It reasons across that graph the way a scientist reasons over thin data, carrying full provenance and a confidence on every link, pruning to the low-magnitude signals that actually decide whether a drug works.

This is where durable value sits and the frontier labs cannot reach it by doing what they do best. Their weapon is scale on data, and the data here, how good scientists reason, does not exist in capturable form, because expert judgement is not written down and is not revealed to a chatbot. To get it they would have to build a judgement layer of their own, which is to say build this. The Judgement Engine runs on top of the frontier models rather than inside them, so it is model-agnostic, we already move between Anthropic, OpenAI and Google by subtask, and a better base model is one we drop in rather than one that obsoletes us.

With the Allen Institute, the engine has already produced a result of exactly the kind this thesis describes. Their multi-region Alzheimer’s atlas shows one population of visual-cortex neurons dying while its near neighbours survive, and from that atlas the engine proposed why those specific cells are the ones lost, that they carry a molecular profile leaving them chronically over-excited, a hypothesis it generated from the single-cell data and traced back to the evidence, now in a paper submitted to Cell with the preprint public[9].

Beyond that result, the proof we have committed to is the one most of the field avoids because it is the most difficult to nail, a clinical trial outcome prediction benchmark. A clinical trial is the one place biology gives a straight answer, so the fairest test of whether the engine predicts a human outcome, rather than recognising a drug it has already read about is a benchmark scored on trials. It asks the engine to put a trustworthy number on a specific trial failing and to name the biological reason it fails, the kind a sponsor could act on, from evidence frozen before the trial read out. Building the benchmark itself is difficult, the outcome can leak both from a model’s memory of the result and from fetching it online, so the plan is to screen out any trial a model already recognises (leaked into the training data) and date-gate the retrieval to before the readout.

Sources

[1] Hay M, Thomas DW, Craighead JL, Economides C, Rosenthal J. Clinical development success rates for investigational drugs. Nature Biotechnology, 32(1):40–51, 2014. Likelihood of approval from Phase 1 ≈ 10%, i.e. roughly 90% of drugs entering trials never reach approval.

[9] Elman × Allen Institute. Collaboration preprint (mechanism of Alzheimer’s-vulnerable neurons). bioRxiv, 2026. Submitted to Cell.