Elman

Thesis

Our thesis

Therapeutic development will become predictable when we can learn which evidence actually translates to humans. The missing training data is the link between evidence available at a decision, the judgement it justified, and what happened next. Elman is building that dataset and the model that learns from it.

Everywhere else in AI, scale has won: more data beating hand-built expertise across vision, language, code, game play and even protein structure. The pattern that won those domains rests on conditions therapeutic development cannot supply: it needs a cheap way to generate near-unlimited labelled experience, a simulator or a game an agent can play against itself billions of times, it needs a reset so that a failure costs nothing and the next attempt is free, and above both it needs that experience to be drawn from the same space as the question rather than from a cheaper proxy for it. Therapeutic development denies the first two outright, because no simulator is faithful to a living human and building one would mean solving the biology first, and because when a prediction fails in the clinic a person dies rather than resetting, and it starves the third for every question that matters, because representative human outcomes are far too sparse to train on and what is left beneath them is mostly adjacent, a mouse that is not a human, a cell measured outside its niche, an organoid that is neither.

The data that would let scale win was therefore never collected at the breadth the problem needs and for almost all questions it can never be. A representative sample would mean measuring every biological layer, DNA, RNA, protein, cell, tissue, organ and whole system, inside living humans at once and without perturbing them, which is impossible. What exists instead is sparse evidence scattered across those layers, each source unfaithful in its own way, a cell measured outside its niche, a mouse that is not a human, an organoid that is neither. Roughly 90% of drugs that enter clinical trials never reach approval[1]. The causes are many, from efficacy to safety to commercial choice, but the deeper problem is not that these proxies are useless. It is that their faithfulness to the eventual human answer has never been measured systematically, so the field has accumulated enormous amounts of biological evidence without accumulating the labels that say how far any particular kind of evidence should carry.

The winning move is therefore a system that reasons across the unfaithful layers and makes that translation explicit. That system is the Judgement Engine. It does not wait for a scientist to hand it a perfectly curated evidence package: its agents identify the gaps in an evidential argument, search the literature, fetch relevant datasets, run analyses and specialist tools, and incorporate new experimental results as they become available. It extracts and structures evidence before it reasons, separating the experimental fact from the bias baked into a model and the bias carried in the prompt and input. Each result is placed into a persistent, provenance-tracked evidence graph carrying what was tested, in which biological context, under which intervention or comparison, with what effect, statistics and source. The engine can then explore competing explanations, validate the weak links between them and make an explicit judgement about what the evidence actually justifies, where the argument breaks and what would change the answer.

That representation matters because it preserves something therapeutic development normally throws away: the state of the evidence at the moment a judgement is made. The graph records not only what was eventually known but what was knowable then, and the judgement records how far that evidence was allowed to carry. Link that state to the experiment, programme decision or clinical outcome that follows and the problem changes shape. What was once a one-off scientific judgement becomes a labelled example of evidence, judgement and reality.

This is where the new training signal comes from. Better frontier models will make literature search, data analysis, coding, hypothesis generation and scientific agents better, and the Judgement Engine is deliberately model-agnostic so that we can route between Anthropic, OpenAI, Google and specialist models as they improve. Their progress makes Elman stronger. What it does not do is make the biological layers line up. Better prediction of a molecular state does not tell you how faithfully that state carries into a cell, a tissue, an organism or a patient, and scaling on more mice, more cells or more papers does not manufacture the missing human outcomes against which those proxies should have been calibrated. Nobody can train their way to representative human experience that was never generated.

The distinction is therefore not that frontier labs cannot build agents, graphs, provenance or scientific workflows; they plainly can. It is that scaling a model on the biological data that already exists does not, by itself, produce the supervision the problem needs. That supervision is the history of evidence available at a real decision → the judgement justified by that evidence → what happened afterwards. Creating it means deploying against consequential therapeutic questions, preserving the evidence state before the answer is known, and collecting the downstream experimental, programme and clinical result. Elman is built so that every real use can generate another example of exactly that form.

With the Allen Institute, the engine has already produced a result of the kind the first half of this thesis requires. Their multi-region Alzheimer’s atlas showed an unexpected human observation: one population of visual-cortex neurons is selectively vulnerable while its near neighbours survive. The atlas showed what was happening; it did not contain the experiments needed to explain why. Elman went beyond the atlas, generated competing explanations and tested them against experimental results gathered from the scientific literature, datasets and complementary analyses. In the resulting workflow, ~170,000 experimental facts were structured into a common evidential representation and 225 hypotheses were investigated. 51% of hypotheses generated by the initial state-of-the-art multi-agent workflow were disproven by Elman validation on the same evidence, with 95% agreement with expert review. The surviving evidence converged on hyperexcitability, with receptor balance, calcium handling and ion-channel activity providing independent mechanistic routes to a specific physiological prediction, now in a paper submitted to Cell with the preprint public[2] and the findings inspectable here.

Beyond that result, the proof we have committed to is the one most of the field avoids because it is the most difficult to nail, a clinical trial outcome prediction benchmark. A clinical trial is the one place biology gives a straight answer, so the fairest test of whether the engine predicts a human outcome, rather than recognising a drug it has already read about, is a benchmark scored on trials. It asks the engine to put a trustworthy prediction on a specific trial outcome and to name the biological reason behind it, the kind a sponsor could act on, from evidence frozen before the trial read out.

Building the benchmark itself is difficult. The outcome can leak both from a model’s memory of the result and from fetching it online, so the plan is to screen out any trial a model already recognises — leaked into the training data — and date-gate retrieval to before the readout. Each benchmark case therefore starts from a time-locked evidence graph containing only what could have been known at the time, asks Elman what that evidence justifies, and reveals the outcome only afterwards. Held-out historical cases test whether the system generalises beyond the examples it has seen; prospective readouts test the same thing without hindsight at all.

The objective, however, is not one universal accuracy number. Faithfulness is conditional. Evidence can carry differently by therapeutic area, modality, biological context, evidence regime and decision type, and a system that is excellent at judging one combination may be materially weaker at another. The benchmark is therefore designed to produce an accuracy map: where does Elman predict therapeutic translation well, where does it not, and what kinds of evidence have to be present before a judgement deserves confidence?

Those outcome-linked cases then become more than a benchmark. They become the training data for models that learn which patterns in the evidence graph predict successful translation. Clinical trials provide the most consequential human labels, but they are not the only ones: wet-lab experiments and programme outcomes return faster, narrower answers along the way. Over time, every real deployment can add another example of evidence at decision time → judgement → outcome, turning the act of using Elman into the process that builds the dataset therapeutic development has never systematically collected.

Sources

[1] Hay M, Thomas DW, Craighead JL, Economides C, Rosenthal J. Clinical development success rates for investigational drugs. Nature Biotechnology, 32(1):40–51, 2014. Likelihood of approval from Phase 1 ≈ 10%, i.e. roughly 90% of drugs entering trials never reach approval.

[2] Elman × Allen Institute. Collaboration preprint (mechanism of Alzheimer’s-vulnerable neurons). bioRxiv, 2026. Submitted to Cell.