What this is
A proof of concept for one decision in early drug discovery: the commit-to-animal gate, where a candidate molecule either earns a program of animal studies or does not. The science is locked to a single question -- blocking a protein called NLRP3 as a treatment for Alzheimer's disease -- and the sections below explain that question from the beginning, for a reader who has never worked on it.
SIM reads a fixed body of evidence and, for each question a committee would need to weigh, argues the case both ways: one run builds the strongest honest case that the evidence is sufficient, a second builds the case that it is not, and a third adjudicates between them. That is repeated several times per question, and where the repetitions disagree the disagreement is reported rather than averaged away.
It does not decide. The output is adjudicated reasoning and the evidence behind it, assembled for a human committee that makes the call.
Every page here is generated ahead of publication and the data files it was built from are served alongside it, each with its digest recorded. Every figure can be checked against those files.
Two things on this page are not derived that way, and both say so where they appear. The plain-English background is authored: no served file states what NLRP3 is or why brain penetrance is hard. And the answer box is the one live component -- its replies are generated by a language model when you ask, are not fingerprinted, and will not be identical if you ask twice.
What holds for the answer box is narrower and worth stating exactly: it is given the same evidence published here and nothing else, and every claim, run or precedent identifier it cites is checked against that evidence before you see the reply. An identifier that does not resolve is reported to you rather than quietly dropped.
Ask a question about this evidence
The whole of the gate below -- every claim, the reasoning behind each one, the numbers and where they came from, the citations, and the stated limitations -- is also published as a plain-text document. This box puts your question to a language model that is given that document and nothing else.
It answers only from the published evidence for gate_20260901T100219Z. It does not search, does not reason beyond that record, and does not make the decision -- SIM produces adjudicated reasoning and a human committee decides.
Answers are model-generated. They are not part of the published artifact, carry no digest, and are not reproducible: the same question can return differently-worded replies. Every identifier a reply cites is checked against the evidence document before you see it, and any that does not resolve is shown to you.
The first reply takes ten to thirty seconds while the model session starts.
What is on this site
Three pages and the files they were generated from. Everything below is reachable from here.
On this page
- What this is -- the gate, the science, and what the system does not do.
- The decision being demonstrated -- what a commit-to-animal gate is and why it is expensive to get wrong.
- The science, in plain terms -- NLRP3, Alzheimer's, the candidate molecule, and the flaw built into it. Includes a glossary.
- What goes in -- the records the system reads, which of them are real, and which are written for the demonstration.
- What was asked, and what was done -- the claims put to the engine and the loop each one went through.
- What each of these files holds -- what is actually inside each served file.
- The source data -- every figure recomputed from the baked files, with the seed corpus and gate censuses.
- The end-to-end workflow -- the generated diagram of the three processors and the file-based data layer.
- Findings from the reasoning engine -- deliberately not written yet, and why.
- Carried limitations -- the full set, repeated on every surface rather than linked from them.
The two apps
- Science Reporting -- The gate committee's readable evidence view.
- Tracer -- Traceability surface. Per-claim verdict distribution, surfaced precedent, numeric provenance, and citations.
The files these pages were generated from
Served alongside the pages. Fetch any of them and check it against the digests published in the source-data section above.
- evidence_package.json -- 1112290 bytes
- manifest.json -- 65348 bytes
- seed_pin.json -- 580 bytes
- seed_manifest.json -- 7547 bytes
- numeric_index.json -- 27180 bytes
- citation_metadata.json -- 43530 bytes
- qa_context.md -- 487700 bytes
- bake_manifest.json -- the digests every file above is checked against. It is listed separately because it cannot appear in its own file list: the digest would have to be taken over a document that already contained it. It is served, and it is not self-attested.
The decision being demonstrated
A candidate drug spends years in laboratory glassware before it ever reaches a living animal. Work at that stage is comparatively cheap and quick: cells in a dish, purified proteins, chemical assays. At some point a team has to decide whether a molecule has earned the next stage -- dosing live animals to find out whether it does anything for the disease itself.
That is the commit-to-animal gate. It matters because it is expensive in every currency a research organization has. Animal studies cost real money, take months to read out, absorb the attention of the people who would otherwise be improving the molecule, and carry an ethical weight that no budget line captures. Committing to a molecule that was never going to work burns a year. Declining one that would have worked can end the program.
It is worth being precise about the kind of question this is. It is not a verdict on whether the underlying biology is right, and it is not a prediction that the drug will succeed. It is a judgment about one molecule at one moment, made on incomplete evidence, by a committee of people who will disagree with each other and who have to justify the call afterwards.
SIM's job is to assemble what that committee reads: for each question bearing on the decision, the evidence found, the case for and the case against, the precedents that are genuinely comparable rather than merely similar, and the citations behind all of it. The committee decides. SIM does not, and nothing on these pages is a recommendation.
Everything in the source-data section further down is computed from the served files and can be checked against a published digest. This section cannot be: it is authored rather than derived, written to make the rest of the page readable by someone who does not work in this field. No figure anywhere on this page rests on it.
The science, in plain terms
Alzheimer's disease is usually described through two hallmarks: amyloid plaques and tau tangles. A third strand of research concerns inflammation. The brain has its own immune cells, called microglia, and in Alzheimer's they are persistently switched on. The open question is whether that inflammation is merely a consequence of the damage or part of what drives it.
NLRP3 sits at the center of that question. It is a protein inside immune cells that behaves like an alarm: when it detects a danger signal it assembles, and the assembled form releases potent inflammatory messengers, chiefly IL-1beta. Amyloid can trip that alarm. If chronic NLRP3 firing is part of what pushes the disease forward, then a molecule that blocks NLRP3 ought to slow it down.
That is the therapeutic bet. It is a live hypothesis rather than a settled result, and it is a popular one -- which is a hazard in its own right, and one this demonstration carries as a standing limitation rather than a footnote.
Why getting into the brain is the hard part
A molecule can shut NLRP3 down perfectly in a dish and still be useless. The brain is protected by the blood-brain barrier, a tight lining of blood vessels that keeps most substances out, reinforced by pumps -- P-glycoprotein chief among them -- that actively eject foreign molecules back into the bloodstream.
Two measurements decide whether a candidate clears it. Kp,uu compares the free drug concentration in brain tissue against the free concentration in blood; only unbound drug can act, so this is the figure that matters rather than the total amount present. The P-gp efflux ratio says how hard the pump is working against the molecule. A weak Kp,uu on its own is a concern. A weak Kp,uu together with a strong efflux ratio is the combination that experienced teams have learned not to wave through.
The candidate on these pages
The candidate is a fictional molecule: an oral, small-molecule, direct NLRP3 inhibitor designed to reach the brain. It was built for this demonstration with a specific weakness deliberately designed into it, because a decision-support system is only interesting on a hard case. Its weakness is exactly the one described above, and the gate states it in its own words:
The candidate's borderline brain exposure (Kp,uu ~0.22 with P-gp efflux) is sufficient to support advancing to in-vivo AD-efficacy studies, on the strength of the dapansutrile APP/PS1 cognitive-rescue precedent.
Quoted from the evidence package this page was built from, not written into this section.
Glossary
Every term used above and on the two evidence surfaces.
| term | in plain English |
|---|---|
| Alzheimer's disease | A progressive brain disease. Usually described through two hallmarks, amyloid plaques and tau tangles; a third strand of research, the one this demonstration sits in, concerns chronic inflammation. |
| Microglia | The brain's own immune cells. They carry NLRP3, and in Alzheimer's disease they are persistently switched on. |
| NLRP3 | A protein inside immune cells that acts as an alarm. When it detects a danger signal it assembles, and the assembled form triggers inflammation. It is the target this candidate is designed to block. |
| Inflammasome | The assembled machine NLRP3 forms part of. Putting it together is what converts a danger signal into active inflammation. |
| IL-1beta | One of the inflammatory messengers released when NLRP3 fires. Measuring it is the standard way to tell whether a molecule has shut NLRP3 off. |
| Blood-brain barrier | A tight lining of the brain's blood vessels that keeps most substances out. Any drug for a brain disease has to get past it. |
| Kp,uu | The ratio of free drug in brain tissue to free drug in blood. Only unbound drug can act, so this is the number that matters rather than the total. Around 1 means the barrier is no obstacle; well below 1 means most of the drug never reaches the target. |
| P-gp efflux | P-glycoprotein is a pump in the blood-brain barrier that actively throws foreign molecules back into the bloodstream. An efflux ratio well above 1 means the pump is working hard against the drug. |
| In vivo | In a living animal, as opposed to in a dish. Entering this stage is what the commit-to-animal gate decides. |
| APP/PS1 | A strain of mouse engineered to develop Alzheimer's-like brain pathology. It is one of several such models, and they do not always agree with each other -- which is itself one of the claims below. |
| ADMET / PK | Absorption, distribution, metabolism, excretion and toxicity, and the time course of drug levels in the body. What the body does to the drug, rather than what the drug does to the body. |
| Developability | Whether a molecule can realistically be turned into a manufacturable, stable, dosable medicine. |
| Precedent | A published molecule whose record can be set beside this candidate. On these pages the precedents are real and carry real identifiers. |
| Comparability basis | A written statement of what makes a precedent comparable: the shared mechanism, the shared readout, and what differs. Without one, a precedent is a resemblance rather than evidence. |
Everything in the source-data section further down is computed from the served files and can be checked against a published digest. This section cannot be: it is authored rather than derived, written to make the rest of the page readable by someone who does not work in this field. No figure anywhere on this page rests on it.
What goes in
SIM reads one fixed, committed body of records. It is not trained or fine-tuned on any of this material -- the records are opened and read at run time, the way an analyst opens a folder rather than the way a model absorbs a training set. That curated corpus is fixed before the run and digest-pinned, and the run does not add to it: what is listed and fingerprinted here is exactly what it was when the run started, and the counts are in the next section.
The corpus is not the only thing the reasoning sees. The inference service issues its own literature searches while it reasons. SIM does not decide what is searched for, does not choose when to search, and does not filter what comes back -- it records what was asked and what returned. That means search queries leave this machine. A query naming a target, an indication and a model system discloses program intent to an external service. Everything else stays local: the corpus, the composition, the expert key, the outputs and the logs.
A citation on these pages therefore has one of three origins. It came from the curated corpus; or it was returned by a search issued during reasoning; or it is neither, and appears in no record this run can point at. The third is the one that matters, and this build does not yet carry that origin as a per-citation field you can inspect. Read what is fingerprinted here as what was supplied to the run, not as the full set of what the run read.
The records come in two kinds, and the distinction is the single most important thing to understand about this demonstration.
- Real public records. The precedents -- published molecules and what was reported about them -- are genuine, drawn from public sources and carrying real identifiers: PubMed IDs, DOIs, PubMed Central IDs and clinical trial registrations. Any of them can be looked up.
- Hand-authored synthetic records. The candidate molecule, every one of its measurements, and the notes describing how this fictional organization reasons are written for the demonstration. No real compound is being described, and no measurement here was ever taken.
Both kinds are made to sit on a common footing before they are compared, on a basis that is written down rather than assumed. What that basis covers, and what it does not, is stated in the next section.
The fingerprint the run recorded when it started matches the fingerprint of the corpus index served with this page. The counts below describe the records the engine actually read.
The kinds of record held
| bucket | what it holds |
|---|---|
| (unclassified) | Entries carrying no bucket label: the adversarial-position instructions the engine argues under, the organizational reasoning note, and one expert-key file. Shown rather than dropped. |
| expert_key | A separately authored record of which precedents ought to count as relevant. It is used to check the system, not to run it, and is kept apart from what the engine reads. |
| precedent | A published molecule and what was reported about it, together with a written statement of why it is comparable to the candidate. Real, with real identifiers that can be looked up. |
| project_science | The candidate molecule itself and its readouts. Synthetic -- written for this demonstration, describing no real compound. |
| tacit_note | Judgment of the kind that normally lives in a team's heads rather than in a report: rules of thumb about what has been waved through before and what has not. Synthetic. |
Everything in the source-data section further down is computed from the served files and can be checked against a published digest. This section cannot be: it is authored rather than derived, written to make the rest of the page readable by someone who does not work in this field. No figure anywhere on this page rests on it.
What was asked, and what was done
The gate decision is broken into separate claims, each one a question about a single body of evidence. Every claim is phrased the same way: given what this readout shows, and given the liability it carries, is that sufficient to justify going into animal studies? Asking about sufficiency rather than about quality is deliberate. It forces an answer relative to the decision actually being taken, instead of a verdict in the abstract that nobody has to act on.
| readout | what this claim asks | how this gate resolved it |
|---|---|---|
| target_engagement | Does the molecule actually switch NLRP3 off? Read out as suppression of IL-1beta, the inflammatory signal NLRP3 releases when it fires. | resolved |
| selectivity | Does it hit NLRP3 and leave related proteins alone? A molecule that also hits the neighbors carries side effects that have nothing to do with the disease. | resolved |
| cns_penetrance | Does enough free drug actually reach brain tissue? For a brain disease this is the make-or-break question, and it is where this candidate is weak. | resolved_low_conf |
| in_vivo_efficacy | In a live mouse bred to develop Alzheimer's-like pathology, does treatment improve anything measurable? | resolved |
| admet_pk_hepatic | Is the molecule absorbed, distributed, broken down and cleared in a workable way, and does it leave the liver alone? | resolved |
| developability | Could this realistically be made into an oral medicine -- stable, manufacturable, with properties that suit a brain target? | resolved |
The readouts and their statuses are read from the gate this page renders. The middle column is authored.
What the engine did with each one
Every claim goes through the same loop. One run argues the strongest honest case that the evidence is sufficient. A second, working independently, argues that it is not. A third reads both and adjudicates between them. The instructions each side works under are themselves committed records, so what the advocate and the skeptic were told to do is inspectable rather than implicit.
The loop is then repeated, and the repetitions are not averaged. Where they disagree, the disagreement is the finding: a claim that comes out four-to-one is a different object from one that comes out five-to-nothing, and both are reported as distributions. A separate calibration step reads that spread and states how strong the evidence is -- keyed on the strength of the evidence itself, not on how many runs happened to agree.
Where a run found no case to make at all, that is recorded as a one-sided result rather than counted as agreement. An absent position is not a defeated one, and a count that looks unanimous because nobody argued the other side is not unanimity. Both evidence surfaces carry that distinction per claim.
Everything in the source-data section further down is computed from the served files and can be checked against a published digest. This section cannot be: it is authored rather than derived, written to make the rest of the page readable by someone who does not work in this field. No figure anywhere on this page rests on it.
What each of these files holds
The section below proves these files are the ones this page was built from, by publishing a digest for each. It does not say what any of them contains. This does.
| file | what it holds |
|---|---|
| evidence_package.json | The gate's output, and the artifact a committee would actually read. One entry per claim: the question put to the engine, how the repetitions came out, which precedent was surfaced, and what the calibration step found. |
| manifest.json | The run's own record of what it was asked to do -- which claims were requested and what status each reached. Written by the engine while it ran, not assembled afterwards by this page. |
| seed_pin.json | A fingerprint of the seed manifest, taken when the run started and stamped into the gate. It is what ties the corpus described on this page to the corpus the run actually read; if they had diverged, the figures above would not be rendered at all. |
| seed_manifest.json | An index of every record in the seed corpus, each with its bucket, its provenance and whether it is synthetic. This is the list of what the system was given. |
| numeric_index.json | Every number the run touched, with how it was arrived at: replicates used, values dropped, the aggregation applied, and whether the stated value matches the one recomputed from the raw record. Regenerated on each build -- a derived view rather than a source of truth. |
| citation_metadata.json | The identifiers behind each cited record -- PubMed IDs, DOIs, PubMed Central IDs and trial registrations -- and what happened when each was resolved. Resolution means the identifier was reachable, not that the source supports the claim citing it. |
| qa_context.md | The whole of this gate written out as plain text rather than JSON: the claims, the adjudication behind each one, the numbers and how they were derived, the citations, and the carried limitations. It is generated by the same bake and carries its own digest. Nothing on this site reads it yet; it exists so that this evidence can be put to a language model without handing it raw JSON, and so that any answer can be checked against identifiers appearing verbatim here. |
| bake_manifest.json | The digest list every file above is checked against, with the gate id and the build timestamps. It cannot appear in its own file list, so it is served and is not self-attested. |
The file list is read from the bake manifest. The descriptions are authored.
Everything in the source-data section further down is computed from the served files and can be checked against a published digest. This section cannot be: it is authored rather than derived, written to make the rest of the page readable by someone who does not work in this field. No figure anywhere on this page rests on it.
The source data
Every figure in this section is computed from the files this page was built from, listed at the foot of the section. None of them is written into the text. The Concept Demo is rebuilt continuously, and a number typed into prose is wrong the moment a new gate lands.
The seed corpus
A committed, fingerprinted set of records that the reasoning engine reads. It is the whole of what the system was given.
| bucket | records |
|---|---|
| expert_key | 1 |
| precedent | 16 |
| project_science | 7 |
| tacit_note | 4 |
| (unclassified) | 5 |
| total | 33 |
| provenance | records |
|---|---|
| graded_class lifted from corpus/seed/precedents/*.json (authoritative) | 1 |
| public | 16 |
| synthetic | 11 |
| (unclassified) | 5 |
| total | 33 |
Five of the seed manifest entries carry no bucket, no provenance and no synthetic flag: the three adversarial-position clauses, the organizational reasoning note, and one of the two expert key files. They are shown as an unclassified row rather than omitted. This is not repairable here -- the manifest is fingerprinted by the seed pin and that fingerprint is stamped into the gate being shown, so completing the labels would break the correspondence between this page and the run it describes.
Records held, and records used
The counts below are what the seed holds and what this gate drew on. They are stated separately and deliberately: the second is not a selection made from the first.
One committed comparability basis, authored for the centerpiece (cns_penetrance) claim. Claims marked basis_alignment=shared_basis_not_claim_specific were run against it for engine coverage; their surfaced precedent is NOT independently comparability-justified for that claim. Per-claim bases require SME authorship (the committed basis is itself sme_validated=false).
| basis alignment | records |
|---|---|
| centerpiece | 1 |
| shared_basis_not_claim_specific | 5 |
| total | 6 |
This gate
| claim status | records |
|---|---|
| resolved | 5 |
| resolved_low_conf | 1 |
| total | 6 |
The data layer
Local JSON and text files. There is no database server and no query tier between these pages and the corpus: each page is generated ahead of publication and the files it was generated from are served beside it, so the exact bytes behind every figure above can be fetched and checked.
| file | bytes | sha256 |
|---|---|---|
| evidence_package.json | 1112290 | 3b6b8d2cf3de5fda |
| manifest.json | 65348 | ba2f53b2ec4c9d51 |
| seed_pin.json | 580 | 7c075107204474d0 |
| seed_manifest.json | 7547 | 0ef606729451025d |
| numeric_index.json | 27180 | ef22500fa4aa3c99 |
| citation_metadata.json | 43530 | ffce5dea4f0552df |
| qa_context.md | 487700 | 0b828f02e6538936 |
The end-to-end workflow
Raw inputs are curated into a seed corpus, the engine reasons over it and writes back, and these pages are generated from what it wrote. The three processors never hand data to one another: each reads and writes through the file layer, so every intermediate state is on disk and inspectable.
Generated when this page was built, from the same constants the application runs on, and drawn inline so it follows the page theme and inverts when printed. The repetition count shown is read from the gate rather than written into the picture. An earlier committed diagram is superseded: it drew a database layer this deployment does not have, and two applications where there is now one.
Findings from the reasoning engine
An account of what the engine actually did, written against the gate rendered on this page. Deferred deliberately. An earlier account described a different run, and rather than publish it beside data it contradicts, this section waits for findings grounded in the gate being shown.
Carried limitations
- SIM emits no go/no-go
Nothing in this interface is a recommendation. The commit-to-animal decision is made by a human committee. - Read-vs-collapse divergence is universal
fred's in-trace self-assessment and C1's distribution-level collapse disagree on every claim, across two independent gates. This is a named calibration question, not a defect. - C1 keys on more than one ground, and the band is only one of them
Under this gate every claim's comparability call tracks its evidence band exactly: four claims band weak and all four take counts-with-caveat, two band moderate and both take counts. THAT AGREEMENT IS NOT EVIDENCE THAT THE BAND IS THE WHOLE STORY. Grouped on the full five-key verdict counts, this gate has four distinct groups and ONE COLLISION: cns_penetrance, developability and target_engagement all return A=0, B=5 with nothing both, neither or unresolved -- 3 of 6 claims sharing one distribution. THEY DO NOT SHARE A CALL. Two band weak and take counts-with-caveat; the third bands moderate and takes counts. A reader inferring the call from the verdict counts is therefore wrong on at least one member of that group, and comparing members of it is comparing draws rather than independent evidence. THE GROUNDS ARE NAMED IN THE PACKAGE AND ARE NOT ONE. Four claims are caveated on majority_quality_weak_or_absent; cns_penetrance carries minority_low_confidence in addition, on two repetitions flagged low confidence, and is the only claim that does. Each claim's caveat_reason is the authority on its own ground rather than this paragraph. - Unanimity here is reproducibility, not corroboration
THREE of six claims return the same verdict in all five repetitions under this gate -- target_engagement, cns_penetrance and developability, each at B=5, A=0. The five draws share one model, one prompt and one evidence base, so agreement measures how reproducibly that evidence is read, not how independently it is corroborated. A fourth claim, in_vivo_efficacy, returned A=3 with two repetitions unresolved and no repetition taking the skeptical side, which is a second reason a count is not a tally of independent judgements: NOT EVERY DRAW RESOLVED. - Breadth is bounded
One precedent against one comparability basis for all six claims. That basis is itself not SME-validated, and five of six claims are not independently comparability-justified. - in_vivo_efficacy is a trend, not a significant result
p = 0.099 against alpha = 0.05. This is material at a commit-to-animal gate, and it was concealed behind a no-numerics flag until the classifier was fixed. - Citation coverage is bounded
The two opening arms cite 12 distinct PubMed identifiers across all six claims. THAT 12 IS THE ONLY FIGURE IN THIS ENTRY AN INSTRUMENT DERIVES: build_citation_metadata.py owns it, the bake stamps it, and a probe checks this page against it. The arbitration cites further identifiers that neither arm cited, and the package holds more again than are cited anywhere. THOSE TWO COUNTS ARE NOT STATED HERE. They were stated for an earlier gate and have NOT been re-derived for this one, and a figure carried across a gate boundary is a figure about the wrong run. The package also carries PubMed Central, DOI, ISRCTN and ClinicalTrials.gov identifiers. NO COUNT ACROSS THOSE NAMESPACES IS STATED HERE, BECAUSE NO INSTRUMENT DERIVES ONE; an earlier version of this entry stated four such counts on no authority. Two spellings of one PubMed Central accession render as distinct, which is a normalization defect and it renders. THESE ARE IDENTIFIERS, NOT RECORDS: a DOI and a PubMed identifier can name the same paper, and nothing in this pipeline resolves that. What bounds coverage is what the repetitions reached for, not what retention kept. No count here is a review of the field. - Resolution is not verification
A resolving identifier proves a record exists. It does not prove the record supports the proposition it was cited for. - Engine value claimed is capability presence
The expert key is grade-lifted rather than firewalled by an independent SME, and the outstanding SME ruling is externally blocked. Any comparative headline remains circular until that ruling lands. - The convergence-amplifier critique is load-bearing
The Scannell and Rogozinska argument that repeated model agreement amplifies rather than corrects error applies to this architecture and is carried on its own merits. - Two configurations reached the same call on all six claims
This gate and gate_20260824T235603Z were run on the same candidate and the same six claims, and their calls agree in direction on all six with no reversal. ONE claim is supported in both -- in_vivo_efficacy. THAT IS STABILITY, NOT REPLICATION. The two runs share a candidate, a precedent set and a reasoning engine, so agreement bounds how much the reading moved under a configuration change and establishes nothing about whether either reading is right. TWO THINGS DIFFER BETWEEN THEM, NOT ONE: the precedent records gained measured potency, and both arguing sides were separately told such data might be present. NO DIFFERENCE BETWEEN THE TWO RUNS CAN BE ASSIGNED TO EITHER CHANGE ALONE. - The measured-potency tier was reasoned over, not quoted
This gate's precedent records carry measured potency for four of twelve programs, and both arguing sides were told the paired record may carry such a block. The gate was scored against a threshold fixed in writing BEFORE it launched: at least three distinct values from a declared fourteen appearing in model output, contributed by at least two claims. IT RETURNED ONE VALUE IN ONE CLAIM. THE THRESHOLD WAS NOT MET. The control run returned zero and a negative control of build-machinery strings returned zero on every claim, so the instrument was working. What the arms did instead was restate the content in their own words -- naming the human macrophage stratum and its species, declining to pool it across assay systems, describing the mouse values as heterogeneous. A PARAPHRASE SCORES ZERO ON A QUOTATION TEST, WHICH IS WHAT THE THRESHOLD MEASURES. The null is reported as it stands and is not reinterpreted afterwards. - An objection that carried every repetition may be wrong
On cns_penetrance, five repetitions of five found for the skeptical case on the ground that the anchor precedent supplies no quantified brain exposure, so there is no common scale against which to judge this candidate's Kp,uu. That absence holds in the publication the reasoning searched. IT DOES NOT HOLD IN THE PUBLIC RECORD: a 2023 report describes the same compound crossing the blood-brain barrier and reaching therapeutic brain concentrations in a different model, and a separate 2026 study assessed plasma and brain exposure after oral dosing. Neither was returned by the searches this gate ran. The adjudication is published here as it was written, because this artifact records what the reasoning produced -- NOT BECAUSE IT IS CORRECT. A reader should treat the objection as untested against the full literature. - Elapsed time measures something outside this system
Generation happens on a remote inference service whose behavior is not observable from the machine running the composition, and no artifact this pipeline writes counts tokens. Any timing figure therefore measures the service and the network as much as the reasoning, and no comparison against an earlier gate is offered here. A run that took longer did not thereby think harder, and nothing in this package can tell the difference. - Two reasoning stages were lost, and the loss is visible
Two constructive openings on cns_penetrance exceeded the inference service's time ceiling and were abandoned. Those repetitions are flagged low confidence and the claim is recorded as resolved with that flag rather than silently. SIXTY OF THE NINETY STAGES IN A GATE ARE OPENINGS; the arbitration stage never receives the evidence bundle directly and reasons over the two arms' arguments instead, so any statement about what the adjudicator did with a precedent record has to be read with that in mind. - The seed pin covers manifested files only
It fingerprints the files listed in the committed manifest, not a whole-tree scan. Staleness is detectable within that scope and not beyond it. The manifest lists 33 of the 33 files the seed holds. Only 8 of the cited identifiers appear in that seed; the rest appear nowhere in it, so the pin bounds the corpus and bounds nothing about the evidence the repetitions actually reached. - Candidate readouts are synthetic; precedents are real
The candidate molecule and its measurements are fabricated for this demonstration. The surfaced precedents are real public records. Reading the candidate data as real measurements misreads the entire artifact. - Arbiter reasoning is retained in full, not audited
The package keeps per-repetition arbiter prose for 30 of the 30 repetitions run, across all six claims and regardless of confidence. What the Tracer shows is the gate's reasoning as it was written, not a sample that survived retention. Completeness of retention says nothing about the quality of the reasoning retained, and nothing here checks it. - Between one in ten and one in five arbitration citations do not resolve to the curated corpus
Measured across four EARLIER runs and NOT RE-MEASURED FOR THIS GATE: 90.1, 88.2, 79.2 and 86.3 percent of the identifiers the arbitration cites resolve to one of the twelve curated precedent records at source level. The remainder do not. SIM's own composition has no retrieval or ranking stage -- both arms are handed the same fixed menu on every repetition -- but the inference service issues its own literature searches while it reasons. An identifier outside the curated corpus may have been returned by one of those searches or produced from model weights, and this pipeline does not resolve which. The citation sidecar resolves the identifiers the PACKAGE CITES -- nine on this gate, nine of nine resolved -- which is a different set from the twelve curated records above. The arbitration also introduces identifiers NEITHER opening arm cited on roughly one repetition in four, and none of those resolved to the corpus. These are UPPER BOUNDS on what sits outside it: an alternate registry spelling of a curated record fails a syntactic membership test, and at least one does so here. - This demonstration is n=1
One corpus, one model, one candidate. What is shown here is a case study with internal validity only: it records what this configuration did on this evidence, and carries no claim about how SIM would behave on another program, another corpus or another model. Generalization requires a different study, not a better run.
Why the limitations are repeated
The same set is carried on all three surfaces, stated beside the claims and the provenance it qualifies rather than linked from them, because a limitation a reader has to go looking for is a limitation that will not be read. It is repeated rather than summarized: what is bounded, what is contingent, what is not yet independently validated, and what is a named open question rather than a result.
This page carried a pointer to the other two surfaces until the convergence-amplifier entry made the case against it. That entry qualifies the architecture described here rather than any single claim, so a reader who stopped at this page would never have met it.