The SIM Framework

Concept Demo

A collaboration between Lenovo and Anduril Limited

This gate keeps the demo from being read by accident. It is not a security control: the underlying data files are served directly and do not pass through it.

raminderpal@hitchhikersai.org

The SIM Framework

Concept Demo -- decision support for the commit-to-animal gate in early drug discovery.

A collaboration between Lenovo and Anduril Limited · raminderpal@hitchhikersai.org

Science Reporting

The gate committee's readable evidence view.

Open Science Reporting

Tracer

Traceability surface. Per-claim verdict distribution, surfaced precedent, numeric provenance, and citations.

Open Tracer

Scope of this deployment

The SIM Framework is a federated AI agentic system. This Concept Demo runs entirely on a single machine; the federated deployment is architectural and is not exercised here.

SIM emits no go/no-go

These interfaces present adjudicated reasoning and the evidence behind it. The commit-to-animal decision is made by a human committee. Nothing in this interface is a recommendation.

The candidate is synthetic; the precedents are real

The candidate molecule and every one of its readouts are fabricated for this demonstration. The surfaced precedents are real public records. Reading the candidate data as real measurements misreads the entire artifact.

gate renderedgate_20260901T100219Z

Every figure on this page comes from that gate. More than one gate exists on disk, and the Concept Demo is rebuilt continuously, so which one is rendered is stated rather than assumed.

What this is

A proof of concept for one decision in early drug discovery: the commit-to-animal gate, where a candidate molecule either earns a program of animal studies or does not. The science is locked to a single question -- blocking a protein called NLRP3 as a treatment for Alzheimer's disease -- and the sections below explain that question from the beginning, for a reader who has never worked on it.

SIM reads a fixed body of evidence and, for each question a committee would need to weigh, argues the case both ways: one run builds the strongest honest case that the evidence is sufficient, a second builds the case that it is not, and a third adjudicates between them. That is repeated several times per question, and where the repetitions disagree the disagreement is reported rather than averaged away.

It does not decide. The output is adjudicated reasoning and the evidence behind it, assembled for a human committee that makes the call.

What a reader can check

Every page here is generated ahead of publication and the data files it was built from are served alongside it, each with its digest recorded. Every figure can be checked against those files.

Two things on this page are not derived that way, and both say so where they appear. The plain-English background is authored: no served file states what NLRP3 is or why brain penetrance is hard. And the answer box is the one live component -- its replies are generated by a language model when you ask, are not fingerprinted, and will not be identical if you ask twice.

What holds for the answer box is narrower and worth stating exactly: it is given the same evidence published here and nothing else, and every claim, run or precedent identifier it cites is checked against that evidence before you see the reply. An identifier that does not resolve is reported to you rather than quietly dropped.

Ask a question about this evidence

The whole of the gate below -- every claim, the reasoning behind each one, the numbers and where they came from, the citations, and the stated limitations -- is also published as a plain-text document. This box puts your question to a language model that is given that document and nothing else.

What this box is, and what it is not

It answers only from the published evidence for gate_20260901T100219Z. It does not search, does not reason beyond that record, and does not make the decision -- SIM produces adjudicated reasoning and a human committee decides.

Answers are model-generated. They are not part of the published artifact, carry no digest, and are not reproducible: the same question can return differently-worded replies. Every identifier a reply cites is checked against the evidence document before you see it, and any that does not resolve is shown to you.

The first reply takes ten to thirty seconds while the model session starts.

What is on this site

Three pages and the files they were generated from. Everything below is reachable from here.

On this page

The two apps

  • Science Reporting -- The gate committee's readable evidence view.
  • Tracer -- Traceability surface. Per-claim verdict distribution, surfaced precedent, numeric provenance, and citations.

The files these pages were generated from

Served alongside the pages. Fetch any of them and check it against the digests published in the source-data section above.

The decision being demonstrated

A candidate drug spends years in laboratory glassware before it ever reaches a living animal. Work at that stage is comparatively cheap and quick: cells in a dish, purified proteins, chemical assays. At some point a team has to decide whether a molecule has earned the next stage -- dosing live animals to find out whether it does anything for the disease itself.

That is the commit-to-animal gate. It matters because it is expensive in every currency a research organization has. Animal studies cost real money, take months to read out, absorb the attention of the people who would otherwise be improving the molecule, and carry an ethical weight that no budget line captures. Committing to a molecule that was never going to work burns a year. Declining one that would have worked can end the program.

It is worth being precise about the kind of question this is. It is not a verdict on whether the underlying biology is right, and it is not a prediction that the drug will succeed. It is a judgment about one molecule at one moment, made on incomplete evidence, by a committee of people who will disagree with each other and who have to justify the call afterwards.

SIM's job is to assemble what that committee reads: for each question bearing on the decision, the evidence found, the case for and the case against, the precedents that are genuinely comparable rather than merely similar, and the citations behind all of it. The committee decides. SIM does not, and nothing on these pages is a recommendation.

This section is background, not data

Everything in the source-data section further down is computed from the served files and can be checked against a published digest. This section cannot be: it is authored rather than derived, written to make the rest of the page readable by someone who does not work in this field. No figure anywhere on this page rests on it.

The science, in plain terms

Alzheimer's disease is usually described through two hallmarks: amyloid plaques and tau tangles. A third strand of research concerns inflammation. The brain has its own immune cells, called microglia, and in Alzheimer's they are persistently switched on. The open question is whether that inflammation is merely a consequence of the damage or part of what drives it.

NLRP3 sits at the center of that question. It is a protein inside immune cells that behaves like an alarm: when it detects a danger signal it assembles, and the assembled form releases potent inflammatory messengers, chiefly IL-1beta. Amyloid can trip that alarm. If chronic NLRP3 firing is part of what pushes the disease forward, then a molecule that blocks NLRP3 ought to slow it down.

That is the therapeutic bet. It is a live hypothesis rather than a settled result, and it is a popular one -- which is a hazard in its own right, and one this demonstration carries as a standing limitation rather than a footnote.

Why getting into the brain is the hard part

A molecule can shut NLRP3 down perfectly in a dish and still be useless. The brain is protected by the blood-brain barrier, a tight lining of blood vessels that keeps most substances out, reinforced by pumps -- P-glycoprotein chief among them -- that actively eject foreign molecules back into the bloodstream.

Two measurements decide whether a candidate clears it. Kp,uu compares the free drug concentration in brain tissue against the free concentration in blood; only unbound drug can act, so this is the figure that matters rather than the total amount present. The P-gp efflux ratio says how hard the pump is working against the molecule. A weak Kp,uu on its own is a concern. A weak Kp,uu together with a strong efflux ratio is the combination that experienced teams have learned not to wave through.

The candidate on these pages

The candidate is a fictional molecule: an oral, small-molecule, direct NLRP3 inhibitor designed to reach the brain. It was built for this demonstration with a specific weakness deliberately designed into it, because a decision-support system is only interesting on a hard case. Its weakness is exactly the one described above, and the gate states it in its own words:

The candidate's borderline brain exposure (Kp,uu ~0.22 with P-gp efflux) is sufficient to support advancing to in-vivo AD-efficacy studies, on the strength of the dapansutrile APP/PS1 cognitive-rescue precedent.
claim carrying the flawcns_penetrance_sufficiency
how this gate resolved itresolved_low_conf

Quoted from the evidence package this page was built from, not written into this section.

Glossary

Every term used above and on the two evidence surfaces.

termin plain English
Alzheimer's diseaseA progressive brain disease. Usually described through two hallmarks, amyloid plaques and tau tangles; a third strand of research, the one this demonstration sits in, concerns chronic inflammation.
MicrogliaThe brain's own immune cells. They carry NLRP3, and in Alzheimer's disease they are persistently switched on.
NLRP3A protein inside immune cells that acts as an alarm. When it detects a danger signal it assembles, and the assembled form triggers inflammation. It is the target this candidate is designed to block.
InflammasomeThe assembled machine NLRP3 forms part of. Putting it together is what converts a danger signal into active inflammation.
IL-1betaOne of the inflammatory messengers released when NLRP3 fires. Measuring it is the standard way to tell whether a molecule has shut NLRP3 off.
Blood-brain barrierA tight lining of the brain's blood vessels that keeps most substances out. Any drug for a brain disease has to get past it.
Kp,uuThe ratio of free drug in brain tissue to free drug in blood. Only unbound drug can act, so this is the number that matters rather than the total. Around 1 means the barrier is no obstacle; well below 1 means most of the drug never reaches the target.
P-gp effluxP-glycoprotein is a pump in the blood-brain barrier that actively throws foreign molecules back into the bloodstream. An efflux ratio well above 1 means the pump is working hard against the drug.
In vivoIn a living animal, as opposed to in a dish. Entering this stage is what the commit-to-animal gate decides.
APP/PS1A strain of mouse engineered to develop Alzheimer's-like brain pathology. It is one of several such models, and they do not always agree with each other -- which is itself one of the claims below.
ADMET / PKAbsorption, distribution, metabolism, excretion and toxicity, and the time course of drug levels in the body. What the body does to the drug, rather than what the drug does to the body.
DevelopabilityWhether a molecule can realistically be turned into a manufacturable, stable, dosable medicine.
PrecedentA published molecule whose record can be set beside this candidate. On these pages the precedents are real and carry real identifiers.
Comparability basisA written statement of what makes a precedent comparable: the shared mechanism, the shared readout, and what differs. Without one, a precedent is a resemblance rather than evidence.
This section is background, not data

Everything in the source-data section further down is computed from the served files and can be checked against a published digest. This section cannot be: it is authored rather than derived, written to make the rest of the page readable by someone who does not work in this field. No figure anywhere on this page rests on it.

What goes in

SIM reads one fixed, committed body of records. It is not trained or fine-tuned on any of this material -- the records are opened and read at run time, the way an analyst opens a folder rather than the way a model absorbs a training set. That curated corpus is fixed before the run and digest-pinned, and the run does not add to it: what is listed and fingerprinted here is exactly what it was when the run started, and the counts are in the next section.

The corpus is not the only thing the reasoning sees. The inference service issues its own literature searches while it reasons. SIM does not decide what is searched for, does not choose when to search, and does not filter what comes back -- it records what was asked and what returned. That means search queries leave this machine. A query naming a target, an indication and a model system discloses program intent to an external service. Everything else stays local: the corpus, the composition, the expert key, the outputs and the logs.

A citation on these pages therefore has one of three origins. It came from the curated corpus; or it was returned by a search issued during reasoning; or it is neither, and appears in no record this run can point at. The third is the one that matters, and this build does not yet carry that origin as a per-citation field you can inspect. Read what is fingerprinted here as what was supplied to the run, not as the full set of what the run read.

The records come in two kinds, and the distinction is the single most important thing to understand about this demonstration.

  • Real public records. The precedents -- published molecules and what was reported about them -- are genuine, drawn from public sources and carrying real identifiers: PubMed IDs, DOIs, PubMed Central IDs and clinical trial registrations. Any of them can be looked up.
  • Hand-authored synthetic records. The candidate molecule, every one of its measurements, and the notes describing how this fictional organization reasons are written for the demonstration. No real compound is being described, and no measurement here was ever taken.

Both kinds are made to sit on a common footing before they are compared, on a basis that is written down rather than assumed. What that basis covers, and what it does not, is stated in the next section.

The seed this gate read is the seed published here

The fingerprint the run recorded when it started matches the fingerprint of the corpus index served with this page. The counts below describe the records the engine actually read.

seed fingerprinta5a931e79613a2fa...

The kinds of record held

bucketwhat it holds
(unclassified)Entries carrying no bucket label: the adversarial-position instructions the engine argues under, the organizational reasoning note, and one expert-key file. Shown rather than dropped.
expert_keyA separately authored record of which precedents ought to count as relevant. It is used to check the system, not to run it, and is kept apart from what the engine reads.
precedentA published molecule and what was reported about it, together with a written statement of why it is comparable to the candidate. Real, with real identifiers that can be looked up.
project_scienceThe candidate molecule itself and its readouts. Synthetic -- written for this demonstration, describing no real compound.
tacit_noteJudgment of the kind that normally lives in a team's heads rather than in a report: rules of thumb about what has been waved through before and what has not. Synthetic.
This section is background, not data

Everything in the source-data section further down is computed from the served files and can be checked against a published digest. This section cannot be: it is authored rather than derived, written to make the rest of the page readable by someone who does not work in this field. No figure anywhere on this page rests on it.

What was asked, and what was done

The gate decision is broken into separate claims, each one a question about a single body of evidence. Every claim is phrased the same way: given what this readout shows, and given the liability it carries, is that sufficient to justify going into animal studies? Asking about sufficiency rather than about quality is deliberate. It forces an answer relative to the decision actually being taken, instead of a verdict in the abstract that nobody has to act on.

readoutwhat this claim askshow this gate resolved it
target_engagementDoes the molecule actually switch NLRP3 off? Read out as suppression of IL-1beta, the inflammatory signal NLRP3 releases when it fires.resolved
selectivityDoes it hit NLRP3 and leave related proteins alone? A molecule that also hits the neighbors carries side effects that have nothing to do with the disease.resolved
cns_penetranceDoes enough free drug actually reach brain tissue? For a brain disease this is the make-or-break question, and it is where this candidate is weak.resolved_low_conf
in_vivo_efficacyIn a live mouse bred to develop Alzheimer's-like pathology, does treatment improve anything measurable?resolved
admet_pk_hepaticIs the molecule absorbed, distributed, broken down and cleared in a workable way, and does it leave the liver alone?resolved
developabilityCould this realistically be made into an oral medicine -- stable, manufacturable, with properties that suit a brain target?resolved

The readouts and their statuses are read from the gate this page renders. The middle column is authored.

What the engine did with each one

Every claim goes through the same loop. One run argues the strongest honest case that the evidence is sufficient. A second, working independently, argues that it is not. A third reads both and adjudicates between them. The instructions each side works under are themselves committed records, so what the advocate and the skeptic were told to do is inspectable rather than implicit.

The loop is then repeated, and the repetitions are not averaged. Where they disagree, the disagreement is the finding: a claim that comes out four-to-one is a different object from one that comes out five-to-nothing, and both are reported as distributions. A separate calibration step reads that spread and states how strong the evidence is -- keyed on the strength of the evidence itself, not on how many runs happened to agree.

claims adjudicated in this gate6
repetitions per claim5
adversarial repetitions run in total30

Where a run found no case to make at all, that is recorded as a one-sided result rather than counted as agreement. An absent position is not a defeated one, and a count that looks unanimous because nobody argued the other side is not unanimity. Both evidence surfaces carry that distinction per claim.

This section is background, not data

Everything in the source-data section further down is computed from the served files and can be checked against a published digest. This section cannot be: it is authored rather than derived, written to make the rest of the page readable by someone who does not work in this field. No figure anywhere on this page rests on it.

What each of these files holds

The section below proves these files are the ones this page was built from, by publishing a digest for each. It does not say what any of them contains. This does.

filewhat it holds
evidence_package.jsonThe gate's output, and the artifact a committee would actually read. One entry per claim: the question put to the engine, how the repetitions came out, which precedent was surfaced, and what the calibration step found.
manifest.jsonThe run's own record of what it was asked to do -- which claims were requested and what status each reached. Written by the engine while it ran, not assembled afterwards by this page.
seed_pin.jsonA fingerprint of the seed manifest, taken when the run started and stamped into the gate. It is what ties the corpus described on this page to the corpus the run actually read; if they had diverged, the figures above would not be rendered at all.
seed_manifest.jsonAn index of every record in the seed corpus, each with its bucket, its provenance and whether it is synthetic. This is the list of what the system was given.
numeric_index.jsonEvery number the run touched, with how it was arrived at: replicates used, values dropped, the aggregation applied, and whether the stated value matches the one recomputed from the raw record. Regenerated on each build -- a derived view rather than a source of truth.
citation_metadata.jsonThe identifiers behind each cited record -- PubMed IDs, DOIs, PubMed Central IDs and trial registrations -- and what happened when each was resolved. Resolution means the identifier was reachable, not that the source supports the claim citing it.
qa_context.mdThe whole of this gate written out as plain text rather than JSON: the claims, the adjudication behind each one, the numbers and how they were derived, the citations, and the carried limitations. It is generated by the same bake and carries its own digest. Nothing on this site reads it yet; it exists so that this evidence can be put to a language model without handing it raw JSON, and so that any answer can be checked against identifiers appearing verbatim here.
bake_manifest.jsonThe digest list every file above is checked against, with the gate id and the build timestamps. It cannot appear in its own file list, so it is served and is not self-attested.

The file list is read from the bake manifest. The descriptions are authored.

This section is background, not data

Everything in the source-data section further down is computed from the served files and can be checked against a published digest. This section cannot be: it is authored rather than derived, written to make the rest of the page readable by someone who does not work in this field. No figure anywhere on this page rests on it.

The source data

Every figure in this section is computed from the files this page was built from, listed at the foot of the section. None of them is written into the text. The Concept Demo is rebuilt continuously, and a number typed into prose is wrong the moment a new gate lands.

The seed corpus

A committed, fingerprinted set of records that the reasoning engine reads. It is the whole of what the system was given.

seed manifest entries33
entry count recorded by the seed pin33
seed manifest digesta5a931e79613a2fab04091b323daa716bbd708f5ea20754e34da840302ee2ea5
bucketrecords
expert_key1
precedent16
project_science7
tacit_note4
(unclassified)5
total33
provenancerecords
graded_class lifted from corpus/seed/precedents/*.json (authoritative)1
public16
synthetic11
(unclassified)5
total33
Five entries are unlabeled

Five of the seed manifest entries carry no bucket, no provenance and no synthetic flag: the three adversarial-position clauses, the organizational reasoning note, and one of the two expert key files. They are shown as an unclassified row rather than omitted. This is not repairable here -- the manifest is fingerprinted by the seed pin and that fingerprint is stamped into the gate being shown, so completing the labels would break the correspondence between this page and the run it describes.

Records held, and records used

The counts below are what the seed holds and what this gate drew on. They are stated separately and deliberately: the second is not a selection made from the first.

precedent records in the seed12
comparability basis records in the seed4
distinct precedents surfaced by this gate1
precedent surfacedprecedent_pc7b071
claims carrying a surfaced precedent6 of 6
comparability basis usedcomparability_basis.json
What the comparability basis does and does not cover

One committed comparability basis, authored for the centerpiece (cns_penetrance) claim. Claims marked basis_alignment=shared_basis_not_claim_specific were run against it for engine coverage; their surfaced precedent is NOT independently comparability-justified for that claim. Per-claim bases require SME authorship (the committed basis is itself sme_validated=false).

basis alignmentrecords
centerpiece1
shared_basis_not_claim_specific5
total6

This gate

gategate_20260901T100219Z
claims adjudicated6
repetitions per claim5
modelive_fred
run created2026-09-01T10:02:20.060589+00:00
evidence assembled2026-09-01T16:50:39.343259+00:00
elapsed between them0.3 days
page baked2026-09-01T20:50:40.386917+00:00
claim statusrecords
resolved5
resolved_low_conf1
total6

The data layer

Local JSON and text files. There is no database server and no query tier between these pages and the corpus: each page is generated ahead of publication and the files it was generated from are served beside it, so the exact bytes behind every figure above can be fetched and checked.

filebytessha256
evidence_package.json11122903b6b8d2cf3de5fda
manifest.json65348ba2f53b2ec4c9d51
seed_pin.json5807c075107204474d0
seed_manifest.json75470ef606729451025d
numeric_index.json27180ef22500fa4aa3c99
citation_metadata.json43530ffce5dea4f0552df
qa_context.md4877000b828f02e6538936

The end-to-end workflow

Raw inputs are curated into a seed corpus, the engine reasons over it and writes back, and these pages are generated from what it wrote. The three processors never hand data to one another: each reads and writes through the file layer, so every intermediate state is on disk and inspectable.

SIM Framework Concept Demo: end-to-end workflowRaw public and hand-authored inputs are curated under an authorship firewall into a seed corpus held as local JSON and text files. There is no database server. The agentic engine runs an adversarial loop at R = 5 per claim, collapses the distribution through C1 evidence-strength banding, and assembles a per-claim evidence package. fred is a frozen vendor inference service. One app with three surfaces publishes the result, and a human committee makes the commit-to-animal call.RAW INPUTSPublic sourcesPubMed, ClinicalTrials.gov,ISRCTN and DOI recordsHand-authored syntheticcandidate readouts, tacitnotes, org reasoningDATA CURATION -- batch, run once before the demoIngest and mapclean, normalize, map tocommitted record shapesAuthorship firewallopaque precedent ids,class labels stripped,key authored apartBoundary scanasserts no class labelreaches the key authorTHE DATA LAYER -- local JSON and text files, no database servercorpus/seed/precedents, bases, readouts,tacit notes, clausescorpus/MANIFEST.jsonpinned paths and sha256,covered by the seed pincorpus/runtime/run traces, gates,evidence packages, logsTHE AGENTIC ENGINE -- adversarial loop, R = 5 repetitions per claimPrecedent surfacingthe precedent bucketminus the paired itemno ranking, no cutoffA: constructive caseargues the claim issufficiently supportedB: skeptical caseargues that it is notArbitrationadjudicates A against B,records a verdictLOOP OUTCOMESSentinel: one-sidedone opening carried nocase; recorded, andnever silently droppedC1 calibrationcollapses by evidence-strength banding, notby verdict agreementGate assemblerper-claim evidencepackage with citationsfred -- FROZEN127.0.0.1:8102vendor service;no editsPUBLICATION -- every step fails closedbakecopies inputs verbatim,records md5 and sha256rendergenerates all threesurfaces ahead of timeprobeasserts served bytes,locally and at the edgepublishone app,three surfacesTHE HUMAN GATEGate committeereads the evidence and makesthe commit-to-animal call.SIM emits no go/no-go.

Generated when this page was built, from the same constants the application runs on, and drawn inline so it follows the page theme and inverts when printed. The repetition count shown is read from the gate rather than written into the picture. An earlier committed diagram is superseded: it drew a database layer this deployment does not have, and two applications where there is now one.

Findings from the reasoning engine

Not written yet

An account of what the engine actually did, written against the gate rendered on this page. Deferred deliberately. An earlier account described a different run, and rather than publish it beside data it contradicts, this section waits for findings grounded in the gate being shown.

Carried limitations

  1. SIM emits no go/no-go
    Nothing in this interface is a recommendation. The commit-to-animal decision is made by a human committee.
  2. Read-vs-collapse divergence is universal
    fred's in-trace self-assessment and C1's distribution-level collapse disagree on every claim, across two independent gates. This is a named calibration question, not a defect.
  3. C1 keys on more than one ground, and the band is only one of them
    Under this gate every claim's comparability call tracks its evidence band exactly: four claims band weak and all four take counts-with-caveat, two band moderate and both take counts. THAT AGREEMENT IS NOT EVIDENCE THAT THE BAND IS THE WHOLE STORY. Grouped on the full five-key verdict counts, this gate has four distinct groups and ONE COLLISION: cns_penetrance, developability and target_engagement all return A=0, B=5 with nothing both, neither or unresolved -- 3 of 6 claims sharing one distribution. THEY DO NOT SHARE A CALL. Two band weak and take counts-with-caveat; the third bands moderate and takes counts. A reader inferring the call from the verdict counts is therefore wrong on at least one member of that group, and comparing members of it is comparing draws rather than independent evidence. THE GROUNDS ARE NAMED IN THE PACKAGE AND ARE NOT ONE. Four claims are caveated on majority_quality_weak_or_absent; cns_penetrance carries minority_low_confidence in addition, on two repetitions flagged low confidence, and is the only claim that does. Each claim's caveat_reason is the authority on its own ground rather than this paragraph.
  4. Unanimity here is reproducibility, not corroboration
    THREE of six claims return the same verdict in all five repetitions under this gate -- target_engagement, cns_penetrance and developability, each at B=5, A=0. The five draws share one model, one prompt and one evidence base, so agreement measures how reproducibly that evidence is read, not how independently it is corroborated. A fourth claim, in_vivo_efficacy, returned A=3 with two repetitions unresolved and no repetition taking the skeptical side, which is a second reason a count is not a tally of independent judgements: NOT EVERY DRAW RESOLVED.
  5. Breadth is bounded
    One precedent against one comparability basis for all six claims. That basis is itself not SME-validated, and five of six claims are not independently comparability-justified.
  6. in_vivo_efficacy is a trend, not a significant result
    p = 0.099 against alpha = 0.05. This is material at a commit-to-animal gate, and it was concealed behind a no-numerics flag until the classifier was fixed.
  7. Citation coverage is bounded
    The two opening arms cite 12 distinct PubMed identifiers across all six claims. THAT 12 IS THE ONLY FIGURE IN THIS ENTRY AN INSTRUMENT DERIVES: build_citation_metadata.py owns it, the bake stamps it, and a probe checks this page against it. The arbitration cites further identifiers that neither arm cited, and the package holds more again than are cited anywhere. THOSE TWO COUNTS ARE NOT STATED HERE. They were stated for an earlier gate and have NOT been re-derived for this one, and a figure carried across a gate boundary is a figure about the wrong run. The package also carries PubMed Central, DOI, ISRCTN and ClinicalTrials.gov identifiers. NO COUNT ACROSS THOSE NAMESPACES IS STATED HERE, BECAUSE NO INSTRUMENT DERIVES ONE; an earlier version of this entry stated four such counts on no authority. Two spellings of one PubMed Central accession render as distinct, which is a normalization defect and it renders. THESE ARE IDENTIFIERS, NOT RECORDS: a DOI and a PubMed identifier can name the same paper, and nothing in this pipeline resolves that. What bounds coverage is what the repetitions reached for, not what retention kept. No count here is a review of the field.
  8. Resolution is not verification
    A resolving identifier proves a record exists. It does not prove the record supports the proposition it was cited for.
  9. Engine value claimed is capability presence
    The expert key is grade-lifted rather than firewalled by an independent SME, and the outstanding SME ruling is externally blocked. Any comparative headline remains circular until that ruling lands.
  10. The convergence-amplifier critique is load-bearing
    The Scannell and Rogozinska argument that repeated model agreement amplifies rather than corrects error applies to this architecture and is carried on its own merits.
  11. Two configurations reached the same call on all six claims
    This gate and gate_20260824T235603Z were run on the same candidate and the same six claims, and their calls agree in direction on all six with no reversal. ONE claim is supported in both -- in_vivo_efficacy. THAT IS STABILITY, NOT REPLICATION. The two runs share a candidate, a precedent set and a reasoning engine, so agreement bounds how much the reading moved under a configuration change and establishes nothing about whether either reading is right. TWO THINGS DIFFER BETWEEN THEM, NOT ONE: the precedent records gained measured potency, and both arguing sides were separately told such data might be present. NO DIFFERENCE BETWEEN THE TWO RUNS CAN BE ASSIGNED TO EITHER CHANGE ALONE.
  12. The measured-potency tier was reasoned over, not quoted
    This gate's precedent records carry measured potency for four of twelve programs, and both arguing sides were told the paired record may carry such a block. The gate was scored against a threshold fixed in writing BEFORE it launched: at least three distinct values from a declared fourteen appearing in model output, contributed by at least two claims. IT RETURNED ONE VALUE IN ONE CLAIM. THE THRESHOLD WAS NOT MET. The control run returned zero and a negative control of build-machinery strings returned zero on every claim, so the instrument was working. What the arms did instead was restate the content in their own words -- naming the human macrophage stratum and its species, declining to pool it across assay systems, describing the mouse values as heterogeneous. A PARAPHRASE SCORES ZERO ON A QUOTATION TEST, WHICH IS WHAT THE THRESHOLD MEASURES. The null is reported as it stands and is not reinterpreted afterwards.
  13. An objection that carried every repetition may be wrong
    On cns_penetrance, five repetitions of five found for the skeptical case on the ground that the anchor precedent supplies no quantified brain exposure, so there is no common scale against which to judge this candidate's Kp,uu. That absence holds in the publication the reasoning searched. IT DOES NOT HOLD IN THE PUBLIC RECORD: a 2023 report describes the same compound crossing the blood-brain barrier and reaching therapeutic brain concentrations in a different model, and a separate 2026 study assessed plasma and brain exposure after oral dosing. Neither was returned by the searches this gate ran. The adjudication is published here as it was written, because this artifact records what the reasoning produced -- NOT BECAUSE IT IS CORRECT. A reader should treat the objection as untested against the full literature.
  14. Elapsed time measures something outside this system
    Generation happens on a remote inference service whose behavior is not observable from the machine running the composition, and no artifact this pipeline writes counts tokens. Any timing figure therefore measures the service and the network as much as the reasoning, and no comparison against an earlier gate is offered here. A run that took longer did not thereby think harder, and nothing in this package can tell the difference.
  15. Two reasoning stages were lost, and the loss is visible
    Two constructive openings on cns_penetrance exceeded the inference service's time ceiling and were abandoned. Those repetitions are flagged low confidence and the claim is recorded as resolved with that flag rather than silently. SIXTY OF THE NINETY STAGES IN A GATE ARE OPENINGS; the arbitration stage never receives the evidence bundle directly and reasons over the two arms' arguments instead, so any statement about what the adjudicator did with a precedent record has to be read with that in mind.
  16. The seed pin covers manifested files only
    It fingerprints the files listed in the committed manifest, not a whole-tree scan. Staleness is detectable within that scope and not beyond it. The manifest lists 33 of the 33 files the seed holds. Only 8 of the cited identifiers appear in that seed; the rest appear nowhere in it, so the pin bounds the corpus and bounds nothing about the evidence the repetitions actually reached.
  17. Candidate readouts are synthetic; precedents are real
    The candidate molecule and its measurements are fabricated for this demonstration. The surfaced precedents are real public records. Reading the candidate data as real measurements misreads the entire artifact.
  18. Arbiter reasoning is retained in full, not audited
    The package keeps per-repetition arbiter prose for 30 of the 30 repetitions run, across all six claims and regardless of confidence. What the Tracer shows is the gate's reasoning as it was written, not a sample that survived retention. Completeness of retention says nothing about the quality of the reasoning retained, and nothing here checks it.
  19. Between one in ten and one in five arbitration citations do not resolve to the curated corpus
    Measured across four EARLIER runs and NOT RE-MEASURED FOR THIS GATE: 90.1, 88.2, 79.2 and 86.3 percent of the identifiers the arbitration cites resolve to one of the twelve curated precedent records at source level. The remainder do not. SIM's own composition has no retrieval or ranking stage -- both arms are handed the same fixed menu on every repetition -- but the inference service issues its own literature searches while it reasons. An identifier outside the curated corpus may have been returned by one of those searches or produced from model weights, and this pipeline does not resolve which. The citation sidecar resolves the identifiers the PACKAGE CITES -- nine on this gate, nine of nine resolved -- which is a different set from the twelve curated records above. The arbitration also introduces identifiers NEITHER opening arm cited on roughly one repetition in four, and none of those resolved to the corpus. These are UPPER BOUNDS on what sits outside it: an alternate registry spelling of a curated record fails a syntactic membership test, and at least one does so here.
  20. This demonstration is n=1
    One corpus, one model, one candidate. What is shown here is a case study with internal validity only: it records what this configuration did on this evidence, and carries no claim about how SIM would behave on another program, another corpus or another model. Generalization requires a different study, not a better run.

Why the limitations are repeated

Limitations travel with the evidence, not in a footnote

The same set is carried on all three surfaces, stated beside the claims and the provenance it qualifies rather than linked from them, because a limitation a reader has to go looking for is a limitation that will not be read. It is repeated rather than summarized: what is bounded, what is contingent, what is not yet independently validated, and what is a named open question rather than a result.

This page carried a pointer to the other two surfaces until the convergence-amplifier entry made the case against it. That entry qualifies the architecture described here rather than any single claim, so a reader who stopped at this page would never have met it.