Data Foundation

The Knowledge Graph Behind Every Answer

ClinicaLister isn't a list of databases bolted together — it's a substance-level knowledge graph. Canonical molecules, diseases, and devices unify 18+ sources into one queryable layer, so every link is consistent, confidence-scored, and traceable to its source.

397,000+
Verified Links
9,100+
Canonical Substances
3-Tier
Confidence Scoring

Raw databases don't agree with each other

The Problem

ClinicalTrials.gov, FDA, EMA, and GSRS each name the same drug differently — brand vs. generic, salt vs. base, dev code vs. INN, with typos and synonyms throughout. Joining them by hand means hours of reconciliation, and the joins break every time a source updates.

The Foundation

We resolve every record to a canonical identity — a molecule, a disease, a device — and connect them with typed, confidence-scored edges. Reconciliation happens once, in the graph, so every search, chart, and AI answer reads from the same consistent layer.

Three canonical identities unify everything

Drugs, diseases, and devices each get one stable identity that every source maps into.

Substances → Molecules

Every FDA, EMA, and GSRS drug record collapses to one canonical molecule, anchored on its UNII substance code. 9,100+ molecules, 83% UNII-anchored — so a brand, a generic, a salt, and a development code all resolve to the same entity.

Diseases → MONDO

Every trial condition resolves to a canonical MONDO disease — 33,700+ terms and 111,600+ synonyms — collapsing 'NSCLC', 'Non-Small Cell Lung Cancer', and every subtype into one queryable identity. 500,000+ trials (~85%) are mapped.

Devices → GUDID

Device submissions (510(k)/PMA) link to their GUDID identifiers and FDA classifications, with sponsor-gated trial matching so device links carry the same rigor as drugs.

Explore the graph

Every node type and every edge type, exactly as the graph is modelled. Rotate it, filter to one dimension, and select any node to see what resolves it.

Pick a node

Drag to rotate the graph, then pick a dimension or select any node to see what it is and how each of its edges is resolved.

  • SubstanceThe molecule golden record — one canonical identity per substance, anchored on its UNII, with its aliases, classes, targets and salt/biosimilar parent.
  • Drug productThe commercial and regulatory record: FDA and EMA products, and everything that hangs off them — NDC, pricing, shortages, labels and FAERS safety.
  • TrialA first-class hub. A trial ties drugs, devices, diseases, geography and published evidence together, so one traversal answers a question that would otherwise take five joins.
  • DiseaseCanonical MONDO identities with the full is-a hierarchy, so a query for a parent disease can pull in every subtype.
  • DeviceFDA-regulated devices with their GUDID identity, classification, recalls and MAUDE adverse events.
SubstanceDrug productTrialDiseaseDeviceTrusted linkGhost link

Node types (29)

NodeDimensionWhat it is
MoleculeSubstanceThe golden record and the primary key of the whole graph. Every FDA, EMA and GSRS drug record collapses into one canonical substance identity, anchored on its UNII — so a brand, a generic, a salt and a development code all resolve to the same entity. Trial linking happens here, at the substance level, never at the product level.
Moiety parentSubstanceThe active-moiety parent a salt or biosimilar rolls up to. This is what makes “all trials for this substance” mean exactly that — while keeping the salt a distinct record. A salt rolls up to its drug base, never to its counterion.
AliasSubstanceEvery other name a molecule goes by — development codes, INN, official and common names, plus brand names backfilled from FDA and EMA. Brand aliases are deliberately excluded from the deterministic auto-link path, because an unfiltered first cut produced a class of false positives.
ClassificationSubstanceATC, MEDRT, FDA_SPL and SNOMED_CT class memberships, aggregated from all of a molecule's FDA products. This is what makes class-level questions possible — “every trial testing any PD-1 inhibitor” rather than a drug-by-drug list.
TargetSubstanceThe pharmacological target a molecule acts on, typed by interaction (inhibitor, agonist, antagonist, substrate). Sourced from GSRS substance relationships, which are sparse — this edge covers a small subset of molecules, and we'd rather say so than imply full coverage.
GSRS substanceSubstanceFDA's Global Substance Registration System — the registry that supplies the UNII, INN, CAS number, molecular formula and structure keys. It is the anchor that makes substance identity deterministic rather than a fuzzy name match.
FDA drugDrug productAn approved US drug product — the NDA, ANDA or BLA record, with its brand and generic names, active ingredients, route and marketing status. Every enrichment below hangs off this record by exact key.
EMA drugDrug productThe European marketing-authorisation record, resolved to the same canonical molecule as its US counterpart — which is what lets one question span both regulators.
NDC productDrug productThe packaged, marketed presentations of a drug product, joined on application number. The NDC is also the hinge that pricing data attaches to.
FAERS safetyDrug productPost-market adverse-event signal from the FDA Adverse Event Reporting System, summarised per drug product and joined on application number. It is the safety half of the picture a trial record alone can't give you.
REMS & designationsDrug productRisk-management programmes and the special designations that change a development path — orphan, fast-track and breakthrough — carried as flags on the product record.
ShortageDrug productCurrent and resolved supply shortages, linked by generic name — the signal that a marketed product is under supply pressure.
PricingDrug productAcquisition cost (NADAC, by NDC), physician-administered reimbursement (ASP, resolved through an HCPCS→NDC crosswalk) and Medicare Part D spend. Three different pricing questions, each with its own resolution path.
Open PaymentsDrug productCMS industry-to-physician transfers of value, linked to the drug by NDC and to devices by name and manufacturer.
DailyMed labelDrug productThe structured product label — the authoritative prescribing text — joined on application number, falling back to RxCUI.
TrialTrialA registered clinical study, keyed by its NCT ID, with phase, status, sponsor, enrolment, arms and interventions. The trial is a hub, not a leaf: it is what ties drugs, devices, diseases, sites and evidence into one traversal.
EU trialTrialThe CTIS registration for European studies, cross-referenced to its NCT counterpart where the registry declares one — so a molecule's trial set doesn't stop at the US border.
Extracted entityTrialA drug or device name pulled out of a trial's free text, tagged with the role it plays — the raw material the linker resolves into a molecule edge. The original string is preserved, so you can always see what the pipeline actually read.
Sites & geographyTrialWhere a trial actually runs — the facility, city and country records behind the map view and any geographic filter.
PublicationsTrialPublished evidence tied back to the trial by NCT ID, from PubMed, Europe PMC and the registry itself — the bridge from “a trial exists” to “here is what it found”.
Change historyTrialField-level change tracking captured by database trigger on every refresh, so a status flip, an enrolment cut or a completion-date slip is a queryable event rather than something you had to be watching for.
Disease (MONDO)DiseaseA canonical disease identity from the Monarch Disease Ontology, which unifies ICD-10, MeSH, DOID, Orphanet, NCI Thesaurus, OMIM and UMLS under one stable ID. This is what collapses “NSCLC” and “Non-Small Cell Lung Cancer” into a single queryable thing.
Parent & ancestorsDiseaseThe is-a hierarchy, with the full transitive closure precomputed. Asking for “lung cancer” can therefore return every subtype beneath it in one query rather than a hand-maintained synonym list.
Disease aliasDiseaseMONDO synonyms — exact, related, broad and narrow — which are what let a free-text trial condition find its canonical identity in the first place.
DeviceDeviceAn FDA-regulated device: its 510(k) or PMA submission, the applicant that filed it, and its three-letter product code. Devices are company-specific, which shapes how their trial edges are scored.
GUDID identityDeviceThe Global UDI Database record, matched to its premarket submission by submission number plus product code and company — the closest thing a device has to a stable identity.
ClassificationDeviceThe FDA product code and risk class (I, II or III) that determine a device's regulatory pathway.
RecallsDeviceDevice recall events, joined by product code — the clearest single signal that a device family is in trouble.
MAUDE eventsDeviceAdverse-event reports from the MAUDE database, joined by product code. Sampled rather than exhaustive, which is a limit of the source, not a choice we made quietly.

Edge types (34)

EdgeTrustResolved by
GSRS substance anchors MoleculeIdentity — exact keyUNII — the registry identifier both sides share
Molecule also known as AliasSource factGSRS names + FDA/EMA brand backfill
Molecule classified as ClassificationSource factRxClass ATC / FDA_SPL / MEDRT / SNOMED_CT
Molecule interacts with TargetSource factGSRS substance relationships
Molecule rolls up to Moiety parentHierarchy — is-a / rollupGSRS active-moiety, salt and biosimilar rollup
FDA drug resolves to MoleculeIdentity — exact keyUNII, falling back to normalised generic name
EMA drug resolves to MoleculeIdentity — exact keyUNII, falling back to normalised medicine name
FDA drug registered as GSRS substanceIdentity — exact keyUNII, else ingredient name
FDA drug marketed as NDC productSource factapplication number
FDA drug safety signal FAERS safetySource factapplication number
FDA drug designated REMS & designationsSource factapplication number
FDA drug in shortage ShortageSource factgeneric name
NDC product priced at PricingSource factNDC, segment-aware; ASP via HCPCS→NDC crosswalk
FDA drug payments for Open PaymentsSource factNDC
FDA drug labelled by DailyMed labelSource factapplication number, else RxCUI
Molecule involved in TrialTrusted link — scored ≥ 0.95alias-first exact → fuzzy with signal boosting → AI verification, scoring ≥ 0.95
Molecule possibly involved in TrialGhost link — 0.85 to 0.94scored 0.85–0.94 — held as a ghost link until verified
Molecule involved in EU trialTrusted link — scored ≥ 0.95the same pipeline, mirrored onto CTIS trials
Extracted entity resolves to MoleculeTrusted link — scored ≥ 0.95alias-exact on the safe set, else fuzzy + AI curation
Trial mentions Extracted entitySource factAI extraction from trial free text, audit-gated
Trial studies Disease (MONDO)Source factexact → alias → AI fallback on the free-text condition
Trial runs at Sites & geographySource factregistry-supplied site records
Trial evidenced by PublicationsSource factNCT ID across PubMed, Europe PMC and the registry
Trial changes tracked Change historySource factfield-level database trigger on every refresh
EU trial cross-references TrialIdentity — exact keyNCT ID declared in the CTIS secondary identifiers
Disease (MONDO) is a Parent & ancestorsHierarchy — is-a / rollupMONDO is-a graph, ancestors precomputed
Disease (MONDO) also known as Disease aliasSource factMONDO synonyms
Device studied in TrialTrusted link — scored ≥ 0.95fuzzy name match; sponsor agreement is required to reach this tier
Device possibly studied in TrialGhost link — 0.85 to 0.94no sponsor agreement, a generic token, or a score of 0.85–0.94
Device identified by GUDID identityIdentity — exact keysubmission number + product code and company
Device classified as ClassificationSource factFDA product code
Device recalled RecallsSource factproduct code
Device adverse events MAUDE eventsSource factproduct code (sampled)
Device payments for Open PaymentsSource factdevice name and manufacturer

Every link is scored and traceable

We don't dump unvalidated matches on you. Each edge carries a confidence tier and full provenance.

Three confidence tiers

High-confidence links surface by default; Medium-confidence (AI-suggested) links wait for validation; Low-confidencecandidates are re-checked before they're promoted. You always know how sure we are.

Reviewed before verified

Every active molecule↔trial link carries a curation verdict — an automated pre-check and an AI reviewer sign off before a link is shown as verified.

Full provenance

Every edge records where it came from and how it was matched — so a regulatory or competitive-intelligence claim can always be traced back to its source.

Data quality you can audit

A knowledge graph is only as good as its edges. We model the hard cases correctly — and guard against them coming back.

Placebo arms are modeled as a dedicated concept — never mistaken for a real drug

Generic terms (“chemotherapy”, “active comparator”) never resolve to a specific molecule

Distinct drugs are kept distinct — no rolling bupivacaine into ropivacaine

A salt rolls up to its drug base, never to its counterion (no quetiapine → fumaric acid)

Each guardrail is enforced at the source and regularly re-checked to stay at zero

The same canonical identities also roll salts and metabolites up to their parent molecule, so “all trials for this substance” means exactly that.

Why our AI answers are grounded

The knowledge graph is what makes the AI useful

When you connect an AI assistant through the Model Context Protocol, it isn't guessing across raw tables — it reads the knowledge graph. Canonical identities and typed links let it answer “every trial for this molecule and its biosimilars, with the regulatory and pricing context” in one hop, with provenance attached.

Intelligence is only as good as the data underneath it.