A generative system is only as good as what it is allowed to reason over. Ours reasons over a single evidence-typed graph fused from more than ninety sources: at the current snapshot, 193M+ entities across 26 types and 484M+ relations across 80 predicates, with gene-level coverage spanning 20k+ organisms.
The scale is the least interesting part. Three design choices matter more.
Negatives are first-class
Most graphs record what binds. This one records what does not, from curated, experimentally-confirmed non-interactions and inactive screening outcomes, which outnumber the positive interaction edges. None of it is sampled or inferred to pad a training set.
That matters for generation. A model that has only ever seen hits has no idea what a near-miss looks like, and a red-team step with no record of failed chemistry has nothing to push back with.
The caveat travels with the data: an inactive result records inactivity under the reported assay conditions. It does not establish non-binding in general, and the graph does not pretend otherwise.
Every edge says how it was established
Each relationship carries its source, an evidence type (experimental, computational, predicted, text-mined or curated), a native score, a negation flag and an evidence count. Per-edge provenance and affinity measurements live in their own store, so both can be queried without denormalising the core graph.
This is what makes a traced claim meaningful. "Supported by four sources" is worth little if three are text-mined from the same review. Typed evidence lets a reviewer discount what they do not trust without discarding the whole answer.
Disagreement is preserved
Sources contradict each other constantly. When they do, the conflicting edges are all retained with their own evidence rather than collapsed into a majority view. Positive evidence takes precedence only where a query has to resolve to a single sign, and that is a query-time decision, not a deletion.
The alternative is tidier and worse: a graph that has already decided for you, in a step you cannot inspect.
Finding everything like it
Proteins connect to proxy nodes for what they share: family, bound ligands, modifications, functional keywords, domains. Two proteins in the same family, or binding the same cofactor, meet at that shared node, two edges apart, in one traversal and with no custom query.
That is how a conservation question or a homolog search stops being a pipeline and becomes a walk. The full schema, the entity and predicate breakdowns, and real example edges are on the knowledge graph page.



