Skip to content

About

Natural product knowledge base — MIBiG, ChEBI and LOTUS harmonized onto one evidence-backed record per chemical structure, carrying producer organisms, biosynthetic gene clusters and mechanism

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

418 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

NaturalProductMech

Knowledge base of individual natural product structures — one record per chemical structure made by a living organism, carrying who makes it, from which gene cluster, what it does, and the evidence for all three.

NaturalProductMech is the biosynthetic-origin counterpart of AntibioticMech (antimicrobial compounds and their mechanisms), TraitMech (traits), CultureMech (growth media), MediaIngredientMech (ingredients), HabitatMech (habitats), CellStructureMech (cell structures), ProteinTraitsMech (proteins) and CommunityMech (communities), and follows the curation pattern established by dismech: one YAML per entity, ontology-grounded, evidence-backed, schema-validated, curated incrementally.

Status: seeded from nine sources, pathway curation underway

2026-09-09. The corpus reproduces offline from committed inventories: MIBiG (structures, producers, gene clusters), ChEBI (grounding, origins), LOTUS and CyanoMetDB (cited occurrences), NCBI Taxonomy (name resolution), NPClassifier (the filing pathway), PubChem BioAssay and BindingDB (measured activities and targets), and AntibioticMech (antimicrobial classification and cross-corpus links). Every record remains SEEDED until a curator signs off identity, structure, filing class, and producer claims; partial biosynthetic_pathway and causal_graphs curation is now underway. Owed work is in NEXT_TASKS.md.

Curated mechanisms can use inline graphs or exact-owner components referenced by causal_graph_refs. Use the resolving reader for the complete mechanism; see curation for the data contract.

records: 3115
unique InChIKeys: 3115
structures with undefined stereocentres: 1124

by pathway (the filing decision, computed):
  ALKALOIDS                             509
  AMINO_ACIDS_AND_PEPTIDES              655
  CARBOHYDRATES                         103
  FATTY_ACIDS                            82
  POLYKETIDES                           869
  SHIKIMATES_AND_PHENYLPROPANOIDS        64
  TERPENOIDS                            213
  UNCLASSIFIED                          620

by curation status:
  SEEDED                               3115

by grounding status:
  EXACT                                 362
  MINTED                               2744
  REVIEW_NEEDED                           9

field coverage (records carrying at least one item):
  biosynthesis_origin                  3115
  producer_taxon_groups                3076
  occurrence_taxon_groups              2342
  producer_organisms                   3076
  occurrences                          2342
  biosynthetic_gene_clusters           3115
  biosynthetic_pathway                  157
  bioactivities                         176
  bioactivity_summary                   317
  molecular_targets                      14
  causal_graphs                         241
  related_records                       209
  discussions                           325

producer claims: 3407 (808 causal, 162 correlational)
  BGC_CHARACTERIZED                     808
  BGC_CORRELATED                        162
  SOURCE_ASSERTION                     2437

cluster link claims: 3449 (1359 demonstrated)
  CLUSTER_CORRELATED                    110
  CLUSTER_DEMONSTRATED                 1359
  CLUSTER_PREDICTED                       2
  CLUSTER_UNSTATED                     1978

records where a producer taxon is independently corroborated by a cited occurrence: 732
records whose only origin evidence is a gene cluster with no named host: 27

These grade two different questions and they come apart. A producer
claim says this TAXON makes the compound; a cluster link says this
LOCUS does. Heterologous expression settles the second and leaves the
first exactly where the isolation report left it.

Development uses Python 3.13 via .python-version; CI selects the same minor explicitly and runs the full quality gate once. The package compatibility floor remains declared in pyproject.toml.

just install        # uv sync --locked --extra dev --extra chemistry
just qc             # every local and CI quality gate
just seed           # dry run: what would be written, per pathway
just report         # corpus, grounding and origin-evidence coverage
just source-queue   # the ranked data-source queue

Nothing above touches the network.

Document What it is
PLAN.md Scope, schema design, the AntibioticMech join, milestones M0–M6
CLAUDE.md Operational guidance for editing agents: commands, boundaries, invariants
NEXT_TASKS.md Owed work, by issue, in the order it is worth doing
docs/HARMONIZATION.md, docs/CURATION.md Identity, merging and scope; decision semantics and evidence rules
research/2026-09-07-natural-product-data-sources.md The verified source landscape, with focused follow-up reports beside it
curation/source_queue.tsv 43 sources, ranked — 9 adopted — with licence status and the gap each closes
.claude/skills/ Five repo-local curation workflows
src/naturalproductmech/schema/ The LinkML schema, closed-validated
conf/ Scope and the record budget, the producer-evidence grading, the adopted and staged sources, the sibling pin

The gap it fills

AntibioticMech answers what a compound does to a microbe and how the microbe resists it. It has nowhere to say who makes erythromycin, from which gene cluster, by which pathway. TraitMech says an organism produces antibiotics; ProteinTraitsMech says what a polyketide synthase domain is. None of them says which structure a given Streptomyces strain makes, from which locus, with what evidence.

Section of a record Question it answers
producer_organisms Which taxon biosynthesizes it, and how do we know?
occurrences Where has it been found, and in whose paper?
biosynthetic_gene_clusters Which cluster, in which genome, at which coordinates?
biosynthetic_pathway Which enzymes, in which order?
bioactivities / molecular_targets What does it do, measured how?
causal_graphs How is it made, and how does it act? Every edge cited.

The two distinctions the corpus is built on

One record is one chemical structure, keyed on a Standard InChIKey. A compound class is not a record. Neither is an extract, a fraction, an essential oil or a herbal preparation — activity measured on a mixture is not evidence about any constituent of it.

Occurrence is not production. A compound isolated from a sponge may be made by its bacterial symbiont; one "found in" a plant extract may come from a fungal endophyte. So producer_organisms requires biosynthesis-grade evidence — a gene cluster with experimental support, heterologous expression, isotope feeding, or production by an axenic culture — while occurrences requires a cited isolation or detection report. The seeder never promotes one to the other. That distinction is why LOTUS's 674,454 structure–organism–reference triples are occurrences, and why MIBiG's per-locus evidence vocabulary is what grades a producer claim.

Sources

The first release seeds from six sources whose licences were verified against their own terms pages on 2026-09-07:

Source Licence Contributes
MIBiG 4.0 CC BY 4.0 producers with NCBI taxid, gene clusters, cluster class, references
ChEBI CC BY 4.0 identity, structures, specialized-metabolite roles
LOTUS via Wikidata CC0 referenced occurrences
NPClassifier CC0 the filing classification, computed from structure
PubChem public domain structures the others lack
CyanoMetDB CC BY 4.0 curated cyanobacterial metabolites

Measured bioactivity and molecular targets come from a fifth and sixth source, both verified the same day: BindingDB's own-curated subset (CC BY 3.0, and separable from its ChEMBL-derived records by a provenance column) and PubChem BioAssay (public domain, recorded per depositor).

Excellent resources that cannot be seeded into a CC BY 4.0 corpus, and why: NPAtlas (CC BY-NC from release 2024_09), Norine (CC BY-NC-SA), NP-MRD (CC BY-NC), CMNPD (CC BY-NC-SA), NPBS Atlas (CC BY-NC), IMPPAT and DrugBank (non-commercial), KNApSAcK (redistribution prohibited), ChEMBL, DrugCentral and the Guide to PHARMACOLOGY (share-alike), ClassyFire/ChemOnt (bespoke terms), MetaCyc (subscription), KEGG (not a public database). NPASS, the Therapeutic Target Database and StreptomeDB state no licence at all, which is not the same as permission, and CO-ADD reserves all rights while calling itself open-access. COCONUT advertises CC0 over a collection that demonstrably contains rows from several of the restricted sources above.

The full reasoning, with licence text quoted from each primary page, is in the research report.

Contributing

See the native merge queue guide for PR checks, queue validation, and recovery when a queued change fails.

Licence

Project-authored data, records, annotations, mappings, data exports and narrative documentation are licensed under CC BY 4.0. Project-authored code, scripts, tests, schemas and website templates are licensed under BSD-3-Clause. See LICENSE for scope and attribution.

Third-party material retains its own licenses and notices. Preserve upstream attribution and source-specific terms when redistributing a record or subset. Previously released material remains available under its original license.

About

Natural product knowledge base — MIBiG, ChEBI and LOTUS harmonized onto one evidence-backed record per chemical structure, carrying producer organisms, biosynthetic gene clusters and mechanism

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages