Executive Summary
MethaNet is building the molecular-attestation knowledge graph for climate-sensitive blue-carbon monitoring. The current warehouse contains 7,965 registered MAG/proteome units, 7,710 ESM-2 embeddings, 7,717 gLM2 payloads, and 7,484 data-complete tri-views. MBAG makes each relationship reviewable by carrying its evidence type, provenance, comparability state, and next validation action alongside the molecular unit.
The tri-view contract gives the release a durable scientific structure. The 625-unit POC core carries a common curated mechanism-feature contract. The 4,358 MSM and Futian mangrove tri-views provide complete annotation outputs and await common feature aggregation. The 2,501-unit MUCC v1 wetland lane adds source DRAM and gene annotations plus processed expression detection. These states serve different analytical roles and remain explicit throughout the report.
That structure creates a compelling climate-tech product. A partner can inspect a candidate or a monitoring context, see direct molecular evidence separately from representation context, identify the claim currently supported, and receive the next highest-value measurement. The result is an evidence card and validation plan that supports molecular diligence, sampling design, and study prioritization today while building the paired evidence required for future calibrated methane-risk intelligence.
MBAG As A Molecular Attestation Knowledge Graph
MBAG is the connective tissue of MethaNet. It links each MAG or proteome to molecular representations, direct functional observations, genomic-context evidence, QC and provenance guardrails, sample-linkage readiness, and field-validation requirements. Relationships preserve their evidence class. The graph therefore supports transparent synthesis without converting proximity or annotation availability into a biological conclusion.
The graph produces five operational outputs. It supports candidate review, measurement design, project-data diligence, validation-portfolio selection, and the evidence ledger required for future MRV deployment. The first four are available as molecular intelligence capabilities. Calibrated sample-level methane-risk estimates follow after abundance, environmental covariates, uncertainty, and field or process validation enter the same graph.
MBAG evidence architecture. Solid relationships join present molecular evidence and reliability guardrails to a MAG or proteome record. Dashed amber relationships identify the validation pathway from exact sample linkage to abundance and environmental context, then to field or process evidence. The visual expresses an evidence model and a decision workflow. Causal assertions require direct mechanism and field-validation evidence.
Evidence Integrity And Current Scope
The table records the consequential findings from reconciling the current warehouse against the prior V9 ledger, per-MAG outputs, embedding protocols, taxonomy fields, MUCC expression tables, and staged ecological evidence. This reconciliation establishes the evidence states that MBAG carries forward. It protects partner decisions from numerical comparisons across unlike feature contracts.
| Audit class | Finding | Observed result | Report action |
|---|---|---|---|
| contract correction | Data-complete and mechanism-comparable tri-views are distinct evidence states. | 7,484 data-complete; 625 mechanism-comparable; 4,358 annotation-complete/harmonization-pending; 2,501 source-scaffold. | Use the four-state evidence contract everywhere and retire the former 4,980 canonical/mechanism-equivalent claim. |
| blocking defect | The former cross-lane methane density mixed incompatible numerators. | MSM/Futian used raw many-to-many annotation rows plus all HMM rows; POC used curated accepted/present features; MUCC used source DRAM terms. | Cross-lane methane/sulfur/substrate densities and the universal attestation ranking are quarantined. Raw annotation counts remain only as provenance diagnostics until the common accepted/present feature rebuild is complete. |
| ranking impact | The former combined index was materially coupled to pipeline-specific methane counts. | Pearson r=0.605; 99.4% of the legacy top 500 were mangrove rows. | Mangrove candidates use geometry and QC while the cross-lane mechanism rank remains in quarantine. |
| geometry boundary | ESM-2 supports target-domain neighborhood navigation and routes source transfer into validation. | Raw reciprocal unique pairs include 15,727 mangrove↔wetland, 1 rumen↔wetland, 0 mangrove↔rumen. After dimension z-scoring the counts are 15,064, 0, and 0, respectively. | Describe ESM-2 links as latent neighborhoods and route source-independent ecological or mechanism transfer to validation. |
| confounding | Bridge taxonomy is structured and GTDB release is source-confounded. | 79.2% of usable reciprocal mangrove↔wetland pairs match phylum after conservative synonym normalization; release metadata is absent outside POC. | Treat bridge continuity as partly taxonomic homophily and require harmonized taxonomy/phylogeny-aware nulls. |
| orthogonal evidence | MUCC expression and field-observation lanes are valuable, with exact joins pending. | 1,358 MAGs have processed methane-gene detection and 1,868 sulfur-associated rows. Exact sample-environment-flux links are 0/133. | Surface expression as detection/occupancy support and staged flux as a validation lane. Activity magnitude and flux attribution await the authoritative join. |
Interpretation follows four evidence stages. The sequence begins with payload availability, advances through within-protocol signal and cross-lane mechanism comparability, then reaches sample and ecological validation.
The underlying warehouse retains detailed audit records and source provenance. This public report presents the resulting evidence contract, decision logic, and validation agenda without exposing raw technical bundles.
The Tri-View Evidence Contract
A formal tri-view row carries ESM-2, gLM2, and a functional payload. Its evidence state records whether those payloads support a common quantitative interpretation. MBAG carries this distinction at row level, in the freeze manifest, in candidate cards, and in release validation gates.
| Lane | Registered | ESM-2 | gLM2 | Functional payload | Data-complete tri-view | Mechanism-comparable tri-view | Functional contract |
|---|---|---|---|---|---|---|---|
| POC reference core | 625 | 625 | 625 | 625 | 625 | 625 | Curated accepted/present mechanism features (comparable) |
| MSM mangrove | 1,428 | 1,428 | 1,428 | 1,427 | 1,427 | 0 | Annotation-complete; common feature aggregation pending |
| Futian mangrove | 3,404 | 3,156 | 3,156 | 2,931 | 2,931 | 0 | Annotation-complete; common feature aggregation pending |
| MUCC v1 wetland | 2,508 | 2,501 | 2,508 | 2,508 | 2,501 | 0 | DRAM/gene/expression source scaffold (non-equivalent) |
Static fallback

gLM2 remains protocol-stratified. 5,209 units use one native and one shuffled window, while 2,508 MUCC units use 10 native and 10 shuffled windows. The shared model family supports context availability across the atlas. Numerical comparisons remain within protocol class. ESM-2 uses one 650M model family with a 6,000-protein cap, and 121 capped rows remain explicit.
Source Provenance And Environmental Readiness
Environmental metadata gives MBAG a provenance-aware route into sample and site rollups. The report shows where each evidence lane originates, its current resolution, and the next required link. This turns metadata gaps into a practical partner agenda for abundance mapping, environmental context, and field validation.
| Evidence lane | Report units | Metadata universe | Primary source | Resolution now | Use now | Blocking gap |
|---|---|---|---|---|---|---|
| Rumen reference | 518 | 555 embedded POC rumen proteomes | Stewart et al. 2019 | 555/555 exact ERZ analysis-accession matches | Methane-domain reference provenance and source-aware bridge comparison | Animal and sample environmental metadata remain cohort-level; blue-carbon interpretation requires a target-domain sample context |
| Wetland/MUCC target | 107 | 107 embedded wetland/MUCC Methanoregula proteomes | Bechtold et al. 2025 | 20 exact NCBI assembly/BioSample; 23 OWC bin plus site/project; 64 source-bucket rows | Target-domain provenance, wetland source-bucket context, and metadata-readiness triage | Uniform MAG-to-sample BioSample mapping for JGI/PPR/STM/source-bucket rows |
| Mangrove/Futian target | 3,404 | 3,404 phase-1 rMAGs; 3,156 ready payload rows; 248 explicit gap rows | Qi et al. 2026 | 2,931/3,404 tri-view units in the current functional snapshot; 65 exact sediment sample metadata rows | Interim mangrove/mudflat molecular niche expansion and time/depth/habitat readiness design | Bacteria functional completion, depth-resolved MAG-to-sample assignment, abundance/read coverage, and flux/process validation |
| Mangrove/MSM target | 1,428 | 1,428 local MAG candidates; source paper reports 966 final MAGs | Pan et al. 2025 | 1,427/1,428 tri-view units; 82 sediment sample rows; 71 exact BioSample rows | Target-domain molecular screening and sample-readiness prioritization | MAG-to-sample assignment and 966-vs-1428 denominator reconciliation before sample/site rollups |
| MUCC v1 Old Woman Creek wetland reference | 2,508 | 2,508 checksum-validated archive MAGs; 2,502 meet the paper-defined HQ/MQ screen; 7 lack direct source protein payload | Borton et al. 2026, mSystems | 2,501/2,508 data-complete source-scaffold tri-views; processed expression supports 1,948 MAGs across 133 source sample columns; 275 chamber-flux, 5,280 porewater, and 29,280 tower-flux rows are staged | Wetland molecular-reference screening and source-aware candidate review under its source-scaffold mechanism contract | Canonical MethaNet curated mechanism annotations and an authoritative sample/date/depth/environment/flux crosswalk. Exact ecological validation joins are 0/133, and expression normalization units remain unresolved |
The POC crosswalk spans a broader 662-proteome context. This report renders 625 MAG or bin comparable POC units plus registered mangrove and MUCC v1 wetland rows. The embedding map contains 7,710 registered ESM-2-bearing units. Pending, source-gap, mixed-resolution, and unlinked rows remain visible as explicit readiness states.
Three Molecular Views In One Evidence Graph
The MethaNet Bridge Attestation Graph organizes molecular similarity into a reviewable evidence trail. A 2D embedding bridge provides discovery context. Functional mechanism claims require convergent evidence from the appropriate view and protocol. Each view therefore carries its own eligible comparison set and validation gaps.
The public report exposes evidence availability, protocol class, numerator provenance, and authorized claim wording. A common cross-lane mechanism score becomes eligible after the shared feature contract is rebuilt and validated.
ESM-2 Geometry With Measured Limitations
ESM-2 defines a high-dimensional proteome-neighborhood surface for 7,710 units. The current raw cosine space is strongly anisotropic. Random-pair cosine has mean 0.9912 and median 0.9936. Median similarity to the global centroid is 0.9969. Raw cross-domain kNN edges therefore occupy a saturated range with median 0.999486. MBAG uses this geometry for neighborhood navigation and carries functional and validation evidence separately.
The graph contains a reproducible target-domain pattern. Raw space contains 15,727 unique reciprocal mangrove↔wetland pairs, 1 rumen↔wetland pair, and 0 rumen↔mangrove pairs. Per-dimension z-scoring retains 15,064 mangrove↔wetland reciprocal pairs, while both rumen transfer categories fall to zero. The release therefore supports target-domain neighborhood continuity and routes source-independent transfer questions into the validation agenda.
Taxonomy explains an important fraction of that continuity. Among reciprocal mangrove↔wetland pairs with usable phylum labels, 60.7% are exact raw-name matches and 79.2% match after conservative synonym normalization. Because GTDB release metadata is recorded only for the POC lane, source and taxonomy-release effects are confounded. Harmonized taxonomy and phylogeny-aware source nulls are required before interpreting neighborhood enrichment as functional convergence.
Diffusion coordinates are the primary navigation view because they are built from the same neighborhood graph used for inspection. UMAP, t-SNE, and PCA remain sensitivity views; no projection is treated as proof.
Scientific anchors include ESM-2 protein language model · medium-sized protein language models for transfer learning · dimension-reduction evaluation principles · Diffusion maps · PHATE for biological manifolds · similarity network fusion · graph ML for integrated multi-omics · UMAP documentation · Old Woman Creek wetland microbiome study. Recent dimensionality-reduction benchmarks reinforce this design. Visual methods differ in local and global preservation, so the report exposes the high-dimensional kNN substrate and candidate evidence cards alongside each projection.
Molecular Niche-Space Bridge Map
The bridge map provides a navigation layer for every embedding-bearing MAG or proteome unit in the release payload. It overlays auditable high-dimensional evidence links from the original 1,280-dimensional ESM-2 space. Source-lane gap records remain visible in status tables as explicit evidence states.
Gold links connect selected case-study candidates to their nearest POC reference neighbor. Gray and teal links show cross-domain kNN evidence from the ESM-2 neighborhood graph. Points encode source ecosystem and functional-payload availability. Select a halo to inspect the row's evidence contract, gLM2 protocol, numerator provenance, QC, taxonomy, expression detection, and authorized claim.
Static fallback

Candidate Evidence Cards
The candidate layer asks which evidence exists for each review hypothesis and which comparisons can support an authorized review. POC cards retain their historical internal bridge ordering. Mangrove cards use ESM-2 neighborhood geometry and QC. MUCC cards carry source-scaffold review evidence with processed expression detection where present.
The matrix and wheel display ESM-2, gLM2, functional payload, common mechanism contract, expression, QC, taxonomy, and sample context. Filled cells record evidence availability or eligibility. Mechanism strength, activity, and flux causality require their own direct supporting evidence.
Candidate evidence-eligibility matrix
Evidence coverage wheel
Static fallback

Functional Metric Harmonization
The prior report applied the label “methane marker density per 1,000 proteins” to a lane-dependent row aggregate. That interpretation belongs to the curated POC feature contract. In MSM and Futian, the numerator combines raw MCycDB hit rows with METABOLIC HMM output rows. A protein can contribute multiple hit rows, so the ratio can exceed one while representing a single marker gene. MUCC carries a third source-term contract.
| Lane | Functional units | Numerator provenance | Median raw methane rows | Median proteins | Raw rows/protein >1 | Public status |
|---|---|---|---|---|---|---|
| Futian mangrove | 2,931 | Raw many-to-many annotation-hit rows plus all HMM rows | 3,757.0 | 2,431.0 | 94.2% | Quarantined raw hit-row numerator requires feature harmonization |
| MSM mangrove | 1,427 | Raw many-to-many annotation-hit rows plus all HMM rows | 4,072.0 | 2,886.0 | 93.1% | Quarantined raw hit-row numerator requires feature harmonization |
| MUCC v1 wetland | 2,508 | Source DRAM term rows plus processed expression detection | 18.0 | 2,568.0 | 0.0% | Source-scaffold term density with a separate cross-lane contract |
| POC reference core | 625 | Curated accepted/present mechanism features | 2.0 | 1,945.0 | 0.0% | Comparable curated-feature density within the POC contract |
The legacy methane component correlated with the former combined index at Pearson r=0.605; 99.4% of the former top 500 were mangrove rows. The release keeps universal ranking in quarantine. Raw counts remain provenance diagnostics in the internal warehouse. Cross-lane mechanism densities and ranks await a shared event contract.
MUCC v1 Adds Expression Evidence And A Field-Validation Lane
The Old Woman Creek lane adds processed metatranscriptome detection across 133 source sample columns. Expression support is present for 1,948 MAGs. 1,358 carry at least one processed methane-associated expressed-gene row, and 1,868 carry sulfur-associated rows. These are detection and occupancy signals from deposited processed tables. Expression normalization remains the next requirement for activity-magnitude comparison.
The warehouse also stages 275 chamber-flux rows (188 source-valid), 5,280 porewater rows (1,563 source-valid), and 29,280 half-hourly gap-filled tower rows. It includes 694 exploratory FlashWeave associations, of which 126 pass the current stability filter, plus 3 non-grey descriptive WGCNA modules.
The decisive evidence gap is linkage. 0/133 sequencing samples currently have an authoritative exact sample, depth, environment, and flux join. 133 remain in an explicit ecological-validation block. Flux records therefore form staged validation context. MAG, expression-signature, and network-edge attribution await the authoritative join.
Sample-Linkage Readiness
This panel organizes MAG and proteome evidence at the strongest environmental context available today. Futian rows are grouped by site and month with chemistry metadata across multiple depth samples. MSM rows are grouped by source sample and BioSample sets. The dark bar overlay shows the share of each context carrying ESM-2, gLM2, and functional annotations.
The readiness layer guides abundance mapping, metadata reconciliation, and field validation. Sample-level methane-risk estimates enter the product after per-MAG abundance, exact sample assignment, environmental permissiveness, and flux or process validation become available.
From Molecular Evidence To Environmental Readiness
The atlas becomes operationally stronger when validated MAG and proteome features roll up to physical samples, metagenomes, sites, and monitoring periods. The POC lane currently supplies mechanism-comparable methane, sulfur, and substrate features. MSM and Futian await a common feature rebuild, while MUCC contributes source-scaffold and expression-detection evidence. Environmental methane-risk modeling becomes eligible when comparable molecular features receive abundance weights, exact sample provenance, environmental covariates, uncertainty, and field or process validation.
Graphical abstract. MethaNet's current evidence layer operates at MAG and proteome grain. ESM-2 and gLM2 context, methane, sulfur, and substrate annotations, QC, taxonomy, and provenance define molecular fingerprints for bridge-candidate review. The next product layer links those fingerprints to physical samples or metagenomes, weights them by MAG or read abundance and unbinned marker evidence, adds environmental permissiveness covariates, uncertainty, and flux or process validation status, then emits readiness labels such as blocked, needs metadata, needs abundance, needs environment, needs flux validation, monitor more, or scoreable provisional. Calibrated MRV outputs enter after the relevant validation gates pass.
A defensible sample score combines four gated layers. The molecular layer uses one common validated mechanism-feature contract across MAGs and unbinned marker evidence. The community layer weights those features by read coverage, relative or absolute abundance, pathway redundancy, and unassembled signal. The environmental layer captures salinity, sulfate, redox or oxygen proxies, pH, temperature, organic carbon, depth, vegetation, hydrology, season, and management. The validation layer anchors predictions against chamber fluxes, dissolved methane, porewater chemistry, incubations, or repeated field observations with explicit temporal and spatial joins.
1. Link molecules to samples
Resolve MAG-to-sample and MAG-to-site provenance, retain resolution tiers, and preserve unlinked MAGs as explicit readiness states.
2. Weight by community abundance
Turn genome potential into sample capacity using MAG coverage, marker abundance, unbinned functional reads, and uncertainty from incomplete assembly.
3. Add environmental permissiveness
Use measured metadata first, modeled covariates second, and mark every salinity, sulfate, redox, substrate, depth, and vegetation field by evidence tier.
4. Calibrate with field evidence
Use flux, porewater, geochemistry, and temporal resampling to learn which molecular signatures predict methane risk under real blue-carbon conditions.
Field work is the learning engine that turns MBAG from a molecular atlas into a progressively stronger risk system. Dense sampling across mangroves, salt marshes, freshwater wetlands, restored sites, degraded sites, salinity gradients, depth profiles, seasons, and management regimes will expand the molecular niche map, reveal source-specific blind spots, and calibrate bridge signatures in blue-carbon settings. Every new sample strengthens the atlas when it arrives with clean provenance, abundance, environmental measurements, and a validation target.
The immediate product output is a sample-risk readiness layer. A sample can be labeled scoreable, monitor more, needs metadata, needs abundance, needs environmental covariates, or needs flux validation. This gives partners a concrete sampling and diligence plan while building the evidence base for calibrated methane-risk scoring.
Strategic Readout
The durable achievement is a queryable, provenance-rich warehouse spanning 7,965 registered units and multiple evidence lanes. It already supports payload auditing, latent-neighborhood exploration, protocol-aware candidate review, expression-detection queries, metadata-gap prioritization, and validation-study design. MBAG consolidates those capabilities into a coherent climate-tech decision system.
The current release carries three explicit evidence states. The 625-unit POC core is mechanism-comparable. The 4,358-unit expansion awaits common feature aggregation. The 2,501-unit MUCC scaffold carries substantial processed expression evidence. Keeping these states visible makes future improvements measurable and protects downstream partner decisions from pipeline artifacts.
The highest-value next build produces one lane-independent mechanism-feature table, harmonized taxonomy with phylogeny-aware nulls, calibrated gLM2 protocols, exact sample and abundance mappings, and field or process validation with uncertainty. Those gates will enable cross-lane mechanism ranking and calibrated sample-risk modeling on a sound scientific foundation.