Skip to main content
Genomics· 20-page report· 10 figures

Tara Oceans MAG–Nutrient Associations

Model 957 metagenome-assembled genomes across 93 ocean samples to identify 10 region-robust nutrient associations and design a confirmatory microcosm experiment.

What this research found

Which ocean microbes genuinely track dissolved nutrients, and which apparent associations are artefacts of where the water was sampled? Abundances for 957 metagenome-assembled genomes across 93 Tara Oceans surface and deep-chlorophyll samples were regressed against nutrient measurements using compositional transforms and region-aware mixed models, producing 314 associations that survive false-discovery control. Holding out each of the 12 ocean regions in turn narrowed that to 168 doubly stable associations, and the 10 most robust genomes were converted into a falsifiable microcosm experiment.

  • Across 5,550 mixed models (925 genomes by 6 predictors), 314 associations passed a 5% false-discovery threshold: 125 for phosphate, 124 for nitrate-plus-nitrite, 39 for the log nitrogen-to-silicon ratio, 19 for silicate, 7 for log phosphorus-to-silicon, and none at all for log nitrogen-to-phosphorus.
  • Spatial cross-validation discarded roughly half the hits. Refitting the entire model set with each of 12 ocean regions held out, 66,600 fits in total, left 297 associations directionally stable, 179 significance-stable, and 168 meeting both criteria.
  • All 10 shortlisted genomes were significant in 12 of 12 folds. The strongest is a South Pacific genome negatively associated with phosphate (median fold coefficient -2.52, full-model q of 1.0e-03), and 7 of the 10 associations are negative, an oligotroph signature, while 3 track nitrate upward.
  • The shortlist spans 6 Proteobacteria, 1 Verrucomicrobia, 1 Bacteroidetes, 1 Euryarchaeota from the Marine Group II and III archaea, and 1 Candidatus Marinimicrobia, with genomes of 0.89 to 3.25 megabases, completion estimates of 50.7% to 94.8%, and redundancy of 0.18% to 8.09%.
  • The proposed test is a 2x2 nitrate-by-phosphate microcosm: ambient controls against additions of 10 micromolar nitrate and 1 micromolar phosphate, 4 bottles per treatment, sampled at six timepoints over 96 hours, with each genome assigned marker genes such as pstS, glnA, and amtB and a directional prediction to falsify.

How it was done

A mean-coverage abundance matrix of 957 genomes across 93 metagenomes was joined to nutrient and sensor records held in a separate archive. Because the identifier schemes differ, samples were matched on station number and depth layer, resolving all 93 metagenomes and all 63 stations with no mismatches. Rare genomes were filtered out, leaving 925, abundances were centred-log-ratio transformed to respect their compositional nature, and below-detection zeros in the nutrient readings were replaced with half the smallest observed non-zero value before ratios were formed. Since the nutrients are strongly inter-correlated, each of the six predictors was modelled separately with temperature, salinity, and depth as covariates and ocean region as a random intercept, and Huber robust regressions confirmed the leading effects were not driven by extreme samples. The work is reported as a 20-page paper with 9 figures and 39 references.

Data sources

  • Tara Oceans metagenome-assembled genome abundance matrix from Delmont et al. 2018 (Figshare 10.6084/m9.figshare.4902938.v3), 957 genomes by 93 metagenomes
  • PANGAEA nutrient measurements, DOI 10.1594/PANGAEA.875575, 34,516 sample records across 202 stations
  • PANGAEA depth-resolved sample registry, DOI 10.1594/PANGAEA.875582, 34,730 records across 203 stations

Limitations

The associations are correlational: leave-one-region-out testing establishes spatial robustness rather than causation, iron, light, and dissolved organic matter go unmeasured, and biotic interactions were not modelled at all. None of the 10 shortlisted genomes appears in the source supplement's gene-level functional tables, so nutrient-metabolism capacity is inferred from lineage physiology and would need direct re-annotation before any primers could be designed.

Figures from this analysis

How this research was produced

K-Dense Web planned and ran this genomics investigation end to end — gathering the sources, carrying out the analysis, producing the figures, and drafting the report. The full session transcript, including every intermediate step, is available to view.

Share:
Genomics

Longevity Gene Analysis

Map model organism longevity genes to human orthologs and evaluate their validation through GWAS studies.

Genomics

T2-Low Asthma Endotype Discovery

Identify molecular endotypes in 679 asthma patients using nasal airway RNA-Seq and pathway-based clustering.

Genomics

NGS Variant Callers in Clinical Practice

Comprehensive review of germline and somatic variant callers for clinical genomics with ACMG/AMP classification frameworks.

Run this kind of analysis on your own question

Try K-Dense Web free and see how an AI co-scientist accelerates your research.