What this research found
Which ocean microbes genuinely track dissolved nutrients, and which apparent associations are artefacts of where the water was sampled? Abundances for 957 metagenome-assembled genomes across 93 Tara Oceans surface and deep-chlorophyll samples were regressed against nutrient measurements using compositional transforms and region-aware mixed models, producing 314 associations that survive false-discovery control. Holding out each of the 12 ocean regions in turn narrowed that to 168 doubly stable associations, and the 10 most robust genomes were converted into a falsifiable microcosm experiment.
- Across 5,550 mixed models (925 genomes by 6 predictors), 314 associations passed a 5% false-discovery threshold: 125 for phosphate, 124 for nitrate-plus-nitrite, 39 for the log nitrogen-to-silicon ratio, 19 for silicate, 7 for log phosphorus-to-silicon, and none at all for log nitrogen-to-phosphorus.
- Spatial cross-validation discarded roughly half the hits. Refitting the entire model set with each of 12 ocean regions held out, 66,600 fits in total, left 297 associations directionally stable, 179 significance-stable, and 168 meeting both criteria.
- All 10 shortlisted genomes were significant in 12 of 12 folds. The strongest is a South Pacific genome negatively associated with phosphate (median fold coefficient -2.52, full-model q of 1.0e-03), and 7 of the 10 associations are negative, an oligotroph signature, while 3 track nitrate upward.
- The shortlist spans 6 Proteobacteria, 1 Verrucomicrobia, 1 Bacteroidetes, 1 Euryarchaeota from the Marine Group II and III archaea, and 1 Candidatus Marinimicrobia, with genomes of 0.89 to 3.25 megabases, completion estimates of 50.7% to 94.8%, and redundancy of 0.18% to 8.09%.
- The proposed test is a 2x2 nitrate-by-phosphate microcosm: ambient controls against additions of 10 micromolar nitrate and 1 micromolar phosphate, 4 bottles per treatment, sampled at six timepoints over 96 hours, with each genome assigned marker genes such as pstS, glnA, and amtB and a directional prediction to falsify.
How it was done
A mean-coverage abundance matrix of 957 genomes across 93 metagenomes was joined to nutrient and sensor records held in a separate archive. Because the identifier schemes differ, samples were matched on station number and depth layer, resolving all 93 metagenomes and all 63 stations with no mismatches. Rare genomes were filtered out, leaving 925, abundances were centred-log-ratio transformed to respect their compositional nature, and below-detection zeros in the nutrient readings were replaced with half the smallest observed non-zero value before ratios were formed. Since the nutrients are strongly inter-correlated, each of the six predictors was modelled separately with temperature, salinity, and depth as covariates and ocean region as a random intercept, and Huber robust regressions confirmed the leading effects were not driven by extreme samples. The work is reported as a 20-page paper with 9 figures and 39 references.
Data sources
- Tara Oceans metagenome-assembled genome abundance matrix from Delmont et al. 2018 (Figshare 10.6084/m9.figshare.4902938.v3), 957 genomes by 93 metagenomes
- PANGAEA nutrient measurements, DOI 10.1594/PANGAEA.875575, 34,516 sample records across 202 stations
- PANGAEA depth-resolved sample registry, DOI 10.1594/PANGAEA.875582, 34,730 records across 203 stations
Limitations
The associations are correlational: leave-one-region-out testing establishes spatial robustness rather than causation, iron, light, and dissolved organic matter go unmeasured, and biotic interactions were not modelled at all. None of the 10 shortlisted genomes appears in the source supplement's gene-level functional tables, so nutrient-metabolism capacity is inferred from lineage physiology and would need direct re-annotation before any primers could be designed.
Figures from this analysis
How this research was produced
K-Dense Web planned and ran this genomics investigation end to end — gathering the sources, carrying out the analysis, producing the figures, and drafting the report. The full session transcript, including every intermediate step, is available to view.


