What this research found
Which Arabidopsis thaliana lines belong in a combined drought-and-heat stress experiment? Instead of picking the strongest association hits, all 1,135 accessions in the 1001 Genomes panel were scored on genotype, provenance climate, and genetic ancestry, and a maximin algorithm nominated 24 lines that are as mutually different as possible across all three spaces at once. The nominated panel covers all 9 ancestry groups, reaches the climatic extremes of the candidate pool, and arrives with a validated 384-plant factorial layout.
- The 24 nominated accessions span all 9 genetic groups of the 1001 Genomes release, including 2 lines from the divergent Relict lineage, with a minimum pairwise joint distance of 0.384 and a mean of 0.564.
- Selecting for diversity rather than association strength captured the climatic edges of the pool: the chosen accessions match the full-pool minimum and maximum exactly for minimum temperature of the coldest month and for both drought variables, and come close for maximum temperature of the warmest month.
- Two-tier structure control worked. Stratifying association tests within ancestry groups and adjusting for the top 5 principal components cut genomic inflation 6 to 15-fold against a naive pooled model, from lambda values of 20.9 to 40.1 down to 2.53 to 2.75 across the four climate variables.
- Heat signals fell on canonical genes: the strongest associations with maximum temperature of the warmest month sit at HSFA2 (heat shock factor A2) and FT (flowering locus T). At a 5% false-discovery threshold, 901 SNPs were significant for that variable, alongside 995, 1,131, and 1,146 for cold and the two precipitation measures.
- The proposed experiment is a 2x2 factorial randomised complete block design: well-watered at 60% soil water content versus drought at 30%, crossed with a 22 C control and a 6-hour 38 C daily heat treatment, giving 384 plants across 4 blocks and 16 replicates per accession, with every balance and integrity check passing.
How it was done
Accession coordinates and ADMIXTURE group labels for the 1,135-line 1001 Genomes panel were joined to WorldClim v2.1 climate normals, sampling maximum temperature of the warmest month, minimum of the coldest month, and precipitation of the driest month and driest quarter at each collection site; 1,131 accessions returned complete values. A curated 55-gene drought, heat, and phenology panel was resolved to TAIR10 coordinates, and SNPs falling within 10 kb windows were sliced from the imputed genotype matrix, leaving 15,686 sites after biallelic, missingness, and minor-allele-frequency filtering. Genotype-climate associations were fitted within each ancestry group with five structure components as covariates and pooled by inverse-variance weighted meta-analysis. A deterministic greedy maximin search over equally weighted genetic, climatic, and ancestry distance matrices then chose the 24-line panel, and the whole design was documented in a 25-page report with 12 figures and 41 references.
Data sources
- 1001 Genomes GMI-MPI v3.1 imputed SNP matrix (10,709,949 SNPs by 1,135 accessions)
- 1001 Genomes accession metadata with ADMIXTURE assignments from Alonso-Blanco et al., Cell 166:481-491 (2016)
- WorldClim v2.1 bioclimatic normals, 1970-2000, 2.5-minute resolution (BIO5, BIO6, BIO14, BIO17)
- Ensembl Plants REST API for TAIR10 coordinates under the Araport11 annotation
Limitations
Four laboratory strains have no wild collection coordinates and were dropped from climate-based selection, and the 25 Relict accessions were excluded from the association tests for lack of power. Because every SNP sits inside one of 55 candidate genes in strong local linkage disequilibrium, the structure axes and inflation statistics describe those loci rather than a genome-wide background, and the nominated panel is a design proposal that has not yet been grown.
Figures from this analysis
How this research was produced
K-Dense Web planned and ran this agriculture investigation end to end — gathering the sources, carrying out the analysis, producing the figures, and drafting the report. The full session transcript, including every intermediate step, is available to view.


