Skip to main content
Proteomics· 16-page report

TMT Proteomics Spike-in Validation

Validate TMT 6-plex quantification accuracy using Erwinia carotovora spike-in proteins with 97% correlation to ground truth.

What this research found

The PXD000001 benchmark — a tandem mass tag (TMT) 6-plex sample in which four foreign proteins are spiked at known ratios into a constant bacterial background — was used to check whether an automated proteomics pipeline recovers the correct abundance ratios. Observed and expected log2 ratios correlated at a mean Pearson r of 0.9667 across the four spike-ins, and all four surfaced in the top 12 of 399 quantified proteins by variance. The work validates the pipeline against a deliberately simple benchmark; the kidney fibrosis samples it was built for remain untested.

  • Observed and expected log2 ratios correlated at a mean Pearson r of 0.9667 (p < 0.001), clearing the threshold of r > 0.9 set in advance. Individual values ran from 0.925 for bovine cytochrome c to 0.993 for yeast enolase.
  • Root-mean-square error on log2 ratios stayed between 0.225 for rabbit muscle glycogen phosphorylase and 0.492 for bovine serum albumin.
  • All four spike-ins landed in the top 3% of proteins by variance, at ranks 1, 2, 10 and 12 of 399. Yeast enolase led with a variance of 1.144 and bovine serum albumin followed at 1.090.
  • PubMed searches pairing each spike-in protein with renal fibrosis returned 0 results, as did searches for the three highest-variance background proteins, supporting the reading that the detected variance was technical rather than biological.
  • Median normalization turned channel-specific intensity differences into uniform distributions across all six TMT channels, and principal component analysis found no outlier samples, with under 5% of variance falling in components 2 through 6.

How it was done

The PXD000001 benchmark from ProteomeXchange was processed with a PyOpenMS pipeline: a target-decoy database search with false discovery rate control, reporter ion intensity extraction from MS2 spectra across TMT channels 126 to 131, and median normalization to correct channel loading differences. Quantification accuracy was measured as the Pearson correlation and root-mean-square error between expected and observed log2 ratios for the four spike-in proteins, while sensitivity was judged by whether those proteins separated from the 399-protein background on variance ranking. Principal component analysis and pre- and post-normalization distribution plots served as quality control, and PubMed searches checked that the spike-ins had no published link to renal fibrosis. The results were written up as a 16-page grant report of roughly 4,500 words with five figures and 22 verified references.

Data sources

  • ProteomeXchange PXD000001 — TMT 6-plex benchmark with four spike-in proteins against a constant Erwinia carotovora background, 399 proteins quantified
  • Spike-in reference proteins by UniProt accession: P00924 yeast enolase, P02769 bovine serum albumin, P00489 rabbit glycogen phosphorylase, P62894 bovine cytochrome c
  • PubMed — targeted searches for spike-in and background proteins in the renal fibrosis literature
  • 22 peer-reviewed references, including Thompson et al. 2003 on TMT and Röst et al. 2016 on OpenMS

Limitations

Validation rests on a single benchmark dataset with a deliberately simplified design in which the background is held constant, and the analysis included no technical replicates and no confidence interval on the mean correlation. Three background proteins nonetheless ranked third to fifth by variance, which the internal review flagged as needing explanation since background variance can signal noise or incomplete normalization.

Outputs produced

  • TMT_Erwinia_1uLSike_Top10HCD_isol2_45stepped_60min_01.mzML

    Raw mass spectrometry data in mzML format from PRIDE PXD000001

    430 MB
  • erwinia_carotovora.fasta

    Erwinia carotovora proteome reference database (4,499 sequences)

    1.6 MB
  • augmented_erwinia_spikeins.fasta

    Target-decoy database (8,998 sequences: 4,499 target + 4,499 decoy)

    3.2 MB
  • experimental_design.json

    Complete experimental design including TMT channel mapping and spike-in ratios

  • README.txt

    PRIDE repository file listing and metadata

  • PXD000001_mztab.txt

    mzTab format protein identification and quantification results

  • 01_generate_decoy_database.py

    Script to generate target-decoy database by reversing sequences

  • 02_parse_mztab_quantification.py

    Script to parse mzTab file and extract TMT quantification data

How this research was produced

K-Dense Web planned and ran this proteomics investigation end to end — gathering the sources, carrying out the analysis, producing the figures, and drafting the report. The full session transcript, including every intermediate step, is available to view.

Share:
Proteomics

Heart Failure Proteomics Module Discovery

Reanalyze LFQ proteomics from 34 human hearts to identify stable co-abundance modules across ischemic, dilated, and hypertrophic heart-failure etiologies.

Drug Discovery

Natural Products Chemical Space

Cluster 715K+ compounds from the COCONUT database to identify bacterial-derived drug candidates using ML dimensionality reduction.

Genomics

Longevity Gene Analysis

Map model organism longevity genes to human orthologs and evaluate their validation through GWAS studies.

Run this kind of analysis on your own question

Try K-Dense Web free and see how an AI co-scientist accelerates your research.