What this research found
The PXD000001 benchmark — a tandem mass tag (TMT) 6-plex sample in which four foreign proteins are spiked at known ratios into a constant bacterial background — was used to check whether an automated proteomics pipeline recovers the correct abundance ratios. Observed and expected log2 ratios correlated at a mean Pearson r of 0.9667 across the four spike-ins, and all four surfaced in the top 12 of 399 quantified proteins by variance. The work validates the pipeline against a deliberately simple benchmark; the kidney fibrosis samples it was built for remain untested.
- Observed and expected log2 ratios correlated at a mean Pearson r of 0.9667 (p < 0.001), clearing the threshold of r > 0.9 set in advance. Individual values ran from 0.925 for bovine cytochrome c to 0.993 for yeast enolase.
- Root-mean-square error on log2 ratios stayed between 0.225 for rabbit muscle glycogen phosphorylase and 0.492 for bovine serum albumin.
- All four spike-ins landed in the top 3% of proteins by variance, at ranks 1, 2, 10 and 12 of 399. Yeast enolase led with a variance of 1.144 and bovine serum albumin followed at 1.090.
- PubMed searches pairing each spike-in protein with renal fibrosis returned 0 results, as did searches for the three highest-variance background proteins, supporting the reading that the detected variance was technical rather than biological.
- Median normalization turned channel-specific intensity differences into uniform distributions across all six TMT channels, and principal component analysis found no outlier samples, with under 5% of variance falling in components 2 through 6.
How it was done
The PXD000001 benchmark from ProteomeXchange was processed with a PyOpenMS pipeline: a target-decoy database search with false discovery rate control, reporter ion intensity extraction from MS2 spectra across TMT channels 126 to 131, and median normalization to correct channel loading differences. Quantification accuracy was measured as the Pearson correlation and root-mean-square error between expected and observed log2 ratios for the four spike-in proteins, while sensitivity was judged by whether those proteins separated from the 399-protein background on variance ranking. Principal component analysis and pre- and post-normalization distribution plots served as quality control, and PubMed searches checked that the spike-ins had no published link to renal fibrosis. The results were written up as a 16-page grant report of roughly 4,500 words with five figures and 22 verified references.
Data sources
- ProteomeXchange PXD000001 — TMT 6-plex benchmark with four spike-in proteins against a constant Erwinia carotovora background, 399 proteins quantified
- Spike-in reference proteins by UniProt accession: P00924 yeast enolase, P02769 bovine serum albumin, P00489 rabbit glycogen phosphorylase, P62894 bovine cytochrome c
- PubMed — targeted searches for spike-in and background proteins in the renal fibrosis literature
- 22 peer-reviewed references, including Thompson et al. 2003 on TMT and Röst et al. 2016 on OpenMS
Limitations
Validation rests on a single benchmark dataset with a deliberately simplified design in which the background is held constant, and the analysis included no technical replicates and no confidence interval on the mean correlation. Three background proteins nonetheless ranked third to fifth by variance, which the internal review flagged as needing explanation since background variance can signal noise or incomplete normalization.
Outputs produced
- 430 MB
TMT_Erwinia_1uLSike_Top10HCD_isol2_45stepped_60min_01.mzML
Raw mass spectrometry data in mzML format from PRIDE PXD000001
- 1.6 MB
erwinia_carotovora.fasta
Erwinia carotovora proteome reference database (4,499 sequences)
- 3.2 MB
augmented_erwinia_spikeins.fasta
Target-decoy database (8,998 sequences: 4,499 target + 4,499 decoy)
Complete experimental design including TMT channel mapping and spike-in ratios
PRIDE repository file listing and metadata
mzTab format protein identification and quantification results
01_generate_decoy_database.py
Script to generate target-decoy database by reversing sequences
02_parse_mztab_quantification.py
Script to parse mzTab file and extract TMT quantification data
How this research was produced
K-Dense Web planned and ran this proteomics investigation end to end — gathering the sources, carrying out the analysis, producing the figures, and drafting the report. The full session transcript, including every intermediate step, is available to view.


