Jump to ContentJump to Main Navigation
Show Summary Details

Statistical Applications in Genetics and Molecular Biology

Editor-in-Chief: Stumpf, Michael P.H.

6 Issues per year


IMPACT FACTOR increased in 2015: 1.265
5-year IMPACT FACTOR: 1.423
Rank 42 out of 123 in category Statistics & Probability in the 2015 Thomson Reuters Journal Citation Report/Science Edition

SCImago Journal Rank (SJR) 2015: 0.954
Source Normalized Impact per Paper (SNIP) 2015: 0.554
Impact per Publication (IPP) 2015: 1.061

Mathematical Citation Quotient (MCQ) 2015: 0.06

Online
ISSN
1544-6115
See all formats and pricing
Volume 7, Issue 2 (Feb 2008)

Application of the Random Forest Classification Method to Peaks Detected from Mass Spectrometric Proteomic Profiles of Cancer Patients and Controls

Jennifer H Barrett
  • Section of Epidemiology and Biostatistics, Leeds Institute of Molecular Medicine
/ David A Cairns
  • Section of Oncology and Clinical Research, Leeds Institute of Molecular Medicine
Published Online: 2008-02-08 | DOI: https://doi.org/10.2202/1544-6115.1349

The random forest classification method was applied to classify samples from 76 breast cancer patients and 77 controls whose proteomic profile had been obtained using mass spectrometry. The analysis consisted of two stages, the detection of peaks from the profiles and the construction of a classification rule using random forests. Using a peak detection method based on finding common local maxima in the smoothed sample spectra, 444 peaks were detected, reducing to 365 robust peaks found in at least 7 out of 10 random subsets of samples. Subjects were classified as cases or controls using the random forest algorithm applied to the 365 peaks. Based on the prediction of the status of out-of-bag samples, the total error rate was 16.3%, with a sensitivity of 81.6% and a specificity of 85.7%. Measures of importance of each of the peaks were calculated to identify regions of the spectrum influencing the classification, and the four most important peaks were identified as mz3863_13, mz2943_12, mz3193_44 and mz8925_94. Combining initial peak detection with the random forest algorithm provides a high-performance classification system for proteomic data, with unbiased estimates of future performance.

About the article

Published Online: 2008-02-08


Citation Information: Statistical Applications in Genetics and Molecular Biology, ISSN (Online) 1544-6115, DOI: https://doi.org/10.2202/1544-6115.1349. Export Citation

Citing Articles

Here you can find all Crossref-listed publications in which this article is cited. If you would like to receive automatic email messages as soon as this article is cited in other publications, simply activate the “Citation Alert” on the top of this page.

[1]
P. J. Watkins, D. Clifford, G. Rose, D. Allen, R. D. Warner, F. R. Dunshea, and D. W. Pethick
Animal Production Science, 2010, Volume 50, Number 8, Page 782

Comments (0)

Please log in or register to comment.
Log in