Statistical Applications in Genetics and Molecular Biology
Editor-in-Chief: Sanguinetti, Guido
6 Issues per year
IMPACT FACTOR 2017: 0.812
5-year IMPACT FACTOR: 1.104
CiteScore 2017: 0.86
SCImago Journal Rank (SJR) 2017: 0.456
Source Normalized Impact per Paper (SNIP) 2017: 0.527
Mathematical Citation Quotient (MCQ) 2017: 0.04
Weighted Lasso with Data Integration
The lasso is one of the most commonly used methods for high-dimensional regression, but can be unstable and lacks satisfactory asymptotic properties for variable selection. We propose to use weighted lasso with integrated relevant external information on the covariates to guide the selection towards more stable results. Weighting the penalties with external information gives each regression coefficient a covariate specific amount of penalization and can improve upon standard methods that do not use such information by borrowing knowledge from the external material. The method is applied to two cancer data sets, with gene expressions as covariates. We find interesting gene signatures, which we are able to validate. We discuss various ideas on how the weights should be defined and illustrate how different types of investigations can utilize our method exploiting different sources of external data. Through simulations, we show that our method outperforms the lasso and the adaptive lasso when the external information is from relevant to partly relevant, in terms of both variable selection and prediction.
Keywords: adaptive lasso; cervix cancer; copy number alterations; data integration; gene expressions; head and neck cancer; Lasso; p»n; penalized regression; prediction; variable selection; weighted lasso
Here you can find all Crossref-listed publications in which this article is cited. If you would like to receive automatic email messages as soon as this article is cited in other publications, simply activate the “Citation Alert” on the top of this page.