Jump to ContentJump to Main Navigation

Online

99,00 € / $149.00*

* Prices subject to change. Shipping costs will be added if applicable.
Publication Date:
June 2006
ISSN:
1544-6115
DOI:
10.2202/1544-6115.1164

See all formats and pricing

Online
Individual Subscription Online only
Euro [D] 99.00
RRP for USA, Canada, Mexico
US$ 149.00 *
Print
Individual Subscription Online only
Euro [D] 285.00
RRP for USA, Canada, Mexico
US$ 384.00 *
Print + Online
Individual Subscription Online only
Euro [D] 342.00
RRP for USA, Canada, Mexico
US$ 461.00 *
*Prices subject to change. Shipping costs will be added if applicable.

Editor-in-Chief: Stumpf, Michael P.H.

Editorial Board Member: Beaumont, Mark / Binder, Harald / Gupta, Mayetri / Hubbard, Alan E. / Husmeier, Dirk / Ji, Hongkai / Keles, Sunduz / Kerr, Kathleen / Lazzeroni, Laura / Lin, Shili / Ma, Ping / Marjoram, Paul / Mertens, Bart / Nerman, Olle / G. Petretto, Enrico / Plagnol, Vincent / Purdom, Elizabeth / Robin, Stéphane / Rzhetsky, Andrey / Sanguinetti, Guido / van der Laan, Mark J. / von Haeseler, Arndt / Weeks, Daniel E. / Wiuf, Carsten / Zhao, Hongyu

6 Issues per year

IMPACT FACTOR 2011: 1.517
5-year IMPACT FACTOR: 1.704
Rank 27 out of 116 in category Statistics & Probability in the 2011 Thomson Reuters Journal Citation Report/Science Edition

Model Selection for Mixtures of Mutagenetic Trees

Junming Yin / Niko Beerenwinkel / Jörg Rahnenführer / Thomas Lengauer

1Department of EECS, University of California, Berkeley

1Department of Mathematics, University of California, Berkeley

1Max-Planck-Institute for Informatics, Saarbrücken, Germany

1Max-Planck-Institute for Informatics, Saarbrücken, Germany

Citation Information: Statistical Applications in Genetics and Molecular Biology. Volume 5, Issue 1, Pages –, ISSN (Online) 1544-6115, DOI: 10.2202/1544-6115.1164, June 2006

Publication History:
Published Online:
2006-06-23

The evolution of drug resistance in HIV is characterized by the accumulation of resistance-associated mutations in the HIV genome. Mutagenetic trees, a family of restricted Bayesian tree models, have been applied to infer the order and rate of occurrence of these mutations. Understanding and predicting this evolutionary process is an important prerequisite for the rational design of antiretroviral therapies. In practice, mixtures models of K mutagenetic trees provide more flexibility and are often more appropriate for modelling observed mutational patterns.Here, we investigate the model selection problem for K-mutagenetic trees mixture models. We evaluate several classical model selection criteria including cross-validation, the Bayesian Information Criterion (BIC), and the Akaike Information Criterion. We also use the empirical Bayes method by constructing a prior probability distribution for the parameters of a mutagenetic trees mixture model and deriving the posterior probability of the model. In addition to the model dimension, we consider the redundancy of a mixture model, which is measured by comparing the topologies of trees within a mixture model. Based on the redundancy, we propose a new model selection criterion, which is a modification of the BIC.Experimental results on simulated and on real HIV data show that the classical criteria tend to select models with far too many tree components. Only cross-validation and the modified BIC recover the correct number of trees and the tree topologies most of the time. At the same optimal performance, the runtime of the new BIC modification is about one order of magnitude lower. Thus, this model selection criterion can also be used for large data sets for which cross-validation becomes computationally infeasible.

Keywords: model selection; mixtures of mutagenetic trees; BIC; empirical bayes

Comments (0)

Please log in or register to comment.