1. Multiple haplotype reconstruction from allele frequency data
- Author
-
Merle Behr, Housen Li, Axel Munk, Marta Pelizzola, and Andreas Futschik
- Subjects
Multivariate statistics ,Computational complexity theory ,Computer Networks and Communications ,Haplotype ,Computer Science (miscellaneous) ,Design matrix ,Coefficient matrix ,Allele frequency ,Algorithm ,Computer Science Applications ,Mathematics - Abstract
Because haplotype information is of widespread interest in biomedical applications, effort has been put into their reconstruction. Here, we propose an efficient method, called haploSep, that is able to accurately infer major haplotypes and their frequencies just from multiple samples of allele frequency data. Even the accuracy of experimentally obtained allele frequencies can be improved by re-estimating them from our reconstructed haplotypes. From a methodological point of view, we model our problem as a multivariate regression problem where both the design matrix and the coefficient matrix are unknown. Compared to other methods, haploSep is very fast, with linear computational complexity in the haplotype length. We illustrate our method on simulated and real data focusing on experimental evolution and microbial data. haploSep is a computationally efficient method to infer major haplotypes and their frequencies from multiple samples of allele frequency data, and to provide improved estimates of experimentally obtained allele frequencies.
- Published
- 2021
- Full Text
- View/download PDF