Back to Search Start Over

Reliable genomic strategies for species classification of plant genetic resources.

Authors :
van Bemmelen van der Plaat, Artur
van Treuren, Rob
van Hintum, Theo J. L.
Source :
BMC Bioinformatics. 3/31/2021, Vol. 22 Issue 1, p1-18. 18p.
Publication Year :
2021

Abstract

Background: To address the need for easy and reliable species classification in plant genetic resources collections, we assessed the potential of five classifiers (Random Forest, Neighbour-Joining, 1-Nearest Neighbour, a conservative variety of 3-Nearest Neighbours and Naive Bayes) We investigated the effects of the number of accessions per species and misclassification rate on classification success, and validated theirs generic value results with three complete datasets. Results: We found the conservative variety of 3-Nearest Neighbours to be the most reliable classifier when varying species representation and misclassification rate. Through the analysis of the three complete datasets, this finding showed generic value. Additionally, we present various options for marker selection for classification taks such as these. Conclusions: Large-scale genomic data are increasingly being produced for genetic resources collections. These data are useful to address species classification issues regarding crop wild relatives, and improve genebank documentation. Implementation of a classification method that can improve the quality of bad datasets without gold standard training data is considered an innovative and efficient method to improve gene bank documentation. [ABSTRACT FROM AUTHOR]

Details

Language :
English
ISSN :
14712105
Volume :
22
Issue :
1
Database :
Academic Search Index
Journal :
BMC Bioinformatics
Publication Type :
Academic Journal
Accession number :
149572439
Full Text :
https://doi.org/10.1186/s12859-021-04018-6