Back to Search
Start Over
ORFcor: Identifying and Accommodating ORF Prediction Inconsistencies for Phylogenetic Analysis
- Source :
- PLoS ONE, Vol 8, Iss 3, p e58387 (2013), PLoS ONE
- Publication Year :
- 2013
- Publisher :
- Public Library of Science (PLoS), 2013.
-
Abstract
- The high-throughput annotation of open reading frames (ORFs) required by modern genome sequencing projects necessitates computational protocols that sometimes annotate orthologous ORFs inconsistently. Such inconsistencies hinder comparative analyses by non-uniformly extending or truncating 5′ and/or 3′ sequence ends, causing ORFs that are in fact identical to artificially diverge. Whereas strategies exist to correct such inconsistencies during whole-genome annotation, equivalent software designed to correct subsets of these data without genome reannotation is lacking. We therefore developed ORFcor, which corrects annotation inconsistencies using consensus start and stop positions derived from sets of closely related orthologs. ORFcor corrects inconsistent ORF annotations in diverse test datasets with specificities and sensitivities approaching 100% when sufficiently related orthologs (e.g., from the same taxonomic family) are available for comparison. The ORFcor package is implemented in Perl, multithreaded to handle large datasets, includes related scripts to facilitate high-throughput phylogenomic analyses, and is freely available at www.currielab.wisc.edu/downloads.html.
- Subjects :
- Gene prediction
Sequence Databases
lcsh:Medicine
Computational biology
Biology
Sensitivity and Specificity
Genome
DNA sequencing
Open Reading Frames
03 medical and health sciences
Annotation
0302 clinical medicine
Predictive Value of Tests
Genome Analysis Tools
Genome Databases
Evolutionary Systematics
ORFS
Gene Prediction
lcsh:Science
Phylogeny
030304 developmental biology
computer.programming_language
Genetics
Evolutionary Biology
0303 health sciences
Multidisciplinary
Models, Genetic
lcsh:R
Computational Biology
Molecular Sequence Annotation
Genomics
Comparative Genomics
Phylogenetics
Open reading frame
lcsh:Q
Perl
Sequence Analysis
computer
Algorithms
Software
030217 neurology & neurosurgery
Research Article
Subjects
Details
- ISSN :
- 19326203
- Volume :
- 8
- Database :
- OpenAIRE
- Journal :
- PLoS ONE
- Accession number :
- edsair.doi.dedup.....49412a279db6d230232353e2b160e7db