| dc.contributor.author | Pérez Rodriguez, Javier | |
| dc.contributor.author | García-Pedrajas, Nicolás | |
| dc.date.accessioned | 2024-03-18T14:41:45Z | |
| dc.date.available | 2024-03-18T14:41:45Z | |
| dc.date.issued | 2016-03 | |
| dc.identifier.citation | Pérez-Rodríguez, Javier & García-Pedrajas, Nicolás. (2016). Stepwise approach for combining many sources of evidence for site-recognition in genomic sequences. BMC Bioinformatics. 17. 10.1186/s12859-016-0968-y. | es |
| dc.identifier.issn | 1471-2105 | |
| dc.identifier.uri | https://hdl.handle.net/20.500.12412/5471 | |
| dc.description.abstract | Background: Recognizing the different functional parts of genes, such as promoters, translation initiation sites,
donors, acceptors and stop codons, is a fundamental task of many current studies in Bioinformatics. Currently, the
most successful methods use powerful classifiers, such as support vector machines with various string kernels.
However, with the rapid evolution of our ability to collect genomic information, it has been shown that combining
many sources of evidence is fundamental to the success of any recognition task. With the advent of next-generation
sequencing, the number of available genomes is increasing very rapidly. Thus, methods for making use of such large
amounts of information are needed.
Results: In this paper, we present a methodology for combining tens or even hundreds of different classifiers for an
improved performance. Our approach can include almost a limitless number of sources of evidence. We can use the
evidence for the prediction of sites in a certain species, such as human, or other species as needed. This approach can
be used for any of the functional recognition tasks cited above. However, to provide the necessary focus, we have
tested our approach in two functional recognition tasks: translation initiation site and stop codon recognition. We
have used the entire human genome as a target and another 20 species as sources of evidence and tested our
method on five different human chromosomes. The proposed method achieves better accuracy than the best
state-of-the-art method both in terms of the geometric mean of the specificity and sensitivity and the area under the
receiver operating characteristic and precision recall curves. Furthermore, our approach shows a more principled way
for selecting the best genomes to be combined for a given recognition task.
Conclusions: Our approach has proven to be a powerful tool for improving the performance of functional site
recognition, and it is a useful method for combining many sources of evidence for any recognition task in
Bioinformatics. The results also show that the common approach of heuristically choosing the species to be used as
source of evidence can be improved because the best combinations of genomes for recognition were those not
usually selected. Although the experiments were performed for translation initiation site and stop codon recognition,
any other recognition task may benefit from our methodology. | es |
| dc.language.iso | eng | es |
| dc.rights | Attribution-NonCommercial-NoDerivatives 4.0 Internacional | * |
| dc.rights.uri | http://creativecommons.org/licenses/by-nc-nd/4.0/ | * |
| dc.title | Stepwise approach for combining many sources of evidence for site-recognition in genomic sequences | es |
| dc.type | article | es |
| dc.identifier.doi | 10.1186/s12859-016-0968-y | |
| dc.issue.number | 1 | es |
| dc.journal.title | BMC Bioinformatics | es |
| dc.relation.projectID | This work has been financed in part by Project TIN-2011-22967 of the Spanish Ministry of Science and Innovation and Excellence in Research Projects P09-TIC-4623 and P07-TIC-2682 of the Junta de Andalucía | es |
| dc.rights.accessRights | openAccess | es |
| dc.subject.keyword | Site recognition | es |
| dc.subject.keyword | Combination of evidence | es |
| dc.subject.keyword | Translation initiation site recognition | es |
| dc.subject.keyword | Stop codon recognition | es |
| dc.volume.number | 17 | es |