| dc.contributor.author | García-Pedrajas, Nicolás | |
| dc.contributor.author | Haro-García, Aida de | |
| dc.contributor.author | Pérez Rodriguez, Javier | |
| dc.date.accessioned | 2024-03-18T14:41:09Z | |
| dc.date.available | 2024-03-18T14:41:09Z | |
| dc.date.issued | 2013-04 | |
| dc.identifier.citation | García-Pedrajas, Nicolás & de Haro Garcia, Aida & Pérez-Rodríguez, Javier. (2013). A Scalable Memetic Algorithm for Simultaneous Instance and Feature Selection. Evolutionary computation. 22. 10.1162/EVCO_a_00102. | es |
| dc.identifier.issn | 1530-9304 | |
| dc.identifier.uri | https://hdl.handle.net/20.500.12412/5469 | |
| dc.description.abstract | Instance selection is becoming increasingly relevant due to the huge amount of data that
is constantly produced in many fields of research. At the same time, most of the recent
pattern recognition problems involve highly complex datasets with a large number of
possible explanatory variables. For many reasons, this abundance of variables significantly harms classification or recognition tasks. There are efficiency issues, too, because
the speed of many classification algorithms is largely improved when the complexity of
the data is reduced. One of the approaches to address problems that have too many features or instances is feature or instance selection, respectively. Although most methods
address instance and feature selection separately, both problems are interwoven, and
benefits are expected from facing these two tasks jointly. This paper proposes a new
memetic algorithm for dealing with many instances and many features simultaneously
by performing joint instance and feature selection. The proposed method performs four
different local search procedures with the aim of obtaining the most relevant subsets
of instances and features to perform an accurate classification. A new fitness function
is also proposed that enforces instance selection but avoids putting too much pressure
on removing features. We prove experimentally that this fitness function improves the
results in terms of testing error. Regarding the scalability of the method, an extension of
the stratification approach is developed for simultaneous instance and feature selection.
This extension allows the application of the proposed algorithm to large datasets. An extensive comparison using 55 medium to large datasets from the UCI Machine Learning
Repository shows the usefulness of our method. Additionally, the method is applied
to 30 large problems, with very good results. The accuracy of the method for classimbalanced problems in a set of 40 datasets is shown. The usefulness of the method is
also tested using decision trees and support vector machines as classification methods. | es |
| dc.language.iso | eng | es |
| dc.rights | Attribution-NonCommercial-NoDerivatives 4.0 Internacional | * |
| dc.rights.uri | http://creativecommons.org/licenses/by-nc-nd/4.0/ | * |
| dc.title | A Scalable Memetic Algorithm for Simultaneous Instance and Feature Selection | es |
| dc.type | article | es |
| dc.identifier.doi | 10.1162/EVCO_a_00102 | |
| dc.issue.number | 1 | es |
| dc.journal.title | Evolutionary Computation | es |
| dc.page.initial | 1 | es |
| dc.page.final | 45 | es |
| dc.relation.projectID | This work was supported in part by the Project TIN2011-22967 of the Spanish Ministry of Science and Innovation and the project P09-TIC-4623 of the Junta de Andalucía | es |
| dc.rights.accessRights | openAccess | es |
| dc.subject.keyword | Memetic algorithms | es |
| dc.subject.keyword | Instance selection | es |
| dc.subject.keyword | Feature selection | es |
| dc.subject.keyword | Scaling-up | es |
| dc.volume.number | 22 | es |