Adapting linear discriminant analysis to the paradigm of learning from label proportions
Author:
Pérez Ortiz, María; Gutiérrez, Pedro Antonio; Carbonero Ruz, Mariano
; Hervás Martínez, César
ISBN:
978-150904240-1DOI:
10.1109/SSCI.2016.7850150Date:
2017Abstract:
The recently coined term “learning from label pro-portions” refers to a new learning paradigm where training datais given by groups (also denoted as “bags”), and the only knowninformation is the label proportion of each bag. The aim is thento construct a classification model to predict the class label of anindividual instance, which differentiates this paradigm from theone of multi-instance learning. This learning setting presents verydifferent applications in political science, marketing, healthcareand, in general, all fields in relation with anonymous data. Inthis paper, two new strategies are proposed to tackle this kind of problems. Both proposals are based on the optimisation of patternclass memberships using the data distribution in each bag and theknown label proportions. To do so, linear discriminant analysishas been reformulated to work with non-crisp class memberships.The experimental part of this paper sets different objetives: 1)study the difference in performance, comparing our proposalsand the fully supervised setting, 2) analyse the potential benefitsof refining class memberships by the proposed approaches, and 3)test the influence of other factors in the performance, such as thenumber of classes or the bag size. The results of these experimentsare promising, but further research should be encouraged forstudying more complex data configurations.
The recently coined term “learning from label pro-portions” refers to a new learning paradigm where training datais given by groups (also denoted as “bags”), and the only knowninformation is the label proportion of each bag. The aim is thento construct a classification model to predict the class label of anindividual instance, which differentiates this paradigm from theone of multi-instance learning. This learning setting presents verydifferent applications in political science, marketing, healthcareand, in general, all fields in relation with anonymous data. Inthis paper, two new strategies are proposed to tackle this kind of problems. Both proposals are based on the optimisation of patternclass memberships using the data distribution in each bag and theknown label proportions. To do so, linear discriminant analysishas been reformulated to work with non-crisp class memberships.The experimental part of this paper sets different objetives: 1)study the difference in performance, comparing our proposalsand the fully supervised setting, 2) analyse the potential benefitsof refining class memberships by the proposed approaches, and 3)test the influence of other factors in the performance, such as thenumber of classes or the bag size. The results of these experimentsare promising, but further research should be encouraged forstudying more complex data configurations.
Collections
Files in this item



