On Exhaustive Enumeration of Redescriptions
par
S3 351
Sciences 3
Redescription mining is a data mining task aiming at finding multiple descriptions of the same set of objects. Many methods and algorithms to mine redescriptions have been developed, most of them making use of heuristics in order to get few rules quickly. However, for certain tasks which require not only a good precision but also a high recall, it can be interesting to mine redescriptions exhaustively. In this paper, we take advantage of the Charm-L algorithm to perform this task. The Charm-L algorithm allows the exhaustive extraction of exact redescriptions. We make use of Formal Concept Analysis (FCA) to extend the Charm-L algorithm in order to mine both exact and approximate redescriptions in an exhaustive manner (with regards to given frequency and similarity thresholds), introducing pruning methods to handle the mining of approximate redescriptions without compromising on exhaustivity. We show that this task is possible in a reasonable time even for large datasets, and discuss how it can be extended.