Discovering frequent behaviors: time is an essential element of the context
Knowledge and Information SystemsPublished 20 November 2010Open access
Bashar Saleh, Florent Masséglia
Citations31
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
The definition of solid itemsets are introduced, which represent coherent and compact behaviors over specific periods, and an algorithm for their extraction is proposed, called Sim, which is proposed for their extraction.
Abstract
International audience
Keywords
Computer Science
Mining association rules between sets of items in large databases
14,720 Citations1993Rakesh Agrawal, Tomasz Imieliński +1 more
An efficient algorithm is presented that generates all significant association rules between items in the database of customer transactions and incorporates buffer management and novel estimation and pruning techniques.
ACM SIGMOD RecordMining frequent patterns without candidate generation
6,360 Citations2000Jiawei Han, Jian Pei +1 more
ACM SIGMOD RecordMining association rules between sets of items in large databases
4,475 Citations1993Rakesh Agrawal, Tomasz Imieliński +1 more
Mining frequent patterns without candidate generation
3,195 Citations2000Jiawei Han, Jian Pei +1 more
This study proposes a novel frequent pattern tree (FP-tree) structure, which is an extended prefix-tree structure for storing compressed, crucial information about frequent patterns, and develops an efficient FP-tree-based mining method, FP-growth, for mining the complete set of frequent patterns by pattern fragment growth.
Lecture notes in computer scienceDiscovering Frequent Closed Itemsets for Association Rules
1,361 Citations1999Nicolas Pasquier, Yves Bastide +2 more
This paper proposes a new algorithm, called A-Close, using a closure mechanism to find frequent closed itemsets, and shows that this approach is very valuable for dense and/or correlated data that represent an important part of existing databases.
Efficiently mining long patterns from databases
1,297 Citations1998Roberto J. Bayardo
A pattern-mining algorithm that scales roughly linearly in the number of maximal patterns embedded in a database irrespective of the length of the longest pattern, compared with previous algorithms that scale exponentially with longest pattern length.
Sampling Large Databases for Association Rules
1,079 Citations1996Hannu Toivonen
New algorithms that reduce the database activity considerably by picking a Random sample, to find using this sample all association rules that probably hold in the whole database, and then to verify the results with the rest of the database.
CLOSET+
582 Citations2003Jianyong Wang, Jiawei Han +1 more
The VLDB JournalOverview of multidatabase transaction management
572 Citations1992Yuri Breitbart, Héctor García-Molina +1 more
It is argued that the multidatabase research will become increasingly important in the coming years and basic research issues in multid atabase transaction management are outlined, followed by a discussion of open problems and practical implications.
Mining Frequent Patterns in Data Streams at Multiple Time Granularities
499 Citations2002Chris Giannella, Jiawei Han +2 more
This paper proposes computing and maintaining all the frequent patterns and dynamically updating them with the incoming data streams and incrementally maintain tilted-time windows for each pattern at multiple time granularities.
A clustering-based approach for discovering interesting places in trajectories
421 Citations2008Andrey Tietbohl Palma, Vânia Bogorny +2 more
The proposed solution is a spatio-temporal clustering method, based on speed, to work with single trajectories, and it is shown that the computation of stops using the concept of speed can be interesting for several applications.
Cyclic association rules
404 Citations2002B. Ozden, Sridhar Ramaswamy +1 more
This work devise a new technique called cycle pruning, which reduces the amount of time needed to find cyclic association rules by studying the interaction between association rules and time, and presents two new algorithms for discovering such rules.
IEEE Transactions on Knowledge and Data EngineeringA survey of temporal knowledge discovery paradigms and methods
393 Citations2002John F. Roddick, Myra Spiliopoulou
The confluence of temporal databases and data mining is investigated, the work to date is surveyed, and the issues involved and the outstanding problems in temporal data mining are explored.
Finding recent frequent itemsets adaptively over online data streams
325 Citations2003Joong Hyuk Chang, Won Suk Lee
This paper proposes a data mining method for finding recent frequent itemsets adaptively over an online data stream by decaying the old occurrences of each itemset as time goes by.
Moment: Maintaining Closed Frequent Itemsets over a Stream Sliding Window
297 Citations2005Yün Chi, Haixun Wang +2 more
A compact data structure, the closed enumeration tree (CET), is introduced, to maintain a dynamically selected set of item-sets over a sliding-window that consists of a boundary between closed frequent itemsets and the rest of the itemsets.
Group-by skyline query processing in relational engines
281 Citations2009Ming-Hay Luk, Man Lung Yiu +1 more
The composition of a query plan for a group-by skyline query is examined and the missing cost model for the BBS algorithm is developed and Experimental results show that the techniques are able to devise the best query plans for a variety of group- by skyline queries.
Efficient elastic burst detection in data streams
277 Citations2003Yunyue Zhu, Dennis Shasha
A general data structure for detecting interesting aggregates over elastic windows in near linear time and applications of the algorithm for detecting Gamma Ray Bursts in large-scale astrophysical data are presented.
IEEE Transactions on Knowledge and Data EngineeringMAFIA: a maximal frequent itemset algorithm
265 Citations2005Doug Burdick, Manuel Calimlim +3 more
The experiments show that MAFIA performs best when mining long itemsets and outperforms other algorithms on dense data by a factor of three to 30.
IEEE Transactions on Knowledge and Data EngineeringFast and memory efficient mining of frequent closed itemsets
229 Citations2006Claudio Lucchese, Salvatore Orlando +1 more
A new scalable algorithm for discovering closed frequent itemsets, a lossless and condensed representation of all the frequent itemets that can be mined from a transactional database that outperforms other state-of-the-art algorithms like CLOSET+ and FP-CLOSE, in some cases by more than one order of magnitude.
An approach to discovering temporal association rules
223 Citations2000Juan M. Ale, Gustavo Rossi
The notion of association rules incorporating time incorporating time to the frequent itemsets discovered is expanded and the concept of temporal support is introduced.
Data & Knowledge EngineeringDiscovering calendar-based temporal association rules
189 Citations2002Yingjiu Li, Peng Ning +2 more
Knowledge and Information SystemsCatch the moment: maintaining closed frequent itemsets over a data stream sliding window
173 Citations2006Yün Chi, Haixun Wang +2 more
Elsevier eBooksA Regression-Based Temporal Pattern Mining Scheme for Data Streams
134 Citations2003Wei‐Guang Teng, Ming-Syan Chen⋆ +1 more
A regression-based algorithm to mine frequent temporal patterns for data streams, called algorithm FTP-DS (Frequent Temporal Patterns of Data Streams), which is able to not only conduct mining with variable time intervals but also perform trend detection effectively.
On mining general temporal association rules in a publication database
81 Citations2002Chang-Hung Lee, Cheng-Ru Lin +1 more
An innovative algorithm, progressive-partition-miner (PPM), is proposed, to discover general temporal association rules in a publication database and is designed to employ a filtering threshold in each partition to prune out those cumulatively infrequent 2-itemsets at an early stage.
Mining Frequent Itemsets in a Stream
68 Citations2007Toon Calders, Nele Dexters +1 more
Lecture notes in computer scienceMining Temporal Features in Association Rules
50 Citations1999Xiaohong Chen, Ilias Petrounias
The major concerns in this paper are the identification of the valid period and periodicity of patterns and more specifically association rules.
Knowledge and Information SystemsComputing the minimum-support for mining frequent patterns
41 Citations2007Shichao Zhang, Xindong Wu +2 more
This paper proposes a computational strategy for identifying frequent itemsets, consisting of polynomial approximation and fuzzy estimation, which automatically generate actual minimum-supports according to users’ mining requirements.
Data Mining and Knowledge DiscoveryWeb usage mining: extracting unexpected periods from web logs
39 Citations2007Florent Masséglia, Pascal Poncelet +2 more
This paper proposes a specific data mining process (in particular, to extract frequent behaviour patterns) in order to reveal the densest periods automatically and extracts the frequent sequential patterns related to the extracted periods.
Choice Reviews OnlineSuccesses and new directions in data mining
35 Citations2008
Capturing defining research on topics such as fuzzy set theory, clustering algorithms, semi-supervised clustering, modeling and managing data mining patterns, and sequence motif mining, this book is an indispensable resource for library collections.
Lecture notes in computer scienceFast Burst Correlation of Financial Data
18 Citations2005Michail Vlachos, Kun‐Lung Wu +2 more
The proposed methods and data-structures can find applications for anomaly or novelty detection in telecommunications and network traffic, as well as in medical data.
Knowledge and Information SystemsCharacterizing pattern preserving clustering
17 Citations2008Hui Xiong, Michael Steinbach +2 more
A new approach for clustering—pattern preserving clustering—which produces more easily interpretable and usable clusters and illustrates how patterns, if preserved, can aid cluster interpretation.
Efficient itemset generator discovery over a stream sliding window
14 Citations2009Chuancong Gao, Jianyong Wang
An efficient algorithm called StreamGen is devised to mine frequent itemset generators over a stream sliding window to generate simple classification rules according to the MDL principle, and outperforms other state-of-the-art algorithms which perform similar tasks in terms of both runtime and memory usage efficiency.
Maintenance of maximal frequent itemsets in large databases
12 Citations2007Wang Lian, David W. Cheung +1 more
This paper clearly addresses the relationships between old and new maximal frequent itemsets and proposes an algorithm IMFI, which is based on these relationships to reuse previously discovered knowledge, which follows a top-down mechanism rather than traditional bottom-up methods to produce fewer candidates.
Lecture notes in computer scienceMining Time-Profiled Associations: An Extended Abstract
6 Citations2005Jin Soung Yoo, Pusheng Zhang +1 more
A novel one-step algorithm to unify the generation of statistical parameter sequences and sequence retrieval and substantially reduces the itemset search space by pruning candidate itemsets based on the monotone property of the lower bounding measure of the sequence of statistical parameters.
Lecture notes in computer scienceFalse-Negative Frequent Items Mining from Data Streams with Bursting
6 Citations2005Zhihong Chong, Jeffrey Xu Yu +3 more
This paper proposes a new false-negative frequent items mining algorithm, called Loss-Negative, for handling bursting in data streams, and presents theoretical bound of the new algorithm, and analyzes the possibility of minimization of missing frequent items, in terms of two possibilities.
Time Aware Mining of Itemsets
2 Citations2008Bashar Saleh, Florent Masséglia
The definition of solid itemsets are introduced, which represent a coherent and compact behavior over a specific period, and an algorithm for their extraction is proposed, which may find many applications in sensitive domains such as fraud or intrusion detection.
