Detecting pattern-based outliers
Pattern Recognition LettersPublished 8 August 2003Open access
Tianming Hu, Sam Yuan Sung
Citations58
SJR quartileQ4
SJR score0.11
SNIP0.06
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
This work proposes two techniques: one to identify the two patterns called low density regularity and the other to detect the corresponding outliers.
Abstract
Outlier detection targets those exceptional data that deviate from the general pattern. Besides high density clustering, there is another pattern called low density regularity. Thus, there are two types of outliers w.r.t. them. We propose two techniques: one to identify the two patterns and the other to detect the corresponding outliers.
Keywords
Computer Science
A density-based algorithm for discovering clusters in large spatial Databases with Noise
19,114 Citations1996Martin Ester, Hans‐Peter Kriegel +2 more
DBSCAN, a new clustering algorithm relying on a density-based notion of clusters which is designed to discover clusters of arbitrary shape, is presented which requires only one input parameter and supports the user in determining an appropriate value for it.
Journal of the Royal Statistical Society Series A (Statistics in Society)Statistics for Spatial Data.
5,160 Citations1993Mike Rees, Noel Cressie
LOF
3,862 Citations2000Markus Breunig, Hans‐Peter Kriegel +2 more
This paper contends that for many scenarios, it is more meaningful to assign to each object a degree of being an outlier, called the local outlier factor (LOF), and gives a detailed formal analysis showing that LOF enjoys many desirable properties.
Identification of Outliers
2,808 Citations1980D. M. Hawkins
A computer normalizes the one or more sets of historical data points and creates a first visual representation corresponding to the first set of the oneor more sets and the second set of additional points.
Automatic subspace clustering of high dimensional data for data mining applications
2,386 Citations1998Rakesh Agrawal, Johannes Gehrke +2 more
CLIQUE is presented, a clustering algorithm that satisfies each of these requirements of data mining applications including the ability to find clusters embedded in subspaces of high dimensional data, scalability, end-user comprehensibility of the results, non-presumption of any canonical data distribution, and insensitivity to the order of input records.
Journal of the Royal Statistical Society Series A (Statistics in Society)Outliers in Statistical Data.
1,698 Citations1995Anthony C. Atkinson, V. Barnett +1 more
Outliers in Statistical Data, 3rd edition by V. Barnett and T. Lewis.
Journal of the American Statistical AssociationA First Course in Probability.
1,361 Citations1984Malcolm J. Sherman, Sheldon M. Ross
The VLDB JournalDistance-based outliers: algorithms and applications
1,196 Citations2000Edwin M. Knorr, Raymond T. Ng +1 more
Outlier detection can be done efficiently for large datasets, and for k-dimensional datasets with large values of k, and it is shown that outlier detection is a meaningful and important knowledge discovery task.
IEEE Transactions on Knowledge and Data EngineeringCLARANS: a method for clustering objects for spatial data mining
1,184 Citations2002Raymond T. Ng, Jiawei Han
A new clustering method is proposed, called CLARANS, whose aim is to identify spatial structures that may be present in the data, and two spatial data mining algorithms that aim to discover relationships between spatial and nonspatial attributes are developed.
Outlier detection for high dimensional data
1,034 Citations2001Charų C. Aggarwal, Philip S. Yu
New techniques for outlier detection are discussed which find the outliers by studying the behavior of projections from the data set by studying the behavior of projections from the data set.
Journal of the Royal Statistical Society Series C (Applied Statistics)A Kernel Method for Smoothing Point Process Data
606 Citations1985Peter J. Diggle
Journal of the Royal Statistical Society Series B (Statistical Methodology)Multivariate Discriminant Analysis and Maximum Penalized Likelihood Density Estimation
16 Citations1995Vincent Granville, J.P. Rasson
The final classification algorithm referred to as APML for approximate penalized maximum likelihood compares favourably in terms of error rate and time efficiency with other algorithms tested, including multinormal, nearest neighbour and convex hull classifiers.
