Probabilistic estimation-based data mining for discovering insurance risks
IEEE Intelligent Systems and their ApplicationsPublished 1 November 1999
Chid Apte, Emiliano Grossman, Edwin Pednault, Barry K. Rosen, Fateh A. Tipu, Brandyn White
Citations36
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
IBM's underwriting profitability analysis application mines property and casualty insurance policy and claims data to construct predictive models for insurance risks and UPA uses the ProbE data-mining kernel to discover risk-characterization rules by analyzing large, noisy data sets.
Abstract
IBM's underwriting profitability analysis application mines property and casualty insurance policy and claims data to construct predictive models for insurance risks. UPA uses the ProbE data-mining kernel to discover risk-characterization rules by analyzing large, noisy data sets.
Keywords
Computer Science
The Nature of Statistical Learning Theory
39,279 Citations1995Vladimir Vapnik
TechnometricsStatistical Learning Theory
26,913 Citations1999Yuhai Wu, Vladimir Vapnik
Presenting a method for determining the necessary and sufficient conditions for consistency of learning process, the author covers function estimates from small data pools, applying these estimations to real-life problems, and much more.
BiometricsClassification and Regression Trees.
23,841 Citations1984Alexander Gordon, Leo Breiman +3 more
C4.5: Programs for Machine Learning
23,665 Citations1992J. R. Quinlan
A complete guide to the C4.5 system as implemented in C for the UNIX environment, which starts from simple core learning methods and shows how they can be elaborated and extended to deal with typical problems such as missing data and over hitting.
Programs for Machine Learning
5,794 Citations1994Steven L. Salzberg, Alberto M. Segre
In his new book, C4.5: Programs for Machine Learning, Quinlan has put together a definitive, much needed description of his complete system, including the latest developments, which will be a welcome addition to the library of many researchers and students.
Journal of the Royal Statistical Society Series C (Applied Statistics)An Exploratory Technique for Investigating Large Quantities of Categorical Data
2,841 Citations1980Gordon V. Kass
The technique set out in the paper, CHAID, is an offshoot of AID (Automatic Interaction Detection) designed for a categorized dependent variable with built-in significance testing, multi-way splits, and a new type of predictor which is especially useful in handling missing information.
Elsevier eBooksIntroduction to Robust Estimation and Hypothesis Testing
2,109 Citations2012Rand R. Wilcox
Journal of the American Statistical AssociationLoss Models: From Data to Decisions
1,288 Citations1999James D. Broffitt, Stuart A. Klugman +2 more
SPRINT: A Scalable Parallel Classifier for Data Mining
781 Citations1996John Shafer, Rakesh Agrawal +1 more
A new decision-tree-based classification algorithm, called SPRINT, is presented that removes all of the memory restrictions, and is fast and scalable, and designed to be easily parallelized, allowing many processors to work together to build a single consistent model.
Large datasets lead to overly complex models: an explanation and a solution
68 Citations1998Tim Oates, David Jensen
This paper proposes a general solution to the problem of large datasets and compact models based on a statistical technique known as randomization testing, and empirically evaluates its utility.
Future Generation Computer SystemsA statistical perspective on data mining
57 Citations1997J. R. M. Hosking, Edwin Pednault +1 more
Comparing three approaches to machine learning that have developed largely independently: classical statistics, Vapnik's statistical learning theory, and computational learning theory concludes that statisticians and data miners can profit by studying each other's methods and using a judiciously chosen combination of them.
