Text analytics in industry: Challenges, desiderata and trends
Computers in IndustryPublished 30 December 2015Open access
Ashwin Ittoo, Le-Minh Nguyen, Antal van den Bosch
Citations127
SJR quartileQ1
SJR score2.21
SNIP3.09
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
A systematic review of the current state of the art in the application of text analytics in industry and a set of desiderata that text analytics techniques should satisfy in order to alleviate challenges and to ensure their successful deployment in industry are provided.
Abstract
\n Contains fulltext :\n 159021.pdf (Publisher’s version ) (Closed access)\n
Keywords
Computer ScienceBusiness, Management and Accounting
Glove: Global Vectors for Word Representation
33,769 Citations2014Jeffrey Pennington, Richard Socher +1 more
A new global logbilinear regression model that combines the advantages of the two major model families in the literature: global matrix factorization and local context window methods and produces a vector space with meaningful substructure.
Neural NetworksDeep learning in neural networks: An overview
18,033 Citations2014Jürgen Schmidhuber
This historical survey compactly summarizes relevant work, much of it from the previous millennium, review deep supervised learning, unsupervised learning, reinforcement learning & evolutionary computation, and indirect search for short programs encoding deep and large networks.
ACM SIGKDD Explorations NewsletterThe WEKA data mining software
17,849 Citations2009Mark Hall, Eibe Frank +4 more
This paper provides an introduction to the WEKA workbench, reviews the history of the project, and, in light of the recent 3.6 stable release, briefly discusses what has been added since the last stable version (Weka 3.4) released in 2003.
arXiv (Cornell University)Sequence to Sequence Learning with Neural Networks
13,362 Citations2014Ilya Sutskever, Oriol Vinyals +1 more
Recursive Deep Models for Semantic Compositionality Over a Sentiment Treebank
6,774 Citations2013Richard Socher, Alex Perelygin +5 more
A Sentiment Treebank that includes fine grained sentiment labels for 215,154 phrases in the parse trees of 11,855 sentences and presents new challenges for sentiment compositionality, and introduces the Recursive Neural Tensor Network.
Information Systems ResearchDeterminants of Perceived Ease of Use: Integrating Control, Intrinsic Motivation, and Emotion into the Technology Acceptance Model
6,408 Citations2000Viswanath Venkatesh
This work presents and tests an anchoring and adjustment-based theoretical model of the determinants of system-specific perceived ease of use, and proposes control, intrinsic motivation, and emotion as anchors that determine early perceptions about the ease ofuse of a new system.
The MIT Press eBooksFast Training of Support Vector Machines Using Sequential Minimal Optimization
5,462 Citations1998John Platt
NLTK
3,299 Citations2002Edward Loper, Steven Bird
The Natural Language Toolkit has been rewritten, simplifying many linguistic data structures and taking advantage of recent enhancements in the Python language.
Linguistic Regularities in Continuous Space Word Representations
2,882 Citations2013Tomáš Mikolov, Wen-tau Yih +1 more
The vector-space word representations that are implicitly learned by the input-layer weights are found to be surprisingly good at capturing syntactic and semantic regularities in language, and that each relationship is characterized by a relation-specific vector offset.
Accurate methods for the statistics of surprise and coincidence
2,688 Citations1993Ted Dunning
A Fast and Accurate Dependency Parser using Neural Networks
1,879 Citations2014Danqi Chen, Christopher D. Manning
This work proposes a novel way of learning a neural network classifier for use in a greedy, transition-based dependency parser that can work very fast, while achieving an about 2% improvement in unlabeled and labeled attachment scores on both English and Chinese datasets.
European Journal of Information SystemsMeasuring information systems success: models, dimensions, measures, and interrelationships
1,860 Citations2008Stacie Petter, William DeLone +1 more
This work builds on the prior research related to IS success by summarizing the measures applied to the evaluation of IS success and by examining the relationships that comprise the D&M IS success model in both individual and organizational contexts.
IEEE Transactions on Systems Man and CyberneticsDevelopment and application of a metric on semantic nets
1,810 Citations1989Roy Rada, Hafedh Mili +2 more
Experiments in which distance is applied to pairs of concepts and to sets of concepts in a hierarchical knowledge base show the power of hierarchical relations in representing information about the conceptual distance between concepts.
The MIT Press eBooksCombining Local Context and WordNet Similarity for Word Sense Identification
1,638 Citations1998Claudia Leacock
This chapter contains sections titled: Introducfion, Training and Testing Data, Experiment 1: The Local Context Classifier, Experiment 2: Measuring Word Similarity In Wordnet, and Combining Local Context and Wordnet Similarity Measures.
Proceedings of the 40th Annual Meeting on Association for Computational Linguistics
1,597 Citations2002Pierre Isabelle
Corpus-based and knowledge-based measures of text semantic similarity
1,189 Citations2006Rada Mihalcea, Courtney D. Corley +1 more
This paper shows that the semantic similarity method out-performs methods based on simple lexical matching, resulting in up to 13% error rate reduction with respect to the traditional vector-based similarity metric.
Automatically Constructing a Corpus of Sentential Paraphrases.
1,114 Citations2005William B. Dolan, Chris Brockett
The creation of the recently-released Microsoft Research Paraphrase Corpus, which contains 5801 sentence pairs, each hand-labeled with a binary judgment as to whether the pair constitutes a paraphrase, is described.
Semantically Conditioned LSTM-based Natural Language Generation for Spoken Dialogue Systems
846 Citations2015Tsung-Hsien Wen, Milica Gašić +4 more
A statistical language generator based on a semantically controlled Long Short-term Memory (LSTM) structure that can learn from unaligned data by jointly optimising sentence planning and surface realisation using a simple cross entropy training criterion, and language variation can be easily achieved by sampling from output candidates.
IEEE Transactions on Knowledge and Data EngineeringSentence similarity based on semantic nets and corpus statistics
821 Citations2006Yuhua Li, David McLean +3 more
Experiments demonstrate that the proposed method provides a similarity measure that shows a significant correlation to human intuition and can be used in a variety of applications that involve text knowledge representation and discovery.
A Hierarchical Neural Autoencoder for Paragraphs and Documents
518 Citations2015Jiwei Li, Thang Luong +1 more
This paper introduces an LSTM model that hierarchically builds an embedding for a paragraph from embeddings for sentences and words, then decodes this embedding to reconstruct the original paragraph and evaluates the reconstructed paragraph using standard metrics to show that neural models are able to encode texts in a way that preserve syntactic, semantic, and discourse coherence.
Proceedings of the 32nd annual meeting on Association for Computational Linguistics
470 Citations1994James Pustejovsky
arXiv (Cornell University)Verb Semantics and Lexical Selection
460 Citations1994Zhibiao Wu, Martha Palmer
GATE
385 Citations2001Hamish Cunningham, Diana Maynard +2 more
GATE, a framework and graphical development environment which enables users to develop and deploy language engineering components and resources in a robust fashion, and can be used to develop applications and resources in multiple languages, based on its thorough Unicode support.
Elsevier eBooksInformation retrieval
326 Citations2003Sean Breen, Michel Manago +2 more
Case-Based Reasoning (CBR) technology has introduced intelligent sales support to e-commerce applications at Analog Devices, and CBR-related techniques for indexing and clustering a case base could be used to help customers refine their queries through a systematic analysis of their needs.
Minerva Access (University of Melbourne)NLTK: The Natural Language Toolkit
244 Citations2002Steven Bird, Edward Loper
Research portal (Tilburg University)Sentence Simplification by Monolingual Machine Translation
239 Citations2012Sander Wubben, Antal van den Bosch +1 more
By relatively careful phrase-based paraphrasing this model achieves similar simplification results to state-of-the-art systems, while generating better formed output, and argues that text readability metrics such as the Flesch-Kincaid grade level should be used with caution when evaluating the output of simplification systems.
arXiv (Cornell University)A Hierarchical Neural Autoencoder for Paragraphs and Documents
207 Citations2015Jiwei Li, Minh-Thang Luong +1 more
IEEE SoftwareUnderstanding service-oriented software
201 Citations2004Nicolas Gold, A. Mohan +2 more
The problems software engineers still face when working with service-oriented software are discussed and some new issues that they must consider, including how to address service provision difficulties and failures are introduced.
Proceedings of the 2006 Conference on Empirical Methods in Natural Language Processing
196 Citations2006Dan Jurafsky, Éric Gaussier
Experimenting with Distant Supervision for Emotion Classification
180 Citations2012Matthew Purver, Stuart Battersby
The method is suitable for some emotions (happiness, sadness and anger) but less able to distinguish others; and that different labelling conventions are more suitable forSome emotions than others.
Computers in IndustryToward a cloud-based manufacturing execution system for distributed manufacturing
178 Citations2014Petri Helo, Mikko Suorsa +2 more
A proposal for the core of architecture for next generation of MES solution is made and a pilot software tool has been developed to support the needs related to real time, cloud-based, light weight operation.
Computers in IndustryNatural language processing for aviation safety reports: From classification to interactive analysis
176 Citations2015Ludovic Tanguy, Nikola Tulechki +3 more
The different NLP techniques designed and used in collaboration between the CLLE-ERSS research laboratory and the CFH/Safety Data company to manage and analyse aviation incident reports are described.
Modeling Interestingness with Deep Neural Networks
175 Citations2014Jianfeng Gao, Patrick Pantel +3 more
The results on large-scale, real-world datasets show that the semantics of documents are important for modeling interest-ingness and that the DSSM leads to significant quality improvement on both tasks, outperforming not only the classic document models that do not use semantics but also state-of-the-art topic models.
A framework for analysis of data freshness
154 Citations2004Mokrane Bouzeghoub
This paper proposes a taxonomy based upon the nature of the data, the type of application and the synchronization policies underlying the multi-source information system, and analyzes the way freshness is defined and used in several types of systems.
Journal of Artificial Intelligence ResearchText Relatedness Based on a Word Thesaurus
147 Citations2010George Tsatsaronis, Iraklis Varlamis +1 more
Experimental evaluation shows that the proposed method outperforms every lexicon-based method of semantic relatedness in the selected tasks and the used data sets, and competes well against corpus-based and hybrid approaches.
Data & Knowledge EngineeringSyMSS: A syntax-based measure for short-text semantic similarity
130 Citations2011Jesús Oliva, J. Ignacio Serrano +2 more
The results show that SyMSS outperforms state-of-the-art methods in terms of rank correlation with human intuition, thus proving the importance of syntactic information in sentence semantic similarity computation.
Paraphrase recognition via dissimilarity significance classification
104 Citations2006Long Qiu, Min‐Yen Kan +1 more
Experimental results show that while being accurate at discerning non-paraphrasing dissimilarities, the implemented system is able to achieve higher paraphrase recall (93%), at an overall performance comparable to the alternatives.
Latent Structure Perceptron with Feature Induction for Unrestricted Coreference Resolution
97 Citations2012Eraldo Rezende Fernandes, Cícero dos Santos +1 more
A machine learning system based on large margin structure perceptron for unrestricted coreference resolution that introduces two key modeling techniques: latent coreference trees and entropy guided feature induction that achieves high performances with a linear learning algorithm.
Proceedings of the COLING/ACL on Interactive presentation sessions -
87 Citations2006
The COLING/ACL 2006 Interactive Presentations allowed developers of implemented computational linguistics software systems and libraries the opportunity to describe the design, development and functionality of their work in an interactive setting, and offered the presenters an opportunity to pro-actively engage more closely with the audience.
Information & ManagementUser interface features influencing overall ease of use and personalization
85 Citations2003Ram L. Kumar, Michael Alan Smith +1 more
This study examines user perceptions of gathering information on, locating, and buying compact discs through Web-based interfaces through part-time MBA students in Hong Kong to identify user interface features that a designer should emphasize and features for which personalization is likely to be beneficial.
arXiv (Cornell University)Dependency-based Convolutional Neural Networks for Sentence Embedding
74 Citations2015Mingbo Ma, Liang Huang +2 more
Proceedings of the 1st Workshop on Vector Space Modeling for Natural Language Processing
60 Citations2015Oren Melamud, Omer Levy +1916 more
GATE: an Architecture for Development of Robust HLT Applications
59 Citations2002Hamish Cunningham, Diana Maynard +2 more
Computers in IndustryTurning user generated health-related content into actionable knowledge through text analytics services
46 Citations2015Paloma Martı́nez, José L. Martínez +4 more
A system that monitors health social media streams is described, based on several text analytics processes supported, among others, by MeaningCloud, a commercial platform which provides meaning extraction from texts in a Software as a Service mode.
Computers in IndustryIntegrating a semantic-based retrieval agent into case-based reasoning systems: A case study of an online bookstore
33 Citations2015Jia Wei Chang, Ming-Che Lee +1 more
A novel framework for a case-based reasoning system that includes a collaborative filtering mechanism and a semantic-based case retrieval agent is presented that outperforms most previously described approaches.
Computers in IndustryA methodology for traffic-related Twitter messages interpretation
30 Citations2015Fábio C. Albuquerque, Marco A. Casanova +6 more
An automatic tweet interpretation tool, based on Machine Learning techniques, that achieves good performance for traffic-related tweets distributed by traffic authorities and news agencies is described.
Data & Knowledge EngineeringMinimally-supervised extraction of domain-specific part–whole relations using Wikipedia as knowledge-base
26 Citations2012Ashwin Ittoo, Gosse Bouma
It is shown that domain-specific part-whole relations cannot be conclusively classified in existing taxonomies, and a mechanism that mitigates the negative impact of semantic-drift on minimally-supervised algorithms is proposed.
Expert Systems with ApplicationsTowards automatic tweet generation: A comparative study from the text summarization perspective in the journalism genre
25 Citations2013Elena Lloret, Manuel Palomar
It was observed that although the original tweets were considered as model tweets with respect to their informativeness, they were not among the most interesting ones from a human viewpoint, and relying only on these tweets may not be the ideal way to communicate news through Twitter, especially if a more personalized and catchy way of reporting news wants to be performed.
Computers in IndustryText classification based filters for a domain-specific search engine
24 Citations2015Sebastian Schmidt, Steffen Schnitzer +1 more
Important aspects that need to be taken into consideration when implementing text classification-based filters in the industrial setting of a domain-specific search engine are presented.
Expert Systems with ApplicationsTerm extraction from sparse, ungrammatical domain-specific documents
19 Citations2012Ashwin Ittoo, Gosse Bouma
This article addresses the term extraction challenges posed by sparse, ungrammatical texts with domain-specific contents, such as customer complaint emails and engineers' repair notes, and presents ExtTerm, a novel term extraction system that outperforms a state-of-the-art baseline in extracting terms from a domain- specific, sparse and un grammatical real-life text collection.
Data & Knowledge EngineeringMinimally-supervised learning of domain-specific causal relations using an open-domain corpus as knowledge base
19 Citations2013Ashwin Ittoo, Gosse Bouma
It is shown that open-domain corpora can be exploited as knowledge bases to overcome data sparsity issues posed by domain-specific relation extraction, and that they enable substantial performance gains.
SentiKLUE: Updating a Polarity Classifier in 48 Hours
15 Citations2014Stefan Evert, Thomas Proisl +2 more
SentiKLUE is an update of the KLUE polarity classifier – which achieved good and robust results in SemEval-2013 with a simple feature set – implemented in 48 hours.
A Proactive Application to Monitor Truck Fleets
11 Citations2013Fábio da Costa Albuquerque, Marco A. Casanova +3 more
Basic requirements for proactive real-time monitoring applications are discussed, and how to structure and geo-reference unstructured text information available on the Internet with a focus on road conditions change and using available geocoding services is described.
PolibitsCombining Active and Ensemble Learning for Efficient Classification of Web Documents
9 Citations2014Steffen Schnitzer, Sebastian Schmidt +2 more
This article presents an approach that minimizes the human annotation effort by interactively incorporating human annotators into the training process via active learning of an ensemble learner.
Word Embeddings vs Word Types for Sequence Labeling: the Curious Case of CV Parsing
7 Citations2015Melanie Tosik, Carsten Lygteskov Hansen +2 more
New methods of improving Curriculum Vitae (CV) parsing for German documents are explored by applying recent research on the application of word embeddings in Natural Language Processing (NLP).
Computers in IndustryA distributional approach to open questions in market research
5 Citations2015Stefan Evert, Paul Greiner +2 more
The Klugator Engine is a system for semi-automatic identification, exploration and visualization of topics and sentiment in large collections of such free-text responses or other short text fragments, using state-of-the-art techniques of natural language processing and machine learning to transform textual input into a structured corpus.
Computers in IndustryAnalysing and evaluating the task of automatic tweet generation: Knowledge to business
3 Citations2015Elena Lloret, Manuel Palomar
Results show that automatically informative and interesting natural language tweets can be generated as a result of summarisation approaches and one can characterise good and bad tweets based on specific linguistic features not present in other types of tweets.
