Location extraction from tweets
Information Processing & ManagementPublished 16 November 2017Open access
Thi Bich Ngoc Hoang, Josiane Mothe
Citations72
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
This paper introduces a model to predict whether a tweet contains a location or not and shows that location prediction is a useful pre-processing step for location extraction, and defines a number of new tweet features and conducts an intensive evaluation.
Abstract
International audience
Keywords
Social SciencesComputer Science
ACM SIGKDD Explorations NewsletterThe WEKA data mining software
17,849 Citations2009Mark Hall, Eibe Frank +4 more
This paper provides an introduction to the WEKA workbench, reviews the history of the project, and, in light of the recent 3.6 stable release, briefly discusses what has been added since the last stable version (Weka 3.4) released in 2003.
Incorporating non-local information into information extraction systems by Gibbs sampling
3,035 Citations2005Jenny Rose Finkel, Trond Grenager +1 more
By using simulated annealing in place of Viterbi decoding in sequence models such as HMMs, CMMs, and CRFs, it is possible to incorporate non-local structure while preserving tractable inference.
Feature-rich part-of-speech tagging with a cyclic dependency network
2,851 Citations2003Kristina Toutanova, Dan Klein +2 more
A new part-of-speech tagger is presented that demonstrates the following ideas: explicit use of both preceding and following tag contexts via a dependency network representation, broad use of lexical features, and effective use of priors in conditional loglinear models.
Design challenges and misconceptions in named entity recognition
1,476 Citations2009Lev Ratinov, Dan Roth
Some of the fundamental design challenges and misconceptions that underlie the development of an efficient and robust NER system are analyzed, and several solutions to these challenges are developed.
Microblogging during two natural hazards events
1,415 Citations2010Sarah Vieweg, Amanda Hughes +2 more
Analysis of microblog posts generated during two recent, concurrent emergency events in North America via Twitter, a popular microblogging service, aims to inform next steps for extracting useful, relevant information during emergencies using information extraction (IE) techniques.
Named Entity Recognition in Tweets: An Experimental Study
1,203 Citations2011Alan Ritter, Sam Clark +1 more
The novel T-ner system doubles F1 score compared with the Stanford NER system, and leverages the redundancy inherent in tweets to achieve this performance, using LabeledLDA to exploit Freebase dictionaries as a source of distant supervision.
Artificial IntelligenceUnsupervised named-entity extraction from the Web: An experimental study
1,128 Citations2005Oren Etzioni, Michael Cafarella +6 more
An overview of KnowItAll's novel architecture and design principles is presented, emphasizing its distinctive ability to extract information without any hand-labeled training examples, and three distinct ways to address this challenge are presented and evaluated.
Model-based feedback in the language modeling approach to information retrieval
799 Citations2001ChengXiang Zhai, John Lafferty
This paper proposes and evaluates two different approaches to updating a query language model based on feedback documents, one based on a generative probabilistic model of feedback documents and onebased on minimization of the KL-divergence over feedback documents.
Proceedings of the International AAAI Conference on Web and Social MediaEvent Detection in Twitter
718 Citations2021Jianshu Weng, Bu‐Sung Lee
This paper attempts to tackle the challenges of event detection in Twitter with EDCoW (Event Detection with Clustering of Wavelet-based Signals), which builds signals for individual words by applying wavelet analysis on the frequencybased raw signals of the words.
Find me if you can
712 Citations2010Lars Bäckström, Eric Sun +1 more
Using user-supplied address data and the network of associations between members of the Facebook social network, an algorithm is introduced that predicts the location of an individual from a sparse set of located users with performance that exceeds IP-based geolocation.
Proceedings of the 22nd international conference on World Wide Web
684 Citations2013
Proceedings of the 21st International Conference on World Wide Web
486 Citations2012Alain Mille, Fabien Gandon +3 more
The technical programme of WWW 2012 represents a well-balanced mix over the 13 tracks listed in the call for papers, and neither enforced a proportional share of acceptances among the tracks nor did it favor tracks with lower numbers of submissions.
Recognizing Named Entities in Tweets
372 Citations2011Xiaohua Liu, Shaodian Zhang +2 more
This work proposes to combine a K-Nearest Neighbors classifier with a linear Conditional Random Fields model under a semi-supervised learning framework to tackle the challenges of Named Entities Recognition for tweets.
Measuring geographical regularities of crowd behaviors for Twitter-based geo-social event detection
299 Citations2010Ryong Lee, Kazutoshi Sumiya
This study aims to develop a geo-social event detection system by monitoring crowd behaviors indirectly via Twitter to find out the occurrence of local events such as local festivals, using geographical regularities obtained from a large number of geo-tagged tweets around Japan via Twitter.
Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies - Volume 1
275 Citations2011Dekang Lin
This year's ACL is expected to attract an even larger number of participants than usual, since 2011 happens to be an off-year for COLING, EACL and NAACL.
TwiNER
236 Citations2012Chenliang Li, Jianshu Weng +5 more
A novel 2-step unsupervised NER system for targeted Twitter stream, called TwiNER, which leverages on the global context obtained from Wikipedia and Web N-Gram corpus to partition tweets into valid segments (phrases) using a dynamic programming algorithm.
Simple supervised document geolocation with geodesic grids
215 Citations2011Benjamin Wing, Jason Baldridge
This work investigates automatic geolocation (i.e. identification of the location, expressed as latitude/longitude coordinates) of documents and describes several simple supervised methods for document geolocated using only the document's raw text as evidence.
TwitIE: An Open-Source Information Extraction Pipeline for Microblog Text: slides
179 Citations2014Kalina Bontcheva, Leon Derczynski
This paper introduces each stage of the TwitIE pipeline, which is a modification of the GATE ANNIE open-source pipeline for news text, and an evaluation against some state-of-the-art systems is presented.
An effective two-stage model for exploiting non-local dependencies in named entity recognition
174 Citations2006Vijay Krishnan, Christopher D. Manning
This paper shows that a simple two-stage approach to handle non-local dependencies in Named Entity Recognition (NER) can outperform existing approaches that handleNon- local dependencies, while being much more computationally efficient.
International Conference on Computational LinguisticsGeolocation Prediction in Social Media Data by Finding Location Indicative Words
168 Citations2012Bo Han, Paul Cook +1 more
This paper shows that an information gain ratiobased approach surpasses other methods at LIW selection, outperforming state-of-the-art geolocation prediction methods by 10.6% in accuracy and reducing the mean and median of prediction error distance on a public dataset.
Location extraction from disaster-related microblogs
137 Citations2013John Lingad, Sarvnaz Karimi +1 more
This work investigates the feasibility of applying Named Entity Recognizers to extract locations from microblogs, at the level of both geo-location and point-of-interest, and shows that such tools once retrained on microblog data have great potential to detect the where information, even at the granularity of point- of-interest.
Estimating Twitter User Location Using Social Interactions--A Content Based Approach
133 Citations2011Swarup Chandra, Latifur Khan +1 more
A baseline probability estimate of the distribution of words used by a user is calculated by using the fact that terms used in the tweets of a certain discussion may be related to the location information of the user initiating the discussion, and yields an accuracy higher that the 10% accuracy of the current state of the art estimation.
An Empirical Evaluation of doc2vec with Practical Insights into Document Embedding Generation
111 Citations2016Jey Han Lau, Timothy Baldwin
It is found that doc2vec performs robustly when using models trained on large external corpora, and can be further improved by using pre-trained word embeddings.
Fine-grained location extraction from tweets with temporal awareness
98 Citations2014Chenliang Li, Aixin Sun
The proposed solution, named PETAR, consists of two main components: a POI inventory and a time-aware POI tagger, designed to simultaneously identify the POIs and resolve their associated temporal awareness.
Location inference using microblog messages
96 Citations2012Yohei Ikawa, Miki Enoki +1 more
This paper attempts to discover geolocation information from microblog messages to assess disasters by learning associations between a location and its relevant keywords from past messages, and guesses where a new message came from.
Inducing Gazetteers for Named Entity Recognition by Large-Scale Clustering of Dependency Relations
82 Citations2008Jun’ichi Kazama, Kentaro Torisawa
This work parallelized a clustering algorithm based on expectationmaximization (EM) and thus enabled the construction of large-scale MN clusters and demonstrated with the IREX dataset for the Japanese NER that using the constructed clusters as a gazetteer (cluster gazetteser) is a effective way of improving the accuracy of NER.
Proceedings of the Fifteenth Conference on Computational Natural Language Learning
74 Citations2011Sharon Goldwater, Christopher D. Manning
Proceedings of the 2003 Conference of the North American Chapter of the Association for Computational Linguistics on Human Language Technology - Volume 1
68 Citations2003Marti A. Hearst, Mari Ostendorf
Proceedings of the International AAAI Conference on Web and Social MediaCatching the Long-Tail: Extracting Local News Events from Twitter
60 Citations2021Puneet Agarwal, Rajgopal Vaithiyanathan +2 more
This paper describes how this ‘long tail’ of events can be detected in spite of their sparsity, and achieves success-rates in the 80% range for event detection and 76% on event-correlation.
Joint Recognition and Linking of Fine-Grained Locations from Tweets
56 Citations2016Zongcheng Ji, Aixin Sun +2 more
Experimental results show that the proposed joint learning algorithm outperforms the state-of-the-art solutions, and learning from unlabeled data improves both the recognition and linking accuracy.
Journal of Computational SciencePredicting information diffusion on Twitter – Analysis of predictive features
53 Citations2017Thi Bich Ngoc Hoang, Josiane Mothe
A model based on new features to represent tweets in order to predict their possible propagation evaluates the model built on top of both features from the literature and features defined on three collections and shows the usefulness of the features in the prediction.
CEUR-WS.org eBooksMaking sense of microposts (#MSM2013) concept extraction challenge
49 Citations2013Amparo Elizabeth Cano Basave, Andrea Varga +3 more
The evaluation process is described and the performance of different approaches in different contexts are explained, to help clarify the suitability of concept extraction tools and methods for Micropost data.
Proceedings of the Thirteenth Conference on Computational Natural Language Learning
44 Citations2009Suzanne Stevenson, Xavier Carreras
Subword and Spatiotemporal Models for Identifying Actionable Information in Haitian Kreyol
41 Citations2011Robert Munro
A novel system, drawing on 40,000 emergency text messages sent in Haiti following the January 12, 2010 earthquake, predominantly in Haitian Kreyol, is presented, showing that keyword/ngram-based models using streaming MaxEnt achieve up to F=0.21 accuracy, and current state-of-the-art subword models increase this substantially.
Combining Terminology Resources and Statistical Methods for Entity Recognition: an Evaluation
37 Citations2008Angus Roberts, Robert Gaizauskas +2 more
This work combines lexical lookup with simple filtering of ambiguous terms, to improve precision, and shows that the combined method boosts precision with little loss of recall, and that linkage from recognised entities back to the domain knowledge resources can be maintained.
Information Processing & ManagementEvidential estimation of event locations in microblogs using the Dempster–Shafer theory
28 Citations2016Özer Özdikiş, Halit Oğuztüzün +1 more
This work focuses on the spatio-temporal characteristics of events detected in microblogs, and proposes a method to estimate their locations using the Dempster-Shafer theory, applicable to any event type and does not require training.
2015 International Conference on Computing, Networking and Communications (ICNC)Location-based event search in social texts
7 Citations2015Yan Huang, Zhi Liu +1 more
Based on geotagged users and texts, effective and efficient algorithms need to be developed to integrate key word and spatial search to allow event query in location-event search in social texts.
arXiv (Cornell University)Home Location Identification of Twitter Users
6 Citations2014Jalal Mahmud, Jeffrey Nichols +1 more
