Lost in Space: Automated Geolocation in Event Data

Improving geolocation accuracy in text data has long been a goal of automated text processing. We depart from the conventional method and introduce a two-stage supervised machine learning algorithm that evaluates each location mention to be either correct or incorrect. We extract contextual information from texts, i.e., N-gram patterns for location words, mention frequency, and the context of sentences containing location words. We then estimate model parameters using a training dataset and use this model to predict whether a location word in the test dataset accurately represents the location of an event. We demonstrate these steps by constructing customized geolocation event data at the subnational level using news articles collected from around the world. The results of an evaluation show that the proposed algorithm out-performs existing geocoders in terms of accuracy.

[Manuscript] [Online Appendix] [Replication Data]