A comparison between geostatistical and machine learning models for spatio-temporal prediction of PM2.5 data
Abstract
Ambient air pollution poses significant health and environmental challenges. Exposure to high concentrations of PM2.5 has been linked to increased respiratory and cardiovascular hospital admissions, more emergency department visits and deaths. Traditional air quality monitoring systems such as EPA-certified stations provide limited spatial and temporal data. The advent of low-cost sensors has dramatically improved the granularity of air quality data, enabling real-time, high-resolution monitoring. This study exploits the extensive data from PurpleAir sensors to assess and compare the effectiveness of various statistical and machine learning models in producing accurate hourly PM2.5 maps across California. We evaluate traditional geostatistical methods, including universal kriging, nearest neighbor Gaussian process and fixed rank kriging, against advanced machine learning approaches such as neural network, random forest, and support vector machine, as well as ensemble model. Our findings highlight the synergistic value of combining geostatistical methods with machine learning to achieve higher-accuracy PM2.5 predictions with lower computational burden, while flexibly capturing nonlinear structure when present.
Repository Citation
Mohamed, Zeinab, and Wenlong Gong. 2026. "A comparison between geostatistical and machine learning models for spatio-temporal prediction of PM2.5 data." Spatial Statistics 74: article 100981.
Publisher
Elsevier
Publication Date
8-2026
Publication Title
Spatial Statistics
Document Type
Article
DOI
https://doi.org/10.1016/j.spasta.2026.100981
Keywords
Air pollution, Geostatistical model, Spatio-temporal data, Machine learning model, Ensemble model
Language
English
Format
text
