A comparison between geostatistical and machine learning models for spatio-temporal prediction of PM2.5 data

Abstract

Ambient air pollution poses significant health and environmental challenges. Exposure to high concentrations of PM2.5 has been linked to increased respiratory and cardiovascular hospital admissions, more emergency department visits and deaths. Traditional air quality monitoring systems such as EPA-certified stations provide limited spatial and temporal data. The advent of low-cost sensors has dramatically improved the granularity of air quality data, enabling real-time, high-resolution monitoring. This study exploits the extensive data from PurpleAir sensors to assess and compare the effectiveness of various statistical and machine learning models in producing accurate hourly PM2.5 maps across California. We evaluate traditional geostatistical methods, including universal kriging, nearest neighbor Gaussian process and fixed rank kriging, against advanced machine learning approaches such as neural network, random forest, and support vector machine, as well as ensemble model. Our findings highlight the synergistic value of combining geostatistical methods with machine learning to achieve higher-accuracy PM2.5 predictions with lower computational burden, while flexibly capturing nonlinear structure when present.

Publisher

Elsevier

Publication Date

8-2026

Publication Title

Spatial Statistics

Document Type

Article

DOI

https://doi.org/10.1016/j.spasta.2026.100981

Keywords

Air pollution, Geostatistical model, Spatio-temporal data, Machine learning model, Ensemble model

Language

English

Format

text

Share

COinS