Urban air pollution remains a critical challenge for public health and environmental management, particularly in regions where ground-based monitoring stations are sparse. Previous studies have primarily examined individual pollutants [1,2], while simultaneous contamination by multiple pollutants has received less attention. Effective modeling of multi-pollutant air pollution requires integrating diverse data sources and addressing the spatial heterogeneity of urban areas. Existing approaches often fail to achieve high accuracy in urban settings due to limited data coverage and the complex interplay among built-up areas, green spaces, and pollutant dispersion. This study presents a machine learning-based framework utilizing a Random Forest (RF) model to estimate the multi-pollutant concentrations (NO₂, SO₂, CO) distribution in Vilnius (Fig1). 
It integrates diverse data sources, including satellite retrievals, sparse ground station measurements, meteorological data, urban built-up density, green space metrics, and geographical classification maps, to address urban heterogeneity. Evaluation of the RF model across the study area yielded robust performance metrics: Accuracy = 0.900, Precision = 0.895, Recall = 0.897, and F1 Score = 0.896. This work demonstrates that combining machine learning with multi-source data and urban spatial features provides scalable and accurate multi-pollutant air pollution modeling, offering valuable insights for urban environmental planning and public health interventions. This research was funded by the Lithuanian Research Council LMT and is implemented under the Postdoctoral Fellowship agreement Nr. S-PD-24-137