YTUP
Journals
About
Services
Guides
Sign InSubmit Article
HomeJournalsSigma Journal of Engineering and Natural Sciences10.14744/sigma.2023.00074
SJSigma Journal of Engineering and Natural Sciences
Get Alerted Download PDF
AbstractKeywordsIntroductionMaterials And MethodsResults And DiscussionConclusionData Availability StatementConflict Of InterestEthicsShare and CiteRelated Articles
Article Open Access1 January 2024

Vector autoregressive clustering for redundancy analysis in air pollution monitoring networks at Tür

Order Reprints Cite Share

Aytaç PEKMEZCİ1

1Muğla University

Sigma Journal of Engineering and Natural Sciences 2024, Vol. 42, Issue 2, pp. 399-406; doi.org/10.14744/sigma.2023.00074

Download PDF View DOI record

Abstract

This study proposes a new approach to reduce the information redundancy at Air Pollution Monitoring Networks (APMNs) and costs required for monitoring them. Proposed approach is based on Vector Autoregressive (VAR) model which describes the relationship between multivariate time series and consists of three main steps: In the first step, VAR model between two or more than two time series consisting of air pollutant observations is estimated. This step is repeated as the number of monitoring stations (n) under study and thus, n parameter vectors are obtained. In the second step, parameters vectors are divided into homogenous groups by using clustering analysis. The objective of this step is to identify the similar mon-itoring stations in terms of the relationship. Last step is to calculate the reduced information redundancy and the monitoring costs. To evaluate the efficiency of proposed approach, data sets consisting of PM10 and SO2 time series obtained from 116 APMNs at Türkiye are used. Fuzzy K-Medoids (FKM) as clustering method Xie-Beni (XB) index as cluster validity index are preferred. Experimental results showed that information redundancy and monitoring cost in PM10 and SO2 stations can reduced at the rate of 63.36 by following proposed approach.

Keywords: Air Pollution; Information Redundancy; Vector Autoregressive Models; Time Series Analysis

Introduction

Air pollution is presence of chemicals or compounds, in the atmosphere, at levels that effect negative on human and environment health. These chemicals or compounds are generally called as “air pollutant”. Particulate Matter, Carbon Monoxide, Sulphur Dioxide (SO2), Carbon Monoxide (CO), Carbon Dioxide (CO2) and Nitrogen (N) are the most important air pollutants. Many studies have investigated the effects of the pollutants on human health

and ecosystem (Ghorani -Azam et al., 2016; Kurt Kar et al., 2016; Liu et al., 2018; Landrigan et al., 2019). It is concluded that there is significant correlation between them. In order to minimize these effects, detecting of air pollution rapidly is considerable important. Air pollution monitoring network (APMN) is a main tool used for this objective. It provides an opportunity of giving correct information about air quality to public, evaluating the results of air pollution and taking precaution for protecting the environment and decreasing harmful effects of air pollution on creatures.

*Corresponding author. *E-mail address: aytac0803@yahoo.com This paper was recommended for publication in revised form by Regional Editor Gülhayat Gölbaşı Şimşek Published by Yıldız Technical University Press, İstanbul, Turkey Copyright 2021, Yıldız Technical University. This is an open access article under the CC BY-NC license (http://creativecommons.org/licenses/by-nc/4.0/).

Sigma J Eng Nat Sci, Vol. 42, No. 2, pp. 399−406, April, 2024

But, APMNs require a lot of monitoring costs and need to expensive devices for monitoring. In this case, it becomes extremely important to decrease the costs required for APMNs. So far, many studies have been carried out for this objective. The most of these studies are based on detecting the stations having similar behavior in terms of an air pollutant via clustering analysis (Giri et al., 2006; Gramch et al., 2006; Lu et al., 2006; Morlini, 2007; Ignaccolo et al. 2008; Pires et al., 2008; D’Urso and Maharaj, 2009; D’Urso et al., 2015; Güler et al., 2016a, 2016b, Cotta et all., 2020). But in all of these studies, either one air pollutant is considered or analyses are carried out for each air pollutant separately and the relationship between air pollutants are not taken into account. In this study, an approach is proposed for reducing monitoring cost in APMNs at Türkiye for more than one air pollutant simultaneously. The proposed approach is based on clustering the parameters of the VAR model which indicates the relationship between air pollutants. In this way, it is aimed to get information about all air pollutants in the model by only monitoring medoid (cluster centers) stations of air pollutant(s) selected as independent variable(s) and to decrease more the monitoring. The organization scheme of this study can be given as follows. In Section 2, the material and methods used in this study are explained. Section 3 consists of and Section 4 concludes the study.

The estimating of VAR model consists of several steps. These steps can be given as follows.

Materials And Methods

This section briefly explains the proposed approach. In this study, the relationship between weekly PM10 and SO2 concentrations is considered. The data set are download from the website of http://laboratuvar.cevre.gov.tr/Default. ltr.aspx and each of which involves the period of between January 2018 and September 2021. Estimating Vector Autoregressive Models VAR model is a statistical model used for investigating the relationships between two or more than two time series. VAR model between two number of time series can be defined as below: (1)

Step 1: Testing stationarity VAR model assumes that all-time series (yt, xt) to be analyzed are stationary, i.e, statistical properties of time series such as mean, variance and covariance are all constant over time. In order to test stationary, unit root tests are used. These tests basically examine following hypothesis. H0: Unit root is present in time series H1: Unit root is not present in time series Where hypothesis H0: states that time series is nonstationary. In the literature, there exists many unit root tests. In this study, Augmented Dickey Fuller (ADF) (Dickey and Fuller, 1979) has been used. ADF test statistic is calculated as follows: (3) Where SE(β11) is standard error of β11. To decide whether time series is stationary or not, absolute value of ADF test statistic is compared with critical value (Dickey and Fuller, 1979). If this value is smaller than critical value, it is decided that time series is non-stationary. In that case, first-order difference of original time series is taken in order to make time series stationary. Unit root test is applied to differenced time series again and if differenced time series is still non-stationary, its second-order difference is taken. This process is repeated until time series become stationary. The number of taken difference indicates stationarity order of time series. For estimating VAR model, in the other words, for performing cointegration test, time series in the model must be stationary of same order. Step 2: Determining lag length The second step of estimating VAR model is the determination of lag length that refers to the number of previous values of time series (yt-1, yt-2,…,yt-p1, xt-1, xt-2,…, xt-p2). In this respect, many criteria have been employed in the literature. The most known criteria are Akaike Information Criterion (AIC), Schwarz Information Criterion (SIC) and Hannan-Quinn Criterion (HQC). In this study, AIC and SIC criteria are preferred for selecting lag length. These criteria are calculated as follows:

Where yt and xt are time series relating to different variables, p1 and p2 are lag length and ε1t and ε2t are error terms that follow normal distribution with zero mean and σ2 variance.

(5) Where n is length of time series and p is the lag length. in equations is computed as bellow

Sigma J Eng Nat Sci, Vol. 42, No. 2, pp. 399−406, April, 2024

(6) The working principle of lag length selection is as follows. VAR model is estimated for various lag length. The model errors ( ) and information criteria are calculated for these models. Finally lag length that provides smallest information criteria is selected. Step 3: Cointegration test Cointegration test investigates that whether a longrun relationship between the time series exists. Johansen (JH) and Engle-Granger (EG) are widely used cointegraand tion tests. In JH test, firstly are calculated by using parameters of estimated VAR model. The JH test includes examination of matrix π. Let rank of π be r. r gives the number of cointegrated time series. In here, three possible cases arise: (1) r = n (the number of variables in the model): Time series are stationary at level. (2) r = 0: There are no cointegration between time series. (3) r < n: There exists r cointegrated time series. With aim of detecting the number of cointegrated time series or whether cointegration exists or not, JH test uses two likelihood ratios known as trace test and maximum eigenvalue statistics. These statistics are calculated as follows: (7)

H1: at least one of β2 coef ficients ≠ 0 (Xt series is Granger causes of Yt) H0: α21 = α22 = ... = α2p2 = 0 (Yt series is not Granger causes of Xt) H1: at least one of α2 coef ficients ≠ 0 (Yt series is Granger causes of Xt) In order to test above hypothesis, F test should be calculated: (9) Where RSSR is residual sum of squares relating to regression model consisting of only Y (Model 1), RSSu is residual sum of squares for the regression model consisting of both Y variables and X variables (Model 2) and k is the number of parameters in Model 2. The calculated F statistics is compared to critical value. If calculated F statistics is higher than critical value, it is decided that Xt series is Granger causes of Yt. Then, VAR model defined in Eq. (1) is estimated. If second hypothesis is tested and if F statistics calculated for this hypothesis is higher than critical value, it is decided that Yt series is Granger causes of Xt. Then, VAR model given in Eq. (2) is estimated. Step 5: Estimating VAR model VAR model is estimated by using Ordinary Least Squares (OLS) technique. If it is assumed that VAR model given in Eq. (1) is estimated, following equation is used: (10)

(8) Where λi is the estimate value of characteristics roots of the π matrix. These test statistics are compared to critical values tabulated by Osterwald-Lenum (1992). If the test statistics are larger than critical value, it is decided that cointegration exists. Besides, value of r gives the number of cointegrated time series. Step 4: Granger causality test Although cointegration indicates that time series have the long-run relationship, it does not give information about the direction of this relationship. Granger Causality (GC) test (Granger, 1969) is used for this kind of analysis. According to GC test, if previous value of Xt is useful in forecasting Yt, the Xt series is Granger causes of Yt or if previous values of Yt is useful in forecasting Xt, the Yt series is Granger causes of Xt. Following hypothesis are testing by using GC test: H0: β21 = β22 = ... = β2p2 = 0 (Xt series is not Granger causes of Yt)

Where β is parameter vector of VAR model. Matrix Z is defined as below: (11) After estimating the parameter vector β, predicted values ( ) are calculated by substituting values of β in Eq. (1) Xie-Beni (XB) INDEX Before applying clustering algorithm, optimal number of clusters should be determined. In this study, cluster validity indices proposed by XB (Xie and Beni, 1991) is preferred. XB index is based on two clustering criteria called as the compactness and separation. Let β = {β11, β12,…, β1p1, β21, β22,…, β2p2} be data set, where βs are parameter vector of VAR model, βijs , {i = 1, 2,…, c, j = 1, 2,…, p1 + p2} are j. component of i. cluster, ukis {k = 1, 2,.., n, i = 1, 2,.., c} are fuzzy membership degree of k. monitoring station to i. cluster. The compactness and separation are calculated for XB index.

Sigma J Eng Nat Sci, Vol. 42, No. 2, pp. 399−406, April, 2024

the data set as cluster center. In FKM, medoid is calculated by the following equation: (12)

Where n is number of monitoring stations and c is the is Euclidian distance. number of cluster and

(16) The update equation for membership degree (uij) is obtained as follows:

(13) Based on compactness and separation criteria, XB index is given as follows: (14) XB index is calculated for all cluster numbers until predefined maximum number of clusters is reached and then the number of clusters, providing minimum XB index is set as the optimal number of clusters. Clustering And Fuzzt K-Medoids Algorithm Clustering analysis is a data mining technique used for dividing data set into groups such that data points within the same group are as similar as possible, whereas data points from different groups are as dissimilar as possible. These groups are called as cluster. Many clustering algorithms exist in the literature. This study uses FKM (Joshi and Krishnapuram, 1999) clustering algorithm based on fuzzy clustering. In fact, the objective function in fuzzy clustering can be defined as follows: (15) In FKM, cluster centers are called as the medoid ( ) which corresponds to data point of a cluster whose sum of distance to all the data points in the cluster makes minimal. The reason of choosing the FKM is to find a data point in

The steps of FKM are as in Table 1. Redundancy Analysis For redundancy analysis, below steps are followed. • Stationarity test is separately performed for each of all PM10 and SO2 time series. The first difference of non-stationary PM10 and SO2 series are taken. Unit root test is re-applied and if the differenced time series is still non-stationary, its second-order difference is taken. This process is repeated until PM10 and SO2 series become stationary. The number of taken difference indicates stationarity order of the time series. In order to estimate VAR model, PM10 and SO2 series obtained from the same station must be stationary of same order. The next steps of the analysis are continued with same-order stationary stations. PM10 and SO2 stations which are not same-order stationary are continued to be monitored. When the number of these stations is considered as n1, the number of stations which are continued to be monitored is determined as 2xn1 (n1 number of PM10 and SO2) in this step. • In the second step, PM10 and SO2 time series having long-time relationship are determined by using Cointegration test. If the number of non-cointegrated series is equal to n2, 2*n2 stations are continued to monitor at this step. • Dependent and independent variables are determined by using GC test. 2*n3 stations with no causality relationship are continued to monitor. The parameters of VAR model are estimated for the stations remained.

Table 1. FKM algorithm Step 1: Entering initial values Number of clusters (c), initial medoids ( iteration number iter = 1

Step 2 : Calculating membership degrees (uij i=1,2,..,c k=1,2,…,n) by using Eq. (17)

Step 4 : Calculating new values of medoids ( ) by using Eq. (16) Step 5 : If

Step 6 : Assign data points (βk) to clusters according to maximum membership degrees.

Sigma J Eng Nat Sci, Vol. 42, No. 2, pp. 399−406, April, 2024

The number of cluster c is determined by using these parameters and FKM clustering algorithm is applied. Lastly, below equation is used to calculate the percentage of the decreased monitoring cost (PDR-MC). (18) Where N is total number of monitoring stations, c is the number medoid stations to be monitored from air pollutant determined as independent variable.

Results And Discussion

To identify stations that do not require to be monitored and thus reduce the monitoring cost, the procedure given in Fig1. is followed. When the procedure given in Fig. 1 is followed, below results are found. • According to ADF test, it is found that all PM10 and SO2 time series are stationary of same order. Thus, the value

of n1 is determined as zero and no monitoring station has been eliminated at this step. • The lag lengths are selected between 2 and 12. Since all time series are stationary at the level, no station is removed from the analysis (n2 = 0). • GC test arise that PM10 series are Granger cause of SO2 series for 75 stations. According to this result, SO2 series are selected as dependent variables and PM10 series are selected as independent variables in the VAR model. This means that PM10 concentrations can be used to estimate SO2 concentrations. Besides, 41 PM10 and SO2 stations with no causality relationship are removed from the analysis and thus it is concluded that these stations are required to monitor. Table 2 shows these stations. According to this, analyses are continued with 75 number of PM10 and SO2 stations. VAR model denoted the relationship between SO2 and PM10 is estimated for each of these stations. Table 3 shows the parameters of these models.

Table 2. The stations which continue to be monitored Name of Stations Adana Meteoroloji, Aksaray, Amasya Suluova, Ankara Kayaş, Ankara Keçiören Sanatoryum, Ardahan, Aydın, Balıkesir Bandırma, Bilecik, Bilecik Bozüyük, Bingöl, Bursa Beyazıt, Çankırı, Çorum Mimar Sinan, Denizli Bayramyeri, Elazığ, Gaziantep, Isparta, İstanbul Başakşehir, İstanbul Esenyurt, İzmir Bornova, İzmir Güzelyalı, Karabük Kardemir2, Kayseri OSB Sanayi, Kırklareli, Kocaeli Körfez, Kocaeli Yeniköy, Konya Karatay, Mardin, Muş, Niğde, Sakarya, Samsun Tekkeköy, Şanlıurfa, Tekirdağ, Tekirdağ Çerkezköy, Tekirdağ Çorlu, Tokat Turhal, Tunceli, Yalova Armutlu, Zonguldak Çatalağzı

Sigma J Eng Nat Sci, Vol. 42, No. 2, pp. 399−406, April, 2024

According to Table 4, the optimal number of clusters is found as 3 since it has the smallest XB index. This states that there are three different groups in terms of the relationship between PM10 and SO2 in Türkiye. When FKM clustering algorithm with 3 number of clusters is applied to the parameters given in Table 3, the clusters given in Table 5 are constituted. The results obtained from Table 5 can be interpreted as follows:

Clusters consist of 31, 41 and 3 stations respectively. When the results are interpreted according to cluster 3, it can be said that the stations Edirne Keşan, Hakkari and Tekirdağ Merkez have similar behavior in terms of the relationship between PM10 and SO2. SO2 and PM10 values of all these stations can be estimated by only monitoring Tekirdağ Merkez PM10 station. It is possible to interpret the other clusters similarly.

Sigma J Eng Nat Sci, Vol. 42, No. 2, pp. 399−406, April, 2024

Table 5. Stations for each cluster and medoid stations Cluster Number

Adana Doğankent, Amasya, Amasya Merzifon, Ankara Sıhhıye, Ankara Sincan, Ankara Siteler, Bartın, Batman, Bayburt, Bursa İnegöl, Çanakkale, Diyarbakır, Iğdır Aralık, İstanbul Kandilli, İzmir Çiğli İBB, İzmir Gaziemir, İzmir Şirinyer, Kahramanmaraş Elbistan, Karabük Kardemir 1, Karabük Tören Alanı, Karaman, Kırşehir, Kocaeli Gebze, Konya Karkent Sanayi, Konya Meram, Malatya, Muğla, Ordu Ünye, Siirt, Sinop Boyabat, Sivas Meteroloji

Adana Çatalan, Adana Valilik, Afyon, Ağrı Doğubeyazıt, Ağrı Patnos, Ankara Bahçelievler, Antalya, Balıkesir, Burdur, Çanakkale Biga İçdaş, Çanakkale Çan, Çorum, Denizli Merkezefendi, Düzce, Edirne, Erzincan, Erzincan Trafik, Erzurum Aziziye, Erzurum Palandöken, Gümüşhane, İstanbul Şirinevler, İstanbul Ümraniye, Kars, Kırıkkale, Kırklareli Lüleburgaz, Kocaeli, Kocaeli Alikahya, Nevşehir, Ordu Stadyum, Osmaniye, Rize, Samsun Atakum, Samsun Bafra, Samsun Canik, Trabzon Akçaabat, Trabzon Fatih, Van, Yalova, Yozgat, Zonguldalk Çatalağzı Kuzyaka, Zonguldak Trafik

Medoids are Kahramanmaraş Elbistan, Samsun Canik and Tekirdağ Merkez respectively and PM10 stations relating to the medoids should be continued to monitor. From this, the percentage of the decreased monitoring cost can be determined as follows.

Conclusion

As a consequence of analyses, the results obtained can be summarized as follows. • The analyses are started with 116 number of PM10 and SO2 monitoring stations at Türkiye. • In the first step of building VAR models, it is observed that all pairs of PM10 and SO2 are stationary of same order. Thus, no monitoring station is eliminated in this step. • In the second step of building VAR models, it is concluded that there exists long-term relationship between all pairs of PM10 and SO2. • According to the results of Granger Causality test, causality relationship between 41 pairs of PM10 and SO2 are not found. It is decided that the monitoring should be continued for these stations. These stations are determined as Adana_Meteroloji, Aksaray, Amasya_Suluova, Ankara_Kayaş, Ankara_Keçiören, Ardahan, Aydın, Balıkesir_Bandırma, Bilecik, Bilecik_Bozüyük, Bingöl, Bursa_Beyazıt, Çankırı, Çorum_Mimar Sinan, Denizli_Bayramyeri, Elazığ, Gaziantep, Isparta, İstanbul_ Başakşehir, Kırklareli, İstanbul_Esenyurt, İzmir_Bornova, İzmir_Güzelyalı, Karabük_Kardemir 2, Kayseri_Sanayi,

Kocaeli_Körfez, Kocaeli_Yeniköy, Mardin, Muş, Niğde, Sakarya, Samsun_Tekkeköy, Şanlıurfa, Tekirdağ, Tekirdağ_ Çorlu, Tekirdağ_Çerkezköy, Tokat_Turhal, Tunceli, Yalova_Armutlu, Zonguldak_Çatalağzı Cumayanı. • Granger Causality test revealed that PM10 concentrations are Granger cause of SO2 concentrations. Therefore, VAR models are estimated such that dependent variables are SO2 and independent variables PM10. • Xie-Beni index found that optimal number of clusters is equal to three. This means that there exist three groups which have different behavior in terms of the relationship between PM10 and SO2 concentrations at Türkiye. • The parameters of VAR models estimated for 75 monitoring stations are clustered by using FKM algorithm. As a result of clustering, the stations that represent the clusters are found as Kahramanmaraş Elbistan, Samsun Canik, Tekirdağ Merkez. • The number of stations to be monitored is found as 85 (41 PM10 +3 PM10+41 SO2). Thus, it is concluded that the monitoring cost and information redundancy at Türkiye are decreased at rate of 63.36% by only monitoring 85 of 232 stations. This reduction in air monitoring cost means that the total cost (manpower, money, time, etc.) required for the future prediction values of PM10 and SO2 variables at each station is reduced. So far, many studies have been carried out with aim of optimizing the number of APMNs and reducing the monitoring cost. But, in all of these studies, either only one air pollutant is considered, or analyses are carried out for each air pollutant separately. There is no study that takes into the relationship between air pollutants. This study proposes an approach based on the relationship between air pollutants for decreasing monitoring cost.

Sigma J Eng Nat Sci, Vol. 42, No. 2, pp. 399−406, April, 2024

Thus, it is aimed to get information about the other air pollutants by monitoring medoid stations relating to only one air pollutant and aimed to reduce monitoring cost more. In this study, relationship between SO2 and PM10 air pollutants are taken into account. It is possible that redundancy analysis approach proposed in this study is carried out for the other pollutants and the other regions. • The originality of this study is that while most of the studies in the literature were on a single variable, this study used more than one variable (PM10 and SO2). The contribution of this study to the literature is to propose a new approach based on the relationship between multiple air pollutants to reduce the information redundancy and the cost of monitoring at air pollution monitoring stations. In future studies, it can be compared using other clustering algorithms other than Fuzzy K-Medoids clustering algorithm.

Data Availability Statement

The authors confirm that the data that supports the findings of this study are available within the article. Raw data that support the finding of this study are available from the corresponding author, upon reasonable request.

Conflict Of Interest

The author declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.

Ethics

There are no ethical issues with the publication of this manuscript.

Share and Cite

PEKMEZCİ, A.; YALÇIN, M.O. Vector autoregressive clustering for redundancy analysis in air pollution monitoring networks at Tür. Sigma Journal of Engineering and Natural Sciences 2024, Vol. 42, pp. 399-406. https://doi.org/10.14744/sigma.2023.00074

Export:

Related Articles

Electricity energy consumption forecasting using LSTM and NARParastou BEHGOUY, Abbas UĞURENVER, 1 January 2026Evaluation of the effects of sustainable development goals on improvement of air quality in citiesŞahin KORKMAZ, Damla MUSTAFAOĞLU et al., 1 January 2026A creative combination of time series models for accurate forecastıng of monthly rainfall A case stuD KARTHIKA, K KARTHIKEYAN, 1 January 2025Role of air pollution on COVID-19 in Istanbulİlknur DÖNMEZ, Zafer ASLAN, 1 January 2023
Publication History
Published1 January 2024
Versionv1
AccessOpen Access
10.14744/sigma.2023.00074
Article Figures (1)
Figure 1
Related Articles
Electricity energy consumption forecasting using LSTM and NARParastou BEHGOUY, Abbas UĞURENVERSigma Journal of Engineering and Natural Sciences, 1 January 2026Evaluation of the effects of sustainable development goals on improvement of air quality in citiesŞahin KORKMAZ, Damla MUSTAFAOĞLU et al.Sigma Journal of Engineering and Natural Sciences, 1 January 2026A creative combination of time series models for accurate forecastıng of monthly rainfall A case stuD KARTHIKA, K KARTHIKEYANSigma Journal of Engineering and Natural Sciences, 1 January 2025
Sigma Journal of Engineering and Natural Sciences coverSigma Journal of Engineering and Natural Sciences Download PDF

Subscribe to YTUP

Stay connected and receive the latest research updates directly in your inbox.

YTUP — Yıldız Technical University Publishing

Advancing knowledge and fostering innovation through high-quality, peer-reviewed academic publications.

About YTU

Discover

  • ›Articles
  • ›Journals
  • ›Research Topics
  • ›Open Access Policy

Guidelines

  • ›Author guidelines
  • ›Services for authors
  • ›Policies and publication ethics
  • ›Editor guidelines
  • ›Fee policy

Explore

  • ›Articles
  • ›Research Topics
  • ›Journals
  • ›How we publish

Support

  • ›Help center
  • ›Emails and alerts
  • ›Contact us
  • ›Submit
  • ›Career opportunities
YTU Logo

© 2026 Yıldız Technical University (Istanbul, Turkey)

Terms and ConditionsTerms of UsePrivacy PolicyPrivacy SettingsDisclaimer
Like this platform? Join our teamHave feedback or questions?
Supervisor