<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing with OASIS Tables v3.0 20080202//EN" "journalpub-oasis3.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:oasis="http://docs.oasis-open.org/ns/oasis-exchange/table" xml:lang="en" dtd-version="3.0" article-type="research-article"><?xmltex \makeatother\@nolinetrue\makeatletter?>
  <front>
    <journal-meta><journal-id journal-id-type="publisher">GI</journal-id><journal-title-group>
    <journal-title>Geoscientific Instrumentation, Methods and Data Systems</journal-title>
    <abbrev-journal-title abbrev-type="publisher">GI</abbrev-journal-title><abbrev-journal-title abbrev-type="nlm-ta">Geosci. Instrum. Method. Data Syst.</abbrev-journal-title>
  </journal-title-group><issn pub-type="epub">2193-0864</issn><publisher>
    <publisher-name>Copernicus Publications</publisher-name>
    <publisher-loc>Göttingen, Germany</publisher-loc>
  </publisher></journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.5194/gi-10-265-2021</article-id><title-group><article-title>Evaluation of multivariate time series clustering for<?xmltex \hack{\break}?> imputation of air pollution data</article-title><alt-title>Evaluation of multivariate time series clustering</alt-title>
      </title-group><?xmltex \runningtitle{Evaluation of multivariate time series clustering}?><?xmltex \runningauthor{W. Alahamade et al.}?>
      <contrib-group>
        <contrib contrib-type="author" corresp="yes" rid="aff1 aff3">
          <name><surname>Alahamade</surname><given-names>Wedad</given-names></name>
          <email>w.alahamade@uea.ac.uk</email>
        <ext-link>https://orcid.org/0000-0001-7647-4374</ext-link></contrib>
        <contrib contrib-type="author" corresp="no" rid="aff2">
          <name><surname>Lake</surname><given-names>Iain</given-names></name>
          
        </contrib>
        <contrib contrib-type="author" corresp="no" rid="aff2">
          <name><surname>Reeves</surname><given-names>Claire E.</given-names></name>
          
        <ext-link>https://orcid.org/0000-0003-4071-1926</ext-link></contrib>
        <contrib contrib-type="author" corresp="no" rid="aff1">
          <name><surname>De La Iglesia</surname><given-names>Beatriz</given-names></name>
          
        </contrib>
        <aff id="aff1"><label>1</label><institution>School of Computing Sciences, University of East Anglia, Norwich NR4 7TJ, UK</institution>
        </aff>
        <aff id="aff2"><label>2</label><institution>School of Environmental Sciences, University of East Anglia, Norwich NR4 7TJ, UK</institution>
        </aff>
        <aff id="aff3"><label>3</label><institution>School of Computing Sciences, Taibah University, Medina 42353, Saudi Arabia</institution>
        </aff>
      </contrib-group>
      <author-notes><corresp id="corr1">Wedad Alahamade (w.alahamade@uea.ac.uk)</corresp></author-notes><pub-date><day>3</day><month>November</month><year>2021</year></pub-date>
      
      <volume>10</volume>
      <issue>2</issue>
      <fpage>265</fpage><lpage>285</lpage>
      <history>
        <date date-type="received"><day>26</day><month>April</month><year>2021</year></date>
           <date date-type="rev-request"><day>17</day><month>May</month><year>2021</year></date>
           <date date-type="rev-recd"><day>22</day><month>September</month><year>2021</year></date>
           <date date-type="accepted"><day>4</day><month>October</month><year>2021</year></date>
      </history>
      <permissions>
        <copyright-statement>Copyright: © 2021 Wedad Alahamade et al.</copyright-statement>
        <copyright-year>2021</copyright-year>
      <license license-type="open-access"><license-p>This work is licensed under the Creative Commons Attribution 4.0 International License. To view a copy of this licence, visit <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link></license-p></license></permissions><self-uri xlink:href="https://gi.copernicus.org/articles/10/265/2021/gi-10-265-2021.html">This article is available from https://gi.copernicus.org/articles/10/265/2021/gi-10-265-2021.html</self-uri><self-uri xlink:href="https://gi.copernicus.org/articles/10/265/2021/gi-10-265-2021.pdf">The full text article is available as a PDF file from https://gi.copernicus.org/articles/10/265/2021/gi-10-265-2021.pdf</self-uri>
      <abstract><title>Abstract</title>

      <p id="d1e124">Air pollution is one of the world's leading risk factors for death, with 6.5 million deaths per year worldwide attributed to air-pollution-related diseases. Understanding the behaviour of certain pollutants through air quality assessment can produce improvements in air quality management that will translate to health and economic benefits. However, problems with missing data and uncertainty hinder that assessment.</p>

      <p id="d1e127">We are motivated by the need to enhance the air pollution data available. We focus on the problem of missing air pollutant concentration data either because a limited set of pollutants is measured at a monitoring site or because an instrument is not operating, so a particular pollutant is not measured for a period of time.</p>

      <p id="d1e130">In our previous work, we have proposed models which can impute a whole missing time series to enhance air quality monitoring. Some of these models are based on a multivariate time series (MVTS) clustering method. Here, we apply our method to real data and show how different graphical and statistical model evaluation functions enable us to select the imputation model that produces the most plausible imputations. We then compare the Daily Air Quality Index (DAQI) values obtained after imputation with observed values incorporating missing data. Our results show that using an ensemble model that aggregates the spatial similarity obtained by the geographical correlation between monitoring stations and the fused temporal similarity between pollutant concentrations produces very good imputation results. Furthermore, the analysis enhances understanding of the different pollutant behaviours and of the characteristics of different stations according to their environmental type.</p>
  </abstract>
    </article-meta>
  </front>
<body>
      

      <?xmltex \hack{\newpage}?>
<sec id="Ch1.S1" sec-type="intro">
  <label>1</label><title>Introduction</title>
      <p id="d1e144">Time series (TS) analysis has received much attention in recent decades due to its importance in many real-world applications such as earthquake prediction <xref ref-type="bibr" rid="bib1.bibx14" id="paren.1"/>, weather forecasting <xref ref-type="bibr" rid="bib1.bibx5" id="paren.2"/>, air pollution forecasting <xref ref-type="bibr" rid="bib1.bibx15" id="paren.3"/>, and human activity recognition <xref ref-type="bibr" rid="bib1.bibx25" id="paren.4"/>.
Generally speaking, TS data can be described as a sequence of observations that a variable takes over time. When several variables are observed and recorded simultaneously, this becomes a multivariate time series (MVTS).</p>
      <p id="d1e159">The quality of the air in the UK is assessed based on five main pollutants. In this study we focus on the four main pollutants: particulate matter less than 2.5 <inline-formula><mml:math id="M1" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">µ</mml:mi></mml:mrow></mml:math></inline-formula>m in diameter (PM<inline-formula><mml:math id="M2" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2.5</mml:mn></mml:msub></mml:math></inline-formula>) or less than 10 <inline-formula><mml:math id="M3" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">µ</mml:mi></mml:mrow></mml:math></inline-formula>m in diameter (PM<inline-formula><mml:math id="M4" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula>), ozone (O<inline-formula><mml:math id="M5" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">3</mml:mn></mml:msub></mml:math></inline-formula>), and nitrogen dioxide (NO<inline-formula><mml:math id="M6" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula>). These pollutants are measured hourly at various monitoring stations.</p>
      <?pagebreak page266?><p id="d1e215">The main challenge with analysing these pollutant TS is that not all the stations report all the pollutants. Even if a station does, it may not measure a particular pollutant all the time due to instrument downtime. In our previous work <xref ref-type="bibr" rid="bib1.bibx3" id="paren.5"/>, we applied an intermediate fusion approach to fuse the distance between stations using the similarity of the four pollutants. The similarity between pollutant TS was measured using shape-based distance (SBD) between hourly pollutant concentrations (TS), as we found that SBD is better than other measures on our dataset <xref ref-type="bibr" rid="bib1.bibx2" id="paren.6"/>. Then we used the <inline-formula><mml:math id="M7" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula>-means clustering algorithm to cluster the stations based on the fused distance; we called that MVTS clustering. Our initial clustering analysis showed that using the basic <inline-formula><mml:math id="M8" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula>-means with the fused distance gives very compact geographical clustering that enhances our understanding of the UK's air pollutant behaviours. Adding to that, using the fused distance to measure the similarity between the pollutants helped us solve some of the uncertainty problems associated with missing pollutant values as the MVTS clustering enables imputation even when no measurement is available for a given pollutant. This is because the multivariate nature of the clustering enabled a station to be allocated to a cluster based on the value of the other pollutants measured.</p>
      <p id="d1e238">Based on the clustering results and station geographical location, we proposed three models to impute the whole time series for the missing pollutant at a given station. In this paper, we apply multiple model evaluation functions to assess which model gives the best results and to demonstrate the validity of our models.</p>
      <p id="d1e242">Our long-term goal is to reduce the uncertainty in air quality assessment by imputing all missing pollutants in the monitoring stations. This will allow us to calculate new air quality indices that may or may not agree with the previous indices; that is, the observed indices that incorporate missing data. This in turn will help us to identify where more measurements can be beneficial.</p>
      <p id="d1e245">We refer to our approach as time series imputation because we used the observed time series to impute missing time series (whole TS) in stations where one pollutant is not measured but other pollutants are. In this process, we are not filling the missing values within the time series (e.g. interpolating) but imputing a new TS. Also, we do not use predictive models; hence, we do not consider this a prediction task. However, it could be argued that our task is close to spatial interpolation <xref ref-type="bibr" rid="bib1.bibx20" id="paren.7"/> even though it is not completely based on spatial information; that is, we did not use any geographical information within the proposed MVTS  clustering. Geographical information, however, is used in nearest-neighbour approaches, which are used in the ensemble proposed. Nevertheless, the main goal of the spatial interpolation is to fill in the gaps (points and/or locations with unknown measurements) using points with known values to cover a certain geographical area <xref ref-type="bibr" rid="bib1.bibx20" id="paren.8"/>. Our goal is to impute unmeasured pollutants (whole TS) in several stations where they are not measured using the fused similarity between stations of other pollutants or using an ensemble of techniques including the MVTS clustering approach. We would argue that our imputation approach incorporates some uncertainty by using a combination of values (within the clustering process and within the ensemble) to produce the imputed value.</p>
      <p id="d1e254">The paper's structure is as follows: Sect. <xref ref-type="sec" rid="Ch1.S2"/> discusses some of the existing TS clustering methods and their application in the air quality field. Section <xref ref-type="sec" rid="Ch1.S3"/> gives a brief introduction of the air quality assessment in the UK and its challenges. Section <xref ref-type="sec" rid="Ch1.S4"/> discusses all the methods we used in detail to impute the missing pollutants and evaluate our proposed solutions. Finally, in Sect. <xref ref-type="sec" rid="Ch1.S5"/>, we analyse the results of our imputation models. Then, we conclude the work with some final remarks and indication for further developments in Sect. <xref ref-type="sec" rid="Ch1.S6"/>.</p>
</sec>
<sec id="Ch1.S2">
  <label>2</label><title>Related work</title>
      <p id="d1e275">In this section, we briefly review some representative research in clustering techniques and its application in air pollution modelling. Data mining techniques have been widely applied to study  air pollution data; however, most of this research focuses only on a single pollutant (univariate TS), while clustering multivariate time series remains a challenging task <xref ref-type="bibr" rid="bib1.bibx22" id="paren.9"/>. Partitioning algorithms such as <inline-formula><mml:math id="M9" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula>-means and <inline-formula><mml:math id="M10" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula>-medoids are very common among works related to TS clustering and have been applied in many papers (e.g. <xref ref-type="bibr" rid="bib1.bibx18 bib1.bibx4 bib1.bibx27" id="altparen.10"/>)</p>
      <p id="d1e298"><xref ref-type="bibr" rid="bib1.bibx4" id="text.11"/> used the <inline-formula><mml:math id="M11" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula>-means algorithm to identify spatial patterns in air pollution data to cluster US cities based on the similarity of their PM<inline-formula><mml:math id="M12" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2.5</mml:mn></mml:msub></mml:math></inline-formula> composition profiles, then characterize these clusters based on chemical characteristics, emission profiles, geographic locations, and population density. <xref ref-type="bibr" rid="bib1.bibx18" id="text.12"/> transformed the TS of pollutant daily observations into a functional form to smooth the TS, then classified the air quality monitoring network in northern Italy using the partitioning around medoids algorithm (PAM) to cluster three individual pollutants, namely NO<inline-formula><mml:math id="M13" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula>, PM<inline-formula><mml:math id="M14" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula>, and O<inline-formula><mml:math id="M15" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">3</mml:mn></mml:msub></mml:math></inline-formula>.  <xref ref-type="bibr" rid="bib1.bibx27" id="text.13"/> applied different clustering algorithms such as <inline-formula><mml:math id="M16" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula>-means, expectation maximization, and canopy for each air pollutant in the dataset (NO, NO<inline-formula><mml:math id="M17" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula>, SO<inline-formula><mml:math id="M18" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula>, PM<inline-formula><mml:math id="M19" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula>, and O<inline-formula><mml:math id="M20" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">3</mml:mn></mml:msub></mml:math></inline-formula>), then aggregated the clustering results based on majority voting to identify one clustering solution for similar regions in terms of air quality.</p>
      <p id="d1e397">On the other hand, there has been some research into similarity within MVTS. For example, <xref ref-type="bibr" rid="bib1.bibx17" id="text.14"/> proposed an MVTS clustering method based on extracted features from the univariate TS. In their work, principal component analysis (PCA) is used to measure the similarity between MVTS, and fuzzy <inline-formula><mml:math id="M21" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula>-means is used to cluster these TS. This clustering approach was used for fault detection in a gas turbine. <xref ref-type="bibr" rid="bib1.bibx29" id="text.15"/> developed an algorithm for clustering MVTS by discovering each TS's temporal patterns. Their algorithm is based on <inline-formula><mml:math id="M22" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula>-means and aims to groups MVTS with similar temporal patterns together into the same cluster. <xref ref-type="bibr" rid="bib1.bibx16" id="text.16"/> proposed robust fuzzy clustering models for MVTS based on an exponential transformation of the dissimilarities. This algorithm was applied to real-world data on the concentrations of three pollutants (NO, NO<inline-formula><mml:math id="M23" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula>, and PM<inline-formula><mml:math id="M24" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula>) in the Metropolitan City of Rome for the problem of detecting pollution alarms.</p>
      <p id="d1e442">In our previous work <xref ref-type="bibr" rid="bib1.bibx2" id="paren.17"/>, we compared different TS distance measures and imputation techniques to impute missing observations and missing pollutants (TS). We found that using shape-based distance (SBD)<?pagebreak page267?> gives better separated clusters than dynamic time warping (DTW). Also, using MICE to impute the TS missing observations is better than using some single imputation methods such as simple moving average (SMA). We used a univariate TS clustering using <inline-formula><mml:math id="M25" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula>-medoids (PAM) to cluster stations and imputed the missing pollutants using the cluster average. In this work, we use the <inline-formula><mml:math id="M26" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula>-means clustering algorithm and include a number of pollutants in the clustering, which makes it MVTS clustering. This clustering algorithm was proposed in <xref ref-type="bibr" rid="bib1.bibx3" id="text.18"/> where more details can be found.  Here we extend that work by applying the imputation solution to real data and using extensive evaluation methods to demonstrate its effectiveness. This enables us to extend our understanding of pollutant behaviour.</p>
</sec>
<sec id="Ch1.S3">
  <label>3</label><title>Air quality assessment</title>
      <p id="d1e473">We will study air pollution using the concentrations measured at the Automatic Urban and Rural Network (AURN) around the UK. The stations in the network are automatic and produce hourly pollutant concentrations. The data are collected and stored, then made directly available via the Web <xref ref-type="bibr" rid="bib1.bibx10" id="paren.19"/>. There are 167 stations with different environmental types: rural, urban, suburban background, roadside, and industrial.</p>
      <p id="d1e479">The Daily Air Quality Index (DAQI) represents air pollution levels in the UK. This index is reported based on the highest individual DAQI derived for each of the five major air pollutants (O<inline-formula><mml:math id="M27" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">3</mml:mn></mml:msub></mml:math></inline-formula>, NO<inline-formula><mml:math id="M28" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula>, PM<inline-formula><mml:math id="M29" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula>, PM<inline-formula><mml:math id="M30" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2.5</mml:mn></mml:msub></mml:math></inline-formula>, and SO<inline-formula><mml:math id="M31" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula>) based on their concentrations. If concentration data for some of these pollutants are not available, the DAQI is based on those pollutants for which data are available. The DAQI is used to provide an indication of the air quality and some associated information that may be used by at-risk groups as well as the general population <xref ref-type="bibr" rid="bib1.bibx10" id="paren.20"/>. The DAQI is numbered from 1 to 10 and divided into four bands: “low” (1–3), “moderate” (4–6), “high” (7–9), and “very high” (10). The air quality is negatively correlated with the DAQI, meaning that a higher DAQI represents worse air quality.</p>
</sec>
<sec id="Ch1.S4">
  <label>4</label><title>Methods</title>
      <p id="d1e540">The MVTS clustering algorithm and our proposed imputation models were implemented in R version (3.5.2) and are fully explained in previous work <xref ref-type="bibr" rid="bib1.bibx3" id="paren.21"/>. To provide a more robust testing scenario, we separate the “model building” stage from the imputation testing stage. We use an initial data period of 3 years (2015–2017) as a training set to build the clustering and then impute on the next year (2018) of the TS to evaluate the goodness of fit.</p><?xmltex \hack{\newpage}?>
<sec id="Ch1.S4.SS1">
  <label>4.1</label><title>Imputation models of missing pollutant TS</title>
      <p id="d1e554">For evaluation purposes, we assume each pollutant from each station is missing entirely and impute it. For any given station, <inline-formula><mml:math id="M32" display="inline"><mml:mi>j</mml:mi></mml:math></inline-formula>, to impute the values of missing pollutant <inline-formula><mml:math id="M33" display="inline"><mml:mrow><mml:msubsup><mml:mi>P</mml:mi><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:msubsup></mml:mrow></mml:math></inline-formula>, where <inline-formula><mml:math id="M34" display="inline"><mml:mi>i</mml:mi></mml:math></inline-formula> represents the different pollutants (<inline-formula><mml:math id="M35" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>≤</mml:mo><mml:mi>i</mml:mi><mml:mo>≤</mml:mo><mml:mn mathvariant="normal">4</mml:mn></mml:mrow></mml:math></inline-formula>), we use different models under two main similarity criteria: the similarity using clustering solutions and the similarity using geographical distance.</p>
      <p id="d1e600">The <inline-formula><mml:math id="M36" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula>-means clustering algorithm is used to group the stations based on their temporal similarity, which is the similarity in time between the hourly pollutant concentrations using SBD as the temporal distance measure. This distance function is implemented in the “dtwclust” package in R <xref ref-type="bibr" rid="bib1.bibx24" id="paren.22"/>. The geographical distance is used to find the spatial similarity between station locations. Adding to that, we use an ensemble model which calculates the median of all the previous imputation models; this model aggregates the temporal and spatial imputation using both the time series clustering and the geographical location similarity. Then, we evaluate these models to select the one that gives the highest similarity to the real values which are known. We explain these models in detail in the following sections.</p>
<sec id="Ch1.S4.SS1.SSS1">
  <label>4.1.1</label><title>Imputation models using clustering results</title>
      <p id="d1e620">Once a clustering of our stations is obtained, we can use the clustering solution to impute missing TS (pollutants). If station <inline-formula><mml:math id="M37" display="inline"><mml:mi>j</mml:mi></mml:math></inline-formula> belongs to cluster <inline-formula><mml:math id="M38" display="inline"><mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, (<inline-formula><mml:math id="M39" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>≤</mml:mo><mml:mi>x</mml:mi><mml:mo>≤</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:math></inline-formula>, where <inline-formula><mml:math id="M40" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula> is the number of clusters) given the measured pollutants over time, then, to impute pollutant <inline-formula><mml:math id="M41" display="inline"><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> based on the clustering results, we use three models.
<list list-type="order"><list-item>
      <p id="d1e678">We impute the average of pollutant <inline-formula><mml:math id="M42" display="inline"><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> in cluster <inline-formula><mml:math id="M43" display="inline"><mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, which is the hourly average of pollutant <inline-formula><mml:math id="M44" display="inline"><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> in all the stations that fall in this cluster. We call this method cluster average (CA).</p></list-item><list-item>
      <p id="d1e715">We impute the average of pollutant <inline-formula><mml:math id="M45" display="inline"><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> in cluster <inline-formula><mml:math id="M46" display="inline"><mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, but using only stations with the same environment type to station <inline-formula><mml:math id="M47" display="inline"><mml:mi>j</mml:mi></mml:math></inline-formula> within the cluster, such as “background rural”, “background urban”, “traffic”, or “industrial”. We call this method CA<inline-formula><mml:math id="M48" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula>ENV. This is in recognition of the fact that the type of station may be important and result in more similar pollutant concentrations.</p></list-item><list-item>
      <p id="d1e755">We impute the average of pollutant <inline-formula><mml:math id="M49" display="inline"><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> in cluster <inline-formula><mml:math id="M50" display="inline"><mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> for stations that belong to the same region. As defined by DEFRA <xref ref-type="bibr" rid="bib1.bibx10" id="paren.23"/> there are 16 regions in the UK for air quality assessment, such as eastern and northern Wales, the East Midlands, and the other UK regions; this method is called CA<inline-formula><mml:math id="M51" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula>REG.</p></list-item></list></p><?xmltex \hack{\newpage}?>
</sec>
<?pagebreak page268?><sec id="Ch1.S4.SS1.SSS2">
  <label>4.1.2</label><title>Imputation models by similarity using geographical distance</title>
      <p id="d1e799">First, we measure the geographic distance using the Harvison metric, which calculates geographic distance on Earth based on longitude and latitude. We calculate the distance between station <inline-formula><mml:math id="M52" display="inline"><mml:mi>j</mml:mi></mml:math></inline-formula> and all other stations that measure pollutant <inline-formula><mml:math id="M53" display="inline"><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>. Then to impute pollutant <inline-formula><mml:math id="M54" display="inline"><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> for station <inline-formula><mml:math id="M55" display="inline"><mml:mi>j</mml:mi></mml:math></inline-formula> we use the following:
<list list-type="order"><list-item>
      <p id="d1e840">the nearest neighbour (1NN) using the Harvison-based distance to station <inline-formula><mml:math id="M56" display="inline"><mml:mi>j</mml:mi></mml:math></inline-formula> – this method is called 1NN; and</p></list-item><list-item>
      <p id="d1e851">the average of the two nearest neighbours (2NN) to station <inline-formula><mml:math id="M57" display="inline"><mml:mi>j</mml:mi></mml:math></inline-formula> – this method is called 2NN.</p></list-item></list></p>
</sec>
<sec id="Ch1.S4.SS1.SSS3">
  <label>4.1.3</label><title>Imputation model by ensemble</title>
      <p id="d1e870">In this approach, for a given station <inline-formula><mml:math id="M58" display="inline"><mml:mi>j</mml:mi></mml:math></inline-formula>, to impute pollutant <inline-formula><mml:math id="M59" display="inline"><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, we use the median value of all the imputed values from the previous models. Those are cluster average (CA), cluster average considering the station type (CA<inline-formula><mml:math id="M60" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula>ENV), cluster average considering the region (CA<inline-formula><mml:math id="M61" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula>REG), first nearest neighbour (1NN), and the average of the two nearest neighbours (2NN). This method is called Median. This imputation approach may be computationally the most expensive as it needs for all others to be computed, but ensembles have the potential to provide very powerful solutions by combining predictions.</p>
</sec>
</sec>
<sec id="Ch1.S4.SS2">
  <label>4.2</label><title>Imputation model evaluation</title>
      <p id="d1e914">We evaluate how plausible the imputation is using different models by comparing truth values to imputed values. The model evaluations are based on the test dataset, which is the 2018 data. As mentioned earlier we do this by taking each existing TS for which we have values, one at a time, and consider them missing. We impute the whole TS by various models and compare that to the ground truth. We are evaluating our models against the real concentrations which contain missing values; hence, we ignore all the missing values in this evaluation. For each model, we can average the different imputation models' behaviour from all the stations to establish the one that provides imputed values closest to the real values. Hence, for our experimental set-up we take each existing TS for a given pollutant and station, <inline-formula><mml:math id="M62" display="inline"><mml:mrow><mml:msubsup><mml:mi>P</mml:mi><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:msubsup></mml:mrow></mml:math></inline-formula>, in turn and impute it by the various models to obtain an imputed TS, <inline-formula><mml:math id="M63" display="inline"><mml:mrow><mml:mi>P</mml:mi><mml:msubsup><mml:mi>I</mml:mi><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:msubsup></mml:mrow></mml:math></inline-formula>. We compare the real values to the imputed values using different statistical and graphical model evaluation functions. The statistical functions include the fraction of predictions within a factor of 2 (FAC2), mean bias (MB), normalized mean bias (NMB), root mean squared error (RMSE), coefficient of correlation (<inline-formula><mml:math id="M64" display="inline"><mml:mi>R</mml:mi></mml:math></inline-formula>), and index of agreement (IOA). These measures are used to evaluate the temporal variation of air pollutants between imputed–modelled and observed concentrations. The graphical functions include a conditional quantile plot, time variation plot, and Taylor diagram. These are functions within the “openair” package, a freely available air quality data analysis tool in R <xref ref-type="bibr" rid="bib1.bibx6" id="paren.24"/> that presents comparisons between the modelled and measured air pollutant concentrations and their statistics graphically. We use the R packages openair <xref ref-type="bibr" rid="bib1.bibx6" id="paren.25"/> and tidyverse <xref ref-type="bibr" rid="bib1.bibx28" id="paren.26"/> for the evaluation.</p>
      <p id="d1e962">Model evaluation functions are beneficial when more than one model is involved in the comparison and help us in understanding why a model does not perform well. The model that gives the lowest error on average, the highest correlation, and the highest degree of agreement between imputed and observed concentrations for all stations (i.e. imputed TS) is initially considered the best model. However, extensive evaluation with various graphical functions enables us to better assess the model quality and how it reflects uncertainty. Note that the best model may change from one pollutant to another and may be affected by other factors such as station type (e.g. urban background, rural, and roadside) or pollutant lifetime and spread.</p>
</sec>
<sec id="Ch1.S4.SS3">
  <label>4.3</label><title>DAQI calculation</title>
      <p id="d1e974">In the UK, DAQI forecasts are issued on a national scale; they are produced by the Met Office in the morning for the current day as well as for the next 4 d. The forecast is improved by incorporating the recent observations of air quality recorded at the AURN stations.
The overall air pollution index for a site or region is determined by the highest DAQI of the five pollutants. The regional DAQI is the highest index among all the stations in that region.</p>
      <p id="d1e977">For our evaluation, we calculated the daily DAQI value using the observed data for each station. This is because the DAQI value is not saved as part of the historical data available, so we need to calculate it from the downloaded data. DEFRA has published a guide for the implementation of DAQI <xref ref-type="bibr" rid="bib1.bibx8" id="paren.27"/>, which explains how the value is calculated, and we follow that guidance. To calculate DAQI, each air pollutant is calculated as follows.</p>
      <p id="d1e983"><list list-type="bullet">
            <list-item>

      <p id="d1e988"><italic>Ozone</italic>. The O<inline-formula><mml:math id="M65" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">3</mml:mn></mml:msub></mml:math></inline-formula> is measured hourly. To determine the DAQI we need to calculate the daily maximum 8-hourly running mean concentration. First, for each hour we calculate the running 8-hourly mean from the previous hours. Then we find the maximum value of these 8-hourly running means. For this calculation 75 % of the data must be captured to calculate the 8-hourly mean.</p>
            </list-item>
            <list-item>

      <p id="d1e1005"><italic>Nitrogen dioxide</italic>. The NO<inline-formula><mml:math id="M66" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> is measured based on an hourly mean. We calculate the daily NO<inline-formula><mml:math id="M67" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> contribution to the DAQI by taking the maximum observation in 24 h every day from 00:00 to 23:00 GMT.</p>
            </list-item>
            <list-item>

      <p id="d1e1031"><italic>Particle PM</italic><inline-formula><mml:math id="M68" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="italic">10</mml:mn></mml:msub></mml:math></inline-formula> <italic>and PM</italic><inline-formula><mml:math id="M69" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="italic">2.5</mml:mn></mml:msub></mml:math></inline-formula>. These are measured hourly. The DAQI is based on the 24 h mean, which we<?pagebreak page269?> calculate by taking the mean value from the hourly observations. For these pollutants 75 % of the daily observations must be captured to calculate the mean; otherwise, the pollutant is considered  missing that day.</p>
            </list-item>
            <list-item>

      <p id="d1e1058">We define the daily index for each pollutant separately. Then, for a station, we take the highest air pollutant index to be the value of the DAQI at that station.</p>
            </list-item>
          </list></p>
      <p id="d1e1063">We called the DAQI that is calculated based on observation “observed DAQI” and the DAQI that is calculated based on imputation “imputed DAQI”. We use the observed DAQI as a performance tool to evaluate our imputation model on its ability to reproduce the Daily Air Quality Index. Note that although we produce only one imputation and not multiple imputations at this stage, we believe they reflect the underlying uncertainty because they are based on a number of aggregated methods.</p>
</sec>
</sec>
<sec id="Ch1.S5">
  <label>5</label><title>Results</title>
      <p id="d1e1075">In this section, we first analyse the proposed pollutant imputation models using some statistical and graphical air pollution modelling evaluation functions. Then, we evaluate the imputation model performance based on the comparison between the observed and imputed DAQI.</p>
<sec id="Ch1.S5.SS1">
  <label>5.1</label><title>Air pollution imputation modelling evaluation</title>
      <p id="d1e1085">We first evaluate imputation models based on the statistical and then on the graphical analysis.</p>
<sec id="Ch1.S5.SS1.SSS1">
  <label>5.1.1</label><title>Model evaluation based on statistical analysis</title>
      <p id="d1e1095">Table <xref ref-type="table" rid="Ch1.T1"/> shows the statistical analysis results. In this table <inline-formula><mml:math id="M70" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula> is the number of stations that measure each pollutant. The table also shows the fraction of predictions within a factor of 2 (FAC2), mean bias (MB), normalized mean bias (NMB), root mean squared error (RMSE), coefficient of correlation (<inline-formula><mml:math id="M71" display="inline"><mml:mi>R</mml:mi></mml:math></inline-formula>), and index of agreement (IOA).</p>
      <p id="d1e1114">In general, model 6 (Median), which is the model that uses the ensemble technique of other models, gives the lowest error average (RMSE), the highest Pearson correlation coefficient (<inline-formula><mml:math id="M72" display="inline"><mml:mi>R</mml:mi></mml:math></inline-formula>), and the highest agreement between imputed and observed concentrations (IOA) for O<inline-formula><mml:math id="M73" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">3</mml:mn></mml:msub></mml:math></inline-formula>, PM<inline-formula><mml:math id="M74" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2.5</mml:mn></mml:msub></mml:math></inline-formula>, and PM<inline-formula><mml:math id="M75" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula>. However, NO<inline-formula><mml:math id="M76" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> shows different behaviour, with model 2 (CA<inline-formula><mml:math id="M77" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula>ENV) achieving slightly higher performance with an increase in the correlation coefficient (by 0.049) and decrease in error average (by 0.826) compared to model 6 (Median). The model bias (MB) for model 2 is 50 % higher than that of model 6. NO<inline-formula><mml:math id="M78" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> shows local patterns, as it is concentrated where it is emitted in urban areas and near  the roadside. Adding to that, NO<inline-formula><mml:math id="M79" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> is shorter-lived than other pollutants and shows greater spatial variability, with concentrations being strongly influenced by the environment type (e.g. roadside, urban background, rural). This changes the NO<inline-formula><mml:math id="M80" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> concentrations from one location to another based on the environmental type <xref ref-type="bibr" rid="bib1.bibx7" id="paren.28"/>.</p>
      <p id="d1e1198">All the selected models performed well, with 71 %–89 % of their imputations falling within a factor of 2 of the observed concentrations as shown in the FAC2 values in Table <xref ref-type="table" rid="Ch1.T1"/>. According to <xref ref-type="bibr" rid="bib1.bibx12" id="text.29"/>, an air quality model minimum requirement is that the FAC2 value is higher than 0.50 and NMB values should be in the range between <inline-formula><mml:math id="M81" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.2</mml:mn></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M82" display="inline"><mml:mrow><mml:mo>+</mml:mo><mml:mn mathvariant="normal">0.2</mml:mn></mml:mrow></mml:math></inline-formula>. Both are met by our models. NMB measures if the model underpredicts or overpredicts, as it estimates the difference between the mean observed and imputed concentrations. Negative NMB means that the model underpredicts and vice versa. All the models have very small biases.</p>

<?xmltex \floatpos{t}?><table-wrap id="Ch1.T1" specific-use="star"><?xmltex \currentcnt{1}?><label>Table 1</label><caption><p id="d1e1230">Performance of the hourly pollutant concentration imputation models based on statistical measures. Best values are in bold for FAC2, RMSE, <inline-formula><mml:math id="M83" display="inline"><mml:mi>R</mml:mi></mml:math></inline-formula>, and IOA.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="8">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:colspec colnum="5" colname="col5" align="right"/>
     <oasis:colspec colnum="6" colname="col6" align="right"/>
     <oasis:colspec colnum="7" colname="col7" align="right"/>
     <oasis:colspec colnum="8" colname="col8" align="right"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Imputation models</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M84" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3">FAC2</oasis:entry>
         <oasis:entry colname="col4">MB</oasis:entry>
         <oasis:entry colname="col5">NMB</oasis:entry>
         <oasis:entry colname="col6">RMSE</oasis:entry>
         <oasis:entry colname="col7"><inline-formula><mml:math id="M85" display="inline"><mml:mi>R</mml:mi></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col8">IOA</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row rowsep="1">
         <oasis:entry namest="col1" nameend="col8">O<inline-formula><mml:math id="M86" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">3</mml:mn></mml:msub></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Model 1 (CA)</oasis:entry>
         <oasis:entry colname="col2">71</oasis:entry>
         <oasis:entry colname="col3">0.867</oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M87" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.008</oasis:entry>
         <oasis:entry colname="col5">0</oasis:entry>
         <oasis:entry colname="col6">15.267</oasis:entry>
         <oasis:entry colname="col7">0.794</oasis:entry>
         <oasis:entry colname="col8">0.712</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Model 2 (CA<inline-formula><mml:math id="M88" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula>ENV)</oasis:entry>
         <oasis:entry colname="col2">71</oasis:entry>
         <oasis:entry colname="col3">0.877</oasis:entry>
         <oasis:entry colname="col4">1.113</oasis:entry>
         <oasis:entry colname="col5">0.022</oasis:entry>
         <oasis:entry colname="col6">14.627</oasis:entry>
         <oasis:entry colname="col7">0.815</oasis:entry>
         <oasis:entry colname="col8">0.729</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Model 3 (CA<inline-formula><mml:math id="M89" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula>REG)</oasis:entry>
         <oasis:entry colname="col2">71</oasis:entry>
         <oasis:entry colname="col3">0.872</oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M90" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.011</oasis:entry>
         <oasis:entry colname="col5">0</oasis:entry>
         <oasis:entry colname="col6">15.014</oasis:entry>
         <oasis:entry colname="col7">0.807</oasis:entry>
         <oasis:entry colname="col8">0.723</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Model 4 (1NN)</oasis:entry>
         <oasis:entry colname="col2">71</oasis:entry>
         <oasis:entry colname="col3">0.831</oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M91" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>1.179</oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M92" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.024</oasis:entry>
         <oasis:entry colname="col6">17.494</oasis:entry>
         <oasis:entry colname="col7">0.757</oasis:entry>
         <oasis:entry colname="col8">0.681</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Model 5 (2NN)</oasis:entry>
         <oasis:entry colname="col2">71</oasis:entry>
         <oasis:entry colname="col3">0.871</oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M93" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.835</oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M94" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.017</oasis:entry>
         <oasis:entry colname="col6">15.159</oasis:entry>
         <oasis:entry colname="col7">0.808</oasis:entry>
         <oasis:entry colname="col8">0.721</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Model 6 (Median)</oasis:entry>
         <oasis:entry colname="col2">71</oasis:entry>
         <oasis:entry colname="col3"><bold>0.888</bold></oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M95" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.373</oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M96" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.008</oasis:entry>
         <oasis:entry colname="col6"><bold>13.776</bold></oasis:entry>
         <oasis:entry colname="col7"><bold>0.837</bold></oasis:entry>
         <oasis:entry colname="col8"><bold>0.745</bold></oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry namest="col1" nameend="col8">NO<inline-formula><mml:math id="M97" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Model 1 (CA)</oasis:entry>
         <oasis:entry colname="col2">157</oasis:entry>
         <oasis:entry colname="col3">0.628</oasis:entry>
         <oasis:entry colname="col4">0.009</oasis:entry>
         <oasis:entry colname="col5">0</oasis:entry>
         <oasis:entry colname="col6">18.33</oasis:entry>
         <oasis:entry colname="col7">0.514</oasis:entry>
         <oasis:entry colname="col8">0.599</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Model 2 (CA<inline-formula><mml:math id="M98" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula>ENV)</oasis:entry>
         <oasis:entry colname="col2">157</oasis:entry>
         <oasis:entry colname="col3"><bold>0.708</bold></oasis:entry>
         <oasis:entry colname="col4">0.247</oasis:entry>
         <oasis:entry colname="col5">0.01</oasis:entry>
         <oasis:entry colname="col6"><bold>15.989</bold></oasis:entry>
         <oasis:entry colname="col7"><bold>0.665</bold></oasis:entry>
         <oasis:entry colname="col8"><bold>0.661</bold></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Model 3 (CA<inline-formula><mml:math id="M99" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula>REG)</oasis:entry>
         <oasis:entry colname="col2">157</oasis:entry>
         <oasis:entry colname="col3">0.63</oasis:entry>
         <oasis:entry colname="col4">0.171</oasis:entry>
         <oasis:entry colname="col5">0.007</oasis:entry>
         <oasis:entry colname="col6">18.364</oasis:entry>
         <oasis:entry colname="col7">0.527</oasis:entry>
         <oasis:entry colname="col8">0.6</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Model 4 (1NN)</oasis:entry>
         <oasis:entry colname="col2">157</oasis:entry>
         <oasis:entry colname="col3">0.605</oasis:entry>
         <oasis:entry colname="col4">2.277</oasis:entry>
         <oasis:entry colname="col5">0.095</oasis:entry>
         <oasis:entry colname="col6">22.591</oasis:entry>
         <oasis:entry colname="col7">0.464</oasis:entry>
         <oasis:entry colname="col8">0.533</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Model 5 (2NN)</oasis:entry>
         <oasis:entry colname="col2">157</oasis:entry>
         <oasis:entry colname="col3">0.618</oasis:entry>
         <oasis:entry colname="col4">2.774</oasis:entry>
         <oasis:entry colname="col5">0.116</oasis:entry>
         <oasis:entry colname="col6">20.46</oasis:entry>
         <oasis:entry colname="col7">0.494</oasis:entry>
         <oasis:entry colname="col8">0.558</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Model 6 (Median)</oasis:entry>
         <oasis:entry colname="col2">157</oasis:entry>
         <oasis:entry colname="col3">0.675</oasis:entry>
         <oasis:entry colname="col4">0.108</oasis:entry>
         <oasis:entry colname="col5">0.005</oasis:entry>
         <oasis:entry colname="col6">16.815</oasis:entry>
         <oasis:entry colname="col7">0.616</oasis:entry>
         <oasis:entry colname="col8">0.642</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3"/>
         <oasis:entry colname="col4">PM<inline-formula><mml:math id="M100" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2.5</mml:mn></mml:msub></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col5"/>
         <oasis:entry colname="col6"/>
         <oasis:entry colname="col7"/>
         <oasis:entry colname="col8"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Model 1 (CA)</oasis:entry>
         <oasis:entry colname="col2">77</oasis:entry>
         <oasis:entry colname="col3">0.835</oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M101" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.118</oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M102" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.012</oasis:entry>
         <oasis:entry colname="col6">5.265</oasis:entry>
         <oasis:entry colname="col7">0.787</oasis:entry>
         <oasis:entry colname="col8">0.713</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Model 2 (CA<inline-formula><mml:math id="M103" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula>ENV)</oasis:entry>
         <oasis:entry colname="col2">77</oasis:entry>
         <oasis:entry colname="col3">0.814</oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M104" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.064</oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M105" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.006</oasis:entry>
         <oasis:entry colname="col6">5.6</oasis:entry>
         <oasis:entry colname="col7">0.76</oasis:entry>
         <oasis:entry colname="col8">0.695</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Model 3 (CA<inline-formula><mml:math id="M106" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula>REG)</oasis:entry>
         <oasis:entry colname="col2">77</oasis:entry>
         <oasis:entry colname="col3">0.838</oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M107" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.064</oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M108" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.006</oasis:entry>
         <oasis:entry colname="col6">5.056</oasis:entry>
         <oasis:entry colname="col7">0.809</oasis:entry>
         <oasis:entry colname="col8">0.725</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Model 4 (1NN)</oasis:entry>
         <oasis:entry colname="col2">77</oasis:entry>
         <oasis:entry colname="col3">0.791</oasis:entry>
         <oasis:entry colname="col4">0.058</oasis:entry>
         <oasis:entry colname="col5">0.006</oasis:entry>
         <oasis:entry colname="col6">5.536</oasis:entry>
         <oasis:entry colname="col7">0.79</oasis:entry>
         <oasis:entry colname="col8">0.7</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Model 5 (2NN)</oasis:entry>
         <oasis:entry colname="col2">77</oasis:entry>
         <oasis:entry colname="col3">0.823</oasis:entry>
         <oasis:entry colname="col4">0.02</oasis:entry>
         <oasis:entry colname="col5">0.002</oasis:entry>
         <oasis:entry colname="col6">4.952</oasis:entry>
         <oasis:entry colname="col7">0.823</oasis:entry>
         <oasis:entry colname="col8">0.726</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Model 6 (Median)</oasis:entry>
         <oasis:entry colname="col2">77</oasis:entry>
         <oasis:entry colname="col3"><bold>0.854</bold></oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M109" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.144</oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M110" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.014</oasis:entry>
         <oasis:entry colname="col6"><bold>4.745</bold></oasis:entry>
         <oasis:entry colname="col7"><bold>0.831</bold></oasis:entry>
         <oasis:entry colname="col8"><bold>0.743</bold></oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry namest="col1" nameend="col8">PM<inline-formula><mml:math id="M111" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Model 1 (CA)</oasis:entry>
         <oasis:entry colname="col2">75</oasis:entry>
         <oasis:entry colname="col3">0.86</oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M112" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.163</oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M113" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.01</oasis:entry>
         <oasis:entry colname="col6">8.747</oasis:entry>
         <oasis:entry colname="col7">0.668</oasis:entry>
         <oasis:entry colname="col8">0.667</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Model 2 (CA<inline-formula><mml:math id="M114" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula>ENV)</oasis:entry>
         <oasis:entry colname="col2">75</oasis:entry>
         <oasis:entry colname="col3">0.851</oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M115" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.148</oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M116" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.009</oasis:entry>
         <oasis:entry colname="col6">9.031</oasis:entry>
         <oasis:entry colname="col7">0.65</oasis:entry>
         <oasis:entry colname="col8">0.662</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Model 3 (CA<inline-formula><mml:math id="M117" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula>REG)</oasis:entry>
         <oasis:entry colname="col2">75</oasis:entry>
         <oasis:entry colname="col3">0.861</oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M118" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.043</oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M119" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.003</oasis:entry>
         <oasis:entry colname="col6">8.797</oasis:entry>
         <oasis:entry colname="col7">0.673</oasis:entry>
         <oasis:entry colname="col8">0.67</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Model 4 (1NN)</oasis:entry>
         <oasis:entry colname="col2">75</oasis:entry>
         <oasis:entry colname="col3">0.816</oasis:entry>
         <oasis:entry colname="col4">0.113</oasis:entry>
         <oasis:entry colname="col5">0.007</oasis:entry>
         <oasis:entry colname="col6">10.363</oasis:entry>
         <oasis:entry colname="col7">0.608</oasis:entry>
         <oasis:entry colname="col8">0.627</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Model 5 (2NN)</oasis:entry>
         <oasis:entry colname="col2">75</oasis:entry>
         <oasis:entry colname="col3">0.858</oasis:entry>
         <oasis:entry colname="col4">0.106</oasis:entry>
         <oasis:entry colname="col5">0.006</oasis:entry>
         <oasis:entry colname="col6">9.23</oasis:entry>
         <oasis:entry colname="col7">0.661</oasis:entry>
         <oasis:entry colname="col8">0.668</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Model 6 (Median)</oasis:entry>
         <oasis:entry colname="col2">75</oasis:entry>
         <oasis:entry colname="col3"><bold>0.882</bold></oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M120" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.216</oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M121" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.013</oasis:entry>
         <oasis:entry colname="col6"><bold>8.224</bold></oasis:entry>
         <oasis:entry colname="col7"><bold>0.715</bold></oasis:entry>
         <oasis:entry colname="col8"><bold>0.697</bold></oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

</sec>
<sec id="Ch1.S5.SS1.SSS2">
  <label>5.1.2</label><title>Model evaluation based on Taylor diagram analysis</title>
      <p id="d1e2277">We use a Taylor diagram to analyse three main statistics: correlation coefficient <inline-formula><mml:math id="M122" display="inline"><mml:mi>R</mml:mi></mml:math></inline-formula>, the standard deviation (sigma), and the root mean square error (centred). These statistics can be plotted on one (2D) graph, which can be represented through the law of cosines <xref ref-type="bibr" rid="bib1.bibx26" id="paren.30"/>.</p>
      <p id="d1e2290">The standard deviation represents the variability between modelled and observed concentrations. The observed variability is plotted on the <inline-formula><mml:math id="M123" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula> axis. The magnitude of the variability is measured as the radial distance from the plot's origin. The black dashed line shows this for the observed value. The grey lines are isopleths for the correlation coefficient (<inline-formula><mml:math id="M124" display="inline"><mml:mi>R</mml:mi></mml:math></inline-formula>) as indicated by the arc-shaped axis; the correlation increases along the arc towards the <inline-formula><mml:math id="M125" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula> axis. The centred root mean square error (RMSE) is represented by the concentric brown dashed lines. The further the points or models are from the observed value, the worse performance they have <xref ref-type="bibr" rid="bib1.bibx6" id="paren.31"/>. Figure <xref ref-type="fig" rid="Ch1.F1"/> shows Taylor diagram plots for all models with all pollutants.</p>
      <p id="d1e2319">In almost all cases the models exhibit less variability than observed, as indicated by the points being closer to the origin than the black dashed line. In general, model 4 (1NN) followed by model 5 (2NN) show variability that is most similar to the observations, as indicated by their relative closeness to the black dashed line. However, these models tend to have the lowest correlation coefficients, as indicated by the grey lines, and the greatest RMSE, as indicated by the brown dashed lines.
Models 4 and 5 use the concentrations from a single site (i.e. the nearest stations) in the imputation, whereas the other models use a cluster average (CA, CA<inline-formula><mml:math id="M126" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula>REG, CA<inline-formula><mml:math id="M127" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula>ENV) or a model ensemble average (Median), so it is reasonable for models 4 and 5 to have variability fairly similar to the observed concentrations. All the other models display less variability than the observed concentrations (as indicated by their points being further from the black dashed line); this may be consistent with their derivation methods, which may smooth out some of the variability.</p>
      <p id="d1e2336">Model 6 (Median), regardless of its ability to capture<?pagebreak page270?> variability, is confirmed as having the highest correlation coefficient and the lowest centred root means squared with all the pollutants except NO<inline-formula><mml:math id="M128" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula>, for which it is the second-best behind model 2 (CA<inline-formula><mml:math id="M129" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula>ENV).</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F1" specific-use="star"><?xmltex \currentcnt{1}?><?xmltex \def\figurename{Figure}?><label>Figure 1</label><caption><p id="d1e2358">Taylor diagrams comparing modelled and observed concentrations for O<inline-formula><mml:math id="M130" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">3</mml:mn></mml:msub></mml:math></inline-formula>, NO<inline-formula><mml:math id="M131" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula>, PM<inline-formula><mml:math id="M132" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2.5</mml:mn></mml:msub></mml:math></inline-formula>, and PM<inline-formula><mml:math id="M133" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula>.</p></caption>
            <?xmltex \igopts{width=497.923228pt}?><graphic xlink:href="https://gi.copernicus.org/articles/10/265/2021/gi-10-265-2021-f01.png"/>

          </fig>

</sec>
<sec id="Ch1.S5.SS1.SSS3">
  <label>5.1.3</label><title>Model evaluation based on conditional quantile analysis</title>
      <p id="d1e2411">We analyse the spread of the modelled and observed pollutant concentrations using conditional quantile plots. Figures <xref ref-type="fig" rid="Ch1.F2"/> and <xref ref-type="fig" rid="Ch1.F3"/> show the conditional quantile plots for the six imputation models (panels a to f). This visualization splits the concentrations into bins according to values of the modelled concentrations. The median line of these values as well  as the 25th <inline-formula><mml:math id="M134" display="inline"><mml:mo>/</mml:mo></mml:math></inline-formula> 75th and  the 10th <inline-formula><mml:math id="M135" display="inline"><mml:mo>/</mml:mo></mml:math></inline-formula> 90th quantile values are plotted together with a blue line showing a “perfect” model. Also shown are histograms of modelled concentrations (shaded grey bars) and histograms of observed concentrations (blue outline bars).</p>
      <p id="d1e2432">These plots show how the modelled concentrations compare with the observed concentrations and how the models capture the variability in the concentrations. The spread of the modelled concentrations around the perfect model line (blue line) is shown by the shaded portions and quantile intervals. If narrow, it indicates high agreement or precision between the modelled and observed concentrations. The quantile intervals also represent the uncertainty bands. In some cases these intervals do not extend along with the median line due to insufficient concentrations to calculate them. The model with good performance is obtained when the median (red line) coincides with the perfect model (blue line) and when the spread in the percentile is as narrow as possible.</p>
      <p id="d1e2435">From these plots, in general, the histograms indicate that model 4 (1NN) (panel d) has better estimation of the variability between the observed and modelled concentrations, as observed before, even though the median line does not<?pagebreak page271?> match the perfect model. This model is positively biased at high concentrations, as shown by the departure of the median line below the blue line for all pollutants. This result supports our analysis from the Taylor diagram that model 4 (1NN) has the lowest variability between modelled and observed concentrations, but with a lower correlation coefficient, and the highest centred root means squared for all pollutants.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F2" specific-use="star"><?xmltex \currentcnt{2}?><?xmltex \def\figurename{Figure}?><label>Figure 2</label><caption><p id="d1e2441">Conditional quantile plot of modelled and observed pollutant concentrations of  O<inline-formula><mml:math id="M136" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">3</mml:mn></mml:msub></mml:math></inline-formula> (left plot) and NO<inline-formula><mml:math id="M137" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> (right plot) for proposed imputation models: <bold>(a)</bold> model 1 (CA), <bold>(b)</bold> model 2 (CA<inline-formula><mml:math id="M138" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula>ENV), <bold>(c)</bold> model 3 (CA<inline-formula><mml:math id="M139" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula>REG), <bold>(d)</bold> model 4 (1NN), <bold>(e)</bold> model 5 (2NN), <bold>(f)</bold> model 6 (Median).</p></caption>
            <?xmltex \igopts{width=489.387402pt}?><graphic xlink:href="https://gi.copernicus.org/articles/10/265/2021/gi-10-265-2021-f02.png"/>

          </fig>

      <?xmltex \floatpos{p}?><fig id="Ch1.F3" specific-use="star"><?xmltex \currentcnt{3}?><?xmltex \def\figurename{Figure}?><label>Figure 3</label><caption><p id="d1e2503">Conditional quantile plot of modelled and observed pollutant concentrations of PM<inline-formula><mml:math id="M140" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2.5</mml:mn></mml:msub></mml:math></inline-formula> (left plot) and PM<inline-formula><mml:math id="M141" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula> (right plot) for proposed imputation models: <bold>(a)</bold> model 1 (CA), <bold>(b)</bold> model 2 (CA<inline-formula><mml:math id="M142" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula>ENV), <bold>(c)</bold> model 3 (CA<inline-formula><mml:math id="M143" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula>REG), <bold>(d)</bold> model 4 (1NN), <bold>(e)</bold> model 5 (2NN), <bold>(f)</bold> model 6 (Median).</p></caption>
            <?xmltex \igopts{width=489.387402pt}?><graphic xlink:href="https://gi.copernicus.org/articles/10/265/2021/gi-10-265-2021-f03.png"/>

          </fig>

      <p id="d1e2563">In Fig. <xref ref-type="fig" rid="Ch1.F2"/> (left), the O<inline-formula><mml:math id="M144" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">3</mml:mn></mml:msub></mml:math></inline-formula> models show that most modelled concentrations match the observations well for a wide range of values. The histograms indicate underestimation in general at the extreme low and high concentrations. In general, the cluster and Median imputation methods (i.e. that use averaging) will tend to struggle to reproduce the lowest and highest concentrations since they take an average approach. Moreover, the highest concentrations are typically limited to relatively few data points. The cases of high ozone concentrations typically occur during specific meteorological conditions and are episodic in nature, and there may be small differences in timings of the peak concentrations at different sites. Very low ozone concentrations are likely to occur at specific sites (near roads where emissions of nitric oxide are large) and therefore may not be reproduced in the models which take a cluster average or where a nearest-neighbour site is not a similar type of site.</p>
      <p id="d1e2577">Model 6 (Median) (panel f) has the best performance, as indicated by an overlapping median line with the blue line. This model has the lowest mean bias and the highest degree of agreement, as indicated by the narrow spread of the modelled concentration quantile intervals.</p>
      <p id="d1e2580">In the same figure (right), NO<inline-formula><mml:math id="M145" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> models show different behaviours from this analysis. Even though the statistical analysis shows that model 2 (CA<inline-formula><mml:math id="M146" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula>ENV) (panel b) gives the best performance, it is clear that in this model, the modelled concentrations tend to be lower than observations for most<?pagebreak page272?> concentration levels (the medians are under the blue line), and the width of the 10th and 75th as well as the 10th and 90th percentiles is quite broad. The only advantage of using this model is its  ability to capture a wide range of concentrations. Model 4 (1NN) (panel d) compared to other models can reproduce the higher concentrations (higher than 125 <inline-formula><mml:math id="M147" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">µ</mml:mi></mml:mrow></mml:math></inline-formula>g m<inline-formula><mml:math id="M148" display="inline"><mml:msup><mml:mi/><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>) as it does not take an average approach. However, this model is positively biased (NMB <inline-formula><mml:math id="M149" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 0.095), which is shown by the departure of the median line from the blue one.</p>
      <p id="d1e2627">The variation between PM<inline-formula><mml:math id="M150" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2.5</mml:mn></mml:msub></mml:math></inline-formula> models in Fig. <xref ref-type="fig" rid="Ch1.F3"/> (left) shows similar performance for the different models. The quantile intervals are wider within the area of high concentrations <inline-formula><mml:math id="M151" display="inline"><mml:mrow><mml:mo>≤</mml:mo><mml:mn mathvariant="normal">60</mml:mn></mml:mrow></mml:math></inline-formula> <inline-formula><mml:math id="M152" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">µ</mml:mi></mml:mrow></mml:math></inline-formula>g m<inline-formula><mml:math id="M153" display="inline"><mml:msup><mml:mi/><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>, and all models underestimate the high concentrations <inline-formula><mml:math id="M154" display="inline"><mml:mrow><mml:mo>≤</mml:mo><mml:mn mathvariant="normal">80</mml:mn></mml:mrow></mml:math></inline-formula> <inline-formula><mml:math id="M155" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">µ</mml:mi></mml:mrow></mml:math></inline-formula>g m<inline-formula><mml:math id="M156" display="inline"><mml:msup><mml:mi/><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>; note that these concentrations are very low-frequency events.</p>
      <p id="d1e2702">Model 6 (Median) (panel f) gives better performance, as indicated by the narrow spread of the modelled concentration quantile intervals and minimal bias, which is indicated by the overlaps between the red and blue lines compared to other models. Models for PM<inline-formula><mml:math id="M157" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula> (right) show performance similar to PM<inline-formula><mml:math id="M158" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2.5</mml:mn></mml:msub></mml:math></inline-formula>.</p>
</sec>
<sec id="Ch1.S5.SS1.SSS4">
  <label>5.1.4</label><title>Model evaluation based on conditional quantile analysis and station environmental types</title>
      <p id="d1e2732">In this analysis, we focus on the performance of model 6 (Median) and model 2 (CA<inline-formula><mml:math id="M159" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula>ENV), as those performed best for the different pollutants in the previous section, but now we break down the analysis for the six environmental types (background rural, background urban, background suburban, and industrial urban, industrial suburban, and traffic urban) to which stations belong. Notice that a pollutant may or may not be measured in all stations and the number of stations of each type is different as shown in Table <xref ref-type="table" rid="Ch1.T2"/>. We also use conditional quantiles to analyse our model's performance within each environmental type.</p>
      <p id="d1e2744">First, we show the monthly average concentrations for each pollutant under each environment type in our test dataset (year 2018) to understand the normal variation of the pollutant concentrations in different environment types. Figures <xref ref-type="fig" rid="Ch1.F5"/>, <xref ref-type="fig" rid="Ch1.F7"/>, <xref ref-type="fig" rid="Ch1.F9"/>, and <xref ref-type="fig" rid="Ch1.F11"/> show conditional quantile plots by the environmental types for the selected models. Table <xref ref-type="table" rid="Ch1.T2"/> shows the statistical measures of performance also broken down by environment type.</p>

      <?xmltex \floatpos{p}?><fig id="Ch1.F4" specific-use="star"><?xmltex \currentcnt{4}?><?xmltex \def\figurename{Figure}?><label>Figure 4</label><caption><p id="d1e2759">Monthly average concentrations of observed NO<inline-formula><mml:math id="M160" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> for each environmental type for the year 2018.</p></caption>
            <?xmltex \igopts{width=341.433071pt}?><graphic xlink:href="https://gi.copernicus.org/articles/10/265/2021/gi-10-265-2021-f04.png"/>

          </fig>

      <?xmltex \floatpos{t}?><fig id="Ch1.F5" specific-use="star"><?xmltex \currentcnt{5}?><?xmltex \def\figurename{Figure}?><label>Figure 5</label><caption><p id="d1e2780">Conditional quantile plot of modelled and observed pollutant concentrations of NO<inline-formula><mml:math id="M161" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> based on model 2 (CA<inline-formula><mml:math id="M162" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula>ENV) for all station environmental types: <bold>(a)</bold> background rural, <bold>(b)</bold> background suburban, <bold>(c)</bold> background urban, <bold>(d)</bold> industrial suburban, <bold>(e)</bold> industrial urban, and <bold>(f)</bold> traffic urban.</p></caption>
            <?xmltex \igopts{width=426.791339pt}?><graphic xlink:href="https://gi.copernicus.org/articles/10/265/2021/gi-10-265-2021-f05.png"/>

          </fig>

      <?pagebreak page274?><p id="d1e2824"><?xmltex \hack{\newpage}?>The most common sources of NO<inline-formula><mml:math id="M163" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> are roads; however, NO<inline-formula><mml:math id="M164" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> concentrations are influenced by traffic density, road locations, and meteorological conditions, which cause variation from one roadside location to another. Figure <xref ref-type="fig" rid="Ch1.F4"/> shows that high NO<inline-formula><mml:math id="M165" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> concentrations are found at traffic urban followed by industrial suburban, then background urban sites, while the background rural sites have the lowest NO<inline-formula><mml:math id="M166" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> concentrations.</p>
      <p id="d1e2866">Figure <xref ref-type="fig" rid="Ch1.F5"/> shows the conditional quantile plots by station type for NO<inline-formula><mml:math id="M167" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> imputation using model 2 (CA<inline-formula><mml:math id="M168" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula>ENV). Here, we see that modelled concentrations are higher than observed concentrations with all environmental types. This is confirmed by all the statistical model quality measures presented in Table <xref ref-type="table" rid="Ch1.T2"/>, where we can observe  a positive mean bias for NO<inline-formula><mml:math id="M169" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula>.</p>
      <p id="d1e2898">As NO<inline-formula><mml:math id="M170" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> distributions in general are skewed to the lower values and our selected model (model 2) (CA<inline-formula><mml:math id="M171" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula>ENV) is based on the average concentrations, the model performs better with lower concentrations.</p>
      <p id="d1e2917">From Table <xref ref-type="table" rid="Ch1.T2"/> based on model RMSE, the model's best performance is associated with background rural stations, while the worst performance is shown for traffic urban stations. Contrasting this with quantile plots, Fig. <xref ref-type="fig" rid="Ch1.F5"/>a shows that for background rural stations the histogram and the median line show better performance with lower concentrations (less than 30 <inline-formula><mml:math id="M172" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">µ</mml:mi></mml:mrow></mml:math></inline-formula>g m<inline-formula><mml:math id="M173" display="inline"><mml:msup><mml:mi/><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>). On the other hand, for traffic urban stations (panel f), the quantile intervals are wider within the area of high concentrations (higher than 25 <inline-formula><mml:math id="M174" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">µ</mml:mi></mml:mrow></mml:math></inline-formula>g m<inline-formula><mml:math id="M175" display="inline"><mml:msup><mml:mi/><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>), and the modelled concentrations tend to be lower than observed concentrations.</p>

      <?xmltex \floatpos{p}?><fig id="Ch1.F6" specific-use="star"><?xmltex \currentcnt{6}?><?xmltex \def\figurename{Figure}?><label>Figure 6</label><caption><p id="d1e2968">Monthly average concentrations of observed O<inline-formula><mml:math id="M176" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">3</mml:mn></mml:msub></mml:math></inline-formula> for each environmental type for the year 2018.</p></caption>
            <?xmltex \igopts{width=341.433071pt}?><graphic xlink:href="https://gi.copernicus.org/articles/10/265/2021/gi-10-265-2021-f06.png"/>

          </fig>

      <?xmltex \floatpos{p}?><fig id="Ch1.F7" specific-use="star"><?xmltex \currentcnt{7}?><?xmltex \def\figurename{Figure}?><label>Figure 7</label><caption><p id="d1e2988">Conditional quantile plot of modelled and observed pollutant concentrations of O<inline-formula><mml:math id="M177" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">3</mml:mn></mml:msub></mml:math></inline-formula> based on model 6 (Median) for station environmental types: <bold>(a)</bold> background rural, <bold>(b)</bold> background suburban, <bold>(c)</bold> background urban, <bold>(d)</bold> industrial suburban, <bold>(e)</bold> industrial urban, and <bold>(f)</bold> traffic urban.</p></caption>
            <?xmltex \igopts{width=426.791339pt}?><graphic xlink:href="https://gi.copernicus.org/articles/10/265/2021/gi-10-265-2021-f07.png"/>

          </fig>

      <p id="d1e3025">For O<inline-formula><mml:math id="M178" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">3</mml:mn></mml:msub></mml:math></inline-formula>, Fig. <xref ref-type="fig" rid="Ch1.F6"/> shows the monthly average  observed O<inline-formula><mml:math id="M179" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">3</mml:mn></mml:msub></mml:math></inline-formula> concentrations in each environment type. From that we can see that ozone in all environment types follows a similar trend. However, background rural stations have the highest concentrations and traffic urban stations have the lowest, consistent with depletion of ozone due to rapid reaction with fresh emissions of nitric oxide from vehicles. Looking at model 6 (Median) performance in Table <xref ref-type="table" rid="Ch1.T2"/> based on the RMSE, the best model performance is associated with industrial urban stations, and its average performance is associated with background rural stations (those with higher concentrations in Fig. <xref ref-type="fig" rid="Ch1.F6"/>), while its worst performance is associated with traffic urban stations (those with lower concentrations in Fig. <xref ref-type="fig" rid="Ch1.F6"/>).</p>
      <?pagebreak page276?><p id="d1e3055">Conditional quantile analysis in Fig. <xref ref-type="fig" rid="Ch1.F7"/> shows the performance of model 6 (Median) for imputing O<inline-formula><mml:math id="M180" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">3</mml:mn></mml:msub></mml:math></inline-formula> for the six environmental types (panels a to f). The model shows similar performance for industrial suburban (panel d) and background rural stations (panel a). For both types, the model is negatively biased (see also Table <xref ref-type="table" rid="Ch1.T2"/>), meaning that the modelled concentrations tend to be lower than observed concentrations (the median lines are above the blue lines).</p>
      <p id="d1e3071">The worst performance based on the RMSE is associated  with traffic urban stations (panel f), which are the stations located at roadsides. With those stations, the modelled concentrations are higher than observed concentrations; i.e. the modelled histogram is shifted to the right. This is indicated by the model positive bias (0.503). The median line also extends beyond the blue line, which means that some modelled concentrations are much higher than observed measurements.</p>
      <p id="d1e3075">The best model performance is associated with industrial urban stations (panel e) according to the RMSE, even though background urban stations (panel c) appear to have the best performance by looking at the conditional quantile plots. The histogram in panel (c) indicates that the distributions of the observed and modelled concentrations tend to be closer to each other for higher concentrations. However, the model overestimates the average concentrations at these stations (between 25 and 70 <inline-formula><mml:math id="M181" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">µ</mml:mi></mml:mrow></mml:math></inline-formula>g m<inline-formula><mml:math id="M182" display="inline"><mml:msup><mml:mi/><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>) and underestimates the very low concentrations.</p>

      <?xmltex \floatpos{p}?><fig id="Ch1.F8" specific-use="star"><?xmltex \currentcnt{8}?><?xmltex \def\figurename{Figure}?><label>Figure 8</label><caption><p id="d1e3100">Monthly average concentrations of observed PM<inline-formula><mml:math id="M183" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2.5</mml:mn></mml:msub></mml:math></inline-formula> for each environmental type for the year 2018.</p></caption>
            <?xmltex \igopts{width=341.433071pt}?><graphic xlink:href="https://gi.copernicus.org/articles/10/265/2021/gi-10-265-2021-f08.png"/>

          </fig>

      <?xmltex \floatpos{p}?><fig id="Ch1.F9" specific-use="star"><?xmltex \currentcnt{9}?><?xmltex \def\figurename{Figure}?><label>Figure 9</label><caption><p id="d1e3120">Conditional quantile plot of modelled and observed pollutant concentrations of PM<inline-formula><mml:math id="M184" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2.5</mml:mn></mml:msub></mml:math></inline-formula> based on model 6 (Median) for station environmental types: <bold>(a)</bold> background rural, <bold>(b)</bold> background suburban, <bold>(c)</bold> background urban, <bold>(d)</bold> industrial suburban, <bold>(e)</bold> traffic urban.</p></caption>
            <?xmltex \igopts{width=441.017717pt}?><graphic xlink:href="https://gi.copernicus.org/articles/10/265/2021/gi-10-265-2021-f09.png"/>

          </fig>

      <p id="d1e3154">Figure <xref ref-type="fig" rid="Ch1.F8"/> shows that PM<inline-formula><mml:math id="M185" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2.5</mml:mn></mml:msub></mml:math></inline-formula> concentrations in rural areas are lower than those in suburban, urban background, and traffic urban areas. That is consistent with the model performance at these sites. Figure <xref ref-type="fig" rid="Ch1.F9"/> shows corresponding conditional quantile plots by station types. Imputing PM<inline-formula><mml:math id="M186" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2.5</mml:mn></mml:msub></mml:math></inline-formula> concentrations using model 6 (Median) gives similar performance for the different station types. In general, the model underestimates the concentrations of PM<inline-formula><mml:math id="M187" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2.5</mml:mn></mml:msub></mml:math></inline-formula>, especially for high concentration levels. Table <xref ref-type="table" rid="Ch1.T2"/> shows that the model underestimates high concentrations in suburban, urban background, and traffic urban areas, as indicated by the model negative biases, while it overestimates the concentrations at industrial urban and background rural sites. The model shows the worst performance for traffic urban (panel e), and this is also indicated by the highest RMSE (5.098) shown in Table <xref ref-type="table" rid="Ch1.T2"/>. The model underestimates the concentrations at these stations, which is confirmed by the model bias (<inline-formula><mml:math id="M188" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.073</mml:mn></mml:mrow></mml:math></inline-formula>) in Table <xref ref-type="table" rid="Ch1.T2"/>. On the other hand, the model's best performance is associated with background suburban sites (Fig. <xref ref-type="fig" rid="Ch1.F9"/>b), even though it underestimates PM<inline-formula><mml:math id="M189" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2.5</mml:mn></mml:msub></mml:math></inline-formula> concentrations with a mean bias of <inline-formula><mml:math id="M190" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.013</mml:mn></mml:mrow></mml:math></inline-formula>.</p>

      <?xmltex \floatpos{p}?><fig id="Ch1.F10" specific-use="star"><?xmltex \currentcnt{10}?><?xmltex \def\figurename{Figure}?><label>Figure 10</label><caption><p id="d1e3230">Monthly average concentrations of observed PM<inline-formula><mml:math id="M191" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula> for each environmental type for the year 2018.</p></caption>
            <?xmltex \igopts{width=341.433071pt}?><graphic xlink:href="https://gi.copernicus.org/articles/10/265/2021/gi-10-265-2021-f10.png"/>

          </fig>

      <?xmltex \floatpos{p}?><fig id="Ch1.F11" specific-use="star"><?xmltex \currentcnt{11}?><?xmltex \def\figurename{Figure}?><label>Figure 11</label><caption><p id="d1e3250">Conditional quantile plot of modelled and observed pollutant concentrations of PM<inline-formula><mml:math id="M192" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula> based on model 6 (Median) for station environmental types: <bold>(a)</bold> background rural, <bold>(b)</bold> background urban, <bold>(c)</bold> industrial urban, <bold>(d)</bold> traffic urban.</p></caption>
            <?xmltex \igopts{width=341.433071pt}?><graphic xlink:href="https://gi.copernicus.org/articles/10/265/2021/gi-10-265-2021-f11.png"/>

          </fig>

      <p id="d1e3280">Finally, PM<inline-formula><mml:math id="M193" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula> levels in background rural and urban areas are lower than those in industrial and traffic urban areas as shown in Fig. <xref ref-type="fig" rid="Ch1.F10"/>. For PM<inline-formula><mml:math id="M194" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula>, imputation  performance shown in Fig. <xref ref-type="fig" rid="Ch1.F11"/> is similar for background urban and background rural sites (panels a and b). The model overestimates the concentrations of PM<inline-formula><mml:math id="M195" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula> that are <inline-formula><mml:math id="M196" display="inline"><mml:mrow><mml:mo>≤</mml:mo><mml:mn mathvariant="normal">10</mml:mn></mml:mrow></mml:math></inline-formula> <inline-formula><mml:math id="M197" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">µ</mml:mi></mml:mrow></mml:math></inline-formula>g m<inline-formula><mml:math id="M198" display="inline"><mml:msup><mml:mi/><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>, while it underestimates the high concentrations of PM<inline-formula><mml:math id="M199" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula> at industrial urban (slightly) and traffic urban sites (panels c and d). That is confirmed by the model mean bias at these sites (<inline-formula><mml:math id="M200" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.002</mml:mn></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M201" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.106</mml:mn></mml:mrow></mml:math></inline-formula>) as shown in Table <xref ref-type="table" rid="Ch1.T2"/>.</p>

<?xmltex \floatpos{t}?><table-wrap id="Ch1.T2" specific-use="star"><?xmltex \currentcnt{2}?><label>Table 2</label><caption><p id="d1e3380">Performance of the hourly pollutant concentration imputation models using model 6 (Median) for O<inline-formula><mml:math id="M202" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">3</mml:mn></mml:msub></mml:math></inline-formula>, PM<inline-formula><mml:math id="M203" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2.5</mml:mn></mml:msub></mml:math></inline-formula>, and PM<inline-formula><mml:math id="M204" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula> and model 2 (CA<inline-formula><mml:math id="M205" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula>ENV) for NO<inline-formula><mml:math id="M206" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> based on statistical measures for all station environment types for all pollutants.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="6">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="left"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:colspec colnum="5" colname="col5" align="right"/>
     <oasis:colspec colnum="6" colname="col6" align="right"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Imputation models</oasis:entry>
         <oasis:entry colname="col2">Environment type</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M207" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">MB</oasis:entry>
         <oasis:entry colname="col5">NMB</oasis:entry>
         <oasis:entry colname="col6">RMSE</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row rowsep="1">
         <oasis:entry namest="col1" nameend="col6">O<inline-formula><mml:math id="M208" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">3</mml:mn></mml:msub></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Model 6 (Median)</oasis:entry>
         <oasis:entry colname="col2">Background rural</oasis:entry>
         <oasis:entry colname="col3">19</oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M209" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>7.648</oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M210" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.130</oasis:entry>
         <oasis:entry colname="col6">14.967</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Model 6 (Median)</oasis:entry>
         <oasis:entry colname="col2">Background suburban</oasis:entry>
         <oasis:entry colname="col3">3</oasis:entry>
         <oasis:entry colname="col4">0.133</oasis:entry>
         <oasis:entry colname="col5">0.003</oasis:entry>
         <oasis:entry colname="col6">12.530</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Model 6 (Median)</oasis:entry>
         <oasis:entry colname="col2">Background urban</oasis:entry>
         <oasis:entry colname="col3">39</oasis:entry>
         <oasis:entry colname="col4">1.578</oasis:entry>
         <oasis:entry colname="col5">0.034</oasis:entry>
         <oasis:entry colname="col6">12.780</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Model 6 (Median)</oasis:entry>
         <oasis:entry colname="col2">Industrial suburban</oasis:entry>
         <oasis:entry colname="col3">2</oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M211" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>1.392</oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M212" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.030</oasis:entry>
         <oasis:entry colname="col6">11.311</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Model 6 (Median)</oasis:entry>
         <oasis:entry colname="col2">Industrial urban</oasis:entry>
         <oasis:entry colname="col3">4</oasis:entry>
         <oasis:entry colname="col4">1.561</oasis:entry>
         <oasis:entry colname="col5">0.033</oasis:entry>
         <oasis:entry colname="col6">11.273</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Model 6 (Median)</oasis:entry>
         <oasis:entry colname="col2">Traffic urban</oasis:entry>
         <oasis:entry colname="col3">3</oasis:entry>
         <oasis:entry colname="col4">16.456</oasis:entry>
         <oasis:entry colname="col5">0.503</oasis:entry>
         <oasis:entry colname="col6">21.580</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry namest="col1" nameend="col6">NO<inline-formula><mml:math id="M213" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Model 2 (CA<inline-formula><mml:math id="M214" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula>ENV)</oasis:entry>
         <oasis:entry colname="col2">Background rural</oasis:entry>
         <oasis:entry colname="col3">15</oasis:entry>
         <oasis:entry colname="col4">0.060</oasis:entry>
         <oasis:entry colname="col5">0.008</oasis:entry>
         <oasis:entry colname="col6">6.699</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Model 2 (CA<inline-formula><mml:math id="M215" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula>ENV)</oasis:entry>
         <oasis:entry colname="col2">Background suburban</oasis:entry>
         <oasis:entry colname="col3">5</oasis:entry>
         <oasis:entry colname="col4">6.590</oasis:entry>
         <oasis:entry colname="col5">0.470</oasis:entry>
         <oasis:entry colname="col6">13.576</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Model 2 (CA<inline-formula><mml:math id="M216" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula>ENV)</oasis:entry>
         <oasis:entry colname="col2">Background urban</oasis:entry>
         <oasis:entry colname="col3">58</oasis:entry>
         <oasis:entry colname="col4">0.025</oasis:entry>
         <oasis:entry colname="col5">0.001</oasis:entry>
         <oasis:entry colname="col6">12.442</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Model 2 (CA<inline-formula><mml:math id="M217" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula>ENV)</oasis:entry>
         <oasis:entry colname="col2">Industrial suburban</oasis:entry>
         <oasis:entry colname="col3">4</oasis:entry>
         <oasis:entry colname="col4">3.929</oasis:entry>
         <oasis:entry colname="col5">0.181</oasis:entry>
         <oasis:entry colname="col6">11.939</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Model 2 (CA<inline-formula><mml:math id="M218" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula>ENV)</oasis:entry>
         <oasis:entry colname="col2">Industrial urban</oasis:entry>
         <oasis:entry colname="col3">11</oasis:entry>
         <oasis:entry colname="col4">0.235</oasis:entry>
         <oasis:entry colname="col5">0.013</oasis:entry>
         <oasis:entry colname="col6">10.481</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Model 2 (CA<inline-formula><mml:math id="M219" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula>ENV)</oasis:entry>
         <oasis:entry colname="col2">Traffic urban</oasis:entry>
         <oasis:entry colname="col3">65</oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M220" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.014</oasis:entry>
         <oasis:entry colname="col5">0.000</oasis:entry>
         <oasis:entry colname="col6">20.500</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry namest="col1" nameend="col6">PM<inline-formula><mml:math id="M221" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2.5</mml:mn></mml:msub></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Model 6 (Median)</oasis:entry>
         <oasis:entry colname="col2">Background rural</oasis:entry>
         <oasis:entry colname="col3">5</oasis:entry>
         <oasis:entry colname="col4">2.167</oasis:entry>
         <oasis:entry colname="col5">0.292</oasis:entry>
         <oasis:entry colname="col6">5.004</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Model 6 (Median)</oasis:entry>
         <oasis:entry colname="col2">Background suburban</oasis:entry>
         <oasis:entry colname="col3">2</oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M222" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.143</oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M223" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.013</oasis:entry>
         <oasis:entry colname="col6">3.434</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Model 6 (Median)</oasis:entry>
         <oasis:entry colname="col2">Background urban</oasis:entry>
         <oasis:entry colname="col3">41</oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M224" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.072</oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M225" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.007</oasis:entry>
         <oasis:entry colname="col6">4.685</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Model 6 (Median)</oasis:entry>
         <oasis:entry colname="col2">Industrial urban</oasis:entry>
         <oasis:entry colname="col3">6</oasis:entry>
         <oasis:entry colname="col4">0.080</oasis:entry>
         <oasis:entry colname="col5">0.009</oasis:entry>
         <oasis:entry colname="col6">3.982</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Model 6 (Median)</oasis:entry>
         <oasis:entry colname="col2">Traffic urban</oasis:entry>
         <oasis:entry colname="col3">23</oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M226" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.781</oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M227" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.073</oasis:entry>
         <oasis:entry colname="col6">5.098</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry namest="col1" nameend="col6">PM<inline-formula><mml:math id="M228" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Model 6 (Median)</oasis:entry>
         <oasis:entry colname="col2">Background rural</oasis:entry>
         <oasis:entry colname="col3">5</oasis:entry>
         <oasis:entry colname="col4">4.205</oasis:entry>
         <oasis:entry colname="col5">0.369</oasis:entry>
         <oasis:entry colname="col6">8.036</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Model 6 (Median)</oasis:entry>
         <oasis:entry colname="col2">Background urban</oasis:entry>
         <oasis:entry colname="col3">26</oasis:entry>
         <oasis:entry colname="col4">1.236</oasis:entry>
         <oasis:entry colname="col5">0.082</oasis:entry>
         <oasis:entry colname="col6">7.097</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Model 6 (Median)</oasis:entry>
         <oasis:entry colname="col2">Industrial urban</oasis:entry>
         <oasis:entry colname="col3">7</oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M229" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.037</oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M230" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.002</oasis:entry>
         <oasis:entry colname="col6">10.027</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Model 6 (Median)</oasis:entry>
         <oasis:entry colname="col2">Traffic urban</oasis:entry>
         <oasis:entry colname="col3">37</oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M231" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>1.939</oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M232" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.106</oasis:entry>
         <oasis:entry colname="col6">8.586</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <p id="d1e4131">Next, we show some examples of our imputed TS compared to the real TS for each pollutant using the selected imputation models in some stations. The following examples in Figs. <xref ref-type="fig" rid="Ch1.F12"/>, <xref ref-type="fig" rid="Ch1.F13"/>, <xref ref-type="fig" rid="Ch1.F14"/>, and <xref ref-type="fig" rid="Ch1.F15"/> show the observed and imputed hourly pollutant concentrations for the four pollutants using the selected imputation model that gave better imputation. We also apply the models for which there is a period of missing values in the observed concentrations to give some idea of how the models work for the whole TS, including when real values do not exist.</p>
      <p id="d1e4143">Figure <xref ref-type="fig" rid="Ch1.F12"/> compares the observed hourly concentrations of PM<inline-formula><mml:math id="M233" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2.5</mml:mn></mml:msub></mml:math></inline-formula> (red) at London Eltham station for 1 to 15 January 2018 with imputed concentrations (black) using model 6 (Median). As we can see, the variation between the imputed and the real TS is very small and the imputed TS reproduces the trend very well, even though there is a period of missing concentrations within the observed TS (red). Similarly, Fig. <xref ref-type="fig" rid="Ch1.F13"/> represents the observed hourly concentrations of PM<inline-formula><mml:math id="M234" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula> (red) at Oxford St Ebbes station for the same period of time with imputed concentrations using model 6. We can see that model 6 underestimates the high concentrations and overestimates the very low concentrations of PM<inline-formula><mml:math id="M235" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula> and PM<inline-formula><mml:math id="M236" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2.5</mml:mn></mml:msub></mml:math></inline-formula>, as mentioned previously in the analysis in Sect. <xref ref-type="sec" rid="Ch1.S5.SS1.SSS3"/>. However, there is still a good match of the trend.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F12" specific-use="star"><?xmltex \currentcnt{12}?><?xmltex \def\figurename{Figure}?><label>Figure 12</label><caption><p id="d1e4191">Imputed (black) and real (red) TS comparison for PM<inline-formula><mml:math id="M237" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2.5</mml:mn></mml:msub></mml:math></inline-formula> at London Eltham station from 1 to 15 January 2018.</p></caption>
            <?xmltex \igopts{width=341.433071pt}?><graphic xlink:href="https://gi.copernicus.org/articles/10/265/2021/gi-10-265-2021-f12.png"/>

          </fig>

      <?xmltex \floatpos{t}?><fig id="Ch1.F13" specific-use="star"><?xmltex \currentcnt{13}?><?xmltex \def\figurename{Figure}?><label>Figure 13</label><caption><p id="d1e4211">Imputed (black) and real (red) TS comparison for PM<inline-formula><mml:math id="M238" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula> at Oxford St Ebbes station from 1 to 15 January 2018.</p></caption>
            <?xmltex \igopts{width=341.433071pt}?><graphic xlink:href="https://gi.copernicus.org/articles/10/265/2021/gi-10-265-2021-f13.png"/>

          </fig>

      <p id="d1e4229">Figure <xref ref-type="fig" rid="Ch1.F14"/> shows a comparison of imputed (black) and observed (red) TS for NO<inline-formula><mml:math id="M239" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> concentrations at Birmingham Acocks Green station for the same period of time (1 to 15 January 2018), but produced by a different imputation model (model 3, CA<inline-formula><mml:math id="M240" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula>ENV) that gives better imputation than others for NO<inline-formula><mml:math id="M241" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula>. It is known that NO<inline-formula><mml:math id="M242" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> has greater spatial variability than other pollutants and it is very complex to impute; the variation between the imputed and the real TS is slightly higher when compared to the previous examples.</p>
      <p id="d1e4268">Figure <xref ref-type="fig" rid="Ch1.F15"/> shows a comparison of the imputed (black) and observed (red) TS for O<inline-formula><mml:math id="M243" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">3</mml:mn></mml:msub></mml:math></inline-formula> concentrations at Birmingham Acocks Green station for the period 16 to 23 January 2018 produced by model 6 (Median). The imputation underestimates the concentrations but represents the trends of high and low values.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F14" specific-use="star"><?xmltex \currentcnt{14}?><?xmltex \def\figurename{Figure}?><label>Figure 14</label><caption><p id="d1e4285">Imputed (black) and real (red) TS comparison for NO<inline-formula><mml:math id="M244" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> at Birmingham Acocks Green station from 1 to 15 January 2018.</p></caption>
            <?xmltex \igopts{width=341.433071pt}?><graphic xlink:href="https://gi.copernicus.org/articles/10/265/2021/gi-10-265-2021-f14.png"/>

          </fig>

      <?xmltex \floatpos{t}?><fig id="Ch1.F15" specific-use="star"><?xmltex \currentcnt{15}?><?xmltex \def\figurename{Figure}?><label>Figure 15</label><caption><p id="d1e4305">Imputed (black) and real (red) TS comparison for O<inline-formula><mml:math id="M245" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">3</mml:mn></mml:msub></mml:math></inline-formula> at Birmingham Acocks Green station from 16 to 23 January 2018.</p></caption>
            <?xmltex \igopts{width=341.433071pt}?><graphic xlink:href="https://gi.copernicus.org/articles/10/265/2021/gi-10-265-2021-f15.png"/>

          </fig>

</sec>
</sec>
<sec id="Ch1.S5.SS2">
  <label>5.2</label><title>Evaluating the imputed concentrations based on the Daily Air Quality Index (DAQI)</title>
      <p id="d1e4332">After imputing the measured pollutants in all the stations, we calculate the DAQI from the imputed data, as explained in Sect. <xref ref-type="sec" rid="Ch1.S4.SS3"/>. Then we compare it with the  DAQI from the observed data to see our selected models' performances. The selected models are model 6 (Median) for O<inline-formula><mml:math id="M246" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">3</mml:mn></mml:msub></mml:math></inline-formula>, PM<inline-formula><mml:math id="M247" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2.5</mml:mn></mml:msub></mml:math></inline-formula>, and PM<inline-formula><mml:math id="M248" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula> and model 2 (CA<inline-formula><mml:math id="M249" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula>ENV) for NO<inline-formula><mml:math id="M250" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula>.</p>
      <?pagebreak page279?><p id="d1e4381">We compare the imputed DAQI with the observed DAQI based on RMSE and the number of days on which there are agreements and disagreements. The total number of days in our dataset is 60 955 d (167 stations <inline-formula><mml:math id="M251" display="inline"><mml:mo>⋅</mml:mo></mml:math></inline-formula> 365 d); there are 2212 d with missing observed DAQI (DAQI <inline-formula><mml:math id="M252" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 0) that have resulted from missing observations on those days. The total number of days to compare is 58 743 d.</p>
      <p id="d1e4398">In general, the total average RMSE from all days in all stations is 0.55. As the station type and the region may affect our imputation, Fig. <xref ref-type="fig" rid="Ch1.F16"/> shows the average RMSE based on air quality regions in the UK (panel a) and station environmental types (panel b); the size of the circles represents the number of stations of each type.  Panel (a) shows that stations classed as traffic urban are associated with the highest RMSE (0.62), while industrial suburban stations have the lowest RMSE (0.36). Panel (b) shows that the north-eastern region  is associated with the lowest RMSE (0.44), while South Wales<?pagebreak page280?> has the highest RMSE (0.74) between imputed and observed DAQI.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F16" specific-use="star"><?xmltex \currentcnt{16}?><?xmltex \def\figurename{Figure}?><label>Figure 16</label><caption><p id="d1e4406">The model performance based on DAQI RMSE: <bold>(a)</bold> the average of the RMSE based on station environmental types and <bold>(b)</bold> the average of the RMSE based on air quality regions.</p></caption>
          <?xmltex \igopts{width=341.433071pt}?><graphic xlink:href="https://gi.copernicus.org/articles/10/265/2021/gi-10-265-2021-f16.png"/>

        </fig>

      <p id="d1e4421">We also study the correlation between the number of measured pollutants in a station and the agreement between modelled and observed DAQI to see if the number of measured pollutants impacts our model's performance.</p>
      <p id="d1e4424">First, we classify stations based on the number of measured pollutants to stations that measured one, two, three, and all four pollutants, as shown in Table <xref ref-type="table" rid="Ch1.T3"/>. Each row in this table represents one group. The second column is the total number of days with associated DAQI from all stations in each group. The RMSE and index of agreement (IOA) are the average of errors and the degree of agreement between observed and modelled DAQI from all stations in each group, then the percentage of each pollutant in each group. Based on this table, we find that stations that measure four pollutants have the lowest RMSE (0.506) and the highest (IOA) (0.806), while stations that measured one pollutant have the worst performance. The majority of stations with one pollutant are stations that measure NO<inline-formula><mml:math id="M253" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula>, with 87 % of the total number of stations in this group (50 stations).</p>

<?xmltex \floatpos{p}?><table-wrap id="Ch1.T3" specific-use="star"><?xmltex \currentcnt{3}?><label>Table 3</label><caption><p id="d1e4441">Comparing observed and modelled DAQI based on the number of measured pollutants in stations.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="9">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:colspec colnum="5" colname="col5" align="right"/>
     <oasis:colspec colnum="6" colname="col6" align="right"/>
     <oasis:colspec colnum="7" colname="col7" align="right"/>
     <oasis:colspec colnum="8" colname="col8" align="right"/>
     <oasis:colspec colnum="9" colname="col9" align="right"/>
     <oasis:thead>
       <oasis:row>
         <oasis:entry colname="col1">Number of measured</oasis:entry>
         <oasis:entry colname="col2">Number of days</oasis:entry>
         <oasis:entry colname="col3">Number of</oasis:entry>
         <oasis:entry colname="col4">RMSE</oasis:entry>
         <oasis:entry colname="col5">IOA</oasis:entry>
         <oasis:entry colname="col6">Percentage</oasis:entry>
         <oasis:entry colname="col7">Percentage</oasis:entry>
         <oasis:entry colname="col8">Percentage</oasis:entry>
         <oasis:entry colname="col9">Percentage</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">pollutants</oasis:entry>
         <oasis:entry colname="col2">in all stations</oasis:entry>
         <oasis:entry colname="col3">stations</oasis:entry>
         <oasis:entry colname="col4"/>
         <oasis:entry colname="col5"/>
         <oasis:entry colname="col6">(O<inline-formula><mml:math id="M254" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">3</mml:mn></mml:msub></mml:math></inline-formula>)</oasis:entry>
         <oasis:entry colname="col7">(NO<inline-formula><mml:math id="M255" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula>)</oasis:entry>
         <oasis:entry colname="col8">(PM<inline-formula><mml:math id="M256" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2.5</mml:mn></mml:msub></mml:math></inline-formula>)</oasis:entry>
         <oasis:entry colname="col9">(PM<inline-formula><mml:math id="M257" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula>)</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">1</oasis:entry>
         <oasis:entry colname="col2">15 684</oasis:entry>
         <oasis:entry colname="col3">50</oasis:entry>
         <oasis:entry colname="col4">0.542</oasis:entry>
         <oasis:entry colname="col5">0.769</oasis:entry>
         <oasis:entry colname="col6">6.1</oasis:entry>
         <oasis:entry colname="col7">87.8</oasis:entry>
         <oasis:entry colname="col8">2.0</oasis:entry>
         <oasis:entry colname="col9">4.1</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">2</oasis:entry>
         <oasis:entry colname="col2">17 581</oasis:entry>
         <oasis:entry colname="col3">48</oasis:entry>
         <oasis:entry colname="col4">0.583</oasis:entry>
         <oasis:entry colname="col5">0.756</oasis:entry>
         <oasis:entry colname="col6">24.5</oasis:entry>
         <oasis:entry colname="col7">48.0</oasis:entry>
         <oasis:entry colname="col8">8.2</oasis:entry>
         <oasis:entry colname="col9">19.4</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">3</oasis:entry>
         <oasis:entry colname="col2">15 443</oasis:entry>
         <oasis:entry colname="col3">43</oasis:entry>
         <oasis:entry colname="col4">0.516</oasis:entry>
         <oasis:entry colname="col5">0.814</oasis:entry>
         <oasis:entry colname="col6">14.0</oasis:entry>
         <oasis:entry colname="col7">31.8</oasis:entry>
         <oasis:entry colname="col8">32.6</oasis:entry>
         <oasis:entry colname="col9">21.7</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">4</oasis:entry>
         <oasis:entry colname="col2">9398</oasis:entry>
         <oasis:entry colname="col3">26</oasis:entry>
         <oasis:entry colname="col4">0.496</oasis:entry>
         <oasis:entry colname="col5">0.814</oasis:entry>
         <oasis:entry colname="col6">25.0</oasis:entry>
         <oasis:entry colname="col7">25.0</oasis:entry>
         <oasis:entry colname="col8">25.0</oasis:entry>
         <oasis:entry colname="col9">25.0</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <p id="d1e4695">We also compare the imputed and the observed DAQI based on the number of days on which the imputed DAQI agrees and disagrees with the observed DAQI. Table <xref ref-type="table" rid="Ch1.T4"/> shows those results and the percentage of time that these situations occurred, which means when agreement or disagreement is found for each DAQI. The total number of days on which the imputed DAQI agrees with the observed DAQI is 43 906 d (75 %), while there are 14 837 d (25 %) of disagreement. We classify the disagreement into two types: the imputed DAQI is higher or lower than the observed DAQI. We find that there are 10 916 d on which the imputed DAQI is lower than the observed DAQI and 3921 d on which the imputed DAQI is higher than the observed DAQI. In most cases, the imputed DAQI is lower than the observed DAQI, in accordance with our analysis of the imputation models that showed underestimation of the pollutant concentrations. From this table, we can see that the highest percentage of disagreement is 42.96 % of the total number of disagreements (14 837) when observed DAQI is 2 and imputed DAQI is 1, followed by 21.35 % of disagreement when observed DAQI is 3 and imputed DAQI is 2.</p>

<?xmltex \floatpos{p}?><table-wrap id="Ch1.T4" specific-use="star"><?xmltex \currentcnt{4}?><label>Table 4</label><caption><p id="d1e4704">Number of days for which imputed DAQI agrees or disagrees with observed DAQI.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="8">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:colspec colnum="4" colname="col4" align="right" colsep="1"/>
     <oasis:colspec colnum="5" colname="col5" align="left"/>
     <oasis:colspec colnum="6" colname="col6" align="right"/>
     <oasis:colspec colnum="7" colname="col7" align="right"/>
     <oasis:colspec colnum="8" colname="col8" align="right"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry namest="col1" nameend="col8">Index agreement </oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">Observed</oasis:entry>
         <oasis:entry colname="col2">Imputed</oasis:entry>
         <oasis:entry colname="col3">Number</oasis:entry>
         <oasis:entry colname="col4">Percentage</oasis:entry>
         <oasis:entry colname="col5">Observed</oasis:entry>
         <oasis:entry colname="col6">Imputed</oasis:entry>
         <oasis:entry colname="col7">Number</oasis:entry>
         <oasis:entry colname="col8">Percentage</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">DAQI</oasis:entry>
         <oasis:entry colname="col2">DAQI</oasis:entry>
         <oasis:entry colname="col3">of days</oasis:entry>
         <oasis:entry colname="col4">of days</oasis:entry>
         <oasis:entry colname="col5">DAQI</oasis:entry>
         <oasis:entry colname="col6">DAQI</oasis:entry>
         <oasis:entry colname="col7">of days</oasis:entry>
         <oasis:entry colname="col8">of days</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">1</oasis:entry>
         <oasis:entry colname="col2">1</oasis:entry>
         <oasis:entry colname="col3">18 920</oasis:entry>
         <oasis:entry colname="col4">43.09</oasis:entry>
         <oasis:entry colname="col5">5</oasis:entry>
         <oasis:entry colname="col6">5</oasis:entry>
         <oasis:entry colname="col7">110</oasis:entry>
         <oasis:entry colname="col8">0.25</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">2</oasis:entry>
         <oasis:entry colname="col2">2</oasis:entry>
         <oasis:entry colname="col3">16 351</oasis:entry>
         <oasis:entry colname="col4">37.24</oasis:entry>
         <oasis:entry colname="col5">6</oasis:entry>
         <oasis:entry colname="col6">6</oasis:entry>
         <oasis:entry colname="col7">19</oasis:entry>
         <oasis:entry colname="col8">0.04</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">3</oasis:entry>
         <oasis:entry colname="col2">3</oasis:entry>
         <oasis:entry colname="col3">7969</oasis:entry>
         <oasis:entry colname="col4">18.15</oasis:entry>
         <oasis:entry colname="col5">7</oasis:entry>
         <oasis:entry colname="col6">7</oasis:entry>
         <oasis:entry colname="col7">5</oasis:entry>
         <oasis:entry colname="col8">0.01</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">4</oasis:entry>
         <oasis:entry colname="col2">4</oasis:entry>
         <oasis:entry colname="col3">525</oasis:entry>
         <oasis:entry colname="col4">1.20</oasis:entry>
         <oasis:entry colname="col5">8</oasis:entry>
         <oasis:entry colname="col6">8</oasis:entry>
         <oasis:entry colname="col7">7</oasis:entry>
         <oasis:entry colname="col8">0.02</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry namest="col1" nameend="col7">Total agreement </oasis:entry>
         <oasis:entry colname="col8">43 906</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry namest="col1" nameend="col7">Total percentage </oasis:entry>
         <oasis:entry colname="col8">74.743</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry namest="col1" nameend="col8">Index disagreement </oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Observed</oasis:entry>
         <oasis:entry colname="col2">Imputed</oasis:entry>
         <oasis:entry colname="col3">Number</oasis:entry>
         <oasis:entry colname="col4">Percentage</oasis:entry>
         <oasis:entry colname="col5">Observed</oasis:entry>
         <oasis:entry colname="col6">Imputed</oasis:entry>
         <oasis:entry colname="col7">Number</oasis:entry>
         <oasis:entry colname="col8">Percentage</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">DAQI</oasis:entry>
         <oasis:entry colname="col2">DAQI</oasis:entry>
         <oasis:entry colname="col3">of days</oasis:entry>
         <oasis:entry colname="col4">of days</oasis:entry>
         <oasis:entry colname="col5">DAQI</oasis:entry>
         <oasis:entry colname="col6">DAQI</oasis:entry>
         <oasis:entry colname="col7">of days</oasis:entry>
         <oasis:entry colname="col8">of days</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">1</oasis:entry>
         <oasis:entry colname="col2">2</oasis:entry>
         <oasis:entry colname="col3">1818</oasis:entry>
         <oasis:entry colname="col4">12.25</oasis:entry>
         <oasis:entry colname="col5">5</oasis:entry>
         <oasis:entry colname="col6">2</oasis:entry>
         <oasis:entry colname="col7">6</oasis:entry>
         <oasis:entry colname="col8">0.04</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">1</oasis:entry>
         <oasis:entry colname="col2">3</oasis:entry>
         <oasis:entry colname="col3">255</oasis:entry>
         <oasis:entry colname="col4">1.72</oasis:entry>
         <oasis:entry colname="col5">5</oasis:entry>
         <oasis:entry colname="col6">3</oasis:entry>
         <oasis:entry colname="col7">54</oasis:entry>
         <oasis:entry colname="col8">0.36</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">1</oasis:entry>
         <oasis:entry colname="col2">4</oasis:entry>
         <oasis:entry colname="col3">11</oasis:entry>
         <oasis:entry colname="col4">0.07</oasis:entry>
         <oasis:entry colname="col5">5</oasis:entry>
         <oasis:entry colname="col6">4</oasis:entry>
         <oasis:entry colname="col7">203</oasis:entry>
         <oasis:entry colname="col8">1.37</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">1</oasis:entry>
         <oasis:entry colname="col2">5</oasis:entry>
         <oasis:entry colname="col3">2</oasis:entry>
         <oasis:entry colname="col4">0.01</oasis:entry>
         <oasis:entry colname="col5">5</oasis:entry>
         <oasis:entry colname="col6">6</oasis:entry>
         <oasis:entry colname="col7">12</oasis:entry>
         <oasis:entry colname="col8">0.08</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">1</oasis:entry>
         <oasis:entry colname="col2">8</oasis:entry>
         <oasis:entry colname="col3">1</oasis:entry>
         <oasis:entry colname="col4">0.01</oasis:entry>
         <oasis:entry colname="col5">6</oasis:entry>
         <oasis:entry colname="col6">7</oasis:entry>
         <oasis:entry colname="col7">4</oasis:entry>
         <oasis:entry colname="col8">0.03</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">2</oasis:entry>
         <oasis:entry colname="col2">8</oasis:entry>
         <oasis:entry colname="col3">1</oasis:entry>
         <oasis:entry colname="col4">0.01</oasis:entry>
         <oasis:entry colname="col5">6</oasis:entry>
         <oasis:entry colname="col6">2</oasis:entry>
         <oasis:entry colname="col7">2</oasis:entry>
         <oasis:entry colname="col8">0.01</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">2</oasis:entry>
         <oasis:entry colname="col2">1</oasis:entry>
         <oasis:entry colname="col3">6374</oasis:entry>
         <oasis:entry colname="col4">42.96</oasis:entry>
         <oasis:entry colname="col5">6</oasis:entry>
         <oasis:entry colname="col6">3</oasis:entry>
         <oasis:entry colname="col7">5</oasis:entry>
         <oasis:entry colname="col8">0.03</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">2</oasis:entry>
         <oasis:entry colname="col2">3</oasis:entry>
         <oasis:entry colname="col3">1479</oasis:entry>
         <oasis:entry colname="col4">9.97</oasis:entry>
         <oasis:entry colname="col5">6</oasis:entry>
         <oasis:entry colname="col6">4</oasis:entry>
         <oasis:entry colname="col7">31</oasis:entry>
         <oasis:entry colname="col8">0.21</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">2</oasis:entry>
         <oasis:entry colname="col2">4</oasis:entry>
         <oasis:entry colname="col3">10</oasis:entry>
         <oasis:entry colname="col4">0.07</oasis:entry>
         <oasis:entry colname="col5">6</oasis:entry>
         <oasis:entry colname="col6">5</oasis:entry>
         <oasis:entry colname="col7">45</oasis:entry>
         <oasis:entry colname="col8">0.30</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">2</oasis:entry>
         <oasis:entry colname="col2">5</oasis:entry>
         <oasis:entry colname="col3">4</oasis:entry>
         <oasis:entry colname="col4">0.03</oasis:entry>
         <oasis:entry colname="col5">7</oasis:entry>
         <oasis:entry colname="col6">8</oasis:entry>
         <oasis:entry colname="col7">1</oasis:entry>
         <oasis:entry colname="col8">0.01</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">3</oasis:entry>
         <oasis:entry colname="col2">1</oasis:entry>
         <oasis:entry colname="col3">337</oasis:entry>
         <oasis:entry colname="col4">2.27</oasis:entry>
         <oasis:entry colname="col5">7</oasis:entry>
         <oasis:entry colname="col6">2</oasis:entry>
         <oasis:entry colname="col7">1</oasis:entry>
         <oasis:entry colname="col8">0.01</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">3</oasis:entry>
         <oasis:entry colname="col2">2</oasis:entry>
         <oasis:entry colname="col3">3168</oasis:entry>
         <oasis:entry colname="col4">21.35</oasis:entry>
         <oasis:entry colname="col5">7</oasis:entry>
         <oasis:entry colname="col6">3</oasis:entry>
         <oasis:entry colname="col7">2</oasis:entry>
         <oasis:entry colname="col8">0.01</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">3</oasis:entry>
         <oasis:entry colname="col2">4</oasis:entry>
         <oasis:entry colname="col3">241</oasis:entry>
         <oasis:entry colname="col4">1.62</oasis:entry>
         <oasis:entry colname="col5">7</oasis:entry>
         <oasis:entry colname="col6">4</oasis:entry>
         <oasis:entry colname="col7">1</oasis:entry>
         <oasis:entry colname="col8">0.01</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">3</oasis:entry>
         <oasis:entry colname="col2">5</oasis:entry>
         <oasis:entry colname="col3">18</oasis:entry>
         <oasis:entry colname="col4">0.12</oasis:entry>
         <oasis:entry colname="col5">7</oasis:entry>
         <oasis:entry colname="col6">5</oasis:entry>
         <oasis:entry colname="col7">10</oasis:entry>
         <oasis:entry colname="col8">0.07</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">3</oasis:entry>
         <oasis:entry colname="col2">6</oasis:entry>
         <oasis:entry colname="col3">2</oasis:entry>
         <oasis:entry colname="col4">0.01</oasis:entry>
         <oasis:entry colname="col5">7</oasis:entry>
         <oasis:entry colname="col6">6</oasis:entry>
         <oasis:entry colname="col7">11</oasis:entry>
         <oasis:entry colname="col8">0.07</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">4</oasis:entry>
         <oasis:entry colname="col2">1</oasis:entry>
         <oasis:entry colname="col3">17</oasis:entry>
         <oasis:entry colname="col4">0.11</oasis:entry>
         <oasis:entry colname="col5">8</oasis:entry>
         <oasis:entry colname="col6">7</oasis:entry>
         <oasis:entry colname="col7">3</oasis:entry>
         <oasis:entry colname="col8">0.02</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">4</oasis:entry>
         <oasis:entry colname="col2">2</oasis:entry>
         <oasis:entry colname="col3">38</oasis:entry>
         <oasis:entry colname="col4">0.26</oasis:entry>
         <oasis:entry colname="col5">8</oasis:entry>
         <oasis:entry colname="col6">9</oasis:entry>
         <oasis:entry colname="col7">1</oasis:entry>
         <oasis:entry colname="col8">0.01</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">4</oasis:entry>
         <oasis:entry colname="col2">3</oasis:entry>
         <oasis:entry colname="col3">598</oasis:entry>
         <oasis:entry colname="col4">4.03</oasis:entry>
         <oasis:entry colname="col5">8</oasis:entry>
         <oasis:entry colname="col6">3</oasis:entry>
         <oasis:entry colname="col7">2</oasis:entry>
         <oasis:entry colname="col8">0.01</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">4</oasis:entry>
         <oasis:entry colname="col2">5</oasis:entry>
         <oasis:entry colname="col3">58</oasis:entry>
         <oasis:entry colname="col4">0.39</oasis:entry>
         <oasis:entry colname="col5">8</oasis:entry>
         <oasis:entry colname="col6">6</oasis:entry>
         <oasis:entry colname="col7">1</oasis:entry>
         <oasis:entry colname="col8">0.01</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">4</oasis:entry>
         <oasis:entry colname="col2">6</oasis:entry>
         <oasis:entry colname="col3">2</oasis:entry>
         <oasis:entry colname="col4">0.01</oasis:entry>
         <oasis:entry colname="col5">9</oasis:entry>
         <oasis:entry colname="col6">8</oasis:entry>
         <oasis:entry colname="col7">2</oasis:entry>
         <oasis:entry colname="col8">0.01</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">5</oasis:entry>
         <oasis:entry colname="col2">7</oasis:entry>
         <oasis:entry colname="col3">1</oasis:entry>
         <oasis:entry colname="col4">0.01</oasis:entry>
         <oasis:entry colname="col5">10</oasis:entry>
         <oasis:entry colname="col6">7</oasis:entry>
         <oasis:entry colname="col7">2</oasis:entry>
         <oasis:entry colname="col8">0.01</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">5</oasis:entry>
         <oasis:entry colname="col2">1</oasis:entry>
         <oasis:entry colname="col3">2</oasis:entry>
         <oasis:entry colname="col4">0.01</oasis:entry>
         <oasis:entry colname="col5">10</oasis:entry>
         <oasis:entry colname="col6">8</oasis:entry>
         <oasis:entry colname="col7">1</oasis:entry>
         <oasis:entry colname="col8">0.01</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry namest="col1" nameend="col7">Total disagreement </oasis:entry>
         <oasis:entry colname="col8">14 837</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry namest="col1" nameend="col7">Total percentage </oasis:entry>
         <oasis:entry colname="col8">25.257</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

</sec>
</sec>
<sec id="Ch1.S6" sec-type="conclusions">
  <label>6</label><title>Discussion and conclusions</title>
      <?pagebreak page282?><p id="d1e5637">In this work, we evaluated our proposed models to impute missing pollutants in a station based on statistical and graphical model evaluation functions (Taylor diagrams and conditional quantile plots) that are designed to evaluate air pollution modelling. We found that the best imputation model based on statistical analysis is model 6 (Median) for O<inline-formula><mml:math id="M258" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">3</mml:mn></mml:msub></mml:math></inline-formula>, PM<inline-formula><mml:math id="M259" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula>, and PM<inline-formula><mml:math id="M260" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2.5</mml:mn></mml:msub></mml:math></inline-formula> and model 2 (CA<inline-formula><mml:math id="M261" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula>ENV) for NO<inline-formula><mml:math id="M262" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> imputation. The station environmental type plays an essential role with NO<inline-formula><mml:math id="M263" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> imputation because NO<inline-formula><mml:math id="M264" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> shows local patterns, as it is concentrated where it is emitted in urban areas and near  the roadside. Adding to that, NO<inline-formula><mml:math id="M265" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> is shorter-lived than other pollutants and shows greater spatial variability, with concentrations being strongly influenced by the environment type (e.g. roadside, urban background, rural). This changes the NO<inline-formula><mml:math id="M266" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> concentrations from one location to another based on the environmental type <xref ref-type="bibr" rid="bib1.bibx7" id="paren.32"/>.</p>
      <p id="d1e5723">On the other hand, the graphical model evaluation functions showed these models' performance based on the distribution of the concentrations and the degree of agreement between imputed–modelled and observed concentrations. These functions help us to understand the relationship between the distributions of the observations and the model's performance. From the histograms in Figs. <xref ref-type="fig" rid="Ch1.F2"/> and <xref ref-type="fig" rid="Ch1.F3"/> we noted that the overall distributions of the observed concentrations of NO<inline-formula><mml:math id="M267" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula>, PM<inline-formula><mml:math id="M268" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2.5</mml:mn></mml:msub></mml:math></inline-formula>, and PM<inline-formula><mml:math id="M269" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula> are skewed to the lower values, while O<inline-formula><mml:math id="M270" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">3</mml:mn></mml:msub></mml:math></inline-formula> has a more normal distribution. From these histograms, we also noticed that the distributions of the modelled O<inline-formula><mml:math id="M271" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">3</mml:mn></mml:msub></mml:math></inline-formula> concentrations are shifted to lower values, while other pollutants' modelled concentrations are shifted to higher values. Hence, the model is not always able to reproduce the edges of the distribution correctly. The skewness in modelled values is mostly associated with model 1 (CA). As a consequence, this shows the greatest difference in skewness between the distributions of the observed and modelled values. However, model 1 (CA) in combination with others, as part of model 6 (Median), reduces the skewness in modelled values and generates better imputation, resulting in the lowest RMSE.</p>
      <p id="d1e5776">Model 6 (Median) is based on the median concentrations from stations with temporal and spatial similarity, so this model's expected performance is to underestimate the highest values and overestimate the lowest values with a normally distributed dataset. We found that the model performance can vary based on the environmental type and the nature of the pollutant, as shown in our analysis of model performance and DAQI RMSE.</p>
      <p id="d1e5779">Through our analysis, we also found that the variation of the model's performance with different environmental types is due to the pollutant behaviour and its emitted sources.</p>
      <?pagebreak page283?><p id="d1e5783">Model 6 (Median) performance with O<inline-formula><mml:math id="M272" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">3</mml:mn></mml:msub></mml:math></inline-formula> imputation changes from one environmental type to another due to the ozone behaviour at these locations. As we know, ozone is not directly emitted into the air, but it is formed as a secondary pollutant by chemistry involving nitrogen oxides (NO<inline-formula><mml:math id="M273" display="inline"><mml:msub><mml:mi/><mml:mi>x</mml:mi></mml:msub></mml:math></inline-formula>): the sum of NO<inline-formula><mml:math id="M274" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula>, nitric oxide (NO), and volatile organic compounds (VOCs) in the presence of sunlight <xref ref-type="bibr" rid="bib1.bibx13" id="paren.33"/>. This chemistry is non-linear, and newly emitted NO can react with O<inline-formula><mml:math id="M275" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">3</mml:mn></mml:msub></mml:math></inline-formula>, leading to reductions in O<inline-formula><mml:math id="M276" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">3</mml:mn></mml:msub></mml:math></inline-formula> concentrations close to sources of NO (e.g. in urban areas and, in particular, close to roads). Consequently, ozone concentrations in urban areas are often lower than those in rural areas <xref ref-type="bibr" rid="bib1.bibx19" id="paren.34"/>, as shown in Fig. <xref ref-type="fig" rid="Ch1.F6"/>.</p>
      <p id="d1e5840">Figure <xref ref-type="fig" rid="Ch1.F7"/>a shows that the model produces a distribution shifted to the left toward lower values, not capturing the ozone for rural areas that are associated with higher concentrations of O<inline-formula><mml:math id="M277" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">3</mml:mn></mml:msub></mml:math></inline-formula>. Similarly, industrial suburban stations (panel d) have a higher frequency of high concentrations (higher than 25 <inline-formula><mml:math id="M278" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">µ</mml:mi></mml:mrow></mml:math></inline-formula>g m<inline-formula><mml:math id="M279" display="inline"><mml:msup><mml:mi/><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>), as shown in the histogram (panel d). Note that the majority of stations measuring O<inline-formula><mml:math id="M280" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">3</mml:mn></mml:msub></mml:math></inline-formula> are background rural or background urban, with few stations in other categories. With traffic urban (panel f), for which the model performs the worst, some modelled concentrations are much higher than observed measurements. This lack of fit may be explained because ozone is suppressed by new emissions of NO close to sources (traffic), which reduces the amount of O<inline-formula><mml:math id="M281" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">3</mml:mn></mml:msub></mml:math></inline-formula> at those station types.</p>
      <p id="d1e5893">From the same figure (panel c), as shown in the histogram for background urban stations, there is a high frequency of low concentrations (less than 10 <inline-formula><mml:math id="M282" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">µ</mml:mi></mml:mrow></mml:math></inline-formula>g m<inline-formula><mml:math id="M283" display="inline"><mml:msup><mml:mi/><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>) at these stations that the model does not capture. This is consistent with the reaction of newly emitted NO from urban roadside that reduces the concentrations of ozone in urban areas. Based on the RMSE and NMB, the model is a middle-performing model. As shown in Table <xref ref-type="table" rid="Ch1.T2"/>, the majority of stations measuring O<inline-formula><mml:math id="M284" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">3</mml:mn></mml:msub></mml:math></inline-formula> belong to this type.</p>
      <p id="d1e5927">As we mentioned earlier, NO<inline-formula><mml:math id="M285" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> is short-lived, so it has large differences between sites near sources (roadside) and those further away. Based on the RMSE, model 2 (CA<inline-formula><mml:math id="M286" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula>ENV) performs better with lower NO<inline-formula><mml:math id="M287" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> concentrations than high values, and since these high NO<inline-formula><mml:math id="M288" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> values exist near  traffic, the model performs the worst with traffic urban stations as shown in Fig. <xref ref-type="fig" rid="Ch1.F5"/>f. In contrast, the model's best performance is associated with background rural stations that have the lowest NO<inline-formula><mml:math id="M289" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> concentrations.</p>
      <p id="d1e5976">PM<inline-formula><mml:math id="M290" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2.5</mml:mn></mml:msub></mml:math></inline-formula> and PM<inline-formula><mml:math id="M291" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula> have many varied sources, so for roads and industrial sites they can be associated with local sources, for example widespread primary sources (direct emissions) and diffused secondary sources (i.e. produced in the atmosphere following emissions of precursor gases). Whilst PM concentrations are often greater at roadsides <xref ref-type="bibr" rid="bib1.bibx11" id="paren.35"/>, the particles can have lifetimes of several days in the atmosphere, meaning that they can be distributed widely. The larger particles are subject to greater loss via sedimentation, so PM<inline-formula><mml:math id="M292" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2.5</mml:mn></mml:msub></mml:math></inline-formula> is more evenly distributed than PM<inline-formula><mml:math id="M293" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula> <xref ref-type="bibr" rid="bib1.bibx23" id="paren.36"/>. This behaviour can also be observed with model 6 (Median) performance, for which there is less variation in the model performance under different environment types compared to the variation of NO<inline-formula><mml:math id="M294" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> and O<inline-formula><mml:math id="M295" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">3</mml:mn></mml:msub></mml:math></inline-formula>, as shown in Table <xref ref-type="table" rid="Ch1.T2"/>.</p>
      <p id="d1e6042">We also observed that the distributions of NO<inline-formula><mml:math id="M296" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula>, PM<inline-formula><mml:math id="M297" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2.5</mml:mn></mml:msub></mml:math></inline-formula>, and PM<inline-formula><mml:math id="M298" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula> are skewed to lower concentrations, which impacts model performance at higher concentrations. All models perform worse for high concentrations of NO<inline-formula><mml:math id="M299" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula>, PM<inline-formula><mml:math id="M300" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2.5</mml:mn></mml:msub></mml:math></inline-formula>, and PM<inline-formula><mml:math id="M301" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula> than O<inline-formula><mml:math id="M302" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">3</mml:mn></mml:msub></mml:math></inline-formula>, as indicated by the width of the quantiles at high values shown in Figs. <xref ref-type="fig" rid="Ch1.F2"/> and <xref ref-type="fig" rid="Ch1.F3"/>. Similarly, for lower concentrations, these models tend to perform better for NO<inline-formula><mml:math id="M303" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula>, PM<inline-formula><mml:math id="M304" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2.5</mml:mn></mml:msub></mml:math></inline-formula>, and PM<inline-formula><mml:math id="M305" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula> than for O<inline-formula><mml:math id="M306" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">3</mml:mn></mml:msub></mml:math></inline-formula>. However, our selected models (model 6, Median;  model 2, CA<inline-formula><mml:math id="M307" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula>ENV) are able to overcome this impact slightly.</p>
      <p id="d1e6158">Our approach enables us to impute and/or estimate plausible concentrations of multiple pollutants at stations across the UK, and the modelled concentrations from the selected models correlated well with the observed concentrations. The performance of these models is very good, with a slight underestimation in model 6 (Median), especially with high concentrations. At the opposite end, model 2 (CA <inline-formula><mml:math id="M308" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula> ENV) slightly overestimates the NO<inline-formula><mml:math id="M309" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> concentrations due to the regional behaviour of this pollutant.</p>
      <p id="d1e6177">We also analysed the performance of these models based on the daily modelled concentrations under different weather types using Lamb weather types (LWTs), which are a synoptic classification of daily weather patterns across the UK <xref ref-type="bibr" rid="bib1.bibx21" id="paren.37"/>. We found that these models work equally well for all LWTs, so we did not include this analysis in this work.</p>
      <p id="d1e6183">In conclusion, MVTS clustering enables imputation even when no measurement is available for a given pollutant since the station can be allocated to a cluster based on the value of the other pollutants measured. Our proposed imputation models, model 6 (Median) for O<inline-formula><mml:math id="M310" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">3</mml:mn></mml:msub></mml:math></inline-formula>, PM<inline-formula><mml:math id="M311" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula>, and PM<inline-formula><mml:math id="M312" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2.5</mml:mn></mml:msub></mml:math></inline-formula> and model 2 (CA <inline-formula><mml:math id="M313" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula> ENV), give the best performance for imputing these pollutants. The advantage of these models is that they aggregate the spatial and temporal imputation. The spatial imputation is obtained from the nearest stations and the temporal imputation is obtained by MVTS clustering that clusters the stations based on similarity in time.</p>
      <p id="d1e6220">In our future work, we aim to improve our imputation by considering more information about the stations, such as station altitude and location in relation to the weather effects. We may also consider the correlation between pollutants in our imputation and include further analysis for the Daily Air Quality Index (DAQI), especially for those days when there is  variation between imputed and observed DAQI. Finally, we need to study all possible uncertainty associated with this type of application, since the pollution level may change from year to year due to some pollution episodes caused by high temperature, wind, wildfire, or other factors.</p>
</sec>

      
      </body>
    <back><notes notes-type="codedataavailability"><title>Code and data availability</title>

      <?pagebreak page284?><p id="d1e6227">Code and data are available at <uri>https://github.com/Wedad-O-A/Modelled-concentrations-/tree/v3.5.2</uri> (last access: 27 October 2021) and <uri>https://doi.org/10.5281/zenodo.5602618</uri> <xref ref-type="bibr" rid="bib1.bibx1" id="paren.38"/>. Data at GitHub are generated by proposed imputation models. The inputs to these models are the original data collected from the DEFRA website (<uri>https://uk-air.defra.gov.uk/data/data_selector_service?show=auto&amp;submit=Reset&amp;f_limit_was=1</uri>, <xref ref-type="bibr" rid="bib1.bibx9" id="altparen.39"/>).</p>
  </notes><notes notes-type="authorcontribution"><title>Author contributions</title>

      <p id="d1e6248">The experimentation and initial draft were produced by WA as part of her PhD. IL and CER contributed ideas, co-supervised the PhD, and revised the draft papers. BDLI was the main supervisor for the work and contributed to the draft revisions.</p>
  </notes><notes notes-type="competinginterests"><title>Competing interests</title>

      <p id="d1e6254">The contact author has declared that neither they nor their co-authors have any competing interests.</p>
  </notes><notes notes-type="disclaimer"><title>Disclaimer</title>

      <p id="d1e6260">Publisher's note: Copernicus Publications remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.</p>
  </notes><ack><title>Acknowledgements</title><p id="d1e6266">We thank the anonymous referees for their useful suggestions and Salvatore Grimaldi for editing.</p></ack><notes notes-type="reviewstatement"><title>Review statement</title>

      <p id="d1e6272">This paper was edited by Salvatore Grimaldi and reviewed by two anonymous referees.</p>
  </notes><ref-list>
    <title>References</title>

      <ref id="bib1.bibx1"><?xmltex \def\ref@label{{Alahamade(2021)}}?><label>Alahamade(2021)</label><?label alahamade2021code?><mixed-citation>
Alahamade, W.: Wedad-O-A/Modelled-concentrations-: Modelled_Concentration_Air_Qaulity (v3.5.2), Zenodo [code and data set], https://doi.org/10.5281/zenodo.5602618, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx2"><?xmltex \def\ref@label{{Alahamade et~al.(2020)}}?><label>Alahamade et al.(2020)</label><?label alahamade2020clustering?><mixed-citation>Alahamade, W., Lake, I., Reeves, C. E., and De La Iglesia, B.: Clustering
Imputation for Air Pollution Data, in: International Conference on Hybrid
Artificial Intelligence Systems, Lecture Notes in Computer Science, 585–597, <ext-link xlink:href="https://doi.org/10.1007/978-3-030-61705-9_48" ext-link-type="DOI">10.1007/978-3-030-61705-9_48</ext-link>, Springer, Cham, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx3"><?xmltex \def\ref@label{{Alahamade et~al.(2021)}}?><label>Alahamade et al.(2021)</label><?label alahamade2021mvts?><mixed-citation>
Alahamade, W., Lake, I., Reeves, C. E., and De La Iglesia, B.: A Multi-variate
Time Series clustering approach based on Intermediate Fusion A case study in
air pollution data imputation, Neurocomputing, in press, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx4"><?xmltex \def\ref@label{{Austin et~al.(2013)}}?><label>Austin et al.(2013)</label><?label austin2013framework?><mixed-citation>Austin, E., Coull, B. A., Zanobetti, A., and Koutrakis, P.: A framework to
spatially cluster air pollution monitoring sites in US based on the PM<inline-formula><mml:math id="M314" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2.5</mml:mn></mml:msub></mml:math></inline-formula>
composition, Environ. Int., 59, 244–254, 2013.</mixed-citation></ref>
      <ref id="bib1.bibx5"><?xmltex \def\ref@label{{Carbajal-Hern{\'{a}}ndez et~al.(2012)}}?><label>Carbajal-Hernández et al.(2012)</label><?label carbajal2012assessment?><mixed-citation>
Carbajal-Hernández, J. J., Sánchez-Fernández, L. P.,
Carrasco-Ochoa, J. A., and Martínez-Trinidad, J. F.: Assessment and
prediction of air quality using fuzzy logic and autoregressive models,
Atmos. Environ., 60, 37–50, 2012.</mixed-citation></ref>
      <ref id="bib1.bibx6"><?xmltex \def\ref@label{{Carslaw and Ropkins(2012)}}?><label>Carslaw and Ropkins(2012)</label><?label carslaw2012Openair?><mixed-citation>
Carslaw, D. C. and Ropkins, K.: Openair – an R package for air quality data
analysis, Environ. Modell. Softw., 27, 52–61, 2012.</mixed-citation></ref>
      <ref id="bib1.bibx7"><?xmltex \def\ref@label{{CenterForCities(2020)}}?><label>CenterForCities(2020)</label><?label centreforcities_report?><mixed-citation>CenterForCities: Cities Outlook 2020 – Air quality in UK cities, available at:
<uri>https://www.centreforcities.org/publication/cities-outlook-2020/</uri> (last access: 22 October 2021), 2020.</mixed-citation></ref>
      <ref id="bib1.bibx8"><?xmltex \def\ref@label{{DEFRA(2013)}}?><label>DEFRA(2013)</label><?label Defra_DAQI?><mixed-citation>DEFRA: Daily Air Quality Index implementation Report, available at:
<uri>https://uk-air.defra.gov.uk/library/reports?report_id=750</uri> (last access: 22 October 2021), 2013.</mixed-citation></ref>
      <ref id="bib1.bibx9"><?xmltex \def\ref@label{{DEFRA(2019)}}?><label>DEFRA(2019)</label><?label Defra_Dataset?><mixed-citation>DEFRA (Department for Environment, Food &amp; Rural Affairs): Data Selector, available at: <uri>https://uk-air.defra.gov.uk/data/data_selector_service?show=auto&amp;submit=Reset&amp;f_limit_was=1</uri>, last access:
1 May 2019.</mixed-citation></ref>
      <ref id="bib1.bibx10"><?xmltex \def\ref@label{{DEFRA(2021)}}?><label>DEFRA(2021)</label><?label defra?><mixed-citation>DEFRA: About Air Pollution,
<uri>https://uk-air.defra.gov.uk/air-pollution</uri>, last access: 22 October 2021.</mixed-citation></ref>
      <ref id="bib1.bibx11"><?xmltex \def\ref@label{{DEFRA LAQM(2016)}}?><label>DEFRA LAQM(2016)</label><?label defralaqum?><mixed-citation>DEFRA LAQM: Public Health Sources and Effects of PM<inline-formula><mml:math id="M315" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2.5</mml:mn></mml:msub></mml:math></inline-formula>, available at:
<uri>https://laqm.defra.gov.uk/public-health/pm25.html</uri> (last access: 20 October 2021), 2016.</mixed-citation></ref>
      <ref id="bib1.bibx12"><?xmltex \def\ref@label{{Derwent et al.(2010)}}?><label>Derwent et al.(2010)</label><?label Derwent_report?><mixed-citation>
Derwent, D., Fraser, A., Abbott, J., Jenkin, M., Willis P., and Murrells, T.: Report:
Evaluating the Performance of Air Quality Models, Department for Environment,
Food and Rural Affairs, London, 2010.</mixed-citation></ref>
      <ref id="bib1.bibx13"><?xmltex \def\ref@label{{Diaz et~al.(2020)}}?><label>Diaz et al.(2020)</label><?label diaz2020ozone?><mixed-citation>Diaz, F. M., Khan, M. A. H., Shallcross, B., Shallcross, E. D., Vogt, U., and
Shallcross, D. E.: Ozone Trends in the United Kingdom over the Last 30 Years,
Atmosphere, 11, 534, <ext-link xlink:href="https://doi.org/10.3390/atmos11050534" ext-link-type="DOI">10.3390/atmos11050534</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx14"><?xmltex \def\ref@label{{Di~Bello et~al.(1996)}}?><label>Di Bello et al.(1996)</label><?label di1996parametric?><mixed-citation>
Di Bello, G., Lapenna, V., Macchiato, M., Satriano, C., Serio, C., and Tramutoli,
V.: Parametric time series analysis of geoelectrical signals: an
application to earthquake forecasting in Southern Italy, 1996.</mixed-citation></ref>
      <ref id="bib1.bibx15"><?xmltex \def\ref@label{{Du et~al.(2020)}}?><label>Du et al.(2020)</label><?label du2020multivariate?><mixed-citation>
Du, S., Li, T., Yang, Y., and Horng, S.-J.: Multivariate time series
forecasting via attention-based encoder–decoder framework, Neurocomputing,
388, 269–279, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx16"><?xmltex \def\ref@label{{D'Urso et~al.(2018)}}?><label>D'Urso et al.(2018)</label><?label d2018robust?><mixed-citation>
D'Urso, P., De Giovanni, L., and Massari, R.: Robust fuzzy clustering of
multivariate time trajectories, Int. J. Approx.
Reason., 99, 12–38, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx17"><?xmltex \def\ref@label{{Fontes and Budman(2017)}}?><label>Fontes and Budman(2017)</label><?label fontes2017hybrid?><mixed-citation>
Fontes, C. H. and Budman, H.: A hybrid clustering approach for multivariate
time series – a case study applied to failure analysis in a gas turbine, ISA
T., 71, 513–529, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx18"><?xmltex \def\ref@label{{Ignaccolo et~al.(2008)}}?><label>Ignaccolo et al.(2008)</label><?label ignaccolo2008analysis?><mixed-citation>
Ignaccolo, R., Ghigo, S., and Giovenali, E.: Analysis of air quality monitoring
networks by functional clustering, Environmetrics, 19, 672–686, 2008.</mixed-citation></ref>
      <ref id="bib1.bibx19"><?xmltex \def\ref@label{{Khan et~al.(2017)}}?><label>Khan et al.(2017)</label><?label h2017estimation?><mixed-citation>
Khan, M. A., Morris, W. C., Galloway, M., A. Shallcross, B. M., Percival,
C. J., and Shallcross, D. E.: An Estimation of the Levels of Stabilized
Criegee Intermediates in the UK Urban and Rural Atmosphere Using the
Steady-State Approximation and the Potential Effects of These Intermediates
on Tropospheric Oxidation Cycles, Int. J. Chem. Kinet.,
49, 611–621, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx20"><?xmltex \def\ref@label{{Lam(1983)}}?><label>Lam(1983)</label><?label lam1983spatial?><mixed-citation>
Lam, N. S.-N.: Spatial interpolation methods: a review,  Am.
Cartographer, 10, 129–150, 1983.</mixed-citation></ref>
      <ref id="bib1.bibx21"><?xmltex \def\ref@label{{Lamb(1972)}}?><label>Lamb(1972)</label><?label lamb1972british?><mixed-citation>
Lamb, H. H.: British Isles weather types and a register of daily sequence of circulation patterns, 1861–1971, Geophysical Memoir 116, HMSO, London, 85 pp., 1972.</mixed-citation></ref>
      <ref id="bib1.bibx22"><?xmltex \def\ref@label{{Liao(2005)}}?><label>Liao(2005)</label><?label liao2005clustering?><mixed-citation>
Liao, T. W.: Clustering of time series data – a survey, Pattern Recogn.,
38, 1857–1874, 2005.</mixed-citation></ref>
      <ref id="bib1.bibx23"><?xmltex \def\ref@label{{National Statistics(2020)}}?><label>National Statistics(2020)</label><?label pm25ANDpm10_repoert?><mixed-citation>National Statistics: National Statistics Concentrations of Particulate
Matter PM<inline-formula><mml:math id="M316" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula> and PM<inline-formula><mml:math id="M317" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">25</mml:mn></mml:msub></mml:math></inline-formula>, available at:
<uri>https://www.gov.uk/government/publications/air-quality-statistics/concentrations-of-particulate-matter-pm10-and-pm25</uri> (last access: 22 October 2021),
2020.</mixed-citation></ref>
      <?pagebreak page285?><ref id="bib1.bibx24"><?xmltex \def\ref@label{{Sarda-Espinosa(2017)}}?><label>Sarda-Espinosa(2017)</label><?label sarda2017package?><mixed-citation>Sarda-Espinosa, A.: Package “dtwclust”, available at: <uri>http://cran.ma.imperial.ac.uk/web/packages/dtwclust/dtwclust.pdf</uri> (last access: 22 October 2021), 2017.</mixed-citation></ref>
      <ref id="bib1.bibx25"><?xmltex \def\ref@label{{Seto et~al.(2015)}}?><label>Seto et al.(2015)</label><?label seto2015multivariate?><mixed-citation>
Seto, S., Zhang, W., and Zhou, Y.: Multivariate time series classification
using dynamic time warping template selection for human activity recognition,
in: 2015 IEEE Symposium Series on Computational Intelligence, 1399–1406,
IEEE, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx26"><?xmltex \def\ref@label{{Taylor(2001)}}?><label>Taylor(2001)</label><?label taylor2001summarizing?><mixed-citation>
Taylor, K. E.: Summarizing multiple aspects of model performance in a single
diagram, J. Geophys. Res.-Atmos., 106, 7183–7192, 2001.</mixed-citation></ref>
      <ref id="bib1.bibx27"><?xmltex \def\ref@label{{Tuysuzoglu et~al.(2019)}}?><label>Tuysuzoglu et al.(2019)</label><?label tuysuzoglu2019majority?><mixed-citation>Tuysuzoglu, G., Birant, D., and Pala, A.: Majority Voting Based Multi-Task
Clustering of Air Quality Monitoring Network in Turkey, Appl. Sci., 9,
1610, <ext-link xlink:href="https://doi.org/10.3390/app9081610" ext-link-type="DOI">10.3390/app9081610</ext-link>, 2019.
</mixed-citation></ref><?xmltex \hack{\newpage}?>
      <ref id="bib1.bibx28"><?xmltex \def\ref@label{{Wickham et al.(2017)}}?><label>Wickham et al.(2017)</label><?label wickham2017package?><mixed-citation>Wickham, H., Averick, M., Bryan, J., Chang, W.,  D'Agostino McGowan, L.,  François, R., Grolemund, G., Hayes, A., Henry, L., Hester, J.,
Kuhn, M., Lin Pedersen, T., Miller, E.,  Milton Bache, S., Müller, K.,  Ooms, J., Robinson, D., Paige Seidel, D., Spinu, V., Takahashi, K., Vaughan, D.,  Wilke, C.,  Woo, K., and Yutani, H.: Package tidyverse, Easily Install and Load the
Tidyverse, Journal of Open Source Software, 4, 1686, <ext-link xlink:href="https://doi.org/10.21105/joss.01686" ext-link-type="DOI">10.21105/joss.01686</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx29"><?xmltex \def\ref@label{{Zhou and Chan(2014)}}?><label>Zhou and Chan(2014)</label><?label zhou2014model?><mixed-citation>
Zhou, P.-Y. and Chan, K. C.: A model-based multivariate time series clustering
algorithm, in: Pacific-Asia Conference on Knowledge Discovery and Data Mining, 805–817, Springer, 2014.</mixed-citation></ref>

  </ref-list></back>
    <!--<article-title-html>Evaluation of multivariate time series clustering for imputation of air pollution data</article-title-html>
<abstract-html/>
<ref-html id="bib1.bib1"><label>Alahamade(2021)</label><mixed-citation>
Alahamade, W.: Wedad-O-A/Modelled-concentrations-: Modelled_Concentration_Air_Qaulity (v3.5.2), Zenodo [code and data set], https://doi.org/10.5281/zenodo.5602618, 2021.
</mixed-citation></ref-html>
<ref-html id="bib1.bib2"><label>Alahamade et al.(2020)</label><mixed-citation>
Alahamade, W., Lake, I., Reeves, C. E., and De La Iglesia, B.: Clustering
Imputation for Air Pollution Data, in: International Conference on Hybrid
Artificial Intelligence Systems, Lecture Notes in Computer Science, 585–597, <a href="https://doi.org/10.1007/978-3-030-61705-9_48" target="_blank">https://doi.org/10.1007/978-3-030-61705-9_48</a>, Springer, Cham, 2020.
</mixed-citation></ref-html>
<ref-html id="bib1.bib3"><label>Alahamade et al.(2021)</label><mixed-citation>
Alahamade, W., Lake, I., Reeves, C. E., and De La Iglesia, B.: A Multi-variate
Time Series clustering approach based on Intermediate Fusion A case study in
air pollution data imputation, Neurocomputing, in press, 2021.
</mixed-citation></ref-html>
<ref-html id="bib1.bib4"><label>Austin et al.(2013)</label><mixed-citation>
Austin, E., Coull, B. A., Zanobetti, A., and Koutrakis, P.: A framework to
spatially cluster air pollution monitoring sites in US based on the PM<sub>2.5</sub>
composition, Environ. Int., 59, 244–254, 2013.
</mixed-citation></ref-html>
<ref-html id="bib1.bib5"><label>Carbajal-Hernández et al.(2012)</label><mixed-citation>
Carbajal-Hernández, J. J., Sánchez-Fernández, L. P.,
Carrasco-Ochoa, J. A., and Martínez-Trinidad, J. F.: Assessment and
prediction of air quality using fuzzy logic and autoregressive models,
Atmos. Environ., 60, 37–50, 2012.
</mixed-citation></ref-html>
<ref-html id="bib1.bib6"><label>Carslaw and Ropkins(2012)</label><mixed-citation>
Carslaw, D. C. and Ropkins, K.: Openair – an R package for air quality data
analysis, Environ. Modell. Softw., 27, 52–61, 2012.
</mixed-citation></ref-html>
<ref-html id="bib1.bib7"><label>CenterForCities(2020)</label><mixed-citation>
CenterForCities: Cities Outlook 2020 – Air quality in UK cities, available at:
<a href="https://www.centreforcities.org/publication/cities-outlook-2020/" target="_blank"/> (last access: 22 October 2021), 2020.
</mixed-citation></ref-html>
<ref-html id="bib1.bib8"><label>DEFRA(2013)</label><mixed-citation>
DEFRA: Daily Air Quality Index implementation Report, available at:
<a href="https://uk-air.defra.gov.uk/library/reports?report_id=750" target="_blank"/> (last access: 22 October 2021), 2013.
</mixed-citation></ref-html>
<ref-html id="bib1.bib9"><label>DEFRA(2019)</label><mixed-citation>
DEFRA (Department for Environment, Food &amp; Rural Affairs): Data Selector, available at: <a href="https://uk-air.defra.gov.uk/data/data_selector_service?show=auto&amp;submit=Reset&amp;f_limit_was=1" target="_blank"/>, last access:
1 May 2019.
</mixed-citation></ref-html>
<ref-html id="bib1.bib10"><label>DEFRA(2021)</label><mixed-citation>
DEFRA: About Air Pollution,
<a href="https://uk-air.defra.gov.uk/air-pollution" target="_blank"/>, last access: 22 October 2021.
</mixed-citation></ref-html>
<ref-html id="bib1.bib11"><label>DEFRA LAQM(2016)</label><mixed-citation>
DEFRA LAQM: Public Health Sources and Effects of PM<sub>2.5</sub>, available at:
<a href="https://laqm.defra.gov.uk/public-health/pm25.html" target="_blank"/> (last access: 20 October 2021), 2016.
</mixed-citation></ref-html>
<ref-html id="bib1.bib12"><label>Derwent et al.(2010)</label><mixed-citation>
Derwent, D., Fraser, A., Abbott, J., Jenkin, M., Willis P., and Murrells, T.: Report:
Evaluating the Performance of Air Quality Models, Department for Environment,
Food and Rural Affairs, London, 2010.
</mixed-citation></ref-html>
<ref-html id="bib1.bib13"><label>Diaz et al.(2020)</label><mixed-citation>
Diaz, F. M., Khan, M. A. H., Shallcross, B., Shallcross, E. D., Vogt, U., and
Shallcross, D. E.: Ozone Trends in the United Kingdom over the Last 30 Years,
Atmosphere, 11, 534, <a href="https://doi.org/10.3390/atmos11050534" target="_blank">https://doi.org/10.3390/atmos11050534</a>, 2020.
</mixed-citation></ref-html>
<ref-html id="bib1.bib14"><label>Di Bello et al.(1996)</label><mixed-citation>
Di Bello, G., Lapenna, V., Macchiato, M., Satriano, C., Serio, C., and Tramutoli,
V.: Parametric time series analysis of geoelectrical signals: an
application to earthquake forecasting in Southern Italy, 1996.
</mixed-citation></ref-html>
<ref-html id="bib1.bib15"><label>Du et al.(2020)</label><mixed-citation>
Du, S., Li, T., Yang, Y., and Horng, S.-J.: Multivariate time series
forecasting via attention-based encoder–decoder framework, Neurocomputing,
388, 269–279, 2020.
</mixed-citation></ref-html>
<ref-html id="bib1.bib16"><label>D'Urso et al.(2018)</label><mixed-citation>
D'Urso, P., De Giovanni, L., and Massari, R.: Robust fuzzy clustering of
multivariate time trajectories, Int. J. Approx.
Reason., 99, 12–38, 2018.
</mixed-citation></ref-html>
<ref-html id="bib1.bib17"><label>Fontes and Budman(2017)</label><mixed-citation>
Fontes, C. H. and Budman, H.: A hybrid clustering approach for multivariate
time series – a case study applied to failure analysis in a gas turbine, ISA
T., 71, 513–529, 2017.
</mixed-citation></ref-html>
<ref-html id="bib1.bib18"><label>Ignaccolo et al.(2008)</label><mixed-citation>
Ignaccolo, R., Ghigo, S., and Giovenali, E.: Analysis of air quality monitoring
networks by functional clustering, Environmetrics, 19, 672–686, 2008.
</mixed-citation></ref-html>
<ref-html id="bib1.bib19"><label>Khan et al.(2017)</label><mixed-citation>
Khan, M. A., Morris, W. C., Galloway, M., A. Shallcross, B. M., Percival,
C. J., and Shallcross, D. E.: An Estimation of the Levels of Stabilized
Criegee Intermediates in the UK Urban and Rural Atmosphere Using the
Steady-State Approximation and the Potential Effects of These Intermediates
on Tropospheric Oxidation Cycles, Int. J. Chem. Kinet.,
49, 611–621, 2017.
</mixed-citation></ref-html>
<ref-html id="bib1.bib20"><label>Lam(1983)</label><mixed-citation>
Lam, N. S.-N.: Spatial interpolation methods: a review,  Am.
Cartographer, 10, 129–150, 1983.
</mixed-citation></ref-html>
<ref-html id="bib1.bib21"><label>Lamb(1972)</label><mixed-citation>
Lamb, H. H.: British Isles weather types and a register of daily sequence of circulation patterns, 1861–1971, Geophysical Memoir 116, HMSO, London, 85 pp., 1972.
</mixed-citation></ref-html>
<ref-html id="bib1.bib22"><label>Liao(2005)</label><mixed-citation>
Liao, T. W.: Clustering of time series data – a survey, Pattern Recogn.,
38, 1857–1874, 2005.
</mixed-citation></ref-html>
<ref-html id="bib1.bib23"><label>National Statistics(2020)</label><mixed-citation>
National Statistics: National Statistics Concentrations of Particulate
Matter PM<sub>10</sub> and PM<sub>25</sub>, available at:
<a href="https://www.gov.uk/government/publications/air-quality-statistics/concentrations-of-particulate-matter-pm10-and-pm25" target="_blank"/> (last access: 22 October 2021),
2020.
</mixed-citation></ref-html>
<ref-html id="bib1.bib24"><label>Sarda-Espinosa(2017)</label><mixed-citation>
Sarda-Espinosa, A.: Package “dtwclust”, available at: <a href="http://cran.ma.imperial.ac.uk/web/packages/dtwclust/dtwclust.pdf" target="_blank"/> (last access: 22 October 2021), 2017.
</mixed-citation></ref-html>
<ref-html id="bib1.bib25"><label>Seto et al.(2015)</label><mixed-citation>
Seto, S., Zhang, W., and Zhou, Y.: Multivariate time series classification
using dynamic time warping template selection for human activity recognition,
in: 2015 IEEE Symposium Series on Computational Intelligence, 1399–1406,
IEEE, 2015.
</mixed-citation></ref-html>
<ref-html id="bib1.bib26"><label>Taylor(2001)</label><mixed-citation>
Taylor, K. E.: Summarizing multiple aspects of model performance in a single
diagram, J. Geophys. Res.-Atmos., 106, 7183–7192, 2001.
</mixed-citation></ref-html>
<ref-html id="bib1.bib27"><label>Tuysuzoglu et al.(2019)</label><mixed-citation>
Tuysuzoglu, G., Birant, D., and Pala, A.: Majority Voting Based Multi-Task
Clustering of Air Quality Monitoring Network in Turkey, Appl. Sci., 9,
1610, <a href="https://doi.org/10.3390/app9081610" target="_blank">https://doi.org/10.3390/app9081610</a>, 2019.

</mixed-citation></ref-html>
<ref-html id="bib1.bib28"><label>Wickham et al.(2017)</label><mixed-citation>
Wickham, H., Averick, M., Bryan, J., Chang, W.,  D'Agostino McGowan, L.,  François, R., Grolemund, G., Hayes, A., Henry, L., Hester, J.,
Kuhn, M., Lin Pedersen, T., Miller, E.,  Milton Bache, S., Müller, K.,  Ooms, J., Robinson, D., Paige Seidel, D., Spinu, V., Takahashi, K., Vaughan, D.,  Wilke, C.,  Woo, K., and Yutani, H.: Package tidyverse, Easily Install and Load the
Tidyverse, Journal of Open Source Software, 4, 1686, <a href="https://doi.org/10.21105/joss.01686" target="_blank">https://doi.org/10.21105/joss.01686</a>, 2017.
</mixed-citation></ref-html>
<ref-html id="bib1.bib29"><label>Zhou and Chan(2014)</label><mixed-citation>
Zhou, P.-Y. and Chan, K. C.: A model-based multivariate time series clustering
algorithm, in: Pacific-Asia Conference on Knowledge Discovery and Data Mining, 805–817, Springer, 2014.
</mixed-citation></ref-html>--></article>
