Evaluation of machine learning algorithms for classification of primary biological aerosol using a new UV-LIF spectrometer



Ruske, Simon, Topping, David O, Foot, Virginia E, Kaye, Paul H, Stanley, Warren R, Crawford, Ian, Morse, Andrew P ORCID: 0000-0002-0413-2065 and Gallagher, Martin W
(2017) Evaluation of machine learning algorithms for classification of primary biological aerosol using a new UV-LIF spectrometer. ATMOSPHERIC MEASUREMENT TECHNIQUES, 10 (2). pp. 695-708.

This is the latest version of this item.

Access the full-text of this item by clicking on the Open Access link.
[img] Text
amt_2016_214.pdf - Published version

Download (2MB)
[img] Text
amt-10-695-2017.pdf - Published version

Download (3MB)

Abstract

<jats:p>Abstract. Characterisation of bio-aerosols has important implications within Environment and Public Health sectors. Recent developments in Ultra-Violet Light Induced Fluorescence (UV-LIF) detectors such as the Wideband Integrated bio-aerosol Spectrometer (WIBS) and the newly introduced Multiparameter bio-aerosol Spectrometer (MBS) has allowed for the real time collection of fluorescence, size and morphology measurements for the purpose of discriminating between bacteria, fungal Spores and pollen. This new generation of instruments has enabled ever larger data sets to be compiled with the aim of studying more complex environments. In real world data sets, particularly those from an urban environment, the population may be dominated by non-biological fluorescent interferents bringing into question the accuracy of measurements of quantities such as concentrations. It is therefore imperative that we validate the performance of different algorithms which can be used for the task of classification. For unsupervised learning we test Hierarchical Agglomerative Clustering with various different linkages. For supervised learning, ten methods were tested; including decision trees, ensemble methods: Random Forests, Gradient Boosting and AdaBoost; two implementations for support vector machines: libsvm and liblinear; Gaussian methods: Gaussian naïve Bayesian, quadratic and linear discriminant analysis and finally the k-nearest neighbours algorithm. The methods were applied to two different data sets measured using a new Multiparameter bio-aerosol Spectrometer which provides multichannel UV-LIF fluorescence signatures for single airborne biological particles. Clustering, in general performs slightly worse than the supervised learning methods correctly classifying, at best, only 72.7 and 91.1 percent for the two data sets respectively. For supervised learning the gradient boosting algorithm was found to be the most effective, on average correctly classifying 88.1 and 97.8 percent of the testing data respectively across the two data sets. </jats:p>

Item Type: Article
Additional Information: This paper is in public online review -- there does not seem to be a category of submission and review?
Uncontrolled Keywords: Generic health relevance
Depositing User: Symplectic Admin
Date Deposited: 03 Mar 2017 11:46
Last Modified: 14 Mar 2024 19:14
DOI: 10.5194/amt-10-695-2017
Open Access URL: http://www.atmos-meas-tech-discuss.net/amt-2016-21...
Related URLs:
URI: https://livrepository.liverpool.ac.uk/id/eprint/3006190

Available Versions of this Item