THE IMPORTANCE OF FOURIER TRANSFORM INFRARED SPECTRA PROCESSING FOR ANALYSIS OF BIOLOGICAL SAMPLES

Tautvydas Taraškevičius1, Rimantė Bandzevičiūtė1, Vytautas Žėkas2, Emilis Gabrielis Byčius2, Dovilė Karčiauskaitė2

1 Institute of Chemical Physics, Vilnius University, Saulėtekio av. 3, 10257 Vilnius, Lithuania

2 Department of Physiology, Biochemistry, Microbiology and Laboratory Medicine, Institute of Biomedical Sciences, Faculty of Medicine, Vilnius University, M. K. Čiurlionio st. 21, 03101 Vilnius, Lithuania

[email protected]

Fourier transform infrared (FTIR) spectroscopy is a versatile vibrational spectroscopy method with great potential for applications in biological and medical research. Its main advantages are the relative ease of use, rapid acquisition of spectra, non-destructivity and non-specificity. However, the accuracy of the technique is hampered by unwanted phenomena introducing deviations of the spectra. This necessitates the application of preprocessing techniques, such as baseline correction and spectra normalization. The goal of this study is to evaluate the impact spectra preprocessing has on infrared spectra classification models based on a case study, as well to identify good practices in FTIR spectra analysis of biological samples.

FTIR spectra of blood derived extracellular vesicles (EVs) from healthy individuals and post myocardial infarction patients were used as a case study. The vesicle samples were prepared at the Vilnius University Faculty of Medicine laboratories. A total of 84 FTIR spectra were used in the analysis, 26 of which were from EVs of infarction patients. The spectra were preprocessed using various techniques and outputted to supervised and unsupervised classification models with the objective of differentiating between vesicles and evaluating the classification accuracy. For unsupervised models, k-means clustering and hierarchical cluster analysis (HCA) were used. For supervised models, logistic regression with and without principal component analysis (PCA) was used. The supervised models were trained and tested using 5-fold cross-validation.

Based on the results of the models (Fig. 1), the following observations were made. First, baseline correction and spectra normalization are vital for cluster analysis. Second, supervised models tend to perform better if the spectral data have their standard deviations normalized to 1. Furthermore, standard deviation normalization also increased the amount of principal components required to achieve a high explained variance. It should be noted that each spectral dataset has its own optimal data processing procedure that usually has to be found by comparing several methods.

Figure 1
Fig. 1. The classification accuracy of extracellular vesicle FTIR spectra preprocessed using various techniques for each of the classification models used in this study.