Westad, Frank and Marini, Federico (2022) Variable Selection and Redundancy in Multivariate Regression Models. Frontiers in Analytical Science, 2. ISSN 2673-9283
pubmed-zip/versions/2/package-entries/frans-02-897605-r1/frans-02-897605.pdf - Published Version
Download (1MB)
Abstract
Variable selection is a topic of interest in many scientific communities. Within chemometrics, where the number of variables for multi-channel instruments like NIR spectroscopy and metabolomics in many situations is larger than the number of samples, the strategy has been to use latent variable regression methods to overcome the challenges with multiple linear regression. Thereby, there is no need to remove variables as such, as the low-rank models handle collinearity and redundancy. In most studies on variable selection, the main objective was to compare the prediction performance (RMSE or accuracy in classification) between various methods. Nevertheless, different methods with the same objective will, in most cases, give results that are not significantly different. In this study, we present three other main objectives: i) to eliminate variables that are not relevant; ii) to return a small subset of variables that has the same or better prediction performance as a model with all original variables; and iii) to investigate the consistency of these small subsets.
Item Type: | Article |
---|---|
Subjects: | European Repository > Chemical Science |
Depositing User: | Managing Editor |
Date Deposited: | 25 Nov 2022 04:31 |
Last Modified: | 23 Feb 2024 03:39 |
URI: | http://go7publish.com/id/eprint/393 |