Variable Selection and Redundancy in Multivariate Regression Models

Westad, Frank and Marini, Federico (2022) Variable Selection and Redundancy in Multivariate Regression Models. Frontiers in Analytical Science, 2. ISSN 2673-9283

[thumbnail of pubmed-zip/versions/2/package-entries/frans-02-897605-r1/frans-02-897605.pdf] Text
pubmed-zip/versions/2/package-entries/frans-02-897605-r1/frans-02-897605.pdf - Published Version

Download (1MB)

Abstract

Variable selection is a topic of interest in many scientific communities. Within chemometrics, where the number of variables for multi-channel instruments like NIR spectroscopy and metabolomics in many situations is larger than the number of samples, the strategy has been to use latent variable regression methods to overcome the challenges with multiple linear regression. Thereby, there is no need to remove variables as such, as the low-rank models handle collinearity and redundancy. In most studies on variable selection, the main objective was to compare the prediction performance (RMSE or accuracy in classification) between various methods. Nevertheless, different methods with the same objective will, in most cases, give results that are not significantly different. In this study, we present three other main objectives: i) to eliminate variables that are not relevant; ii) to return a small subset of variables that has the same or better prediction performance as a model with all original variables; and iii) to investigate the consistency of these small subsets.

Item Type: Article
Subjects: European Repository > Chemical Science
Depositing User: Managing Editor
Date Deposited: 25 Nov 2022 04:31
Last Modified: 23 Feb 2024 03:39
URI: http://go7publish.com/id/eprint/393

Actions (login required)

View Item
View Item