Should data ever be thrown away? Pooling interval-censored data sets with different precision

Tretiak, Krasymyr and Ferson, Scott ORCID: 0000-0002-2613-0650 (2023) Should data ever be thrown away? Pooling interval-censored data sets with different precision. International Journal of Approximate Reasoning, 156. pp. 114-133.

Access the full-text of this item by clicking on the Open Access link.

Official URL: http://dx.doi.org/10.1016/j.ijar.2023.02.007

Abstract

Data quality is an important consideration in many engineering applications and projects. Data collection procedures do not always involve careful utilization of the most precise instruments and strictest protocols. As a consequence, data are invariably affected by imprecision and sometimes sharply varying levels of quality of the data. Different mathematical representations of imprecision have been suggested, including a classical approach to censored data which is considered optimal when the proposed error model is correct, and a weaker approach called interval statistics based on partial identification that makes fewer assumptions. Maximizing the quality of statistical results is often crucial to the success of many engineering projects, and a natural question that arises is whether data of differing qualities should be pooled together or we should include only precise measurements and disregard imprecise data. Some worry that combining precise and imprecise measurements can depreciate the overall quality of the pooled data. Some fear that excluding data of lesser precision can increase their overall uncertainty about results because lower sample size implies more sampling uncertainty. This paper explores these concerns and describes simulation results that show when it is advisable to combine fairly precise data with rather imprecise data by comparing analyses using different mathematical representations of imprecision. Pooling data sets is preferred when the low-quality data set does not exceed a certain level of uncertainty. However, so long as the data are random, it may be legitimate to reject the low-quality data if its reduction of sampling uncertainty does not counterbalance the effect of its imprecision on the overall uncertainty.

Item Type:	Article
Uncontrolled Keywords:	Imprecise data, Censoring, Maximum likelihood, Epistemic uncertainty, Kolmogorov-Smirnov, Descriptive statistics
Divisions:	Faculty of Science and Engineering > School of Engineering
Depositing User:	Symplectic Admin
Date Deposited:	03 Mar 2023 10:52
Last Modified:	05 Apr 2023 12:16
DOI:	10.1016/j.ijar.2023.02.007
Open Access URL:	https://doi.org/10.1016/j.ijar.2023.02.007
Related URLs:	Author Publisher
URI:	https://livrepository.liverpool.ac.uk/id/eprint/3168733