article · Journal of Statistical Sciences and Computational Intelligence
Missing data frequently impairs the reliability and efficiency of statistical estimates drawn from surveys. To address this, an evaluation was conducted on a regression-type estimator designed to calculate population means within a stratified two-stage sampling structure containing incomplete observations. The assessment utilised field survey records consisting of school attendance records as auxiliary data and mathematics test scores as the primary variable. Within this setup, schools formed primary sampling units while students represented secondary units. Missing values were simulated by setting twenty percent of test scores as missing completely at random, which were subsequently replaced using ratio and regression imputation methods. Tested across multiple sample sizes against existing ratio and difference estimators, the regression estimator consistently yielded lower variances and coefficients of variation. Its relative efficiency rose with larger sample sizes, particularly when paired with regression imputation.
Real-world surveys routinely suffer from incomplete responses, which can distort results and weaken statistical confidence. Demonstrating that a regression estimator handles missing values effectively provides survey practitioners with a more dependable method for calculating population averages from multi-stage field data, such as educational testing records.
The method represents early-stage analytical research that could be integrated into survey processing software, educational assessment platforms, or statistical toolkits used by government statistical agencies and polling organisations. While demonstrated on empirical school survey data, the abstract does not indicate that software packages or direct commercial tools have yet been built.
AI-generated from the published abstract. Always read the original work before citing.
Missing data is a recurring challenge in survey sampling, often reducing the efficiency and reliability of estimators. This study proposed and investigated the regression-type estimator of the population mean under a stratified two-stage sampling design in the presence of missing values. The population of the study comprised of Field survey data on students’ school attendance (auxiliary variable) and mathematics test scores (study variable). A stratified two-stage design was adopted, with schools as primary sampling units and students as secondary units. To reflect item nonresponse, 20% of the study variable was declared missing completely at random (MCAR) and handled through both regression and ratio imputation. The performance of the proposed regression estimator was compared with the Bahl-Saini (2011) ratio and difference estimator across sample sizes of 25, 40, 70, and 100, using coefficient of variation (CV), and confidence intervals as evaluation criteria. Results showed that the regression estimator consistently achieved lower variances and CVs than the existing estimators, with efficiency improving as sample size increased. Even in the presence of missing data, the regression estimator maintained superior performance, particularly under regression imputation. The study demonstrated the efficiency of the regression estimator in handling incomplete data and highlights its practical significance for reliable estimation in complex survey designs.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.64497/jssci.128
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.