TL;DR: In this article, an analytical formula is derived to approximate the finite sample bias of the OLS estimator of the autoregressive parameter when the underlying process has a unit root.
Abstract: An analytical formula is derived to approximate the finite sample bias of the ordinary least-squares (OLS) estimator of the autoregressive parameter when the underlying process has a unit root. It is found that the bias is expressible in terms of parabolic cylinder functions which are easy to compute. Numerical evaluation of the formula reveals that the approximation is very accurate. The derived formula inspires a heuristic approximation, obtained by leastsquares fitting of the asymptotic bias. More importantly, the formula proves analytically that the bias declines at a rate which is slower than the consistency rate, thus explaining some previous simulation findings. A case where the bias increases with the sample size is also given.
TL;DR: The early years of the Cote d'Ivoire Standards Survey (CILSS) had a sampling bias, which seriously affected estimates of population statistics such as household size as mentioned in this paper.
Abstract: The sampling aspects of a household data set are important to analysts. The early years of the Cote d'Ivoire Standards Survey (CILSS) had a sampling bias, which seriously affected estimates of population statistics such as household size. The bias arose from sampling procedures that overrepresented larger dwellings. Assuming that samples drawn in later years were unbiased, a correction procedure is applied that uses weights based on household size. Results from the weighted data are then compared with the unweighted findings to assess the seriousness of the bias. Estimates of household expenditure per capita in the early years of the survey are found to be significantly underestimated, resulting in an overestimation of poverty. The sampling bias also resulted in an underestimation of the upward trend in poverty during 1985-88. The CILSS has been a popular and fruitful data set for policy analysis. These findings, however, cast doubt on the robustness of earlier work. Thus, the effort to trace sampling information is particularly worthwhile for policy-oriented applied research.
TL;DR: This paper argued that commitment status does not explain sampling discrepancies in studies on gender differences in schizophrenia, and that the inconsistency in findings most likely can be explained by sampling differences across studies, such as help-seeking behavior and commitment status.
Abstract: Some of the findings on gender differences in schizophrenia have been inconsistent, particularly, those regarding brain abnormalities. The inconsistency in findings most likely can be explained by sampling differences across studies. Walker and Lewine argue that the sampling differences can be explained, in large part, by gender differences in help-seeking behavior and commitment status. This article argues that commitment status does not explain sampling discrepancies in studies on gender differences in schizophrenia .
TL;DR: In this article, the authors use a mathematical model to estimate the magnitude of possible selection and its effects on the mean and variance of GATB validities and find that some of the validity studies conducted were not included in the database and speculated that these studies may have found negligible or even negative validities.
Abstract: Previous analyses have suggested that the database of 755 studies of the validity of the General Aptitude Test Battery (GATB) demonstrates a small but consistent positive correlation with criteria relevant to job performance. Critics have noted that some of the validity studies conducted were not included in the database and have speculated that these studies may have found negligible or even negative validities, so that the extant database is subject to selection bias. The authors use a mathematical model to estimate the magnitude of possible selection and its effects on the mean and variance of GATB validities
TL;DR: In this article, the authors report on household size changes in the Cote d'Ivoire between 1985 and 1988 as evidenced by the Living Standards Survey (CILSS) and identify a change in sampling procedures as the most likely cause of the problem.
Abstract: This paper reports on household size changes in the Cote d'Ivoire between 1985 and 1988 as evidenced by the Cote d'Ivoire Living Standards Survey (CILSS). The decline, from 8.31 to just 6.32, cannot be explained in terms of real-world changes alone, and so must be due also to either sampling bias or non-sampling errors. The paper identifies a change in sampling procedures as the most likely cause of the problem. The over-enumeration of large households in the early years of the survey is also reflected in dwelling size changes. However, observed declines in household size measured within the same sampling arrangement, and even within the same panel of households, suggest that there is also a real-world decline in household size in Cote d'Ivoire. The paper concludes that care must be taken in over-time and cross-section analyses with the CILSS data, and an appropriate re-weighting of the data is called for.
TL;DR: A sampling design and estimation procedures that directly utilize the unequal probabilities with which nests are included in the sample, while acknowledging the biased nature of the data are presented.
Abstract: Encounter-sampling designs are data-collection procedures in which population units are included in the sample as they are detected, or "encountered." Perhaps the most familiar examples of encountersampling designs are line-transect (e.g. Burnham et al. 1980) and line-intercept (e.g. De Vries 1979) procedures. One characteristic of encounter-sampling designs is that a sampling frame (e.g. Cochran 1977) of population units is not required. When individuals are mobile, elusive, or possess other characteristics that make it difficult to construct a sampling frame, encounter-sampling designs often provide the only effective means of sampling a population. A second characteristic of encounter-sampling designs is that they lack control over the subset of the population comprising the sample. As a result, data collected by these procedures often are not representative of the population of interest (i.e. a biased sample), and are best viewed as a probability sample. If variables of interest are correlated with the probability of inclusion, the data cannot be treated as a simple random sample, and estimators based upon random sampling theory are biased (Rao 1965). In these cases, designspecific estimators or bias corrections dependent on the probability structure of the data must be employed. A population of bird nests is one example of a population whose study requires the use of an encountersampling technique. In addition to the obvious lack of a sampling frame, the population is demographically open in that nests are initiated and fail through time. The usual sampling design consists of conducting searches for viable nests, including all detected nests in the sample. Data collected under such a design are biased because longer-lived nests are included in the sample with higher probability than are shorter-lived nests (Mayfield 1961). However, the method by which searches are conducted, typically, is not structured and is not helpful in deriving estimators. The parameter most often of interest is the nestsurvival rate (i.e. probability a nest survives to "success"). Success is often defined as the production of at least one offspring, though other definitions are equally appropriate. Nest-survival rates may be estimated using a number of models. Mayfield (1961, 1975), Johnson (1979), and Bart and Robson (1982) modeled nest survival after detection. Hensler and Nichols (1981), Pollock and Cornelius (1988), and Bromaghin and McDonald (1993) modeled the entire existence of nests by incorporating probabilities of inclusion and partial information of the total lifetime of detected nests in the model. While all of these models attempt to treat the probability structure of the data, only the Pollock-Cornelius and BromaghinMcDonald models do so fully and correctly. Although Heisey and Nordheim (1990) indicated that the Pollock-Cornelius model produces biased estimates of nest survival, recent information (Pollock and Cornelius unpubl. data) suggests that the bias decreases as the time between visits to nests decreases. All of the nest-survival models are designed to produce estimators of a single parameter, the probability of nest-success, while acknowledging the biased nature of the data. To some extent, they have all been successful. However, any number of additional parameters may also be of interest, and their estimation, which is also complicated by the biased nature of the data, has received little attention. We present a sampling design and estimation procedures that directly utilize the unequal probabilities with which nests are included in the sample. The method is developed from a classical sampling approach in that all characteristics of the population of nests are considered as fixed; randomness observed in the data is due solely to the sampling design. The design consists of temporally systematic searches for nests, and nests are included in the sample as they are detected. The model assumes that nests are (approximately) aged at the time of detection and monitored until they either fail or are successful. The systematic design permits probabilities of inclusion to be estimated. The estimates are employed in modified Horvitz-Thompson estimators (Horvitz and Thompson 1952) to obtain estimates of any parameter that can be expressed as a total or a ratio of totals. Such parameters include the probability a nest survives to success, the number of nests which survive to success, and the numbers of nests initiated. The sampling model.-Consider a fixed geographic area in which birds are nesting. Nests are initiated and survive for some period of time. The population of interest consists of all nests which exist within the area for any portion of a specified time frame. For example, the time frame might be constructed to contain all or some interesting portion of the nesting season of the species under consideration. Thus, a geographic area and a time frame are used to define the population. The population size is denoted N and the number of time units in the time frame of interest is denoted D. The time frame of D time units is divided into a
TL;DR: In this paper, the effect of filtering on the distribution of parameters of a dynamic regression model with a lagged dependent variable and a set of exogenous regressors was investigated, using the Census X-11 filter as a specific example.
Abstract: It is common for an applied researcher to use filtered data, like seasonally adjusted series, for instance, to estimate the parameters of a dynamic regression model. In this paper, we study the effect of (linear) filters on the distribution of parameters of a dynamic regression model with a lagged dependent variable and a set of exogenous regressors. So far, only asymptotic results are available. Our main interest is to investigate the effect of filtering on the small sample bias and mean squared error. In general, these results entail a numerical integration of derivatives of the joint moment generating function of two quadratic forms in normal variables. The computation of these integrals is quite involved. However, we take advantage of the Laplace approximations to the bias and mean squared error, which substantially reduce the computational burden, as they yield relatively simple analytic expressions. We obtain analytic formulae for approximating the effect of filtering on the finite sample bias and mean squared error. We evaluate the adequacy of the approximations by comparison with Monte Carlo simulations, using the Census X-11 filter as a specific example
TL;DR: This article showed that the observed changes in household welfare and in Cote d'Ivoire between 1985 and 1988 vanish when corrections are applied to the data for changes in sampling procedures; even the direction of the trend is reversed.
Abstract: Over the years, household surveys have become a popular, valuable data source for empirical research in microeconomics. In developing countries, household survey data have become more available in the past decade, as a result of several international programs. This has spurred interest in the economics of the household in the context of development economics. Many analysts give little attention to the sampling design of the surveys they use, taking the data produced by statisticians and survey practitioners as is. At best, sampling weights are applied to ensure that the results are representative. The authors illustrate the need to pay close attention to the sampling aspects of a household survey used in applied microeconomic analysis - particularly for comparisons over time. This case study shows that observed changes in household welfare and in the incidence of poverty in Cote d'Ivoire between 1985 and 1988 vanish when corrections are applied to the data for changes in sampling procedures; even the direction of the trend is reversed. Similarly, the cross-sectional patterns of welfare and poverty observed in earlier analyses for 1985-86 prove to be incorrect. The Cote d'Ivoire Living Standards Survey, conducted between 1985 and 1988, has provided a popular, fruitful data set for policy analysis. But according to the authors, the recorded decline in mean household size during this period is due to sampling bias in the early years of the survey. If this is true, the robustness of the analyses based on these data is questionable.
TL;DR: Two subsamples of 30 subjects drawn from the same skeletal population were examined in a parallel blind study of dentoalveolar pathology, tooth wear and calculus, pointing to potential effectiveness of sampling bias in archaeological partial samples.
Abstract: Two subsamples of 30 subjects drawn from the same skeletal population (Monte d'Argento, Italy, Medieval age) were examined in a parallel blind study of dentoalveolar pathology, tooth wear and calculus, in order to assess the influence of sampling bias on dietary reconstruction. The effect of interobserver error was controlled and random subdivision was performed on a sample homogeneous for age and social condition. In spite of this, appreciable differences in the frequencies of some pathologies and, therefore, in dietary reconstruction were found, thus pointing to potential effectiveness of sampling bias in archaeological partial samples.