Sampling Survey: Sampling Design, Estimation Methods, and Data Quality
Sample survey basis
On sample surveys
Purpose of the sample survey: The general characteristics of the sample sample sample sample taken from the aggregate, i.e., the acquisition of statistical needs analysis data, were the first stages of the statistical analysis.
The questions in this article can also be addressedDescriptive statistics and visualization: data measurements, distance and related analysisHow the concept of a relatively close read together is developed in different contexts.
In the age of big data, the significance of sample surveys has gradually diminished, and it is more worthwhile to discuss the need to obtain credible conclusions from unstable distributions, although it is also valuable to draw on the technical methods used to obtain better sampling data from big data.
The sample surveys divided the analytical statistics into two categories,Sample and test dataThe former is in the real world, but it needs to be observed to get data. The latter is the data obtained by the test under controlled conditions. The former is required to be obtained using this technology, while the latter is used in the same way as the other one. Pilot design methodology Get...
And we have a comprehensive survey methodology, including a comprehensive census, a statistical statement... that is, a survey of the whole population.
Relatively comprehensive survey, non-exhaustive surveyFocused surveys, typical surveys, sample surveys ; Focusing on research to study units that, although small, account for a large proportion of the research phenomenon. Typical surveys are based on the surveyer’s choice of a number of typical subjects. Sample surveys are the most widely used methods in non-exhaustive surveys, from which the survey is based on the following:Some samples were taken from the entire research audience, in accordance with certain rules/procedures
Sample surveys can reduce costs and time costs. Governments, businesses and a large number of market players need to obtain data through sample surveys and to purchase related services from statistical offices, market survey companies.
Probability and non-probability sampling
Non-probability sampling, also known as non-random sampling, has the core of ** “subjective choice”.** The sample relies on the subjective judgement of the researcher, convenience or voluntary participation of the interviewee. The probability of each unit being drawn in the overall picture is unknown and may even be zero.
Key approaches include:
- I want to take random samples, select the most accessible people as samples
- Purpose/intentional sampling, judging and selecting those units that are “most typical” or “most capable of providing information” based on their own expertise and experience.
- Quota sampling, an attempt to “memolify” a stratification in non-probability samples. The researchers first identified the “quota” (e.g. gender, age) of the sample for certain characteristics, so that it was consistent with the overall ratio.
- Sample of volunteers, based on voluntary selection
- Scramble snowball samples, first a few eligible initial interviewees, then a request for other eligible interviewees
- Advantages:
- Low cost, fast.: Simple to operate, suitable for projects with limited budget and time.
- Wide applicability: For general purposes without sampling frames or for exploratory, qualitative studies.
- Access to specific groups: the only viable way to study “hidden” populations.
- Disadvantages:
- It's impossible to extrapolate the total.This is its fundamental, deadly weakness.
- Highly subjective.: The results are vulnerable to prejudice by researchers or participants.
- Question of representation: Samples are usually subject to systemic deviations.
Probability sample, also known as random sample, with the idea of “equal opportunity” at its coreI'm sorry. Each unit of the total was one unit during the sampling processKnown, non-zero** probability of being drawn.
- The randomity principle: Sample selection is objective and random, excluding subjective interference by researchers.
- Derogability: The results of the sample can be mathematically used to extrapolate the sum as a result of randomness. This is the most valuable feature of a probability sample.
- Calculating Sample Errors: We can calculate the range of errors (between confidence) that may exist between the sample results and the overall real value by statistical formula.
- Scientifically strong.: The results are objective and re-proven, and are the gold standard for quantitative studies and scientific surveys.
The main one is the method below.
- Simple random sampling It's completely randomly extracted from the general population.
- System sampling Calculate a sample interval, choose a random starting point, extract it from the interval
- Scattered Samples First, the grouping of the whole into several non-overlapping “layers” by a particular characteristic, and then the independent and random sampling of each layer (usually simple random or systematic)
- Group sample The group is divided into several “groups” (e.g., classes, communities, streets) and then randomly extracts several “clusters” and conducts a comprehensive survey of all units within the selected groups.
- Multistage sampling mixes the methods in front, and the sampling is phased.
Basic concepts
Before any sample survey can begin, researchers must clearly define several basic elements. These concepts are the cornerstone of the entire survey design, and it is essential to understand the differences and linkages between them. They are: total, sample units, sample frames and samples.
The whole is the ultimate target audience for our research. But in practice, we need to distinguish between two levels of overall: Overall objective It's us.IdeasComplete collection of all units that meet certain conditions that are intended to be studied. It is the “end group” that we want to extend the findings of the study.Overall researchYes.ActualWe can take the whole sample book. It is usually a group represented by a specific “sampling box”.
Sample units are the total, used for samplingBasic modulesI'm sorry. It's the thing we take from the whole thing. All sample units combined must cover the entire study in a complete manner and there can be no overlap (i.e. each element of the total is only a sample unit). The selection of sampling units depends on the purpose of the study, which is not necessarily a single person.
The sample frame isLists, maps or any form of list of all sample units in the studyI'm sorry. It is our “operational manual” or “maps” for actual sampling. The perfect sampling frame should be satisfied: complete, non-duplication, non-duplication, accurate.
The samples were actually extracted from the sampling frame based on sampling methods.The group of sample units.I'm sorry. It is the specific object that we use to carry out data analysis and research.Sample Volume: Number of units included in the sample, usually in letters n - Show.Sample ratio: Sample quantity as a proportion of total research, calculated as f = n / N(where N is the overall scale of the study)
The total and sample in the sample survey are completely different from the mathematical statistical definition, and the overall is unlimited and has a defined distribution for mathematical statistics, but the overall no-distribution assumptions and limited number of samples in the sample survey are not guaranteed by I.I.D.
Sample error versus non-sampling error
In any sample, we extrapolate the sum of the results from the sample, which always differs from the actual situation in the general picture. This general difference is calledGeneral survey errorIt consists of two main categories:Sample errorandNon-sampling error。
The sample error is...The error that came about just because we only looked at part of it, not all of it.I'm sorry. Only in probability samples can we quantify this "lucky" by using random principles. Component. The larger the sample, the more stable the sample is to the general, and the smaller the sampling error.
To measure a good or bad estimate, traditional statistics typically use deviations, differences, and average errors MSE to measure an estimate. Modern statistical learning techniques do not take the latter into account.
Non-sampling error isErrors in all the other components of the survey except the sample itselfI'm sorry. Such errors are more subtle and may exist in both censuses and sample surveys.Source: Investigation of various human or systemic failures in the design, implementation, data processing, etc. It's inevitable and it's hard to measure.
Implementation steps
Step 1: Clear definition of the objectives and overall objectives of the survey
Core issuesThe blogger says: "What are we looking at?" Who do we want to extend the conclusions to?
Step 2: Build or evaluate sampling frames
Core issues: Do we have a map? How good is this map?
Step 3: Designing sampling schemes
Core issues: How do we “draw”? How many?
Step 4: Design and testing questionnaire
Core issues: How do we ask for real and accurate information?
Step 5: Conducting investigations and managing no response
Core issues: How can we ensure that investigations are conducted smoothly and that everyone is involved as far as possible?
Step 6: Data collation, analysis and reporting
Core issues: How can meaningful conclusions be drawn from the data collected?
Simple Random Sampling
SRS
Simple random sampling (Simple Random Sampling, SRS) is the most basic probability sampling method. In this methodology:
- Each sample unitThe same opportunities (whether individual, family, company, etc.) are available. Importability random selection
- Sample selection: randomly sampled fixed sizes from the aggregate, each of which has the same probability of selection as the possible sample combination
- No returned sample: Each of the units drawn will not be selected again. SRS can be achieved by generating random numbers.
Average SRS, variance and standard error
The difference between the sample average and the sample formula is a key tool for estimating the overall average and the overall difference. Sample average: $$\bar{y} = \frac{1}{n} \sum_{i=1}^{n} y_i$$ Sample variance: $$s^2 = \frac{1}{n-1} \sum_{i=1}^{n} (y_i - \bar{y})^2$$ An important feature of simple random sampling is the average sample value and sample size. BadNo bias, that is, their expectations are equal to the overall average and the overall variance, respectively. This means that the averages and squares calculated through sampling do not systematically deviate from the true values of the whole on average.
The standard error is the average deviation of the sample average from the overall average, and the average deviation from the average is measured by the average difference.Sample averageThe error can be passed.Standard ErrorThe Quantification of the Quality of the Quality of the Quality of the Quality of the Quality of the Quality of the Quality of the Quality of the Quality of the Quality of the Quality of the Quality of the Quality of the Quality of the Quality of the Quality of the Quality of the Quality of the Quality of the Quality of the Quality of the Quality of the Quality of the Quality of the Quality of the Quality of the Quality of the Quality of the Quality of theStandard Error: $$\text{S.E.}(\bar{y}) = \frac{s}{\sqrt{n}}$$ of which$s$ It's a standard deviation from the sample.$n$ The smaller the standard error, the closer the average sample is to the overall average, and therefore the larger the sample size reduces the standard error and thus the accuracy of the estimate.
Trust Interval, CI
The range used to estimate the overall parameter (e.g., average) provides a range, indicating the likelihood that the overall parameter will fall into that range ... Standard error (S.E.) and Sample average, we can calculate the confidence interval of the overall average $$置信区间 = \bar{y} \pm Z \times \text{S.E.}$$ Of which:
- $\bar{y}$ is the sample average.
- $Z$ is the fraction of the standard normal distribution, depending on the confidence required degrees
- $\text{S.E.}$ is standard error.
In simple random samples, in addition to estimating the overall average, we can estimate the total amount (e.g., total income, total production, etc.) by way of a sample. For total estimates, use the following formula: $$ \hat{Y} = \frac{N}{n} \sum_{i=1}^{n} y_i $$ I'm not sure.$N$ In general,$n$ It's a sample size.$y_i$ is the value observed in the sample unit.
Standard error in total sample: $$ \text{S.E.}(\hat{Y}) = \frac{N}{n} \times \text{S.E.}(\bar{y}) $$ This formula indicates that the standard error of the total is proportional to the standard error of the sample average.
Estimated ratio
This chapter is structured around the estimation of the proportion in the sample survey and focuses on methods such as single-scale estimation, multi-category analysis, independent sample consolidation and sub-total analysis. Unlike stratification, this chapter highlights the fact that it is not possible to distinguish the whole from the different sub-totals before the sampling. These methods are widely used in the areas of political opinion polls, biostatistics, market research and ecology.
Overall
In the overall scale sample, we are concerned about the proportion of individuals with certain characteristics in the total. Define indicator variables if a simple random sample of n size is taken from the sum of N size: I'm sorry. ♪ I'm not sure I'm gonna be able to do this ♪ The \begin{cases} One. & \text{if i\text{units have target characteristics} \ Photo by Flickr user @un.org & \text{otherwise} The next thing I know, I'm not sure. I'm sorry.
Overall parameters:
- $C = \sum_{i=1}^{N} y_i$ : Total number of units with this characteristic
- $P = \frac{C}{N}$ : Overall
- $Q = 1 - P$ : Percentage not having this characteristic
- $S^2 = \frac{N}{N-1} PQ$ : Total variance
Sample statistics:
- $c = \sum_{i=1}^{n} y_i$ : Units in the sample with this characteristic
- $p = \frac{c}{n} = \bar{y}$ : Sample ratio, yes $P$ No-Effort estimate
Sampling distribution characteristics:
- Number of units in the sample with characteristics $c$ Subscribe to hypergeometric distribution
- When? $N$ When it's big, the hypergeometric distribution is almost two-dimensional.
- $p$ The difference is:$\mathrm{Var}(p) = \left(1 - \frac{n}{N}\right) \frac{PQ}{n}$
Difference estimate:
$$ \hat{v}(p) = \left(1 - \frac{n}{N}\right) \frac{pq}{n-1} $$
Standard error:
$$ \mathrm{SE}(p) = \sqrt{\hat{v}(p)} $$
Application of the overall ratio estimate: Mark the re-capture method, step as follows
- First capture $M$ Animals only. Mark and put back to the original habitat.
- After a while, they catch again. $n$ Animals only
- Calculate the number of animals marked in it $c$
Based on the overall ratio,
$$
p = \frac{c}{n} \text{ 估计了总体中标记动物的比例 } \frac{M}{N}
$$
Equivalent through $\frac{c}{n} = \frac{M}{N}$, obtains a rectangular estimate of the total size:
$$ \hat{N} = \frac{Mn}{c} $$ We'll add amendments to the application. $$ \hat{N} = \frac{(n + 0.5)M}{c + 0.5} $$
Calculations between confidence zones and sample demand
For the overall ratio $P$When the sample is large enough,$p$ The sample distribution is almost subject to the normal distribution. It's built on the very limited logic of the center. $1 - \alpha$ Trust interval: $$ Z = \frac{p - P}{\sqrt{\hat{v}(p)}} \stackrel{d}{\approx} N(0,1) $$ For 95% confidence level: $$ P\left(-1.96 \leq \frac{p - P}{\sqrt{\hat{v}(p)}} \leq 1.96\right) \approx 0.95 $$ The modular is trusted: $$ \left[ p - 1.96 \sqrt{ \left(1 - \frac{n}{N}\right) \frac{pq}{n-1} },; p + 1.96 \sqrt{ \left(1 - \frac{n}{N}\right) \frac{pq}{n-1} } \right] $$
When the total is large ($N \to \infty$In the following: $$ \left[ p - 1.96 \sqrt{ \frac{pq}{n} },; p + 1.96 \sqrt{ \frac{pq}{n} } \right] $$
In designing the survey, an appropriate sample volume needs to be determined to achieve the pre-set estimated accuracy. Sample quantities are usually calculated based on the width of the confidence zone.
For 95% confidence level, require error threshold not to exceed $e$:
$$ 1.96 \times \sqrt{\frac{PQ}{n}} \leq e $$
Because $PQ \leq 0.25$♪ When ♪ $P = 0.5$ The most conservative estimate is:
$$ n \geq \frac{(1.96)^2 \times 0.25}{e^2} $$
For example, the error limit is ±3%:
$$ n \geq \frac{(1.96)^2 \times 0.25}{(0.03)^2} \approx 1067 $$
That's why many polls use about 1100 samples, claiming precision at > 3%.
When the overall level is limited, limited overall correction is required: $$ n = \frac{n_I}{1 + \frac{n_I}{N}} $$ of which $n_I$ The sample size is required on the basis of an unlimited overall assumption.
Why did we introduce the sample count into the scale estimate? He needed a margin. Limit$e$ It is only a more fixed requirement in the scale estimate, and other confidence-between problems can be similarly calculated to calculate sample demand, and Z distribution is very common in sample surveys, so many give estimates averages and variance, and many start to consider the confidence-building and sample volume issues.
Analysis in multi-category situations
Multi-category basic analysis
When the total contains $k$ Each category accounts for a certain proportion of the total, and we need to estimate both those proportions and their interrelationships, when one category is one (e.g., multiple candidates for election). Set$C_1, C_2, \ldots, C_k$ Indicate the number of categories, satisfied $\sum_{j=1}^{k} C_j = N$I'm sorry. The number of categories in the sample is $c_1, c_2, \ldots, c_k$,Fulfilled $\sum_{j=1}^{k} c_j = n$
Sample distribution is subject to multivariant hypergeometric distribution, with a probabilistic mass function:
$$ p(x_1, \ldots, x_k) = \frac{ \binom{C_1}{x_1} \binom{C_2}{x_2} \cdots \binom{C_k}{x_k} }{ \binom{N}{n} } $$
I'm sorry. $j$ Total ratio of category:$P_j = \frac{C_j}{N}$
Sample scale:$p_j = \frac{c_j}{n}$ Sample ratio is the estimate of the overall ratio.
Specimen: $$ \mathrm{Var}(p_j) = \left(1 - \frac{n}{N}\right) \frac{N}{N-1} \frac{P_j(1 - P_j)}{n} $$
When? $N$ In large numbers, the multivariant hypergeometric distribution can be similar to multiple distributions, with the variance being simplified: $$ \mathrm{Var}(p_j) \approx \frac{P_j(1 - P_j)}{n} $$
Difference estimate: $$ \hat{v}(p_j) = \left(1 - \frac{n}{N}\right) \frac{p_j(1 - p_j)}{n-1} $$
Formulas can be used directly when the quantity is estimated instead of the proportion $$ c_i = N p_i $$ His difference is... $$ \mathrm{Var}(c_i) = N^2 \mathrm{Var}(p_i) = N^2 \left(1 - \frac{n}{N}\right) \frac{N}{N-1} \frac{P_i(1 - P_i)}{n}. $$
Multi-category differentials
In multi-category situations, we often need to estimate the difference between the two categories, such as the difference in the support rate for the two candidates. For Category $i$ and $j$ Difference in ratio:
Estimate: $$ p_i - p_j $$
Specimen:
$$ \mathrm{Var}(p_i - p_j) = \frac{N - n}{n(N - 1)} \left[ P_i(1 - P_i) + P_j(1 - P_j) + 2P_i P_j \right] $$
Difference estimate:
$$ \hat{v}(p_i - p_j) = \frac{N - n}{N(n - 1)} \left[ p_i(1 - p_i) + p_j(1 - p_j) + 2p_i p_j \right] $$
Sample volume under multiple categories
We're gonna need to figure out what we're gonna do. $j$ Percentage of categories $P_j$, require confidence level $1 - \alpha$ The estimated error is no more than $e$I'm sorry. That is:
$$ P\left( |p_j - P_j| \leq e \right) = 1 - \alpha $$
When? $N$ The blogger says:$p_j$ Approximate to normal distribution, we can use the following variants (We all use the variance estimates and assume that$N$Largely in calculating confidence interval and sample volumes, when normal distribution assumptions can be used and the denominator is used$n-1$Replace with$n$): $$ z_{\alpha/2} \cdot \sqrt{ \frac{P_j(1 - P_j)}{n} } \leq e $$
of which $z_{\alpha/2}$ It's the top of the standard normal distribution. $\alpha/2$ Bits (e. g. 95% confidence level) $z_{0.025} = 1.96$)
The solution is: $$ n \geq \frac{ z_{\alpha/2}^2 \cdot P_j(1 - P_j) }{ e^2 } $$
When the difference between the two categories is estimated $P_i - P_j$ , the error is not greater than $e$, the sample volume formula is:
$$ z_{\alpha/2} \cdot \sqrt{ \mathrm{Var}(p_i - p_j) } \leq e $$
When? $N$ Very often.$n-1$In exchange for...$n$):
$$ \mathrm{Var}(p_i - p_j) \approx \frac{ P_i(1 - P_i) + P_j(1 - P_j) + 2P_i P_j }{n} = \frac{ P_i + P_j - (P_i - P_j)^2 }{n} $$
Therefore:
$$ n \geq \frac{ z_{\alpha/2}^2 \cdot \left[ P_i(1 - P_i) + P_j(1 - P_j) + 2P_i P_j \right] }{ e^2 } $$
Group of independent samples
In practical studies, there is a constant need to consolidate data from multiple independent surveys to improve the accuracy of estimates or to enhance the reliability of conclusions. Assumptions $k$ An independent survey addresses the same response variables for the same target group.
Set $p_i$ As No. $i$ The percentage of samples surveyed,$n_i$ For its sample volume, the combined estimate is:
$$ \hat{P} = \frac{ \sum_{i=1}^{k} n_i p_i }{ \sum_{i=1}^{k} n_i } = \frac{ \sum_{i=1}^{k} c_i }{ \sum_{i=1}^{k} n_i } $$
This estimate is reasonable because of the sample size of each survey $n_i$ It's not random, it's research design. The combined estimates are essentially weighted averages, with a positive ratio of weight to sample volume.
When the total size is large, the difference in the combined estimate is:
$$ \mathrm{Var}(\hat{P}) = \frac{ \sum_{i=1}^{k} n_i^2 \mathrm{Var}(p_i) }{ n^2 } $$
of which $n = \sum_{i=1}^{k} n_i$
Because $\mathrm{Var}(p_i) \approx \frac{PQ}{n_i}$, received:
$$ \mathrm{Var}(\hat{P}) \approx \frac{ \sum_{i=1}^{k} n_i^2 \cdot \frac{PQ}{n_i} }{ n^2 } = \frac{ PQ \sum_{i=1}^{k} n_i }{ n^2 } = \frac{PQ}{n} $$
This suggests that the difference between the combined sample and the one size is equal to the one size. $n$ the difference of the single sample.
The variance is estimated to be: $$ \hat{v}(\hat{P}) = \frac{ \hat{p} \hat{q} }{ n } $$ of which $\hat{p} = \hat{P} , \hat{q} = 1 - \hat{p}$。
General estimate
On sub-general issues
In sample surveys, we often focus not only on the overall parameters but also on the parameters of the specific part (subtotal) of the overall picture. For example:
- In the population census, we may be concerned about income levels of different age groups
- In agricultural surveys, the production of different cultivation methods may need to be understood
- Health surveys may need to analyse the incidence of diseases by gender or occupational group
The key features of such problems are:The overall composition consists of several non-overlapping parts, of which we are interested in one or more. These are referred to as sub-total or research domains.
The sub-total estimation problem is divided into two main scenarios:
Scattered sampling:
- Each subtotal is known and separated before the sample.
- The sample can be done independently within each subtotal.
- This is part of the stratification sample and will be discussed in a subsequent section.
After-action hierarchy:
- I can't separate the whole sub-total before sampling.
- First, take a simple random sample from the body.
- Samples are then classified according to the characteristics of the sample unit
- This chapter focuses on the size of the sub-group. $N_j$ Unknown, balance the discussion of known
Core differences: In the stratification sample, the subtotal is known and separated in advance; in the case discussed in this chapter, the subtotal is determined after the sampling.
This is common in actual investigations:
- It is not possible to classify the overall unit before the sample (e.g. by income level)
- Lack of subtotal identifier information in sample frames
- Afterward, to analyze the sub-groups that were not foreseen in the survey
- The analytical needs of certain sub-groups were not taken into account in the survey design
Relevant theoretical framework for sub-total
- Yes, I do. $N$ unit, including $J$ A non-overlapping subtotal:
- $N_j$ : No. $j$ Size of the total of individuals (unknown)
- $n$ : Simple random sample volume
- $n_j$ : in sample $j$ Unit numbers for individual aggregates (random variables)
- $y_{ij}$ : No. $j$ The first of the entire population. $i$ Unit response variable values
- $\bar{Y}j = \frac{1}{N_j} \sum{i=1}^{N_j} y_{ij}$ :第 $j$ average for the entire population
- $Y_j = \sum_{i=1}^{N_j} y_{ij}$ : No. $j$ Total number of individual units
Draw size from the total $n$ After a simple random sample:
- Sample is no. $j$ Number of units per unit $n_j$ It's random.
- Significant nature: Prove it, this $n_j$ Unit composition in size $N_j$ Simple random subsamples taken from the total
Estimated subtotal averages and variance
I'm sorry. $j$ The average value of the individual population is estimated at: I'm sorry. ♪ I'm not gonna let you go ♪j = \frac{1}{n_j} \sum{i=1}^{n_j} y_{ij} $$
This is the sub-average. $\bar{Y}_j$ . That is: $$ E(\bar{y}_j) = \bar{Y}_j $$
The difference in the sub-total mean estimate is:
$$ \mathrm{Var}(\bar{y}_j) = \left(1 - \frac{n_j}{N_j}\right) \frac{S_j^2}{n_j} $$
of which $S_j^2 = \frac{1}{N_j - 1} \sum_{i=1}^{N_j} (y_{ij} - \bar{Y}_j)^2$ No. No. $j$ The difference of the individual total.
Because $N_j$ Unknown. We can't use it directly. $n_j / N_j$ As a sample percentage. However, it can be shown that:
$$ E\left( \frac{n_j}{N_j} \right) = \frac{n}{N} $$
It shows that, although... $n_j / N_j$ It's random in itself, but its expectations are equal to the overall sample ratio. $n / N$
Therefore, the margin is estimated to be: $$ \hat{v}(\bar{y}_j) = \left(1 - \frac{n}{N}\right) \frac{s_j^2}{n_j} $$
of which $s_j^2 = \frac{1}{n_j - 1} \sum_{i=1}^{n_j} (y_{ij} - \bar{y}_j)^2$ No. No. $j$ The difference in the total sample.
Attention.: Here the overall sample ratio is used $n/N$ Not the subtotal sample ratio. $n_j / N_j$, because $N_j$ Unknown, if he knows, should be replaced accordingly.
If you want to build the confidence compartment, you can still use the normal distribution method. We have now given estimates of the average and the difference, and you can bring it in.
Estimates of totals
Total Subsize $N_j$ In unknown circumstances, it cannot be simply used $N_j \times \bar{y}_j$ As total estimate, as $N_j$ Unknown.
Introduction of supporting variables to address this problem $U_i$:
$$ U_i = \begin{cases} y_{ij} & \text{if i\text{units belong to the }j\text{total} \ Photo by Flickr user @un.org & \text{otherwise} The next thing I know, I'm not sure. I'm sorry.
There are: $$ \bar{U} = \frac{1}{N} \sum_{i=1}^{N} U_i = \frac{1}{N} \sum_{i=1}^{N_j} y_{ij} = P_j \bar{Y}_j $$ of which $P_j = N_j / N$ No. No. $j$ The overall proportion of each individual in the aggregate.
Based on the supporting variables, the total subtotal is estimated to be: I'm sorry. What's wrong with you?j = N \bar{u} = \frac{N}{n} \sum{i=1}^{n_j} y_{ij} $$ 这是一个无偏估计量,即: $$ E(\hat{Y}_j) = Y_j $$
Supporting variables $U$ The overall variance is:
$$ S_u^2 = \frac{1}{N-1} \left[ \sum_{i=1}^{N} U_i^2 - N \bar{U}^2 \right] = \frac{N_j - 1}{N - 1} S_j^2 + \frac{N}{N - 1} P_j Q_j \bar{Y}_j^2 $$
of which $Q_j = 1 - P_j$。
The difference in the total estimated volume is therefore:
$$ \mathrm{Var}(\hat{Y}_j) = N^2 \mathrm{Var}(\bar{u}) = \frac{N^2(1 - f)}{n} S_u^2 \ = \frac{N^2(1 - f)}{n} \left[ \frac{N_j - 1}{N - 1} S_j^2 + \frac{N}{N - 1} P_j Q_j \bar{Y}_j^2 \right] $$
of which $f = n/N$ For sample ratio.
Sample variance $s_u^2$ This can be expressed as:
$$ s_u^2 = \frac{1}{n-1} \sum_{i=1}^{n} (u_i - \bar{u})^2 = \frac{1}{n-1} \left[ (n_j - 1) s_j^2 + n p_j q_j \bar{y}_j^2 \right] $$
of which $p_j = n_j / n , q_j = 1 - p_j$
Therefore, the variance in the total estimated volume is estimated to be: I'm sorry. ♪ The world is so full of shit ♪j) = \frac{N^2(1 - f)}{n(n-1)} \left[ \sum{i=1}^{n_j} (y_{ij} - \bar{y}_j)^2 + n p_j q_j \bar{y}_j^2 \right] $$ Summing the sum of the subtotals is the sum of the sum of the sum, and the sum of the sum of the sum of the sum of the sum of the sum of the sum of the sum of the sum of the sum of the sum of the sum of the sum of the sum of the sum of the sum of the sum of the sum of the sum of the sum of the sum of the objects of the individual subs, which, without much discussion, can be extended to the sum of the sum of the sum of the objects, and its form before simplification can be used in the case of the known size of the subtotal of the sum of the sum, which is very simple, and we do not need to give any additional formula.
General average value difference
Set $y$ The blogger says that the government is not a party to the law.$\bar{Y}_i$ and $\bar{Y}_j$ Subtotals, respectively $i$ and $j$ average.
- Estimated:$\bar{y}_i - \bar{y}_j$
- Difference:$\text{Var}(\bar{y}_i - \bar{y}_j) = \left(1 - \frac{n_i}{N_i}\right)\frac{s_i^2}{n_i} + \left(1 - \frac{n_j}{N_j}\right)\frac{s_j^2}{n_j}$
- Difference estimate:$v(\bar{y}_i - \bar{y}_j) = \left(1 - \frac{n_i}{N_i}\right)\frac{s_i^2}{n_i} + \left(1 - \frac{n_j}{N_j}\right)\frac{s_j^2}{n_j}$
If confidence is needed to continue using normality, the sample size is considered to be the same.
Calculation of sample quantities
The total mean of the estimated sub-value for the volume of samples is calculated, and the estimated number of samples is calculated as follows: $i$ Average value of individual total $\bar{Y}_i$, is given a margin error $d_i$ And confidence level $1 - \alpha$ in the case of the $n_i$ , and the formula is:
$$ n_i = \frac{N_i \cdot z_{\alpha/2}^2 \cdot S_i^2}{(N_i - 1)d_i^2 + z_{\alpha/2}^2 \cdot S_i^2} $$
Of which:
- $N_i$ No. No. $i$ Size of the total individual (usually unknown)
- $S_i^2$ No. No. $i$ Differences in the total number of individuals (usually unknown, pre-estimated)
- $z_{\alpha/2}$ It's the top of the standard normal distribution. $\alpha/2$ Bits
- $d_i$ is the maximum error allowed
If you want to study the ratio (specify of the second classification), you just need to replace the difference. $$ n_i = \frac{N_i \cdot z_{\alpha/2}^2 \cdot p_i (1 - p_i)}{(N_i - 1) d_i^2 + z_{\alpha/2}^2 \cdot p_i (1 - p_i)} $$
For the case of polykot in general, there are errors.Consider prioritizing the most concerned sub-totals, useMinimum Max Error Method, select the total sample volume below $n$, minimizes the maximum relative error in all subtotal estimates:
$$ \min_{n} \max_{i} \left( \frac{z_{\alpha/2} \cdot S_i}{\sqrt{n_i}} \sqrt{1 - \frac{n_i}{N_i}} \right) $$
of which $n_i$ No. No. $i$ Total individual sample volume, compared to total sample volume $n$ ♪ And the relationship is ♪ $n_i = n \cdot W_i$,$W_i = N_i / N$ No. No. $i$ Total per unit.
Or use the following:Practical approaches
- Estimated overall ratio per sub-unit $W_i$
- Identification of the required sample quantity for each subtotal $n_i$
- Calculated total sample volume:$n = \sum_{i=1}^{k} \frac{n_i}{W_i}$
Example:: Assuming that the total consists of two sub-totals,$W_1 = 0.7$,$W_2 = 0.3$I'm sorry. To achieve the same precision, each of them is required $n_1 = 100$,$n_2 = 150$I'm sorry. The total sample volume is:
$$ n = \frac{100}{0.7} + \frac{150}{0.3} \approx 143 + 500 = 643 $$
Scattered sampling methods and their application
Basic concepts related to stratification sampling
In practical issues, the general body can generally be divided into several non-overlapping sub-totals by characteristics (e.g., different regions, different categories, different time periods, etc.) and these are referred to as the general body of the individual."Layer"(strata). When there is a clear heterogeneity in the overall picture, simple random sampling may result in a lack of representativeness of the sample, thereby reducing the accuracy of the estimate. To address this problem, statisticians have proposed a stratification sampling method.
The basic idea of the stratification sample is to first divide the whole into G layers that do not overlap and cover the whole population (strata), then to conduct the separate (usually simple random) sampling within each layer, and finally to merge the results of the stratification into a layer weight that gives estimates of the overall parameters.
Stratification samples significantly reduce the variance of the estimate when the layer is of high homogeneity and the layer is of high heterogeneity. And it can provide data directly at the research level and, under existing formations, better survey implementation.
Gives the symbol of some sort of stratification sample
- $N$ : Total size
- $G$ : Layers
- $N_g$ : Section$g$The size of the layer, $g = 1, 2, \dots, G$
- $N = \sum_{g=1}^{G} N_g$
- $y_{gi}$ : Section$g$Layer 1$i$Unit response value
- $W_g = N_g / N$ : Section$g$Weight of Layers
- $\bar{Y}g = \frac{1}{N_g} \sum{i=1}^{N_g} y_{gi}$ : 第$G$ level average overall
- $Y_g = N_g \bar{Y}_g$ : Section$g$Total total of layers
- $S_g^2 = \frac{1}{N_g - 1} \sum_{i=1}^{N_g} (y_{gi} - \bar{Y}_g)^2$ : Section$g$Overall variance of layers
- $\bar{Y} = \frac{1}{N} \sum_{g=1}^{G} \sum_{i=1}^{N_g} y_{gi} = \sum_{g=1}^{G} W_g \bar{Y}_g$ : Overall average
- $S^2 = \frac{1}{N - 1} \sum_{g=1}^{G} \sum_{i=1}^{N_g} (y_{gi} - \bar{Y})^2$ : General variance
The overall variance can be broken down into: $$ S^2 = \frac{1}{N - 1} \left[ \sum_{g=1}^{G} (N_g - 1) S_g^2 + \sum_{g=1}^{G} N_g (\bar{Y}_g - \bar{Y})^2 \right] $$
In the upper form, the first is the sum of the squares (SSerror) and the second is the sum of the squares (SStrt). This reflects the rationale for the variance analysis: $$ \text{Var}(Y) = E[\text{Var}(Y|X)] + \text{Var}[E(Y|X)] $$ of which$X$is a layered variable.
We do simple random stratifications, and we do simple random stratifications in each stratification, so we give the mathematical symbols that follow.
- $n_g$ : Section$g$Sample volume of layers
- $\bar{y}g = \frac{1}{n_g} \sum{i \in s_g} y_{gi}$ : 第$g$-level sample average
- $s_g^2 = \frac{1}{n_g - 1} \sum_{i \in s_g} (y_{gi} - \bar{y}_g)^2$ : Section$g$Sample differences for layers
- $\text{Var}(\bar{y}_g) = \left(1 - \frac{n_g}{N_g}\right) \frac{S_g^2}{n_g}$ : Section$g$Difference of the mean of the layer sample
- $v(\bar{y}_g) = \left(1 - \frac{n_g}{N_g}\right) \frac{s_g^2}{n_g}$ : Uneven estimate of the above differences
And here we're just thinking about each floor as a SRSOR, and it's easy to continue the previous section and draw conclusions directly.
Parameter estimation of a layer sample versus confidence interval
Parameter estimation of stratification samples
Mean values for the layered sample (also estimated by Horvitz-Thompson): I'm sorry. ♪ I'm not gonna let you go ♪{st} = \sum{g=1}^{G} W_g \bar{y}g I'm sorry. The estimate is impartial, i.e. $E (\y}{st}) = \bar{Y}$。
Specimen: I'm sorry. \text{Var} (\bar{y}{st}) = \sum{g=1}^{G} W_g^2 \text{Var}(\bar{y}g) = \sum{g=1}^{G} W_g^2 \left(1 - \frac{n_g}{N_g}\right) \frac{S_g^2}{n_g} $$
The difference is estimated in no-sided terms: I'm sorry. v (\bar{y}{st}) = \sum{g=1}^{G} W_g^2 \left(1 - \frac{n_g}{N_g}\right) \frac{s_g^2}{n_g} $$
Estimates of total: I'm sorry. What's wrong with you?{st} = N \bar{y}{st} = \sum_{g=1}^{G} N_g \bar{y}g $$ 方差: $$ \text{Var}(\hat{Y}{st}) = N^2 \text{Var}(\bar{y}{st}) = \sum{g=1}^{G} N_g^2 \left(1 - \frac{n_g}{N_g}\right) \frac{S_g^2}{n_g} $$
The difference is estimated in no-sided terms: I'm sorry. v (hat{Y}{st}) = N^2 v(\bar{y}{st}) = \sum_{g=1}^{G} N_g^2 \left(1 - \frac{n_g}{N_g}\right) \frac{s_g^2}{n_g} $$
Proportional issue of stratification samples
In the survey, when the variable takes only 0 and 1 values (e.g., if it owns a product), we usually focus on the ratio parameter. Will respond in general terms $y_{gi}$ Replace with the binary variable, and get:
- $\bar{Y}_g \to P_g$ , $S_g^2 \to \frac{N_g}{N_g - 1} P_g Q_g$ , of which $Q_g = 1 - P_g$
- $\bar{y}_g \to p_g$ , $s_g^2 \to \frac{n_g}{n_g - 1} p_g q_g$ , of which $q_g = 1 - p_g$
- $\text{Var}(p_g) = \left(1 - \frac{n_g}{N_g}\right) \frac{N_g}{N_g - 1} \frac{P_g Q_g}{n_g}$
- $v(p_g) = \left(1 - \frac{n_g}{N_g}\right) \frac{p_g q_g}{n_g - 1}$
Overall rate: $$ P = \sum_{g=1}^{G} W_g P_g $$
Sample estimate: $$ p_{st} = \hat{P}{st} = \sum{g=1}^{G} W_g p_g $$
Specimen: $$ \text{Var}(p_{st}) = \sum_{g=1}^{G} W_g^2 \text{Var}(p_g) $$
The difference is estimated in no-sided terms: $$ v(p_{st}) = \sum_{g=1}^{G} W_g^2 v(p_g) $$
Trust interval of a stratification sample
All confidence-building issues are ultimately resolved by normal distribution, which is very easy because we have calculated the average of the estimates and the difference between the square and the square.
For the overall average $\bar{Y}$ It's... $100(1 - \alpha)%$ Trust interval: I'm sorry. ♪ I'm not gonna let you go ♪{st} \pm z{\alpha/2} \sqrt{v(\bar{y}_{st})} $$
For total $Y$ It's... $100(1 - \alpha)%$ Trust interval: I'm sorry. What's wrong with you?{st} \pm z{\alpha/2} \sqrt{v(\hat{Y}_{st})} $$
Distribution of sample quantities
Total sample quantity given $n = \sum_{g=1}^{G} n_g$The distribution of sample volumes between layers is a key issue. Common methods of allocation include:
Proportional distribution
Proportional distribution: $n_g \propto N_g$that is $$ n_g = W_g n = \frac{N_g}{N} n $$
The weights are equal in the ratio distribution, with the difference in the mean values of the stratification samples being: I'm sorry. \text{Var} (\bar{y}{st}) = \left(1 - \frac{n}{N}\right) \frac{1}{n} \sum{g=1}^{G} W_g S_g^2 $$
The difference between a stratification sample and a simple random sample is $\text{Var}(\bar{y}) = \left(1 - \frac{n}{N}\right) \frac{S^2}{n}$, so when I'm sorry. ♪ I'm gonna be a little bit more like a little more than a little bit of a ♪ < S2 I'm sorry. When the stratification sample is more accurate than a simple random sample, this variation is easily proven by the use of the squared-difference angle, and they are close to parity only when the mean inter-stage difference is small.
Optimal Distribution / Neyman Distribution
When considering sampling costs, the best-placed target is: at total cost $C = c_0 + \sum_{g=1}^{G} c_g n_g$ (of which) $c_0$ It's a fixed cost.$c_g$ No. No. $g$ ) Under the pressure of the sample cost of the layer) $\text{Var}(\bar{y}_{st})$。
Neyman's solution is: $$ n_g \propto \frac{N_g S_g}{\sqrt{c_g}} $$ Specifically: $$ n_g = n \cdot \frac{N_g S_g / \sqrt{c_g}}{\sum_{j=1}^{G} N_j S_j / \sqrt{c_j}} $$ ♪ When all $c_g$ In the case of equivalence, simplify the text to read: $$ n_g = n \cdot \frac{W_g S_g}{\sum_{j=1}^{G} W_j S_j} $$ At this time, the difference in the average of the stratification samples is: $$ V_{\text{Neyman}} = \frac{1}{n} \left( \sum_{g=1}^{G} W_g S_g \right)^2 - \frac{1}{N} \sum_{g=1}^{G} W_g S_g^2 $$ Compared to the percentage distribution: $$ V_{\text{proportion}} - V_{\text{Neyman}} = \frac{1}{n} \sum_{g=1}^{G} W_g \left( S_g - \sum_{j=1}^{G} W_j S_j \right)^2 \geq 0 $$
This suggests that Neyman distribution is always better than or equal to proportional distribution. When? $S_g$ The greater the difference, the more obvious the advantage is, because the top is about $S_g$ Weights in the market $W_g$ The difference below.
Sample volume determined
Assumptions $N_g$ and $S_g^2$ Known, use ratio distribution, request $\text{Var} (\bar{y}{st}) \leq V$,其中 $V.A. is the given value. Unlocking: I'm sorry. \text{Var} (\bar{y}{st}) = (n^{-1} - N^{-1}) \sum_{g=1}^{G} W_g S_g^2 \leq V $$ 得: $$ ♪ The world is so full of shit ♪ I'm sorry. (Assumptions) $N = \infty$ (samples at the time)
For a limited total, the sample size required is: $$ n = \frac{n_l}{1 + n_l / N} $$
Use Neyman distribution, request $\text{Var} (\bar{y}\st}\leq V$. Square formula: I'm sorry. \text{Var} (\bar{y}{st}) = n^{-1} \left( \sum_{g=1}^{G} W_g S_g \right)^2 - N^{-1} \sum_{g=1}^{G} W_g S_g^2 \leq V $$
When? $N = \infty$ Other Organiser $$ n_l = \frac{\left( \sum_{g=1}^{G} W_g S_g \right)^2}{V} $$
For a limited total: $$ n = \frac{\left( \sum_{g=1}^{G} W_g S_g \right)^2}{V + N^{-1} \sum_{g=1}^{G} W_g S_g^2} $$
Ratio estimate versus regression estimate
In sample surveys, we usually focus on parameters such as the overall average or the ratio. However, in practical application, there are many other important, limited overall parameters, such as median income, percentage of the population living below the poverty line, etc. This chapter will focus on two methods of estimating the ratio of the overall average:
- Rate estimate (Ratio Estimation): Ratio relationship between the use of supporting and target variables
- Return Estimates: Linear relationships between supporting and target variables
These methods are particularly useful in the following cases, where we give examples:
- Need to estimate the ratio of the two overall averages
- Total total needs to be estimated, but total size unknown
- Need to use supporting information to improve the accuracy of the estimates
- The estimates need to be adjusted to reflect demographic characteristics
- Process non-response deviations
Study ratio parameters
Estimated percentage
Consider a limited total, with two variables per sample unit responding $x$ and $y$:
- $X$、$Y$: each indicates the total $x$ and $y$ Average
- Ratio parameters:$R = \frac{Y}{X}$
In simple random sample unreleased (SRSOR), sample size is $n$I'm sorry. Yeah. $R$ Estimates are: $$ \hat{R} = \frac{\bar{y}}{\bar{x}} $$
Difference in the estimated ratio
Ratio estimate $\hat{R}$ The difference is almost as follows: $$ V(\hat{R}) = \left(1 - \frac{n}{N}\right) \frac{S_d^2}{n\bar{X}^2} $$ of which $S_d^2 = S_y^2 + R^2 S_x^2 - 2R S_{xy}$, by definition $d = y - Rx$,$\bar{d} = \bar{y} - R\bar{x}$ Producate.
Noting: $$ \hat{R} - R = \frac{\bar{y}}{\bar{x}} - R = \frac{\bar{d}}{\bar{x}} \approx \frac{\bar{d}}{X} $$
Therefore: $$ V(\hat{R}) \approx \left(1 - \frac{n}{N}\right) \frac{S_d^2}{nX^2} $$
Taking into account relevant factors $\rho = \frac{S_{xy}}{S_x S_y}$, the difference can also be: $$ V(\hat{R}) \approx \left(1 - \frac{n}{N}\right) \frac{S_y^2 + R^2 S_x^2 - 2R\rho S_x S_y}{nX^2} $$
Ratio deviation
Ratio estimate $\hat{R} = \frac{\bar{y}}{\bar{x}}$ The deviation is: $$ B(\hat{R}) = E(\hat{R}) - R = E\left( \frac{\bar{y} - R\bar{x}}{\bar{x}} \right) = E\left[ \frac{\bar{y} - R\bar{x}}{X(1+\epsilon)} \right] $$ of which $\epsilon = \frac{\bar{x} - X}{X}$I'm sorry. Using Taylor to expand, the deviation became:
$$ B(\hat{R}) = \frac{\text{Cov}(\bar{x}, \bar{y}) - R \times V(\bar{x})}{X^2} = (1 - f) R \frac{(C_{xx} - C_{xy})}{n} $$ of which $C_{xx} = \frac{S_x^2}{X^2}$,$C_{xy} = \frac{S_{xy}}{XY}$,$f = \frac{n}{N}$I'm sorry. When a sample $n$ When it's bigger, it's negligible.
Average error in ratio (MSE)
When? $n$ When larger, the deviation is negligible, MSE approximates the difference: $$ \text{MSE}(\hat{R}) = E(\hat{R} - R)^2 = E\left( \frac{\bar{y} - R\bar{x}}{\bar{x}} \right)^2 \approx V(\hat{R}) = E\left( \frac{\bar{y} - R\bar{x}}{X} \right)^2 $$
Definitions $d_i = y_i - Rx_i$,$\bar{D} = 0$, the overall variance is: $$ S_d^2 = \sum_{i=1}^{N} \frac{(y_i - Rx_i)^2}{N-1} = \sum_{i=1}^{N} \frac{[(y_i - Y) - R(x_i - X)]^2}{N-1} = S_y^2 + R^2 S_x^2 - 2R S_{xy} = S_y^2 + R^2 S_x^2 - 2R\rho S_x S_y $$
The natural estimate of the variance is: $$ v(\hat{R}) = (1 - f) \frac{s_d^2}{n\bar{x}^2} $$ of which $d_i = y_i - \hat{R}x_i$,$s_d^2$ These. $d_i$ the sample difference. We can give you the details.$R$ The confidence interval is: $$ \hat{R} \pm 1.96 \sqrt{v(\hat{R})} $$
Application of ratio estimates in average and sum of estimates
In some cases,$x$ and $y$ There is a clear positive correlation, such as land area and production. In addition, sometimes in the aggregate $x$ is known in the average or sum. When? $X$ When known, we can adjust the estimates by: $Y$: $$ \hat{Y}_R = \left( \frac{X}{\bar{x}} \right) \bar{y} $$ When? $x$ and $y$ This estimate may be more effective when it is close to a ratio, which we call the ratio estimate.$\hat{Y}_R$ It's not impartial. It's almost equal to: $$ \text{Var}(\hat{Y}_R) = X^2 \text{Var}(\bar{y}/\bar{x}) = X^2 \times (1 - f) \frac{S_y^2 + R^2 S_x^2 - 2R\rho S_x S_y}{nX^2} = (1 - f) \frac{S_y^2 + R^2 S_x^2 - 2R\rho S_x S_y}{n} $$
The ratio estimate is smaller than the sample average only if: I'm sorry. R^2 S x^2 - 2R\rho S x S y < 0 $$ $$ RS_x < 2\rho S_y $$
Equivalent: I'm sorry. \cdot S x < 2\rho S_y $$ $$ \frac{S_x}{X} < 2\rho \frac{S_y}{Y} $$
Therefore, for the ratio estimates, it is more effective: I'm sorry. \rho > \frac{S_x / X}{2 S_y / Y} = \frac{CV(X)}{2 CV(Y)} $$ $CV(\bar{x})$ and $CV(\bar{y})$ Absolute values do not usually vary significantly. So when $rho > At $1/2 a ratio is estimated to be more effective than SRS.
In short: when $x$ and $y$ The ratior is more suitable for estimation than the sample average when there is a strong positive correlation $y$ Overall average.
$\hat{Y}_R$ The variance is estimated to be: I'm sorry. v (hat{Y}R) = (1 - f) \frac{s_y^2 + \hat{R}^2 s_x^2 - 2\hat{R}s{xy}}{n} $$
Why use ratio estimates
When the parameter is a ratio per se, for example, the average juice content per apple.
As a secondary variable $x$ & Study Variables $y$ Highly relevant and overall $x$ The ratio estimates provide more accurate estimates than simple sample averages when the sum or average of the totals is known.
Total estimated, but unknown
For example: There's a bunch of apples, and we'd like to estimate the total number of apple juices. Set:
- $y_1, \dots, y_n$: the amount of juice per apple in the sample
- $x_i$The weight of each apple in the sample,$\bar{x}$ It's the average mass of the sample.
- $N$ It's hard to count, so... $N\bar{y}$ Hard to get directly
- But the total weight of the whole batch of apples. $X$ Easy to access.
The number of apples can be estimated. $X / \bar{x}$The total volume of apple juice can be estimated at: $$ \hat{Y}_r = \frac{\bar{y}}{\bar{x}} X $$
Adjustment of estimates to reflect demographic aggregates
Example: There are 400 students in a university, taking a sample of 400 SRS:
- 240 women and 160 men in samples
- 84 women and 40 men are planned to work in teaching
Objectives: The total number of students who are planned to become teachers is estimated.
Estimator 1: Use SRS information only $$ N\bar{y} = 4000 \times \frac{124}{400} = 1240 $$
Estimator 2: Integration of demographic information (schools with 2700 women and 1,300 men) $$ \frac{84}{240} \times 2700 + \frac{40}{160} \times 1300 = 1270 $$
Key points:
- Application of estimated rates within each gender group
- Sixty per cent of the samples were female, but 67.5 per cent of the total were female. Estimators adjusted to better reflect demographic ratios
Process non-response deviations
Example: Enterprise sample
- $y_i$: enterprises $i$ Expenditure on health insurance
- $x_i$: enterprises $i$ Number of employees known
- Objectives: Estimated expenditure for general insurance
Estimator 1: $N\bar{y}$
- Companies with fewer employees are unlikely to respond to the survey
- $y_i$ and $x_i$ Proportional
- Estimator 1 overestimates total insurance expenditure. $t_y$
Estimator 2: $X \frac{\bar{y}}{\bar{x}}$
- Because of the greater likelihood of a response from a company with a large number of employees, $X/\bar{x} < N$
- Therefore, the ratio of total health insurance expenditure may be expected to compensate for the lack of impact of companies with small staff Reactions
Regressive estimate in simple random sample
For the return estimate
In survey samples, regression estimates are a statistical technique that uses information from supporting variables to improve the accuracy of estimates. When we study a variable,$y$When (e.g. income, production, etc.) associated supporting variables are often available$x$Information (e.g. population, area, etc.). When x has linear relationships with y and the overall parameters of x (e.g., aggregate average or sum) are known, we can use this relationship to improve the estimates of y.
The core of the regression estimates is the use of the following information:
- Relationship of supporting and target variables: Assumptions $y$ and $x$ Linear relationship exists $E(y) = \beta_0 + \beta_1 x$
- General information on the supporting variable: Known $x$ Overall average $\bar{X}$ Total or total $t_x$
- Sample data: From the sample. $(x_i, y_i)$ Match Data
The key insight of this estimation is that if the auxiliary variable is $x$ & Study Variables $y$ Highly relevant, then. $x$ The overall information we know can help us estimate it more precisely. $y$ General parameters. For example, if we know the total population of a region ($x$and observed population and household income ($y$There is a strong correlation, and then we can use it to obtain more accurate estimates of household income than the simple sample average.
Comparison with simple and ratio estimates
- Simple random sampling estimates: Use only $y$ sample information, ignore any supporting variables
- Ratio estimate: Assumptions $y$ and $x$ Proportional relationship between the point of origin $y = R x$
- Return estimate: Allow $y$ and $x$ General linear relationships between $y = \beta_0 + \beta_1 x$More flexible
The advantage of returning to the estimate is when the relationship does not pass its roots. $x=0$ Time $y \neq 0$It is more relevant than the estimated ratio. For example, even photographs count when estimating the number of dead trees (in the case of trees)$x$) Zero, field count ($y$The return estimate is more appropriate than the rate estimate at this time.
Returning to estimated process
We're re-presenting the process that's needed here. Because there are only two variables, the equation is simple.
Defines the overall regression parameter:
- $\beta_1 = B_1 = \frac{\sum_{i=1}^{N} (x_i - \bar{X})(y_i - \bar{Y})}{\sum_{i=1}^{N} (x_i - \bar{X})^2}$ (General regression rate)
- $\beta_0 = B_0 = \bar{Y} - B_1 \bar{X}$ (Current regression transect) I'm not sure.$\bar{X}$ and $\bar{Y}$ Auxiliary variables $x$ And study variables $y$ Overall average.
From the sample, we can estimate:
- $\hat{\beta}1 = \hat{B}1 = \frac{\sum{i \in S} (x_i - \bar{x})(y_i - \bar{y})}{\sum{i \in S} (x_i - \bar{x})^2}$
- $\hat{\beta}_0 = \hat{B}_0 = \bar{y} - \hat{B}_1 \bar{x}$
$y$ Overall average $\bar{Y}$ The estimated return is: $$ \tilde{y}_{reg} = \hat{B}_0 + \hat{B}_1 \bar{X} = \bar{y} + \hat{B}_1 (\bar{X} - \bar{x}) $$ This estimate can be understood as using the average sample first. $\bar{y}$ estimate $\bar{Y}$, and then by the sample $x$ With known total $x$ Variance $(\bar{X} - \bar{x})$, combined with estimated slope $\hat{B}_1$ Adjustments.
Nature of the estimate
Offset $$ \text{bias}(\tilde{y}_{reg}) = \text{cov}(\hat{B}_1, \bar{x}) $$ The regression estimate is usually biased, but when the sample is larger, the deviation is negligible. If the regression line goes through all the overalls, Points $(x_i, y_i)$, and $\hat{B}_1 = B_1$- No, it's zero.
Average error (MSE) $$ \text{MSE}(\tilde{y}_{reg}) = \left(1 - \frac{n}{N}\right) \frac{S_d^2}{n} $$ of which $d_i = y_i - [\bar{Y} + B_1(x_i - \bar{X})]$,$S_d^2$ These. $d_i$ Overall variance.
Utilization of relevant factors $\rho$MSE can be rewritten as: $$ \text{MSE}(\tilde{y}_{reg}) = \left(1 - \frac{n}{N}\right) \frac{1}{n} S_y^2 (1 - \rho^2) $$ This indicates that:
- When sample amount $n$ Increase or sample ratio $n/N$ MSE decreases when increased
- When? $x$ and $y$ Correlation factors between $\rho$ Close $\pm 1$ MSE is significantly reduced at the time
Total total return estimate $$ \hat{t}{yreg} = \sum{i \in S} y_i + \sum_{i \in S^c} (\hat{B}0 + \hat{B}1 x_i) = \sum{i \in S} y_i + (N - n)\hat{B}0 + \hat{B}1 (t_x - \sum{i \in S} x_i) $$ 当 $n \ll N$ 时,可近似为: $$ \hat{t}{yreg} \approx N(\hat{B}0 + \hat{B}1 \bar{X}) = N \tilde{y}{reg} $$ Standard Error $$ \text{SE}(\tilde{y}{reg}) = \sqrt{\left(1 - \frac{n}{N}\right) \frac{s_d^2}{n}} $$ $$ \text{SE}(\hat{t}=sqrt/2003/left1 - \right)\frac{s d^2} I'm sorry. of which $s_d^2$ It's a cripple. $d_i = y_i - \hat{B}_0 - \hat{B}_1 x_i$ the sample difference.
95% confidence interval: $$ \tilde{y}{reg} \pm t{n-2}(0.025) \sqrt{\left(1 - \frac{n}{N}\right) \frac{s_d^2}{n}} $$ $$ \hat{t}{yreg} \pm t{n-2}(0.025) N \sqrt{\left(1 - \frac{n}{N}\right) \frac{s_d^2}{n}} $$
Summary of estimates of the return to the ratio
Comparison of the three estimation methods
| Estimated methodology | Mean $\bar{Y}$ Estimated amount | Total $T_Y$ Estimated amount |
|---|---|---|
| SRS | $\bar{y}$ | $N\bar{y}$ |
| Ratio | $\hat{B} \bar{X}$ | $\hat{B} t_x$ |
| Return | $\hat{B}_0 + \hat{B}_1 \bar{X}$ | $N(\hat{B}_0 + \hat{B}_1 \bar{X})$ |
Select the situation for the return estimate:
- When? $x$ and $y$ Linear relationships exist, but not necessarily through the original.
- As a secondary variable $x$ and $y$ Heightly Associated ($|rho|) > 0.5$)
- When known $x$ Total information (average or total)
- When there is a systemic deviation, correction is required (e.g., photo counting deviation in dead tree count cases)
Selection of ratio estimates:
- When? $x$ and $y$ Proportional relationship between the point of origin
- When Total Size $N$ Unknown, but the supporting variable is known
- It's particularly useful in the whole group sample.
Select the SRS estimate:
- When no supporting information available
- When the auxiliary variable is relevant to the study variable Weak
- When analysis requires simplicity and ease of interpretation
- Title: Sampling Survey: Sampling Design, Estimation Methods, and Data Quality
- Author: Hyacehila
- Created at : 2025-11-12 08:04:37
- Link: https://hyacehila.github.io//blog/2025/11/12/sampling-survey-notes/
- License: This work is licensed under CC BY-NC-SA 4.0.