Mathematical Statistics: Populations, Samples, and Sampling Distributions
Basic concepts
Statistics are the reverse of probabilism, and in probabilism we have the essence of what happens, to study the results of what happens, and in statistics we reverse the principles by looking at the real world. Mathematical statistics are the foundational knowledge of the entire statistical system, and a lot of the rest goes from here. Mathematical statistics: study of the collection of data with random errors using probabilistic and mathematical methods and analysis of the collected data (statistical analysis) under a set model (statistical analysis) to infer the problems studied (statistical extrapolation) And we'll start with the most basic concept of statistics, and we'll build the foundation of classical statistics, which is the frequency school; the corresponding Bayes school will be presented separately later.
Before embarking on a specific narrative of statistical content, we need to distinguish between another division of statistical science, which is dedicated to the study of observations per se, and statistical inferences, which focuses on the analysis of subjects per se from the point of view.
Overall and sample
Total and individual
The body of the subject is referred to as the sum; the elements that make up the whole are referred to as the individual Indicators that are often of limited interest in statistical research, so the total number of indicators to be studied is at this point in time, and the corresponding number of individuals still present. It's obvious that we can see these numbers as random variables. At this point, the total is a random variable, and his distribution and numerical characteristics are called the total distribution of digital characteristics. The central point is that it's a probability distribution. Because the parameters are unknown, we can also say that overall is a probability distribution group. Sometimes the aggregate cannot be expressed in a parameter distribution, which we call non-parameter aggregate, and at this point applies to the treatment of non-parametric statistics.
Samples and samples
In order to extrapolate the overall distribution and characteristics, a number of individuals are drawn from the general population for observation tests to obtain information about the overall population, which is referred to as “sampling”; The number of individuals included in the sample is called sample capacity. Because the sample is random, each individual is a random variable with a capacity of$n$The sample can be considered as$n$V-random variable Once we get a sample, we get it.$n$A specific number, referred to as an observation value for a sample, a short sample value Those samples that give exact values are called complete samples. Only given the range of sample observations called group samples
Simple random sample
Since the purpose of sampling is to provide statistical inferences to the population as a whole and in order for the sample taken to reflect the overall information well, consideration must be given to the most common sampling method.Simple random sampling He's satisfied.
- All samples taken and total equal distribution
- The samples are independent of each other. A simple random sample is called a simple random sample. In later studies, all samples are considered simple random.
Finally. In fact, the information we obtained from the samples was specific and determined; they were sample values rather than samples; The job of statistics is to look at the general characteristics from the available data, and the sample is our bridge; the reason we're able to do this is because the sample values are determined by the total distribution.
Statistics and their distribution
Statistics
The sample is the basis for statistical extrapolation. But when applied in practice, it is generally not the use of the sample itself, but the construction of appropriate functions, i.e. statistics, that are tailored to the specific issue, using these functions for research.
Definitions:$X_{i}$It's a general sample. $x_{i}$ It's a sample value.$g(X_{1},...X_{n})$For statistics.$g(x_{1}...x_{n})$For the corresponding observation value Note that statistics must not contain any unknown parameters in the overall distribution. Like what? $\frac1{\sigma^2}\sum_{i=1}^n(X_i-\mu)^2$ Whether or not statistics depend on parameters$\sigma,\mu$Known The average of the samples, the difference, the standard difference, the steps of the square, the center of the square are important statistics, and their expressions are no longer repeated.
Fix sample differences
The difference we're using in probability is $$S^{2}=\frac{1}{n}\sum_{i=1}^{n}(X_{i}-\mu)^{2}$$ But in statistics, in order to meet the problem of unbiased estimates that we will point out later, we need to make the following amendments, in addition to replacing the overall average with sample averages. $$S^{2}=\frac{1}{n-1}\sum_{i=1}^{n}(X_{i}-\overline{X})^{2}=\frac1{n-1}[\sum_{i=1}^nX_i^2-n\overline{X}^2]$$
Three sample distributions
The statistics are a function determined entirely by the template, and he's also a random variable, and his distribution is called sample distribution, and here's three very important sample distributions, all of which come from the most classic general pattern of normality.
Carside Distribution$\chi ^{2}$Distribution)
Definition: Setup$X_{1},X_{2}...X_{n}$ It's from the standard normal general.$N(0,1)$The sample is statistical. $$\chi^2=X_1^2+X_2^2+\cdots+X_n^2$$ Call it freedom.$n$I'm not sure I'm going to be able to do that.$\chi^2\sim\chi^2(n).$ We give the probability density function for the distribution of the card directly. The probability density function is $US$\left.f(x)=\left}begin{array}}frac}1}&x>0,\0,&\text Other Organiser.$$ 同时给出一些性质 卡方分布具有可加性(特征函数理论可证明)$$X\sim\chi^2(n 1), Y\sim\chi^2(n 2), \\text{and}X, Y\text{independent}\Rightarrow X+Y\thicksim\chi^2(n 1+n 2)$$ 卡方分布的期望和方差为(中心矩,原点矩理论可证明) $$E(\chi^2)=n,D(\chi^2)=2n.$$ 定义 某点$x$是$\chi^2\sim\chi^2(n).$ 的上$\alpha$ 分位点 当且仅当 $$P{\chi^{2}>\chi_{\alpha}^{2}(n)}=\alpha(0<\alpha<$1 (very basic definition of location)
T distribution
Set$X\sim N(0,1),Y\sim\chi^2(n)$ The two are independent of each other. $$t=\frac X{\sqrt{Y/n}}$$ Satisfied freedom by$n$The T distribution is recorded as$t(n)$ Gives a probability density function for T-distribution $f(x)=\frac{\Gamma [(n+1)/2}sqrt{\pi}\Gamma(n)}left(+\frac{x^2}^right)^-\frac{n+2} (-\info)}<x<+infty It's easy to see that the probability density function of the T-distribution is an even function, and the standard normal probability density is the T-distribution.$n$ Approaching. $\infty$ The limit of time
T-distribution as an even function $$t_{1-\alpha/2}(n)=-t_{\alpha/2}(n)$$
F distribution
Set$X\sim\chi^{2}(n_{1}),Y\sim\chi^{2}(n_{2})$ The two are independent of each other. $$F=\frac{X/n_1}{Y/n_2}$$ Obey freedom.$n_{1},n_{2}$F distribution as recorded$F(n_{1},n_{2})$ Gives a probability density function for F distribution $f(x)=\begin{cases}\Gamma(n_1+n_2)/2^{\frac{n_1}2}x^{\frac{n_1}2-1}\\hline\Gamma(n_1/2)\Gamma(n_2/2)[1+(n_1x/n_2)]^{\frac{n_1+n_2}2}\0,&\text{Other}.\end{cases}$$ 容易知道 F分布有关于交换构成部分分子和分母的性质 $$F\sim F(n_1,n_2)\Rightarrow\frac1F\sim F(n_2,n_1)$$ 关于F分布的分位点有这样的性质 $$F_{1-\alpha}(n_1,n_2)=\frac1{F_\alpha(n_2,n_1)}$$
About freedom.
The original freedom of freedom was originally introduced in the Carabinieri distribution. He expressed the number of stand-alone random variables contained in our statistics. In computing, it's often the sample capacity minus the number of binding equations, and we continue to see many expressions of freedom in many places back there.
Distribution of sample averages and differences in sample size
Theorem I
Assumptions of average and variance overall $$E(X)=\mu,D(X)=\sigma^2,$$ So whatever the distribution of the total, for the sample from this total, There must be a sample average. $$E(\overline{X})=\mu,D(\overline{X})=\frac{\sigma^{2}}{n}.$$ The expectations and differences of the sample (amendment) range can be addressed by the distribution given by Theorem II Combined$\chi^2$ Average and variance of distribution calculated
Theorem II
For a single normal aggregate $ \\mathrm{N)\mu,\sigma2)}$ 其样本均值$\overline{X}$和样本修正方差$Satisfied with $2
- $\overline{X}\sim N(\mu,\frac{\sigma^{2}}{n}).$
- $\frac{(n-1)S^2}{\sigma^2}\sim\chi^2(n-1);$
- If there are no amendments to$\frac{1}{n}$The difference in the coefficient is$\frac{nS^2}{\sigma^2}\sim\chi^2(n-1);$
- If the difference is calculated using a real average instead of a sample average $\chi^2$The freedom of distribution becomes$n$ ;fixing deviations does not involve the use of real averages, the correction is intended to resolve the problem of neutrality, and true averages do not involve this
- $\frac{\overline{X}-\mu}{S^{\color{red}}/\sqrt{n}}\sim t(n-1);$
- Sample average$\overline{X}$and sample correction variance$S^2$Independence
Theorem III
For a double-normal sum of the same difference $ \\mathrm{N)\mu_{1},\sigma^2)}$ $\mathrm{N(\mu_{2},The average difference is
$$\frac{(\overline{X}-\overline{Y})-(\mu_1-\mu_2)}{S_w\cdot\sqrt{\frac1{n_1}+\frac1{n_2}}}\sim t(n_1+n_2-2)$$
of which
$$S_{w}^{2}=\frac{(n_{1}-1)S_{X}^{2}+(n_{2}-1)S_{Y}^{2}}{n_{1}+n_{2}-2}.$$
If not amended
$$S_{w}^{2}=\frac{(n_{1})S_{X}^{2}+(n_{2})S_{Y}^{2}}{n_{1}+n_{2}-2}.$$
Where's that? \overline{Y} ~S X^2}~S Y^2}$ for the average sample and difference
And there is.
$$\frac{S_X^2}{S_Y^2}\thicksim F(n_1-1,n_2-1)$$
If not amended
$$\frac{n_{1}(n_{2}-1)S_X^2}{n_2(n_1-1)S_Y^2}\thicksim F(n_1-1,n_2-1)$$
When the difference is different, but both take the form of correction.
$$\frac{\frac{S_X^2}{\sigma_{x}^2}}{\frac{S_Y^2}{\sigma_{y}^2}}\thicksim F(n_1-1,n_2-1)$$
The key formulas and the three distributions given in this section are very important, and then they are often used to deal with hypothetical tests.
Most of the inducts involved are fundamental deformations and the application of the theorem to the front. Be careful to make a strict distinction between sample differences and sample correction differences.
Examples
Examples of this section Section entitled “Examples of mathematical statistical sampling distribution”
Order statistics and their distribution
Order Count
Assumptions$X_{1}...X_{n}$From the general distribution function$F(x)$If we sort these samples from bottom to big, we get a sequenced sample. $X_{(1)}...X_{(n)}$
We call it No.$i$Order counts from small to large$i$Amount $X_{i}$
Obviously. $X_{1}$ Called Minimum Order Statistics $X_{n}$ Called maximum order statistics
It's very clear that order statistics should be distributed separately, because sample numbers must be limited to a few points, and we can study the distribution of order statistics, but obviously, the distribution of order statistics should not be independent. It's a very good order for the distribution that's separated. Here's what we're going to do.
Distribution of order statistics
General$X$The density function is$p(x)$ Distribution function is$F(x)$ $X_{1}...X_{n}$From the general distribution function$F(x)$And the sample is no.$k$Order statistics$X_{k}$Distribution is $$p_{k}\left(x\right)=\frac{n!}{\left(k-1\right)!\left(n-k\right)!}\left(F\left(x\right)\right)^{k-1}\left(1-F\left(x\right)\right)^{n-k}p(x)$$
For the joint distribution of multiple order statistics, we directly give the binary formula: $00begin{aligned}p (y,z)&=\frac{n!}{\left(i-1\right)!\left(j-i-1\right)!\left(n-j\right)!}\left[F(y)\right]^{i-1}\left[F(z)-F(y)\right]^{j-i-1}\\&\cdot\left[1-F(z)\right]^n-j}p(y)p(z),\cdot\leq z\end{aligned} $ We don't give formulas for the wider picture.
Empirical Distribution Functions
Assumptions$x_{1}...x_{n}$From the general distribution function$F(x)$If we sort these samples from bottom to big, we get a sequenced sample. $x_{(1)}...x_{(n)}$ Define the following functions with an orderly sample $F n(x)=\begin{cases}0,&x<x_{(1)}\k/n,&x_{(k)}\le x<x_{(k+1)},\1,&x (n)\le \end{cases}\quad k=1,2,... n-1 Obviously, this function meets all the conditions of a distribution function, which we call an experience distribution function.$F_{n}(x)$ He's an ordinary jumping function. According to Bernouli's law. $F_{n(x)}$Condense to distribution function by probability$F(x)$
Empirical distribution functions are derived from samples, even if the same sample size varies. For sure.$x$ The empirical distribution function is a random variable. It's an event.<Frequency of occurrence We gave a theory about the empirical distribution function, and he explained that as long as the sample was large enough, the empirical distribution function was a good approximation of the distribution function. Set$x_1,x_2,\cdots,x_n$is the total distribution function as$F(x)$The sample, $F_n(x)$for their experience distribution function, when$n\to+\infty$There is. $$P{ \lim sup\mid F_n( x) - F( x) \mid = 0} =1$$
Compilation and presentation of sample data
References Descriptive statistics and visualization
Parameter estimation
Parameter estimation issues
Definitions
The basic problem with mathematical statistics is to extrapolate the overall distribution and some numerical characteristics of the distribution, based on the information provided by the sample. One of the types of the problem is that the overall distribution is known, while some of its parameters are unknown and, based on the samples obtained, these parameters are extrapolated and referred to as parameter estimates
Estimates and estimates of unknown parameters
Suppose we have a normal distribution overall.$X$ Obey.$N(\mu,\sigma^{2})$ Parameters are unknown How do we estimate the unknown parameters? Yeah.$\mu$ You can use the average sample, median sample, etc.$\sigma$ We can estimate the difference. I can see that we need to construct a sample function. $\hat{\theta}(X_1,X_2,\cdotp\cdotp\cdotp X_n)$ It's obviously a statistical amount that we call the statistical amount used to estimate the parameters. This method of directly giving values for unknown parameters is calledPoint estimate estimate$(0,1)$ Overwrite$\mu$ The probability is 95%. The method to give the value range of unknown parameters is calledEstimates
Common estimation methods
- Rectangular estimation method
- It's a very similar estimate.
- Minimal 2x2 estimation method
- The Bayesian method.
- Scale-neutral method
- Maximum Minimum Estimated Method The usual estimates are basically these. In mathematical statistics, we're going to introduce rectangular estimations and very similar methods.
Rectangular estimate
Theory
The idea of rectangular estimation is a simple alternative thought, proposed by the statistician Pearson. Theoretically, the law of Sinchin. $US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$$US$US$US$US$US$US$US$US$$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$$US$$$$$$US$US$$US$$US$$US$$US$US$<\varepsilon)=1.$$ 样本矩依概率收敛于总体矩 由于 $X_{i}^{k}$ 仍然保证了独立同分布 因此高阶原点矩也可以使用辛钦大数定律 $$\lim_{n\to\infty}P(\frac{1}{n}\sum_{i=1}^{n}X_{i}^{k}-E(X^{k})|<\varepsilon) = $1 This method of estimating the overall rectangular using the corresponding sample rectangular to determine the estimated value of the parameters to be determined is calledRectangular estimation method
Methodology
General$X$Probability function$f(x;\theta_1,\theta_2,...,\theta_l)$Organisation$l$Unknown parameter$\theta_{1}...\theta_{l}$ $X_1,X_2,\cdotp\cdotp,X_n$It's total.$X$ , and in general,$l$Rectangular E (X^k)\left(){k=1,2,..,l}\right)$存在,则它们应是这$Functions for \iota$: $E (X^k}=g{k}(\theta_{1},\cdots,\theta_{l}),\quad k=1,2,\cdots,l$$ 又 样本的$k$ 阶原点矩为$A_{k}=\frac{1}{n}\sum_{i=1}^{n}X_{i}^{k}$ 因此我们可以建立并解以下的方程来确定参数的矩估计值 $$\begin{cases}g 1(\theta 1,\cdots,\theta l)=\frac1n\sum=i,\g 2(theta 1,\cdots,\theta l)=\frac1\sum 1}i^,\cdots\cdots\g l(\theta 1,\cdots,\theta l)=\frac1n\sum i^l,\cdots\ It's very obvious that we should select the number of equations that we have to create, the number of the highest steps in the rectangular.
Examples
Examples of this section Section entitled “Examples of mathematical statistical rectangles”
The advantage of the rectangular approach is that it's simple, and it doesn't need to know what the distribution is.
The disadvantage is that when the overall type is known,Inadequate use of information provided by the distribution .
At the same time, in general, rectangular estimates are not unique, the main reason being that when establishing a rectangular equation, the general rectangles are selected with the corresponding sample rectangles instead of the belt.
Very similar estimates.
A very similar estimate (MLE) is a parameter estimation method used under conditions known for the overall distribution type. It was originally proposed by mathematician Gauss, but the real success is due to statisticians Fisher. It's a very simple idea. Select the one with the greatest probability that the sample will appear in all parameter selections as its estimate. Value It's a very intuitive idea.
It's a very similar principle.
In general, if the probability and parameters of event A occur$\theta\in\Theta$About,$\theta$ The values vary, as do the P(A). the probability of event A occurring$P(A|\theta).$If one test, event A happens, it's considered at this time.$\theta$It's supposed to be a medium P.$\theta$To the biggest one. That's how it works. His advantage is that the information given by the overall distribution when it's known is of better quality, but the difficulty of calculating it has clearly increased. It's very apparent that the assumption is that there's an overall known distribution in the form of parameters, which is also a parameter method.
Methodology
It's a great estimate of the total separation.
If the total$X$It's discrete. $P{X=x}=p(x;\theta)$ Form known but parameters unknown$X_1,\cdots,X_n$It's from$X$the sample; or$X_1,\cdots,X_n$Joint distribution laws $$\prod_{i=1}^np(x_i;\theta)$$ Again.$x_1,\cdots,x_n$Yes.$X_1,\cdots,X_n$A sample value: an easy sample$X_1,...,X_n$Remove$x_1,...,x_n$the probability of an event${X_1=x_1,\cdots,X_n=x_n}$The probability is that: $$L(\theta)=L(x_1,\cdots,x_n;\theta)=\prod_{i=1}^np(x_i;\theta),\theta\in\Theta.$$ of which$L(\theta)$ We just have to find the right parameters to select the right ones under the determined sample values to make the apparent functions extremely significant as the estimate of the parameter, that is, $$L(x_1,\cdots,x_n;\hat{\theta})=\max_{\theta\in\Theta}L(x_1,\cdots,x_n;\theta)$$ $\hat{\theta}(x_1,\cdotp\cdotp,x_n)$ It's called a very similar estimate. $\hat{\theta}(X_1,\cdotp\cdotp\cdotp,X_n)$ It's called a huge estimate of parameters.
It's a long line of estimates.
Use exactly the same principle to get a function. $$L(\theta)=L(x_{1},\cdots,x_{n};\theta)=\prod_{i=1}^{n}f(x_{i};\theta)$$ Or is it the construction that makes the apparition of function great? $$L(x_1,\cdots,x_n;\hat{\theta})=\max_{\theta\in\Theta}L(x_1,\cdots,x_n;\theta)$$ Claims$\hat{\theta}(x_1,\cdots,x_n)$Yes$\theta$It's a huge estimate. Claims$\hat{\theta}(X_1,\cdots,X_n)$Yes$\theta$It's a huge estimate.
Larger methods
If density functions$f(x;\theta),p(x;\theta)$ About$\theta$ Then we can use the analysis's differentials to do a great deal of research. $$\frac{dL(\theta)}{d\theta}=0$$ Solution.$\theta$ That's our final parameter.
Because of the apparent function$L(\theta)$ It's the sum of multiple functions, so it looks like a logarithmic function.$\ln(L(\theta))$It'll simplify the guidance, simplify the operation, and the result will be the same, because...$\ln(x)$ Single
This stems from the theorem: If$\hat\theta$Unknown parameter$\theta$It's a huge estimate, and...$g(\theta)$Yes$\theta$One-to-Tune Functions${g}(\hat\theta)$ Yeah.$g(\theta)$It's a huge estimate.
Examples
Examples of this section and the section on “Suspicions of numerical estimates”
Medium estimate
Let's try to estimate the parameters of a Cauchy distribution.$\theta$ The density function is $$f\left(x,\theta\right)=\frac{1}{\pi\left[1+\left(x-\theta\right)^{2}\right]}$$ We know that Cauchy's rectangles don't exist, so the rectangles don't work. $$\sum_{i=1}^{n}\frac{X_{i}-\theta}{1+\left(X_{i}-\theta\right)^{2}}=0$$ This equation has a lot of roots and it's not easy to take root. But there's a simpler way.
$\theta$It's the median of Cauchy's distribution.
Guidelines for the excellence of estimates
So when there are multiple estimates, which one is better, that's the question that we're going to study now, if we judge the good or the bad, if we get a better estimate.
No bias
Non-selective introduction
We know that the dot estimate is actually a random variable, so as a volatile estimate, the result of many of its fluctuations, if it's around the real value of the parameter, he might be considered a good estimate. That's the system error estimated by the impartial study.
General$X\sim F(x;\theta)(\theta\in\Theta)$As parameter space.$X_1,X_2,...,X_n$Overall$X$The sample,$\hat{\theta}=\hat{\theta}(X_1,X_2,...,X_n)$Unknown parameter $\theta$ I'm just trying to figure it out. If the estimate $\hat\theta$ The mathematical expectations exist and there is. $$E_{\theta}(\hat{\theta})=\theta $$ Name$\hat\theta$ Yes.$\theta$ It's called bias. Claims $$b_{n}(\hat{\theta},\theta)=E_{\theta}(\hat{\theta})-\theta $$ For deviation of estimate
If $b_{n}(\hat{\theta},\theta)\ne0$ Name$\hat\theta$ Yes.$\theta$ Estimated bias If $\pepratorname*lim}{n\to\infty}b{n}({\hat{\theta}})=0$ 则称$\hat\theta$ 是$\theta$.
No bias requires no system error, which is certainly good in theory, but in practical application, the value of neutrality also needs to be determined by specific events.
Theorem
No matter what.$X$Obey what distribution, if $$ \mu\overset{\Delta}{\operatorname*{=}}E(X):,:\sigma^{2}\overset{\Delta}{\operatorname*{=}}D(X) $$ Both exist, then.$\hat{\mu}=\overline{X},\hat{\sigma}^2=S^2$The difference is... $\mu,\sigma^2$ The impartial estimate.
Here.$S^2$ It's a sample variance. It's modified. $\frac{1}{n-1}$is the difference of the coefficient
We introduced it at the beginning of mathematical statistics. $$E(\overline{X})=\mu,D(\overline{X})=\frac{\sigma^{2}}{n}.$$ The former is proof of the neutrality of the mean estimate. $ \begin{aligned} E\left (S{2}right)& =\frac{1}{n-1}E\left[\sum_{i=1}^{n}\left(X_{i}-\overline{X}\right)^{2}\right] \ &=\frac{1}{n-1}E\left[\sum_{i=1}^{n}X_{i}^{2}-n\overline{X}^{2}\right] \ &=\frac{1}{n-1}\left[\sum_{i=1}^{n}E\left(X_{i}^{2}\right)-nE\left(\overline{X}^{2}\right)\right] \ & =\frac{1}{n-1}\left[\sum_{i=1}^{n}\left(\sigma^{2}+\mu^{2}\right)-n\left[\operatorname{Var}(\bar{X})+\left(E(\bar{X})\right)^{2}\right]\right] \ &=\frac{1}{n-1}\left[n\left(\sigma^{2}+\mu^{2}\right)-n\left[\frac{\sigma^{2}}{n}+\mu^{2}\right]\right] \ &=\sigma^{2}. \end{aligned}$$
Note for uncorrected sample variance (in$\frac{1}{n}$ It's only gradual. $$\begin{gathered} {S_{n}}^{2}=\frac{1}{n}\sum_{i=1}^{n}(X_{i}-\bar{X})^{2}=\frac{n-1}{n}S_{n}^{*2} \ E({S_{n}}^{2})=\frac{n-1}{n}E(S_{n}^{*2}) \ =\frac{n-1}{n}\sigma^{2}\rightarrow\sigma^{2}~(n\rightarrow\infty) \end{gathered}$$ Which means, both rectangular and MLE, that their estimates of the overall variance are based directly on the sample center rectangular approach, which is not neutral.
Examples
Examples of this section “Guidelines for the assessment of mathematical statistical estimates” Section
Average error
Assumptions used$T(x)$As Parameter$q(\theta)$The estimates, a natural criterion for evaluating the merits of the estimates, could be defined as follows: $$ MSE_\theta(T)=R(\theta,T)=E(T(x)-q(\theta))^2 $$ Call the top the equation the average error and short the MSE
It's a very natural guideline for estimating the magnitude of errors, and it's common in a lot of statistical places as long as it involves point estimates. Equivalent errors can normally be broken down in the following forms: $$\begin{gathered} MSE=E_\theta\left[(\hat{\theta}-\theta)^2\right]=E_\theta\left[\hat{\theta}^2+\theta^2-2\hat{\theta}\theta\right] \ =E_\theta\left[\hat{\theta}^2\right]-E_\theta\left[\hat{\theta}\right]^2+E_\theta\left[\hat{\theta}\right]^2+\theta^2-2\theta E_\theta\left[\hat{\theta}\right] \ =V_\theta\left[\hat{\theta}\right]+(\theta-E_\theta\left[\hat{\theta}\right])^2 \end{gathered}$$ If we remember the deviation of the bias $$bias=E_\theta[\hat{\theta}]-\theta $$ There is. $$MSE=V_\theta[\hat{\theta}]+bias^2$$ That is the sum of the squares of the average error or deviation estimated
You can see that if the estimate is neutral, the deviation should be 0 MSE equals the method of estimation. In fact, the theory still contains equations of real values, and in actual calculations we need estimates instead of MSEs.
Define if for all$\theta\in\Theta$ There's always.$R(\theta,T)\leq R(\theta,S)$ Then we call it$S$It's an unacceptable estimate.$T$That's right.$S$Better. We don't usually choose unacceptable estimates.
Obviously, we would like to find an estimate with minimal MSE values for all parameters, but that's not possible; the best estimate of uniform error to a minimum does not exist. We usually choose to make some reasonable demands on estimates and select good estimates in estimates that meet the reasonable requirements.
Examples of this section Section on "Equal Errors"
Validity
Definitions$X_1,X_2,\cdot\cdot,X$It's total.$X\sim F(x,\theta);\theta\in\Theta$ The sample, What's that?{1}=\hat{\theta}{1}(X_{1},X_{2},\cdots,X_{n}),:\hat{\theta}{2}=\hat{\theta}{2} (X , X 2, \cdotsX n} Both.$\theta$ {\cHFFFFFF}{\cH00FF00}1)=E(\hat{\theta}2)=\theta\quad$.若$\forall\theta\ta$ $ ♪ And I'm so sorry ♪{1})\leq D(\hat{\theta}{2} $ Name$\hat\theta_{1}$ That's right.$\hat\theta_{2}$ More effective. It can be seen that effectiveness is just a narrower MSE. We introduced the concept of effective estimates in the C-R variant, and in fact the whole C-R variant and UMVUE were derived from the validity.
Compatibility (consistency)
Definitions
This is a good estimate of how the estimates will change as the sample capacity increases. A very natural idea is that with the increase in sample capacity, estimates should be more precise. Define. Set $hat(theta}n=\hat{\theta}(X_1,X_2,...,X_n)$是未知参数 $\theta$ 的点估计, 若 $\forall\theta\in\Theta$ 满足: $\forall\varepsilon>0$ 有 $\lim P{|\hat{\theta}{\cHFFFFFF}{\cH00FF00} Name $\hat{\theta}_n$ Yes.$\theta$ Combined estimates
Theorem
Whatever it is,{X}$服从什么分布,若 $\mu\triangleq E(X), \sigma^2\triangleq D(X)$ Both exist, then.$\hat{\mu}=\overline{X},\hat{\sigma}^2=S_n^{*2}$The difference is... $\mu,\sigma^2$ Combined estimates The core of research compatibility is the Sinchin Big Number Law, which means that the sample rectangles by probability. It's based directly on Sinchin's law. That's right.{i=1}^{n}X_{i}\xrightarrow{P}\mu(n\rightarrow\infty)$$ 研究方差$$\begin{gathered} == sync, corrected by elderman == == sync, corrected by elderman == == sync, corrected by elderman == I'm sorry. Obviously, the result is constricted by probability, because the difference is equal to the two-step square and one-step square. Bad
Conclusions
- Rectangular estimates are consistent estimates.
- Very similar estimates are generally consistent.
- Compatibility estimates may not be neutral, or even consistent.
- If$\hat\theta$ It's an impartial estimate.$\lim_{n\to\infty}D(\hat{\theta})=0$It's a confluence of estimates.It's a good conclusion.I'm not sure. Based on the theory here in Chebishev, there are sufficient conditions to estimate:(progressive) unbiased estimates and close to zero. Proof: Chebishev is different $$P\left(\left|\xi- E(\varepsilon)\right|\geqslant\varepsilon\right)\leqslant\frac{D\left(\xi\right)}{\varepsilon^{2}}$$ If there is.$\lim_{n\to\infty}D(\hat{\theta})=0$ And the difference left is close to zero, based on the pressure theorem.
Progressive normality
A lot of complicated numbers in the$n$Close$\infty$ At the same time, they're moving towards normal distribution. Gradual normality is the extrapolation of the very restrictive rationale; What statistics are gradual and normal, and how to judge is not the focus of our research here. Gradual normality and compatibility are equally large samples.
C-R heterogeneity
One parameter tends to have multiple impartial estimates. We naturally hope that the difference will be smaller, but whether or not the difference is lower, under what conditions. The C-R heterogeneity explains this. Proof that under certain conditions there is no bias in the estimates.$\hat\theta$There's a positive bottom line.
Fisher Information Volume
Definitions
For logarithmic functions$ln(\xi;\theta)$ Define the amount of Fisher's information as follows: $I(\theta)=E (\frac(partial\f(\xi;\theta)\partial\theta}^2}>$0.00 The expectation here is to see the parameters as defined values.$X$ (samples) Seeking expectations, that's...$E_{X|\theta}$ Various properties indicate that the larger the Fisher information, the more information the sample can be considered to contain about unknown parameters.
Examples
Examples of this section "Mathematic statistics Fisher information volume case" section
Conclusions
If (generally most distribution meets) $$\frac\partial{\partial\theta}\int\frac{\partial f(x;\theta)}{\partial\theta}dx=\int\frac{\partial^2f(x;\theta)}{\partial\theta^2}dx,$$ then $$I(\theta)=-E[\frac{\partial^2\ln f(\xi;\theta)}{\partial\theta^2}]$$
C-R heterogeneity
Set$\xi_1,\xi_2,\cdots,\xi_n$To take from a probability function$f(x;\theta),\theta\in\Theta$ $={\theta:a<\theta<b}$的母体的一个子样 ,其中$a,b$为已知常数, 且可设$a=-\infty,b=+\infty.$ 又$\eta=u(\xi_1,\xi_2,\cdots,\xi_n)$是$g (\theta) a neutral estimate and meets normal conditions $\text{collect}{x:f(x;\theta)>\text {not \theta\text}$$ $$\begin{aligned}g^{\prime}(\theta)&\text{and}Partial f(x; \theta)}Partial\theta}\text{and for everything}\theta\Theta,\fra\partial{\theta}&\int f(x;\theta)dx=\int\frac{\partial f(x;\theta)}{\partial\theta}dx\end{aligned}$$ $$\begin{aligned}\frac\partial{\partial\theta}&{\int\cdots\int u(x_1,x_2,\cdots,x_n)f(x_1;\theta)\cdots f(x_n;\theta)dx_1\cdots dx_n}\&=\int\cdots\int u(x_1,x_2,\cdots,x_n)\frac\partial{\partial\theta}[\prod_{i=1}^nf(x_i;\theta)]dx_1\cdots dx_n\end{aligned}$$ 则对于Fisher信息量 $$I(\theta)=E(\frac{\partial\ln f(\xi;\theta)}{\partial\theta})^{2}>0$$ 有 $$D_\theta\eta\geq{\frac{[g^{\prime}(\theta)]^2}{nI(\theta)}}$$ 对于$g(\theta)=\theta$ 的情形 $$D theta\eta\geq\frac1}$$ So, we gave an estimate of the lower CR range, which is also known as the information variable.
We define the estimated amount that meets the normal condition as a formal estimate. It's easy to see that the lower level of the CR is the difference between a formal unbiased estimate. Bottom For other estimates that are not formal or impartial, the difference cannot be given in the CR-I. Border
Application of CR insularity
Definitions$\theta$An impartial estimate$\hat\theta$ Makes the CR indifferent. $$D(\hat{\theta})=\frac1{nE[(\frac{\partial\ln f(\xi;\theta)}{\partial\theta})^2]}=\frac1{nI(\theta)}$$ Yes, it is referred to as a valid estimate (the amount of information here, regardless of multiple sampling, only one sample is taken) Definition$e=\frac{1}{nI(\theta)}/D(\hat\theta)$ It's called neutral efficiency. $e=1$ It's called a valid estimate. Definitions$e\ne1$ If$lim(e)=1$ It's called a progressive and effective estimate.
Examples of this section Section entitled “Mathematical statistics, CR-influences”
Adequate statistics
Adequate statistics
Introduction
The samples come from the general, and they contain the general information, but we often use the function of the tectonic sample -- statistical data -- to extrapolate statistically, how to extract the general information from the sample, whether or not we take the total information from the entire sample, which is the solution here. Here we usually consider only the information on the general distribution parameters contained in the sample, not the general distribution type. We use one example to introduce the concept of adequate statistics. For example, in order to study the fatality rate of an athlete, we tested the athletes 10 times and found that the remaining eight were hit except for the third and sixth. Now we're going to look into the parameters of the hit rate. It's very obvious.$T=x_{1}+x_{2}+...+x_{n}$ The amount of statistics constructed under this scenario will not be lost at all.$\theta$ That's the idea that statisticians Fisher is proposing. Adequate statistics With sufficient statistics, our statistical inferences on this parameter can be converted to statistics that no longer require sample data.
Definitions
Sample$X_{1}...X_{n}$ There's a sample distribution.$F_{\theta}(x)$ It contains all the parameters.$\theta$ Statistics$T$ There's a sample distribution. $T_{\theta}(t)$ The full measure of nature means...$T_{\theta}(t)$ All of it.$F_{\theta}(x)$About parameters$\theta$ The information, the distribution of the samples.$F_\theta(x|T=t)$ Other Organiser$\theta$ Information In that sense, we can give a full statistical definition.
Definitions$X_1,X_2,...,X_n$For Total$X$ The sample,$X$ The distribution function is$F(x;\theta),\quad T{=}T(X_1,X_2,...,X_n)$for a statistical amount, when T=t is given, if the sample$X_1,X_2,...,X_n)$ The distribution of conditions (probability of conditions at the time of separation and density at the time of continuity) and parameters$\theta$It's not relevant, it's called$T$As Parameters$\theta$Adequate statistics
Definition of equivalence$X_1,X_2,..,X_n$For Total$x$ The sample,$X$ The probability function is$f(x,\theta),\quad T{=}T(X_1,X_2,...,X_n)$The probability function for a statistical amount$g(t,\theta)$If$\frac{f(x_1,\theta)f(x_2,\theta)\text{L }f(x_n,\theta)}{g(T(x_1,x_2,\text{L },x_n),\theta)}=h(x_1,x_2,\text{L },x_n)$Establishment; And...$t=T(x_1,x_2,\mathcal{L},x_n)$When taking a fixed value,$T=t$Conditional probabilities under conditions of occurrence$h(x_1,x_2,\mathcal{L},x_n)$Not dependent$\theta$and $T$ As Parameters$\theta$ Adequate statistics
Examples
Examples of this section Section entitled “Question of adequate statistical data”
Factorial Theorem
By definition, it's cumbersome to determine whether a statistical measure is sufficient, so we've given the factor decomposition theorem, which can greatly simplify the search for a full statistical matter.
Factorial Theorem $T$ Yes.$\theta$ The most important condition for a full measure of statistics is that the joint distribution of samples can be broken down into the following forms:$h$ Non-negative and$\theta$ Not relevant $g$ Passes Only$T$ Link to the sample. $$L(\theta)=\prod_{i=1}^nf(x_i;\theta)=h(x_1,x_2,\cdots,x_n)g(T(x_1,x_2,\cdots,x_n);\theta)$$ It can be used very simply. All we have to do is get the combined density function properly deformed and decompose. The core is decomposition.
Theoretically, if$T$ Yes.$\theta$ A full statistical volume $f(t)$ It's a single reversible function.$f(T)$ Yeah.$\theta$ Adequate statistics
Full statistics
First we introduce the concept of a full distribution function. General$X$The distribution function family is${F(x;\theta),\theta\in\Theta}$ For any satisfaction$E_{\theta}[g(X)]=0$♪ To everything ♪$\theta\in\Theta$Random variable$g(X)$Always. $$ P_{\theta}{g(X)=0}=1,:\text{对一切}\theta\in\Theta, $$ Name${F(x;\theta),\theta\in\Theta}$As Full Distribution Functions Definitions$(X_1,X_2,...,X_n)$For Total$F(x;\theta)(\theta\in\Theta)$A sample, if measured$T=T(X_1,X_2,...,X_n)$Distribution Functions${F_{\tau}(x;\theta),\theta\in\Theta}$ is a full distributed function family, which is$T=T(X_1,X_2,\cdots,X_n)\text{}$ To complete statistics You can see the following characteristics of the full statistical data. $US$00begin{aligned}P {\theeta}\big{g}=g=(2}(T)\big}&=1,\quad\forall\theta\in\Theta\\Leftrightarrow E_{\theta}\big[g_{1}(T)\big]&=E_{\theta}\big[g_{2}(T)\big],\forall\theta\in\Theta\text{。}\end{aligned}$$ Examples of this section Section entitled “Question of full and mathematical statistics”
Indicator Distribution
It's a sort of widely used distributed community.
Single Parameter Index Distribution
Definitions: General $X$ or $X|\theta$ Distribution density $p(x|\theta)$ Is: $$ p(x|\theta)=g(x)h(\theta)\exp{t(x)\phi(\theta)} $$ The functions of which are the normal known functions are referred to$p(x|\theta)$ In the single parameter index distribution group Examples Study normal distribution$N(\mu,\sigma^{2})$ When?$\sigma^2$When known $ \begin{aligned} p (x|m)& =\left(2\pi\sigma^2\right)^{-\frac12}\exp\left{-\frac1{2\sigma^2}(x-\mu)^2\right} \ &== sync, corrected by elderman == I'm sorry. A definition that meets a single parameter index distribution group Similar, if the average is thought to be known, the difference is unknown, which is also part of the single-parameter index distribution. It's similar to the porcelain distribution, two distributions, Gamma distributions, Beta distributions, all of which are index distribution groups. In fact, the index is very wide-ranging. But the very basic distribution of evenly distributed forms is not an index distribution group, because his definition of clusters is related to parameters that cannot be summarized as an index. The definition of index distribution is only one of these forms, and it can actually be defined in many different ways, but they're all equal.
Biparameter index distribution
Definitions General $X$ or $X|\theta,\varphi$ The density of the distribution is $p(x|\theta,\varphi),\theta,\varphi$ Unknown parameter if $$ p(x|\theta,\varphi)=g(x)h(\theta,\varphi)\exp{t(x)\phi(\theta,\varphi)+u(x)\chi(\theta,\varphi)} $$ Name $X$ The distribution belongs to the two-parameter index distribution group
A theorem.
Set Random Variables$x$ With a single parameter index distribution,$X_1,X_2$,L ,$X_n$ It's from the general.$x$. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .$\sum_{i=1}^nt(X_i)$is the parameter$\theta$Adequate statistics Here.$t(X_i)$ It's the function inside when defining the index distribution group. If$\Theta^*$As$\mathbb{R}^k$And at this point, the full amount of statistics is complete. That is, for the index population, there's a very close line between full and complete statistics. It's a very good theory to give a full statistical breakdown.
Unanimous Minimal Estimation
The C-R-I-R-I-R-I-I-I told us.
- There's a difference in the amount of statistics.
- In fact, not every parameter has a valid estimate, because not every impartial estimate can reach the lower C-R horizon. So we asked two questions.
- It's known that a neutral estimate can construct a new neutral estimate that is smaller than the original one.
- An impartial estimate, even if it doesn't work, but his equation is minimal.
Rao-Blackwell Theorem
Rao-Blackwell Theorem
Set$X$and$Y$It's two random variables.>\boldsymbol{0}$ 定义$$\varphi(y)=E(X\mid Y=y)$It's okay. There is.$E\varphi(Y)=\mu,\mathrm{Var}(\varphi(Y))\leq\mathrm{Var}(X)$ It's a condition.${:}X\text{和 }\phi(Y)\text{几乎处处相等}.$ He offered a way to reduce the difference. Prove it. Set$X,Y$ combined density is$p(x,y)$ $X$ The condition density is$h(x|y)$ There is. $$\varphi(y)=E\left(X\mid Y=y\right)=\int xh(x\mid y)dx=\int x\frac{p(x,y)}{p_{Y}(y)}dx$$ So you can give it to me. $ \begin{aligned}E\phi(Y)&=\int\varphi(y)p_Y(y)dy\&=\int\int xp(x,y)dxdy=EX=\mu\end{aligned}$$ $$\begin{aligned}\operatorname{Var}(X)&=E\left[(X-\varphi(Y))+(\varphi(Y)-\mu)\right]^2\&=E\left(X-\varphi(Y)\right)^2+E\left(\varphi(Y)-\mu\right)^2\&\color{}{+}2E[(X-\varphi(Y))(\varphi(Y)-\mu)]\end{aligned}$$ 针对后半部分单独计算有 $$\begin{aligned} &E[(X-\varphi(Y))(\varphi(Y)-\mu)] \ &=\int\int[x-\varphi(y)][\varphi(y)-\mu]p(x,y)dxdy \ &=\int\int[x-\varphi(y)][\varphi(y)-\mu]p_Y(y)h(x\mid y)dxdy \ &=\int[\phi(y)-\mu]{\int[x-\phi(y)]h(x|y)dx}p_Y(y)dy=0 \end{aligned}$$ 因此 $$\begin{aligned}\operatorname{Var}(X)&=E\left(X-\varphi(Y)\right)^2+\operatorname{Var}(\varphi(Y))\\operatorname{Var}(X)&=\operatorname{Var}(\varphi(Y))\quad\Leftrightarrow P\left(X-\varphi(Y)=0\right)=1\end{aligned}$$
Use of adequate statistics
Set an overall probability density function as$p(x;\theta),X_1,X_2$, ,$X_{n}$ It's a sample.$T=T(X_{1}...X_{n})$ Yes.$\theta$Adequate statistics $S(X_{1}...X_{n})$ is the parameter$g(\theta)$An impartial estimate $$\varphi(T)=E(S(X)|T(X))$$ Yes.$g(\theta)$ An impartial estimate and $$Var_{\theta}(\varphi(T))\leq Var_{\theta}(S(X)),$$ When and only when$P(\varphi(T)=S(X))=1$ Time is set. This theory tells us that when the expectations are reduced, using a good choice when fully measured, it tells us to choose the conditions.$Y$How More generally: If the unbiased estimate is not a function of sufficient statistical volumes, it is expected that the requirement for sufficient statistical volumes will yield a new neutral estimate that is smaller than the original estimate, thereby reducing the partial estimate. In other words, consider$\theta$The problem of estimating needs only to be performed in a function based on sufficient statistical data, and the statement is correct for all statistical inferences, which is calledThe principle of adequacy
An example
Set$X_{1}...X_{n}$It's from$b(1,p)$ the sample$\overline{X}$ Yes.$p$ Adequacy of statistics$\theta=p^2$Uneven estimates We know that the sample average and the sample range are neutral estimates of the overall average and the range. $$E\overline{X}=E(X)=p;ES^{*^2}=D(X)=p(1-p)$$ There's a difference between the two.$p^2$ This is what we're building for. It's very common to use similar techniques to construct impartial estimates in the later study of UMVUE.
Minimal variance estimation
Now, let's answer another question:
UMVUE Definition
For parameter estimates, set$\hat{\theta}$Yes.$\theta$A neutral estimate, if any$\theta$Uneven estimates$\tilde{\theta}$, in parameter space$\Theta$Both. $$Var(\hat{\theta})\leq Var(\tilde{\theta})$$ It's called the Minimal Specimetric Estimates (UMVUE)
- If UMVUE exists, it must be a function of sufficient statistical weight.
- If the variance reaches the lower C-R limit, it must be UMVUE.
- The variance of the UMVUE does not necessarily reach the lower C-R boundary
UMVUE judgement
Set$X=(X_1,X_2,\cdots,X_n)$It's a sample from a certain group.$\hat{\theta}=\hat{\theta}(X)$Yes.$\theta$An impartial estimate, Var. < + \infty.$如果对任意满足$E(\phi(X))=0$ 的$\phi(X)$ both $$\color{}{\mathrm{Cov}_\theta(\widehat{\theta},\varphi)=0},\quad\forall\theta\in\Theta,$$ then$\hat\theta$Yes.$\theta$UMVUE
Construct UMVUE
Set$T(X)$It's a full-fledged count.$S(X)$Yes.$g(\theta)$, and$\varphi(T)=E_{\theta}(S(X)|T(X))$Yes.$g(\theta)$UMVUE
Further
If for all $\theta\in\Theta$,$Var_{\theta}(\varphi(T))<\infty$, 则$\varphi(T)$是$g(\theta)$唯一的$UMVUE$
Theorem tells us that the UMVUE can be constructed using a full measure of statistics.
And UMVUE is the only one in probability.
In fact, this theorem tells us two ways to find UMVUE, but first we need to find a full measure of statistics.$T(X)$
- Statistical quantum function method: if$\varphi(T(X))$Yes.$g(\theta)$Quantified, then$\varphi(T(X))$ Yeah.$g(\theta)$The UMVUE.$g(\theta)$The impartial estimate.
- Expectations: If available$g(\theta)$An impartial estimate$\varphi(X)$,$E(\varphi(X)|T(X))$Yeah.$g(\theta)$UMVUE.
UMVUE and C-R lower bounds
Some UMVUE can reach the lower C-R level, but others cannot. Normally, the neutral estimate for reaching the lower C-R is UMVUE, provided the lower C-R is present. You can't deny UMVUE without reaching the lower C-R level. It's important that we have an impartial estimate of how to reach the lower C-R level.
Examples
Examples of this section "Mathematic Statistics UMVUE" section
Estimates
The section is expected to introduce some of the elements of the hypothetical tests and to apply some of the elements that will be discussed later in advance, although still within the parameters estimate. We use average error to measure the deviation; but the concept of reliability is still lacking, the average error is much less accurate, and there are no fixed criteria; the inter-area estimate gives a range of parameters according to a certain level of reliability, which is what the entire section of the estimation is about.
Basic concepts of spatial estimates
Definition of confidence interval
General$X$Distribution Functions$F(x;\theta)$Contains an unknown parameter$\theta$, for given value $\alpha\left (0)<\alpha<1\right)$,若由样本$X_1,X_2,\cdots$, $Two statistics determined by X n$ $\underline{\theta}=\underline{\theta}(X_1,X_2,\cdots,X_n)$and$\text{}\bar{\theta}=\overline{\theta}(X_1,X_2,\cdots,X_n)\text{ }$Satisfied $ (X 1, X 2, \cdots, X n)<\theta<== sync, corrected by elderman == @elder man $ Name of random space (%2)$\underline{\theta},\overline{\theta})$Yes.$\theta$The confidence is 1-$\alpha$It's the confidence zone.$\underline{\theta}$ and $\overline{\theta}$It's called confidence 1.$-\alpha$ The two-sided confidence interval is called the upper limit and lower limit.
- The parameters to be determined are certain, but the unknown is random.
- So we can't say there are parameters.$1-\alpha$We should say that there is random space.$1-\alpha$Probability includes parameters
- If there's a lot of areas that are sampled over and over again, then the range that contains the parameters is about as much as$1-\alpha$It's the law of Bernoulie's great numbers.
Solution steps between confidence zones
One.
Find a sample $X_1,X_2,...,X_n$Function: $$ Z=Z(X_1,X_2,\cdotp\cdotp,X_n;\theta) $$ Include only parameters to be assessed $\theta$and $Z$ The distribution is known and does not depend on any unknown parameters (including$\theta$I'm not sure. Obviously, this doesn't fit the statistical definition.
Two.
For given confidence 1-$\alpha$set two constants$a,b$, make $P{a)<Z(X_1,X_2,\cdots,X_n;\theta)<b}=1-\alpha.$$ $a,b$It's all we need to do.
III
If you can get from $a<Z(X_1,X_2,\cdots,X_n;\theta)<b$得到等价的不等式 $\underline{\theta}<\theta<\overline{\theta}$, 其中 $\theta=\theta(X_1,X_2,\cdots,X_n)$, $\overline{\theta}=\overline{\theta}(X_1,X_2,...,X_n)$都是统计量,那么 $It's fine. Yes. $\theta$A confidence is 1- $\alpha$ Other Organiser It's a sort of equation.
A few notes.
- Different levels of confidence$\alpha$ Parameters$\theta$The corresponding confidence interval is different.
- The smaller the confidence interval, the more accurate the estimates, the lower the corresponding confidence level, and vice versa. Yeah.
- If you want to reduce the confidence level, you have to increase the sample capacity.
- The same confidence, the same confidence zone.
- The value of the axis is usually derived from statistical distortion, which is a test of the level of thinking, so how to construct the axis is not the focus we need to have.
Estimates of total normal averages
The overall variance is known.
If the total$X$Obey.$N(\mu,\sigma^2)$ of which$\sigma_{2}$as known$U$Statistics $$U=\frac{\bar{X}-\mu}{\sigma/\sqrt{n}}$$ Parameters$\mu$Make an estimation. For given confidence level$1-\alpha$ Yes. $US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$$US$US$US$US$US$$US$$$US$US$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$...$$$$$$$$...$$$$$$$...$$$$$$...<u_{1-\frac{\alpha}{2}}}=1-\alpha $$ Why is that? We sometimes need theory. And there's a confidence zone. $$\color{}{\left(\bar{X}-u_{\alpha/2}\frac\sigma{\sqrt{n}},\bar{X}+u_{\alpha/2}\frac\sigma{\sqrt{n}}\right)}$$ What happens between the confidence zones is a constant repetition of the problem. Understood.$u$The value means you know how to calculate. Is this the only kind of confidence zone? $$\left(\overline{X}-\frac\sigma{\sqrt{n}}u_{0.01},\overline{X}+\frac\sigma{\sqrt{n}}u_{0.04}\right)$$ The reason we don't do that is because this confidence zone is much longer, so it's not precise enough. High In fact, there are countless options in the confidence zone, but we're only going to choose the one that's the shortest.
The overall equation is unknown
This is the time to choose.$t$Statistics are for testing. It's easy to know. $$\frac{\overline{X}-\mu}{S^{\color{red}}/\sqrt{n}}\sim t(n-1);$$ So give confidence as$1-\alpha$Other Organiser $US$ \\left(x}-\frac{S n^)}{\sqrt{n}}\cdot t_{\alpha/2}(n-1),\quad\bar{X}+\frac{S_n^\cdottt (n-1)\right) Understood.$t$The value means how to calculate.
Estimates of the difference between the two normal overall mean values
Both differences are known.
$$\overline{X}\sim N(\mu_1,\frac{\sigma_1^2}n),\quad\overline{Y}\sim N(\mu_2,\frac{\sigma_2^2}m)$$ It's easy to construct a core number. $$\frac{(\bar{X}-\bar{Y})-(\mu_1-\mu_2)}{\sqrt{\frac{\sigma_1^2}n+\frac{\sigma_2^2}m}}\sim N(0,1)$$ So you can get it.$\mu_1-\mu_2$There's a confidence zone. $$\left((\overline{X}-\overline{Y})-u_{1+\frac a2}\sqrt{\frac{\sigma_1^2}n+\frac{\sigma_2^2}m},\quad(\overline{X}-\overline{Y})+u_{1+\frac a2}\sqrt{\frac{\sigma_1^2}n+\frac{\sigma_2^2}m}\right)$$
The difference is unknown, but the difference is equal.
$$\overline{X}\sim N(\mu_1,\frac{\sigma^2}n),\quad\overline{Y}\sim N(\mu_2,\frac{\sigma^2}m)$$ Give the pivotal amount $$\frac{(\overline{X}-\overline{Y})-(\mu_1-\mu_2)}{\sqrt{\frac1n+\frac1m}\sqrt{\frac{(n-1)S_1^2+(m-1)S_2^2}{n+m-2}}}\sim t(n+m-2)$$ Theoretical reasoning. $P\left(\left|left.\frac{(overline{X}----overline{Y}} }{\sqrt{\rac1n+\frac1m}\sqrt{(n-1)S 1^2+(m-1)S 2^n+m2}}}}right|<t_{1-\frac\alpha2}\right)=1-\alpha\right. $$ 给出置信区间 $$\left((\overline{X}-\overline{Y})\pm t_{1+\frac\alpha2}\sqrt{\frac1n+\frac1m}\sqrt{\frac{(n-1)S_1^2+(m-1)S_2^2}{n+m-2}}\right)$$
The difference is unknown, but the sample is large enough.
When the sample is large enough (which is generally considered to be more than 50) we can replace the difference in the sample correction with the difference in the actual difference and then return to the situation where the difference is known. $$\left((\overline{X}-\overline{Y})\pm u_{1-\frac\alpha2}\sqrt{\frac{S_1^2}n+\frac{S_2^2}m}\right)$$
The difference is unknown, but the sample is equal.
$$\overline{X}\sim N(\mu_1,\frac{\sigma_1^2}n),\quad\overline{Y}\sim N(\mu_2,\frac{\sigma_2^2}m)$$ And there is. $$n=m$$ You!$Z_{i}=X_{i}-Y_{i}$ We can think of the sample now.$Z_{i}$It all comes from $$Z\sim N(\mu_{1}-\mu_{2},\sigma_{1}^{2}+\sigma_{2}^{2})$$ We can assume that at this point we are in the process of estimating the mean of the normal distribution of the unknown difference in the sample. Select$t$Statistics $$\frac{\overline{Z}-\mu}{S^{\color{red}}/\sqrt{n}}\sim t(n-1);$$ The formula before transformation is $US$\left(overline{Z}-\frac{S z^)}{\sqrt{n}}\cdot t_{\alpha/2}(n-1),\quad\overline{Z}+\frac{S_ z^}{\sqrt{n}}\cdot t_{\alpha/2}(n-1)\right)}$$ 换入我们需要的量得到估计区间 $$\left((\overline{X}-\overline{Y})\pm t_{1+\frac\alpha2}(n-1)\frac{S_Z}{\sqrt{n}}\right)$$
Estimated range of normal aggregate differences
Overall$X$Obey.$N(\mu,\sigma^2)$ We just need to introduce.$\mu$Unknown For what is already known about the average, a similar form is given in the section on statistics that can be used to construct the axis, and freedom and coefficients change. One of the caloric distributions that we've described earlier is a pivotal amount. $$\frac{(n-1)S^2}{\sigma^2}\sim\chi^2(n-1)$$ So there is. $P\left{\chi {\alpha/2)^2(n-1)<\frac{(n-1)S^2}{\sigma^2}<\chi_{1-\alpha/2}^2(n-1)\right}=1-\alpha $$ 计算得到置信区间为 $$\left(\frac{(n-1)S^2}{\chi_{1-\alpha/2}^2(n-1)},\frac{(n-1)S^2}{\chi_{\alpha/2}^2(n-1)}\right)$$ 开方就可以得到标准差的置信区间 $$\loft(\1-\alpha/2^2(n-1)},\frac{sqrt{1}(n-1)}} This is the first asymmetric axis we've presented.$F$So is the core. We still picked a symmetrical point on both sides to determine the confidence interval. It's just a matter of choice.
Estimates of the difference between the two normal aggregates
Let's just talk about the fact that the overall average is unknown. Although we can still construct new cores based on the known variance estimates of the average above, it is easy to guess that freedom increases when the average is known. Or do you think you're going to give us a pivotal point based on the conclusions we've presented in the statistics section? Volume $$\frac{\frac{S_X^2}{\sigma_{x}^2}}{\frac{S_Y^2}{\sigma_{y}^2}}\thicksim F(n_1-1,n_2-1)$$ So, give us an inequity. $P\left{F {\alpha/2} (n 1-1, n 2-1)<\frac{S_1{}^2/{\sigma_1}^2}{S_2{}^2/{\sigma_2}^2}<F_{1-\alpha/2}(n_1-1,n_2-1)\right}=1-\alpha $$ 置信区间为 $$\color{}{\left(\frac{S_1^2}{S_2^2}\frac1{F_{1-\alpha/2}(n_1-1,n_2-1)},\frac{S_1^2}{S_2^2}\frac1{F_{\alpha/2}(n_1-1,n_2-1)}\right)}.$$
Single-side confidence interval
At this point, we're sure that the core of the confidence zone will be changed to $ \begin{aligned}P<\theta)&=1-\alpha\quad (\text{or}P(\theta)<\overline(theta} =1-\alpha)\end{aligned} It's understandable. At this point,$\underline{\theta}$ It's called a one-sided lower limit. $\overline{\theta}$It's called a one-sided confidence cap. The manner in which the core amount is constructed will not change the one-sided confidence zone, but the determination of the difference. The single-side confidence zone still meets our definition of confidence. He actually has the same value as the double confidence zone.
Proportional confidence interval (large sample capacity)
As a whole$X$The distribution is unknown, but the sample is very large. $$\overline{X}\sim N(\mu,\frac{\sigma^{2}}{n})$$ Actually, we're back to the whole problem of normality. Use estimate$\overline{X}$To construct the core amount required to finalize the estimated range of parameters
Give an example of a ratio (often attributed to two-point distribution parameters)$p$It's a very common problem in statistics.
Set Distribution$X$Obedience parameter is$p$The two-point distribution of the sample is$X_{1},...,X_{n}$ $n>50$ 求参数$p$置信度为$1-\alpha-dollar confidence interval
It's not a problem with a normal distribution, but it's easy to know that we can be assisted in our research with very limited logic. $$\overline{X}\sim N(p,\frac{p(1-p)}{n})$$ Attention, we're bringing in two-point distribution expectations and differences, not two distributions. Now the question is, is there an unknown variance in the range estimate of the normal distribution mean?
The difference is unknown, but he only contains the parameters we want to estimate.$U$Statistics as pivotal $P\left.<\frac{\sqrt{n}(\overline{X}-p)}{\sqrt{p(1-p)}}<u_{1-\frac\alpha2})\approx1-\alpha $$ 得到关于$p$的方程(和前面的思路还是有点区别的) $$\begin{aligned}0\leq\frac{n(\overline{X}-p)^2}{p(1-p)}<u_{1-\frac{\alpha}2}^2\end{aligned}$$ 化简 $$(n+u_{1-\frac\alpha2}^2)p^2-(2n\overline{X}+u_{1-\frac\alpha2}^2)p+n\overline{X}^2<$0.00 The equation can be estimated at the desired range.
Assumptions test
Introduction of assumptions and assumptions
When we don't know anything about parameters, we usually use the parameter estimation method in the previous chapter. But when the parameter estimates are completed, we have a basic understanding of the parameters, and we want to know if our estimates are correct, and that's what this chapter's hypothetical test is about.
What's hypothetical?
Presentation of specific values for overall parameters For example, the overall average is greater than a certain number, the overall equation is less than a certain amount, and so on.
What's a hypothetical test?
Some assumption of the overall parameter (or distribution form) is first made, and then the determination of whether or not the hypothesis is established is made using sample information
There are two types of tests of parameters and non-parametric tests, the difference between non-parametric and non-parametric statistics.
Logical application of counter-evidence, statistically based on the principle of small probability.
Original and alternative assumptions
null hypothesis alternative hypothesis
The assumption is that we're gathering evidence that we want to object to.
In the hypothetical tests behind us, Thinks it's usually an equation with an equal sign like:$\mu=10,\mu\ge10$
The alternative scenario is the one we want to support. In the hypothetical tests behind us, An equation that does not normally contain equivalents is as follows:>0$
The original and alternative assumptions must be opposed to each other. Group
A hypothetical test example
Let's start with an example of the process of hypothetical testing.
Snails produced at a plant, with a standard strength of 68, and actual production strength$X$ Yeah.$N(m,3.6^2 )$If$E(X)=\mu=68$If the average values are: 69.5 and 67.5 respectively, is the same?
We make two assumptions. $\mu=68$ Alternative assumptions $\mu\ne68$
Now we're going to select one of the two scenarios for a hypothetical test. $\mu=68$ Correct (select original assumption)
Then there is. $$\overline{X}\sim N(68,3.6^2/36)$$ Now we're building a contradiction between a small probability event, and... It's based on a small probability.$\alpha=0.05$ $$P\left(\left|\frac{\overline{X}-68}{3.6/6}\right|>== sync, corrected by elderman == It's possible to decipher the odds and get the acceptance and rejection.
Additional explanation of some terms
Two types of errors
It's very clear that our test results are based on the fact that it's completely incorruptible.
- The assumption is true, but the sample was rejected.
- The original hypothesis was false, but it was accepted for sampling reasons.
We usually write down the probability of a first-class error.$\alpha$ The probability of a second type of error is...$\beta$
Small probability principle
In one trial, the probability of an almost impossible event is called a small probability. Once a small and medium probability event occurs, we have reason to reject the assumption. The small probability is what we decide.
High profile level
That's the small probability we identified as a small probability.$\alpha$
That's what we decided in advance. 0.05 ~0.1 million
And that's why we're using the symbol here because it's the same value.
Test statistics
Test statistics are based on sample observations. The sample statistics used to make decisions about the original and alternative assumptions are also a statistical amount, but we use it for a specific purpose. He's a hypothetical.$H_0$For real, a certain amount of statistical data that is known to be constructed. Core questions on testing classical statistical assumptions when selecting the appropriate statistical volume
Deny Field
Denied domain is a collection of all possible values that can be taken to test the statistical amount as originally assumed, and the boundary of rejected domain is called the threshold
The probability of two types of error
Qualitative analysis
Type two error probability
- Increases as the overall parameters assumed decrease
- Increase with first-class errors
- Increase with overall standard deviation
- Increase with reduced sample capacity
We need to know.
It's impossible to reduce the probability of two types of error at the same time, at the established sub-sample capacity. By increasing the size of the subs, you can reduce the second type of error.
The calculation below can be explained.
Calculating the probability of two errors
Let's just say one example. Example:$X_1,X_2...X_{n}\sim N(\mu,\sigma^2)$ of which$\sigma^2$ Known $H 0: \mu=h 1: \mu>$0.00 Deny Domain As $\bar{x}\ge c_{0}$
- The probability of two types of error.
- Yes.$\mu_0=0.5,\sigma=0.2,\alpha=0.05,n=9,\mu=0.65$Time-based calculations of the probability of not making category two errors
We just need to start with a definitional analysis, and use tests to make some complementary simplicity.
The first type of error is a waiver of the truth, which is the original hypothesis, but a refusal. The second type of error is hypocrisy, which means the original hypothesis is not valid, but acceptance. $$00\ &\alpha=P\left(\c 0H 0真right)\ &=P_{\mu_{0}} \left(\overline{x}\geq c_{0}\right) \ &=P\left(\frac{\overline{x}-\mu_{0}}{\sqrt{\frac{\sigma^2}{n}}}\geq\frac{c_0-\mu_{0}}{\sqrt{\frac{\sigma^{2}}{n}}}\right)=1-\phi\left(\frac{c_{0}-\mu_{0}}{\sqrt{\frac{\sigma^{2}}{n}}}\right) \end{aligned}$$ $$\begin{aligned} &\beta=P\left(\c 0HH 1真right)\ &=P_{\mu} \left(\overline{x}\leq c_{0}\right) \ &=P\left(\frac{\overline{x}-\mu}{\sqrt{\frac{\sigma^2}{n}}}\leq\frac{c_0-\mu}{\sqrt{\frac{\sigma^{2}}{n}}}\right)=1-\phi\left(\frac{c_{0}-\mu}{\sqrt{\frac{\sigma^{2}}{n}}}\right) \end{aligned}$$
From this point on, one pattern of one type of error and two type of error is shown; therefore, the probability of one type of error cannot be reduced indefinitely, corresponding to the probability of another type of error explodes.
We're using the basic definition to calculate the probability of two errors. And then, by testing the amount of statistics to simplify the distribution of some standard, you get the probability of two types of error.
I can see the first and second types of errors.$c_0$It's completely unknown, because the rejection field is determined by the probability of choosing the first type of error, and we control the probability of the first type of error first in the actual hypothetical test.
We don't need to calculate the specific values of two types of errors, for the first type, which is actually an equation, and the probability of the first type of errors is artificially assigned; for the second category of errors, the probability of the first type is assigned.$\alpha$Send$c_0$Then we'll take over.$\beta$expression that you can get the result of
If we give the rejection field directly (or calculate the result in the question before), then the probability of a category I-II error (re-inferred by definition) can be calculated, which is often the case when the probability of a type I error is properly controlled.
Assumed testing of the overall average
A hypothetical test for the known difference average value extraction
Give two hypotheses. $$H_0\colon\mu=\mu_0;\quad H_1\colon\mu\neq\mu_0$$ Construct the number of tests $$U=\frac{\overline{X}-\mu_0}{\sigma/\sqrt{n}}\sim N(0,1)$$ When you accept the original hypothesis, test the statistics.$\overline{X}$ The situation is known.$N(0,1)$
The area of rejection, the area where the small probability of an event occurs. $$P_{H_{0}}(\left|\frac{\overline{X}-\mu_{0}}{\sigma/\sqrt{n}}\right|\geq u_{\frac{\alpha}{2}})=\alpha $$ You can get statistics by deforming.$\overline{X}$Other Organiser
Called$U$Tests, because the numbers are...$U$Statistics
Hypothetical test for the difference unknown
Construct the number of tests $$T=\frac{\overline{X}-\mu_{0}}{S^{*}/\sqrt{n}} \sim t(n-1)$$ It's easy to give a rejection field based on the principle of small probability. Just check if this is in the rejected field. Called$t$Test
Assumed test of equal dual-normal matrix averages
When the difference is unknown but equal
Test the number of statistics to be $$\frac{\overline{X}-\overline{Y}}{\sqrt{\frac1n+\frac1m}\sqrt{\frac{(n-1)S_1^{*2}+(m-1)S_2^{*2}}{n+m-2}}}\sim t(n+m-2)$$
In the case of Big Son
According to the central limits, both values are subject to normal distribution, so they can be constructed.$U$Statistics $$U=\frac{\bar{X}-\bar{Y}}{\sqrt{\frac{S_1^{*2}}n+\frac{S_2^{*2}}m}} \sim N(0,1)$$
Assumptions of overall variance
Hypothetical test of the squared-out value if the average is known
Test statistics $$\chi^2=\frac{\sum_{i=1}^n(X_i-\mu)^2}{\sigma_0^2}\sim\chi^2(n)$$ Here's the difference.$n$And the auxiliary construction.$n$I've got a split.
Hypothetical test for squared-out values in the event of unknown averages
Test statistics $$\chi^{2}=\frac{(n-1)S^{*2}}{\sigma_{0}^{2}}\sim\chi^{2}(n-1)$$
In the case of large samples
For the tests we've been able to construct,$\chi^2(n)$In the case When the sample is large enough, it's very limited by the center. $$\frac{\chi^2-n}{\sqrt{2n}}\overset{\text{}}{ \operatorname* { \sim }}N(0,1)$$ So, the test statistics that are calculated above are brought here to get new tests. All the calorie tests can be converted to large samples.$U$Test
Assuming tests for a two-normal parent-rate difference
Test statistics $$F=\frac{S_1^{*2}/\sigma_1^2}{S_2^{*2}/\sigma_2^2}=\frac{S_1^{*2}}{S_2^{*2}}\sim F(n-1,m-1)$$ Use this to test whether the difference is equal and then return to check the balance; It can be understood that the difference is an essential part of the examination of the mean equivalent.
Match data$t$Test
Match data$t$Testing is one of the most widely applied hypothetical tests in statistics.
How to match
We often want to get two subjects close to the test. The aim is to avoid the influence of factors other than experimental treatment.
Three scenarios with design information
- The same pairing of pairs is given two different treatments to the subject subject, the aim being to infer whether the effects of the two treatments differ.
- The same subject was treated differently, with the aim of insulating that the effects of the two were different
- Comparison of treatment before and after the same subject is tested, the purpose is to infer whether a certain treatment is effective You can see, it's a pair.$t$The tests should be very extensive for processing of experimental data.
$t$Test
Designed with pairs$t$The test was to examine the difference in the average between the two data sets. If we're going to make a difference between two sets of data, then we're going to be a whole. What we're looking at is a hypothetical test of whether the average of the margin values in this group is zero. $$t=\frac{\overline{d}-\mu_{d}}{s_{\overline{d}}}=\frac{\overline{d}-0}{s_{d}/\sqrt{n}}=\frac{\overline{d}}{s_{d}/\sqrt{n}}\quad\sim{{t(n-1)}}$$ It's a hypothetical test of the unknown difference in the overall average. The specific example is not to describe here the idea of understanding this test is enough.
Single-side hypothetical tests
Double and single-sided hypothetical tests
The alternative is no directional.$\ne$ It's called a double-sided test or a double-tail test. The alternative scenario is specific in orientation and contains a symbol>~<$ Called a single-tailed test or a single-tail test
Use symbol<$ 的称为左侧检验 使用$>Called right test
One-sided testing is achieved
The one-sided hypothesis test is very natural, and we need to convert the rejection of the domain to the same level of statistical data.
Whether a single-sided test is used is required to see whether the need for such use exists for the specific subject of the hypothetical test;
Requires that the standard deviation of resistance of a certain conductor should not exceed 0.2. Obviously, we should have a default scenario of $H 1:\sigma^2>0.2 GH2 Our aim is to deny the bias, so we should take it as a hypothetical.
The standard weight of each bag packed with salt-eating automatic packaging is not less than 500, and the machine adjusts to take a subs and tests whether the average weight of the bag is significantly lower. Our goal is to deny the smallness of the dollar.<}500.$
That's the situation where one-sided hypothetical tests are required.
Non-normal overall hypothetical tests
Assumptions of scale (situation of large samples)
Or is it the use of the central, very limited logic that is the subject of a normal general hypothesis test, which we'll present only one scenario, that is, the hypothetical test of the 0-1 distribution parameter? Set Overall as $$P\bigl{X=x\bigr}=p^{x}\left(1-p\right)^{1-x},\quad x=0,1$$ Two scenarios are selected as $$H_{0}:p=p_{0},\quad H_{1}:p\neq p_{0}$$ We know by the very limits of the center when the original hypothesis is established and the sample is sufficiently large. $$U=\frac{\overline{X}-p_{0}}{\sqrt{p_{0}(1-p_{0})/n}}.$$ Approximate to standard normal distribution$N(0,1)$
Available$U$Test export hypothetically tested rejection field The change will give you the same rejection field as the sample average.
Hypothetical testing of the value of the index distribution parameter
The core of the hypothetical test is that the construction of the statistics is not a matter of perception, and first we need to make sure that the distribution of the statistics is known, and then we need to probably at least test the way the statistics are being shifted. $p(x)=\left(begin{array}ll} The \lambda e^, \lambda x, \ambda \ & x\ge0, \ 0, & \text {Other.} I'm sorry.$$ 由于我们已知$E(X)=\frac{1}{\lambda}$ 所以非常自然的思想是 根据均值的偏移方向和均值变形后可能的分布来构造检验统计量量 不妨设原假设为 $H_0:\lambda=\lambda_1$ 那么 均值过大或者过小的时候 拒绝原假设 又 $$2\lambda(1}cdots+x n}) =2\lambda\bar}sim\chi 2 $$ It's not that hard to refuse the field. The construction tests are very dependent on people's experience and knowledge of the distributions.
Practical application of hypothetical tests
Assumed p-test
In the hypothetical tests we conducted before us, we gave permission to make the first type of error and then determined the rejection field;
In the statistical software currently in use, we often give it to the$p$Value for hypothetical tests.
$p$The method is to calculate the theoretical boundary position of the rejected domain at this point, based on the values we observe, and then calculate the probability of error under our control in the first category based on the assumption that the sample value is on the edge of the rejected domain.$\alpha$ And then the value of the calculations is...$p$Value, therefore$p$Less than the required visibility can be judged to be significant.
Under the original assumption, the P-value is subject to the rule.$[0,1]$The flat distribution of the area requires extrapolation based on the testing of the statistics.
On sample capacity
Yes.$p$When the value is getting smaller, the results of the tests are getting more and more obvious.
It's very obvious.$p$The value is smaller as the sample capacity increases.$p$The value sample size is different from what we understand, and the hypothetical tests must be conducted in such a way as to give the size of the sample capacity, and in fact we need to discuss the statistical efficacy of a statistical function that we normally use the probability of a second type of error.$\beta$ . To calculate$1-\beta$ To be statistically effective as a hypothetical test.
Effects and minimum sample amount tested
In the hypothetical test,Minimum Technical Effects, MDE is the level of visibility in the given sample (in the case of the$\alpha$), statistical effectiveness ()$1-\beta$The conditions. Tests can be significant in terms of the minimum effect size.
The minimum detectable effect is: $$\Delta_{\min}=\sqrt{\frac{2\sigma^2}{n}}\cdot(z_{\alpha/2}+z_\beta)$$ The formula is based on the corresponding distribution of the statistically relevant content, the most common two-normal aggregate differences are known, the mean difference Z test is essentially similar to the rest. of which
- $z_{\alpha/2}$is the normal distribution fraction of the profile level. If it's a unilateral study,$z_\alpha$ If other tests are considered, the corresponding distribution will need to be changed.$\alpha$It's a significant level of probability of a first type of error, because it's studied here.$z$The symmetry of distribution will$z_{1-\alpha/2}$Writing$z_{\alpha/2}$
- $z_\beta$is the normal distribution fraction of the statistical effect, and the statistical effect is:$1-\beta$ The second type of error is a difference between probability and probability, because of the research here.$z$The symmetry of distribution will$z_{1-\beta}$Writing$z_{\beta}$
- $\sigma ^{2}$It's the overall variance.$n$is the number of single groups of samples.
If the T test is considered, it needs to be replaced with a T-distribution, which is still short-cut because of its symmetrical nature, but needs to be supplemented by the freedom of the T-distribution, which varies from one degree to another.
The card-side tests are used primarily for the independent testing of disaggregated data or for the test of proposed eugenics, and are not usually defined as minimal detectable effects directly by formulas similar to those of Z and T. - It's a Cramer.'s V, etc. measure the effect size. LikeDescriptive statistics and visualization / Correlation factors in the /Legation Tables $V$Related coefficients” section
The minimum detectable effect is determined by the level of visibility, statistical efficacy, sample size and standard deviation of the sample.
- When the sample volume increases or when the standard deviation decreases,$\Delta_{\min}$And then it drops, which means the experiment is more sensitive.
- The government is not a party to the law, but it is a party to the law.$\Delta_{\min}$The subsequent rise indicates that greater effects are needed to be tested.
- MDE helps researchers to assess the capabilities of the experimental design to ensure a reasonable sample allocation and optimization of resources.
Give us the minimum expected effect we want to detect in advance, based on the minimum detectable effect size formula.$\Delta$ , the formula can be converted to the formula used to calculate the minimum sample amount required, i.e. $$n=\frac{2\sigma^2}{\Delta_{\min}^2}\cdot(z_{\alpha/2}+z_\beta)^2$$
Non-parametric hypothetical tests
Introduction
Non-parametric hypothesis tests are part of non-parametric statistics And not the key point of parameter statistics is that Unknown for the overall distribution We're here to present a few non-parametric hypothetical tests that are very rare.A quote from non-parametric statistics Probability drawings are a classic non-parametric test, and probability drawings make the point from the normal general a straight line on the drawings, and he is the normal QQQ (Quantile-Quantile) map to test whether the sample from the normal aggregate is studied using a fractional number. In addition, we'll present the Pesason Carp's proposed eugenicity test, which is the non-parametric test on which the comparison is based. This is a way to test whether a sample comes from a distribution we've chosen. It's a classic problem in non-parametric statistics.
The idea of Peason's party to match the eugenic test
We divide the results of random experiments into a complete event. Group $A_{1},A_2,...,A_k$ That's the rim up.{i=1}^{\kappa}\mathbf{A}\, \A\A=A=mathbf}, i,j=1,2, \mathbf{L},k.$ We're assuming that.$H_0$Here we go. $$p_i=P(A_i)$$ So there's an analysis of the actual and theoretical frequency. $$\chi^{2}=\sum_{i=1}^{k}\frac{(f_{i}-np_{i})^{2}}{np_{i}}$$ When he was big enough to deviate a lot, then the assumption was fake, the assumption was true.
Theories of the Peason Cartridges to the Presbyterian Presumptuous Test
When?$H_0$For real and$n$When you're big enough to count. $$\chi^{2}=\sum_{i=1}^{k}\frac{(f_{i}-np_{i})^{2}}{np_{i}}=\sum_{i=1}^{k}\frac{f_{i}^{2}}{np_{i}}-n$$ Close to obedience.$\chi^{2}(k-1)$ From this test of the distribution of statistics, we can export the rejection field we need to use. Because we explained earlier that when the original assumption was false when the deviation was large, it was necessary to select a single-side hypothesis test to reject the field as being a non-existent one-sided one-sided test.$\chi^{2}\geq\chi_{\alpha}^{2}(k-1)$
If we're choosing the hypothesis,$H_0$♪ And if ♪$X$The distribution function contains unknown parameters First, use the sample to obtain the maximum semblance of the unknown parameter, using the estimate as the parameter value. We'll come to conclusions.$H_0$For real and$n$When you're big enough to count. $US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US$US${i}\right)^{2}}{n\hat{p}{i}}=\sum_{i=1}^{k}\frac{f_{i}^{2}}{n\hat{p}-$US$$US$US$US$US$US$US$US$US$US$US$$US$$US$$US$$US$$US$$US$$US$$$US$$$US$$$$$$$$$US$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$$ Close to obedience.$\chi^{2}(k-1-r)$ of which$r$is the number of unknown parameters contained in the distribution function Deny domain $\q\chi{\alpha}^{2}(k-1-r)$
Operating habits
Using the Peason Cartrix Probability Test will generally require assurance that the following requirements are used to classify the samples
- Large sample generally considered $n>50$
- Requesting theoretical frequency for groups $np i>5$
- General data are divided into 7-14 groups, which can be smaller than 7 groups to satisfy Article II. The proper grouping of all sample data is in fact the core of the Peason Carpenter's proposed eugenicity test.
Additional elements
We're testing the calculator's proposed merits only for the theoretical distribution with limited values.
If you want to process a continuous variable, we need to divide the compartments, amend them to a limited range and calculate the probability size of the compartments.
The idea of the Person card for the quality test is very well applied in the luminum table, and we will introduce it in the related and other analyses of the luminous table, all of which have a classification, frequency structure.
- Title: Mathematical Statistics: Populations, Samples, and Sampling Distributions
- Author: Hyacehila
- Created at : 2023-03-18 13:28:08
- Link: https://hyacehila.github.io//blog/2023/03/18/mathematical-statistics-notes/
- License: This work is licensed under CC BY-NC-SA 4.0.