Multivariate Statistics Introduction: Random Vectors, Covariance Matrices, and Multivariate Normal Distributions
Additional knowledge of random vectors
In the study of linear statistical models, random vectors appear in various places, where there is some addition to random vectors that are not described in detail in the probability theory.
Random vector numerical feature
For random vectors, we define his average as $$E(X)=(EX_1,\cdots,EX_n)^{\prime}$$ If $Y=AX+b$ Well... $$E(Y)=AE(X)+E(b).$$ And there is. $${E}(АXB)=\mathrm{A}{E}(X)B$$
For random vectors, we define his compromise matrix as $$\mathrm{Cov}(X)=E[(X-EX)(X-EX)^{\prime}].$$ It's easy to see that the matrix is symmetrical, and each dollar is the one that's often seen in probabilities, and that's the difference on the diagonal.$Var(X_{i})$ The balance can be used to determine the relevance. (Only linear correlation can be judged, the difference and the related coefficient do not mean independence from each other) There's something about the matrix. $$\mathrm{trCov}(X)=\sum_{i=1}^{n}\mathrm{Var}(X_{i})$$ The positive properties of the study matrix secondary are: $$协方差矩阵是对称的半正定矩阵$$ If$Y=AX$ $$\operatorname{Cov}(\boldsymbol{Y})=A\operatorname{Cov}(\boldsymbol{X})\boldsymbol{A}^{\prime}.$$ It's extended to the collusive matrix between two random vectors. $$\mathrm{Cov}(X,Y)=E[(X-EX)(Y-EY)^{\prime}].$$
It's common to add to the computational nature of $$\mathrm{Cov}(AX,BY)=A\mathrm{Cov}(X,Y)B^{\prime}.$$
Random Vector Secondary
Defines the default of the random vector as $$X^{\prime}AX=\sum^{n}\sum^{n}a_{ij}X_{i}X_{j}$$ Our random vector replaced the original.$X$, instead of replacing the original symmetric matrix with a collusive matrix$A$ Please note that the recoil of the random vector is essentially a random variable.
For this random variable, we can give the desired formula. $$\text{设 }\operatorname{E}(X)=\boldsymbol{\mu},\operatorname{Cov}(X)=\boldsymbol{\Sigma},\text{则}E(X^{\prime}AX)=\boldsymbol{\mu}^{\prime}A\boldsymbol{\mu}+\mathrm{tr}(A\boldsymbol{\Sigma}).$$ $tr$ The operator is the sum of the diagonal elements used to trace the matrix. Based on this formula, we can provide some inferences. $$\text{f}\mu=0,\text{E}'AX)=\mathrm{tr}(A\Sigma);$$ $$\\textrm=(\bardsymbol{\sigma^I},\text{E^({\prime}A\mu+sigma}(}n}a}ii})=mu^}A\prime^A;$$ $$\text \mathbf(=mathbf} \mathbf},\mathbf(=mathbf{I},\text{E(\bardsymbol{X}'AX})=\mathrm{tr}\boldsymbol{A}.$$
Normal Random Vector
Basic knowledge
We've studied the normal random variable in the front, which means the random variable that follows the normal distribution.$N(\mu,\sigma^{2})$ $$f(x)=\frac{1}{\sqrt{2\pi}\sigma}\mathrm{e}^{-\frac{1}{2\sigma^{2}}(x-\mu)^{2}},-\infty<x<+\infty $$ 他的二维推广应该是 $$f(x_1,x_2)=\frac1{2\pi\sigma_1\sigma_2\sqrt{1-\rho^2}}e^{-\frac1{2(1-\rho^2)}\left(\frac{(x_1-\mu_1)^2}{\sigma_1^2}-2\rho\cdot\frac{x_1-\mu_1}{\sigma_1}\cdot\frac{x_2-\mu_2}{\sigma_2}+\frac{(x_2-\mu_2)^2}{\sigma_2^2}\right)}$$ 现在,我们自然的推广这个一维的形式到随机向量上 定义:$\text{设 }n\text{ 维随机向量 }{X}=(X_1,\cdotp\cdotp\cdotp,X_n)`\text{具有密度函数}$ $$f(x) =\frac^(\det\bardsymbol{\(\mathrm}e}\frac}}(x-\bardsymbol})/\bardsymbol{\Sigma^(x-\bardsymbol{\mbol{\}}, $$$ We usually write it down.$N(\boldsymbol{\mu},\boldsymbol{\Sigma})$ These are their distribution parameters. Theorem: if the normal vector of the top distribution is satisfied $E(X)=\mu,\quad\mathrm{Cov}(X)=\boldsymbol{\Sigma}.$
As can be seen from the top core description, the multiple normal distribution is his average vector.$\mu$A matrix of differences$\Sigma$ Absolutely. The other one.$\mathbf{\mu}=\mathbf{0},\boldsymbol{\Sigma}=\mathbf{I}$ At the time, we called it a multiple standard normal distribution.
His feature function is $$\Phi_{X}\left(t\right)=\exp\left[\mathrm{i}t^{\prime}\mu-\frac{1}{2}t^{\prime}\Sigma t\right]$$ Characteristic functions can be determined with distribution functions (density functions)
Decision Theorem
We know that each weight of the normal vector is a normal variable, but the combined distribution of the two normal variables is not necessarily a normal vector, so we give the following determinations.
Random vector $X=(X_1,\quad X_2,\cdots,\quad X_n)^T$ Obey. $n$ Normal distribution $N(a,B)$ The only necessary condition is any linear combination of it.$Y=\sum_{i=1}^nl_iX_i$Obey the one-dimensional normal distribution.$$N(\sum_{i=1}^nl_ia_i,\sum_{i=1}^n\sum_{k=1}^nl_il_k\operatorname{cov}(X_i,X_k))$$ Based on this theorem, it is easy to see the previously proven conclusion that the joint distribution of the two separate normal variables must be the normal vector.
Decomposition Theorem
We talked about the decomposition of the dichotomy when we studied it, and now we're promoting it. A matrix with multiple normal distributions.$\Sigma$Meet the split diagonal matrix form $$\left.\bardsymbol{11}&\boldsymbol{0}\\boldsymbol{0}&\boldsymbol{\Sigma}{22}\end{matrix}\right.\right],$$ 那么我们可以进行对应的分解 $$\boldsymbol{X}=\begin{pmatrix}\boldsymbol{X}1\\boldsymbol{X}2\end{pmatrix},\quad\boldsymbol{\mu}=\begin{pmatrix}\boldsymbol{\mu}1\\boldsymbol{\mu}2\end{pmatrix},$$ 这时候我们能验证到 $$f(x)=f{1}(x{1})f{2}(x{2}),$$ 当 $$\begin{aligned}f_1(x_1)&=\frac{1}{(2\pi)^{\frac{m}{2}}\det\Sigma_{11}}\mathrm{e}^{-\frac{1}{2}(x_1-\mu_1)/\Sigma_{11}^{-1}(x_1-\mu_1)},\f_2(x_2)&=\frac{1}{(2\pi)^{\frac{n-m}{2}}\det\Sigma_{22}}\mathrm{e}^{-\frac{1}{2}(x_2-\mu_2)\Sigma_{22}^{-1}(x_2-\mu_2)}.\end{aligned}$$ 其实我们在二元正态分布里面提到的相关系数$\rho$ 就是决定协方差矩阵副对角元是否为0的关键 根据上面的推论我们能够给出一个重要的正态分布分解定理 $$\begin{aligned}\text{(a) setx)\sim\(mu,\sigma),\text{and X\text{and \mu\text},\text{and}\boldsymbol{Sigma}\text{syggle},\text{x}\bardsymbol{X} i\sim (\boldsymbol{\mu}i,\boldsymbol{\Sigma}{ii}, i=1,2\text{and independent}. \end{aligned}$$ $$\begin{aligned} (\mathrm{b}\text{\bardsymbol{\Sigma}2\bardsymbol{I},\text{and }\bardsymbol}X}(X 1,\cdots,X n)^prime},\bardsymbol{mu}}(\mu 1,\cdots,\mu^n)^,\text{x}, \text{x^x(i, \sigma^2, i)&=1,\cdots, n\text{and independent} .\end{aligned} This theorem tells us that independence and irrelevance are equal to the weight of a normal vector, and we just need to verify that we have a zero difference to guarantee independence.
Aggregated normal vector
We demand that all normal variables be independent. $$\sum_{r=1}^na_rX_r\sim N{\left(\sum_{r=1}^na_r\mu_r,\sum_{r=1}^na_r^2\sigma_r^2\right)}$$ We want all normal vectors to be independent. $$\sum_{r=1}^nX_rA_r\sim N_m\left(\sum_{r=1}^n\mu_rA_r,\sum_{r=1}^n(A_r^{\prime}\sum_rA_r)\right)$$
The change of the dimension to the normal vector
Let's start with a theory of the core of normal vector variation. $$\begin{gathered} \text{ 设 n 维随机向量}\mathbf{X}\sim N(\boldsymbol{\mu},\boldsymbol{\Sigma}),\boldsymbol{A}\text{ 为}n\times n\text{ 非随机可逆阵},\boldsymbol{b} \ \text{为 }n\times1\text{ 向量,记 }Y=AX+b\text{ ,则} \ \boldsymbol{Y}\sim N(\boldsymbol{A}\boldsymbol{\mu}+\boldsymbol{b},\boldsymbol{A}\boldsymbol{\Sigma}\boldsymbol{A}^{\prime}). \end{gathered}$$ It's obvious that we can use some of the more specific variations to achieve some of the special effects with random reversible matrices. Amending Arguments $$\text{设 }X\sim N(\boldsymbol{\mu},\boldsymbol{\Sigma}),\text{则 }Y=\boldsymbol{\Sigma}^{-\frac{1}{2}}\boldsymbol{X}\sim N(\boldsymbol{\Sigma}^{-\frac{1}{2}}\boldsymbol{\mu},\boldsymbol{I}).$$ Changed to render irrelevant all normal weights that could have been relevant, and the difference was one. Switching $$\text{设 }\boldsymbol{X}\sim N_n(\boldsymbol{\mu},\sigma^2\boldsymbol{I}),\boldsymbol{Q}\text{ 为 }n\times n\text{ 正交阵,则 }\boldsymbol{Q}\boldsymbol{X}\sim N_n(\boldsymbol{Q}\boldsymbol{\mu},\sigma^2I)$$ It's a transformation that guarantees the original independence and equation. Regenerative $$\begin{aligned}\text{}\bardsymbol{x}-\bardsymbol{n(\bardsymbol},\bardsymbol{\Sigma},\bardsymbol{\chi}\bardsymbol{X},\barysymbol{x}, \bardsymbol}Sigma}\text{segmbol}&=\begin{bmatrix}\boldsymbol{X}1\\boldsymbol{X}2\end{bmatrix},\boldsymbol{\mu}=\begin{bmatrix}\boldsymbol{\mu}1\\boldsymbol{\mu}2\end{bmatrix},\boldsymbol{\Sigma}:=\begin{bmatrix}\boldsymbol{\Sigma}{11}&\boldsymbol{\Sigma}{12}\\boldsymbol{\Sigma}{21}&\boldsymbol{\Sigma}\end{bmatrix},\text{it}\boldsymbol{x}The text is called m\timestext, while \bardsymbol{\Sigma}\text{m\times\boldsymbol{m}\text{format, \boldsymbol{x} 1-\boldsymbol{m}1,\boldsymbol{\Sigma}That's right. This theorem tells us that any dimension of a normal vector is also a normal vector.
Change in the normal vector of changes in dimensions
The core theorem is just a small change ahead. $US$\begin{array}\text{set}X\sim n\left(\bardsymbol},\bardsymbol{\right),\bardsymbol{A}text{m\tates n\text{format)}m\lex(m\left}<{\bord0\shad0\alphaH3D}You know, \boldsymbol \boldsymbol \sim\boldsymbol}mbol}mft (\boldsymbol{, \boldsymbol{, \boldsymbol{A} \boldsymbol{\Sigma}\boldsymbol{A^ {\prime}\right.\end{array} $ This theorem applies to the increase in dimensions, which is $m.>Case of $0.00 It makes him look at normal processes at random. Lower the dimension to 1 $$\begin{array}{cc}&\text{x\x\sim{n(\bardsymbol{mu},\bardsymbol{Sigma},\bardsymbol{c}\text{n\text{non-zero vectors} & ^(\prime}\bardsymbol{x}\sim(\pendsymbol^c},\bardsymbol{c}}.\end{array} $ Linear combinations of normal vectors are normal variables And turn the extraction into one. $$\begin{aligned}\quad&\text{x\sim n\left(\bardsymbol{mu},\bardsymbol(sigma}\right),\bardsymbol{mu}=left(,\cdots,\bardsymbol}n\right)^{\prime},\boldsymbol{\Sigma}=\left(\boldsymbol{\sigma}\right, \text{&N\left(\boldsymbol{\mu}i\right.,\\sigma{ii}\left.\right),&i =1,\cdots, n.end{aligned} This theorem tells us that each weight of a normal vector is a normal vector. But please note that the inverse theorem is not valid.
Distribution of conditions and expectations
$ \\begin{aligned}\text{Theoretical 2.3.2]&\text{set}X=begin{bmatrix}X^x}{p-r}\sim N{\rho}(\mu,\Sigma)(\Sigma>0), \\text{, then \\text{to}\text{time}, X^(1)}\text{condition distribution} (X^(1)}XX^(2)}&\sim N_{r}(\mu_{1,2},\Sigma_{11,2}),\end{aligned}$$ 其中 $$\{\mu_{1\cdot2}=\mu^{(1)}+\Sigma_{12}\Sigma_{22}^{-1}(x^{(2)}-\mu^{(2)})}\~~~{\Sigma_{11\cdot2}=\Sigma_{11}-\Sigma_{12}\Sigma_{22}^{-1}\Sigma_{21}.}$$
Corresponds. $$(X^{(1)}|X^{(2)})\sim N_{r}(\mu_{1,2},\Sigma_{11,2})$$ Name $$\mu_{1\cdot2}=\mu^{\left(1\right)}+\sum_{12}\sum_{n2}^{-1}\left(x^{\left(2\right)}-\mu^{\left(2\right)}\right)$$ It's an expectation.$E\left(X^{\left(1\right)}|X^{\left(2\right)}\right)$
Carside distribution
$\begin{aligned}&\\text{order}Z 1, Z 2, \\cdots, Zn\text{and}Z\text{is the standard normal distribution random variable for independent and distributed distribution,}&\text \ and =1^2Z2^cdots+\n^text{, which we call X is subject to the caloric distribution of freedom }end{aligned} It's easy to know that the density function of the calf distribution is $g(x)=\begin{cases}\dfrac{1}Gamma\Big(\dfrac{n}2^ {\mathrm{e}-\frac{2}}&\text \x>0,\\0,&\text \xleqslant0\end{cases}$$ 定理 $$\text{set}\bardsymbol{sim\mathrm{0}bardsymbol{\symbol{sim\chi 2.$$ 定理 $$\begin{aligned}X\sim n(0,I n), A\text{n\times\text{r\text{symmetric},\text{\text{}\matbf{A^mathbf{A}\mathbf{A}\text{second}X^'}AX\sim\chi{r}^{2}\end{aligned}$$ 定理 $$\begin{aligned}&\text{X}\sim n (\mathbf}, \mathbf{I} n), A\text{n\tmes\text{symmetric array}, B\text{m\times\text}. \text{Fang}&Other Organiser$$ 定理 $$\begin{gathered} X\sim n (mathbf{I}), A\text{and B}n\text{symmetrical},\text{and}A\bardsymbol{B}=mathbf,\text{second} \\text{type X'AX and X'BX Independent. ♪ I'm sorry ♪ I'm sorry.
Multiple statistical bases
Parameter estimates for multiple normal distribution
Additional definitions
Common digital features of multiple normals in general have been introduced in probabilistic theory. Primary probability theory It's enough to know the next few.
- Mean vector
- Arguments
- Related coefficient arrays
Additional definitionsSample deviation matrixIs:
$ \begin{aligned}A&=\sum_{a=1}^n\left(X_{(a)}-\overline{X}\right)\left(X_{(a)}-\overline{X}\right)^{\prime}=X^{\prime}X-n\overline{X}\overline{X^{\prime}}\&=X^{\prime}\left[I_n-\frac{1}{n}\mathbf{1}n\mathbf{1}n^{'}\right]X\xrightarrow{\mathrm{def}}\left(a{ij}\right){p\times p}\end{aligned}$$
其中
$$a_{ij}=\sum_{a=1}^{n}\left(x_{ai}-\overline{x}{i}\right)\left(x$-overline{x}
That's where the unit matrix didn't separate $n.or n-1$
Largely specious estimates of averages and alignment matrix
Now, let's figure out how to estimate two core parameters in a multi-state normal analysis. $\mu,\Sigma$
Yes.$\Sigma$Yeah. $ \begin{aligned} L\left (mu, Sigma\right)& =-\frac{np}{2}\ln2\pi-\frac{n}{2}\ln\left|\Sigma\right| \ &-\frac{1}{2}\mathrm{tr}\left[\Sigma^{-1}\sum_{i=1}^{n}\left(x_{\left(i\right)}-\mu\right)\left(x_{\left(i\right)}-\mu\right)^{\prime}\right] \ &=C-\frac{1}{2}\mathrm{tr}\left[\Sigma^{-1}A+n\Sigma^{-1}\left(\overline{X}-\mu\right)\left(\overline{X}-\mu\right)^{\prime}\right] \ &=C-\frac{1}{2}\mathrm{tr}\left(\Sigma^{-1}A\right)-\frac{n}{2}\left[\left(\overline{X}-\mu\right)^{\prime}\Sigma^{-1}\left(\overline{X}-\mu\right)\right] \ &\leqslant C-\frac{1}{2}\mathrm{tr}\left(\Sigma^{-1}A\right). \end{aligned}$$ 等号只在$\mu=\overline{X}$ 的时候取到 也就是 $$\ln L\left(\overline{X},\Sigma\right)=\max_{\mu}\ln L\left(\mu,\Sigma\right).$$
We can prove it with similar ideas. When?$$\hat{\Sigma}=\frac{1}{n}A$$Sometimes. $ \\mathrm}L\left(overline{x),\right.) =max {\{\xXx},\Sigma=>♪ L\left (overline{X}, Sigma\right): $ That's how the parameters are estimated.
The export of other very seemingly significant estimates and the related coefficients are very similar.
We've been telling you how to do this in the course of learning the math. $\mu,\Sigma$ It's a very similar estimate.$\phi(\mu,\Sigma)$ The whole idea is a function of nature.
Well, we can naturally export the MLE. There's a huge estimate of what's going on. What's that?{ij}=\frac{1}{n}\sum{t=1}^{n}\left(x_{ti}-\overline{x}{i}\right)\left(x{tj}-\overline{x}{j}\right)=\frac{1}{n}a{ij}.$$ 根据相关系数定义式可以给出 $$r_{ij}=\frac{\hat{\sigma}{ij}}{\sqrt{\hat{\sigma}{ii}\cdot\hat{\sigma}{jj}}}=\frac{a{ij}}{\sqrt{a_{ii}\cdot a_{jj}}}.$$
Nature of the estimate
It's very apparent that the estimates have the following characteristics.
- $\overline{X}\sim N_{p}\left(\mu,\frac{1}{n}\sum\right)$
- $\overline{X}和A相互独立$
- $A\xrightarrow{d}\sum_{i=1}^{n-1}Z_{i}Z_{i}^{\prime},\text{其中}Z_{1},\cdots,Z_{n-1}\text{独立同 }N_{p}(0,\Sigma)\text{分布}$
- $P\left{A>0\right}=1\Leftrightarrow n>p.$
- The value of the average is very much estimated, neutral, effective, compatible, gradual normality, sufficient statistically.
- It's a very similar estimate of the alignment matrix, the correction is neutral, effective, compatible, gradual normality is a sufficient measure of statistics.
- Estimation of relevant coefficients is gradual and neutral
- The estimate of the simple average against the reconciliation matrix does not require a normal sum
Total sample distribution of multiple normals
In a one-dollar normal general, the hypothetical test involves an overall problem of differential analysis of the two aggregates that have been extended to multiple aggregate averages; we've designed a lot of fine statistics to help us solve these problems; here we extend those sample statistics to multiple normal aggregates, including common ones.$\chi^2,t,F$Multiple normal forms of three statistics
Wishart$W$Distribution
We know that two important statistics are being given. Medium $$\overline{X}\sim N_{p}\left(\mu,\frac{\sum}{n}\right).$$ So, what's the estimate? $S=\frac{1}{n-1}A$ What's the distribution?
Definitions$X(a)\sim ~N_p\left(0,\Sigma\right)\left(\alpha=1,\cdots,n\right)$Independent, remember.$X=$ $\left(X_{\left(1\right)},\cdots,X_{\left(n\right)}\right)^{\prime}为n\times p矩阵,则称随机阵$ $$W=\sum_{a=1}^{n}X_{\left(a\right)}X_{\left(a\right)}^{\prime}=X^{\prime}X$$ It's distributed between the West and the West.$\sim W_{p}\left(n,\Sigma\right).$
When?$p=1$There are times. $$W=\sum_{a=1}^{n}X_{\left(a\right)}^{2}\sim\sigma^{2}\chi^{2}\left(n\right),$$ In other words, this is the spread of the calorie distribution in the total of multiple normals.
Hotelling$T^2$Distribution
We can come straight to the conclusion. It's original.$t$Extension of distribution
Set $X\sim N_{p}\left ( 0, \Sigma \right )$,W\simW p}\left(n,Sigma\right)\left(\Sigma)>0,\right.$ $\gqslant p) and X and W are independent of each other.$为霍特林$T^{2}$统计量,其分布称为服从$n$个自由度的$T^2$分布,记为$I'm sorry. More generally, if$X\sim N_{\rho}\left(\mu,\Sigma\right)\left(\mu\neq0\right)$, or$T^2$The distribution is non-centre-hotrin.$T^2$Distribution, as$T^2\sim T^2(p,n,\mu).$
Wilks.$A$Distribution
When we estimate the parameters, use them. $A$ As an estimate of the coordinated matrix; we defined the broad range in multiple statistics as the array of the coordinated matrix Multiple statistical analysis and the “wide range” section
Set $A \left( , , , , , , , , , )>0,n_{1}\geq\right.$$p),且A_{1}与A_{2}独立,则称广义方差之比$ $$\Lambda = \\right}{\left|A +right}{\right}A}text}Stat{Stat{Stat{Stat{Stat{Stat{Stat{Stat{Specific }Specific }Wilx's Distribution, as }Lambda\sim\Lambda\ft(p, n , n}right}.\ when p=1, \LambdaStat\Species is precisely the parameter in the one-dollar statistics is n {(m}\n}2}text{Spect\text{Spect{(}beta}, n}, {2}).
Special conclusions
$$\Delta\left(p,n,1\right)\frac{d}{1+\frac{1}{n}T^{2}\left(p,n\right)}$$
$$T^{2}\left(p,n\right)=n\cdot\frac{1-\Lambda\left(p,n,1\right)}{\Lambda\left(p,n,1\right)}$$
$$\frac{n-p+1}{np}T^{2}=\frac{n-p+1}{p}\frac{1-A}{\Lambda}=F\left(p,n-p+1\right).$$
Multiple statistical extrapolations
Insulation of individual aggregate mean vectors
Want to test $$H_0:\boldsymbol{\mu}=\boldsymbol{\mu}_0,\quad H_1:\boldsymbol{\mu}\neq\boldsymbol{\mu}_0$$
The AC matrix is known.
Construct statistics $T 0^2=(\overline{x}-\bardsymbol{0}^prime}\left(\frac1n\bardsymbol{\right)^(\overline{bardsymbol{)}n(\overline{\bardsymbol})=(\barysymbol{x}}}0)^{\prime}\boldsymbol{\Sigma}^{-1}(\overline{\boldsymbol{x}}-\boldsymbol{\mu}0)$$ 原假设为真时则有 $$T{0}^{2}\sim\chi^{2}\left(p\right)$$ 使用单侧检验有 $$\text\\alpha^2(p), \\text{rejected}H $0
The Accompany Matrix is unknown
Construct Statistics (Hotlyn)$T^2$Statistics) $$T^{2}=n\left(\bar{x}-\mu_{0}\right)^{\prime}S^{-1}\left(\bar{x}-\mu_{0}\right)$$ When the assumption is true, there is. $$\frac{n-p}{p(n-1)}T^{2}\sim F(p,n-p)$$ One-sided check. $$\text{若}\frac{n-p}{p\left(n-1\right)}T^2\geqslant F_a(p,n-p),\text{则拒绝 }H_0$$
Large sample extrapolation
The previous proofs are based on multiple normal assumptions, but sometimes multiple normal assumptions are not satisfied; but when the sample capacity is large enough, multiple centers can solve problems. When the sample size is large enough, the following approximation relationships occur. $T 0^2=(\overline{x}-\bardsymbol{0}^prime}\left(\frac1n\bardsymbol{\right)^(\overline{bardsymbol{)}n(\overline{\bardsymbol})=(\barysymbol{x}}}0)^{\prime}\boldsymbol{\Sigma}^{-1}(\overline{\boldsymbol{x}}-\boldsymbol{\mu}0)$$ $$T{0}^2=n(\overline{x}-\mu)^{\prime}S^{-1}(\overline{x}-\mu)$$ $$TOther Organiser All hypothetical tests and area estimations follow the formula above.
$$S^{-1}=A^{-1}(n-1)$$
It seems to be more than statistics.
A very large number of important tests of statistics in multiple statistics are derived from the maximum approximation principle, rather than by extending the largest approximation principle in mathematical statistics. And here we're presenting the principle of seemingly statistical and maximum comparison.
Set $p$ The density function of the sum of the elements is$f\left(x,\theta\right)$, where$\theta$is an unknown parameter, and$\theta\in\Theta$ $\left(参数空间\right),又设\Theta_{0}是\Theta 的子集,我们希望对下列假设:$ $$H_{0}:\theta\in\Theta_{0},H_{1}:\theta\in\Theta_{0}$$ To judge, that's a hypothetical test.
From General$X$The extraction capacity is$n$The sample.$X_{(t)}(t=1,\cdots,n).$Use the joint density function of the sample $$L\left(x_{\left(1\right)},\cdots,x_{ \left(n\right)};\theta\right)=\prod_{t=1}^{n}f\left(x_{\left(t\right)};\theta\right)$$ As$L\left(X;\theta\right)$and calls it the apparent function of the sample
Include statistics $$\lambda=\max_{\theta\in\theta_{0}}L\left(X;\theta\right)/\max_{\theta\in\theta}L\left(X;\theta\right),$$ It's a sample.$X_{(t)}\left(t=1,\cdots,n\right)$function, commonly called$\lambda$It's like we're out of statistics.
Known by the principle of maximum approximation if$\lambda$It's too small.$H_0$This sample was observed for real.$X_{(\omega)}(t=1,...,n)$Probability ratio $H_0$ To observe this sample when not true$X_{(\omega)}$ It's much less likely.$H_0$Not established
According to traditional methods, we need to calculate the exact sample distribution that appears to be statistically comparable to the hypothetical test results, but multiple statistics are too complex and we give a large sample that resembles the theorem. When sample size n is large $$-2ln\lambda=-2ln\left[\max_{\theta\in\Theta_{0}}L\left(X;\theta\right)/\max_{\theta\in\Theta}L\left(X;\theta\right)\right]$$ It's like obeying freedom.$f$Yes.$\chi^2$Distribution, where$f=\Theta$Number of dimensions$-\Theta_0$Dimensions (i.e., the margin where freedom is restricted)
Extrapolation of two overall averages
Aligning Matrix
Two totals. $N_{p}(\mu_1,\Sigma),N_{p}(\mu_2,\Sigma)$ Take two separate samples.$x_{n_1},y_{n_2}$ We want tests. $H 0:\bardsymbol}1=\boldsymbol{\mu}2,\quad H_1:\boldsymbol{\mu}1\neq\boldsymbol{\mu}2$$ 我们可以自然的从一元统计的情形中得到霍特林$T$统计量 $$\begin{aligned}T^2=&\left(\frac{1}{n_1}+\frac{1}{n_2}\right)^{-1}\left(\overline{x}-\overline{y}\right)^{\prime}S^{-1}\left(\overline{x}-\overline{y}\right)\=&\frac{n_1n_2}{n_1+n_2}\left(\overline{x}-\overline{y}\right)^{\prime}S^{-1}\left(\overline{x}-\overline{y}\right)\end{aligned}$$ 当原假设成立的时候 $$\frac{n{1}+n{2}-p-1}{p\left(n{1}+n{2}-2\right)}T^{2}\sim F(p,n_{1}+n_{2}-p-1)$$ 可以自然的进行单侧检验量 方向和我们在前面介绍的一样 其中 $$S^{-1}=(\frac{A_1+A_2}{n_1+n_2-2})^{-1}$$ There's a significant difference between the two mean vectors, which doesn't mean there must be a significant difference between them.; that is to say, the equal rejection of the mean vector does not mean that we will be able to detect significant differences when we test each weight separately;
But this difference is still the main reason for the average vector difference, and it is customary for us to examine the significant differences between the weights separately after testing the significant differences in the overall vector.
It's a pair.
It is assumed that two samples are independent in certain circumstances; in a number of experiments, two samples may exist in pairs but are not independent; and pairing data often leads to better statistical extrapolations You! $$d_i=x_i-y_i,\quad i=1,2,\cdots,n$$ There is.$d_{i}$Obey the new distribution. $$N_{p}\left(\delta,\Sigma\right)$$ of which $$\delta=\mu_{1}-\mu_{2}$$ So the original assumption is... $$H_0:\boldsymbol{\mu}_1=\boldsymbol{\mu}_2,\quad H_1:\boldsymbol{\mu}_1\neq\boldsymbol{\mu}_2$$ To $$H_0:\boldsymbol{\delta}=\boldsymbol{0},\quad H_1:\boldsymbol{\delta}\neq0$$ The problem turned into a single whole.
Comparison of multiple aggregate averages (multiple variance analysis)
Assumptions $$H_0:\boldsymbol{\mu}_1=\boldsymbol{\mu}_2=\cdot\cdot\cdot=\boldsymbol{\mu}_k,\quad H_1:\boldsymbol{\mu}_i\neq\boldsymbol{\mu}_j,\text{至少存在一对 }i\neq j$$ One of ours.$\mu$They're all vectors, not single-dollar variables presented in the variance analysis.
Remember $$T=SST=\sum_{i=1}^k\sum_{j=1}^{n_i}(x_{ij}-\overline{x})(x_{ij}-\overline{x})^{\prime}$$ $$E=SSE=\sum_{i=1}^k\sum_{j=1}^{n_i}(x_{ij}-\overline{x_i})(x_{ij}-\overline{x_i})^{\prime}$$ $$H=SSTR=\sum_{i=1}^kn_i\left(x_i-\overline{x}\right)\left(x_i-\overline{x}\right)^{\prime}$$ There is. $$T=E+H$$ Use the approximation test to get Wilks statistics. $$\Lambda=\frac{|E|}{|E+H|}$$ When the original assumption is true, the statistics follow the parameters.$(p,k-1,n-k)$The Wilkes Distribution
The rule for rejection is: $$\text{若}\Lambda\leqslant\Lambda_{1-a}(p,k-1,n-k),\text{则拒绝 }H_0$$ The absence of significant differences in multiple tests does not mean that their weights do not differ significantly; in turn, they are the same; in custom, if multiples detect significant differences, we still have to do a one-dollar variance analysis to see where the differences generally come from.
The Inference of the Arranged Matrix
We don't take into account the hypothetical tests of the single-sum matrix because it's complicated and not unique.
Assumptions $H 0:\bardsymbol{\1=\boldsymbol{\Sigma}2=\cdots=\boldsymbol{\Sigma}k,\quad H_1:\boldsymbol{\Sigma}i\neq\boldsymbol{\Sigma}j,\text{at least one pair exists}i\neq j$$ 修正的似然比统计量为 $$\lambda=\frac{\prod{i=1}^k|S_i|^{(n_i-1)/2}}{|S_p|^{(n-k)/2}}$$ 其中 $$S_i=\frac1{n_i-1}\sum{j=1}^{n_i}{(x{ij}-\bar{x_i})\left(x{ij}-\bar{x_i}\right)}^{\prime}$$ $$S_p=\frac1{n-k}\sum{i=1}^k{(n_i-1)S_i}=\frac1{n-k}E$$ 构造$M$统计量有 $$M=-2\mathrm{ln}\lambda=\left.(n-k)\ln|S_p|-\sum_{i=1}^k{(n_i-1)}\ln|S_i\right. $$ 当原假设为真的时候 $$u=(1-c)M$$ 近似服从自由度为$\frac{1}{2}(k-1)p(p-1)$ 的卡方分布 其中 $$c=\Big(\sum_{i=1}^{k}\frac{1}{n_{i}-1}-\frac{1}{n-k}\Big)\frac{2p^{2}+3p-1}{6\left(p+1\right)\left(k-1\right)}$$ 拒绝规则为 $$\text \text \ if \geqslant\chi\
Preparatory knowledge of multiple statistics
Difference and Broadest
In the one-dollar probability theory, the variance is used to measure the dissegregation of a random variable (variability level) and we explain the difference in the sample in the mathematical statistics; the multiple parts of the probability theory are used to explain the difference before the two random variables, although not explained in the mathematical statistics, but it is not difficult to introduce the difference in the sample (which also needs to be corrected)$n-1$The concept of mathematical software can also serve us; The ACSM is a matrix; how to measure the total variability of random vectors in one number is the question we have to answer;
Total variance
The total variance is defined as $$\mathrm{tr}\left(\boldsymbol{\Sigma}\right)=\sum_{i=1}^{p}\sigma_{ii}$$ He didn't consider the effects of the correlation between the samples; that was one of his flaws, but we would still use it in some places back there. $p=1$And then it turns into a difference.
Broad Difference
The most commonly used definition for wide range is $$\left|\Sigma\right|$$ The broad variance takes into account the correlation between variables; however, it may be misleading that the same broad difference is derived from two completely different matrixes. $p=1$And then it turns into a difference.
O'Hara and Ma's.
Orchid distance.
$p$The distance of two o'clock in Vior's space is $$d(x,y)=\sqrt{(x_{1}-y_{1})^{2}+(x_{2}-y_{2})^{2}+\cdots+(x_{p}-y_{p})^{2}}$$ In order to avoid rooting, we're used to using square distance. $$d^2(x,y)=(x-y)^{\prime}(x-y)=(x_1-y_1)^2+(x_2-y_2)^2+\cdotp\cdotp\cdotp+(x_p-y_p)^2$$ In the case of different weight units, the OSD is of little practical value; Even where units are the same, we need to standardize the data (usually z-scores); otherwise the calculation of the Euros distance is not meaningful; standardization is also the basis for the data pre-processing process.
Ma's distance.
When there is a linear relationship between random variables; the oscillation distance loses its original role and does not allow for the correct judgement of the isolation and proximity;
The Ma's distance was proposed to solve the problem; his essence was to rotate the axis, so that the correlation would disappear and then use the result of the A's distance.
Two vectors are defined as follows: $$d^{2}\left(x,y\right)=\left(x-y\right)^{\prime}\boldsymbol{\Sigma}^{-1}\left(x-y\right)$$ The range of vector to total is defined as follows: $$d^{2}\left(x,\pi\right)=\left(x-\mu\right)^{\prime}\Sigma^{-1}\left(x-\mu\right)$$ of which$x,y$A random vector. $\mu,\Sigma$ The average vector and the all-inclusive matrix, respectively.
Let's give you a brief description of the nature of Mars' distance.
- About change and change$y=Cx+b$ Changes in scale can be expressed as$y=Cx$ Change the value of some plus or minus to$y=x+b$
Ma's distance doesn't mean that the change doesn't affect our judgment about Ma's distance.
Standardized transformations are also a special form of transformations in the front, with no change in the distance between and after standardization When the weights are not relevant, the horse's distance is the standardized distance.
- Title: Multivariate Statistics Introduction: Random Vectors, Covariance Matrices, and Multivariate Normal Distributions
- Author: Hyacehila
- Created at : 2024-09-11 10:36:49
- Link: https://hyacehila.github.io//blog/2024/09/11/multivariate-statistics-introduction-notes/
- License: This work is licensed under CC BY-NC-SA 4.0.