Elementary Probability: Random Events, Probability Models, and Random Variables
Random event and probability
Random phenomena and statistical patterns
Starting with the probabilistic theory, we started looking at random questions that were different from the mathematical branches of his previous studies. And of course, we're here to use a lot of analytical tools and meta-mathematics to study random issues.
Random phenomena
(a) The inevitable event must occur, and it is unlikely that it will never occur under certain conditions; among them, the event may or may not occur at random when the basic conditions remain unchanged;
Frequency stability
In large-scale experiments, sometimes frequency. $$F_{N}(A)=n/N$$ That means random events aren't random, he has one. Statistical regularity And we have the concept of probability. A irrelevant experiment. The inherent value of the system.
Frequency and probability
They have some nature and connections.
- Non-negative $F_{N}(A)\ge 0$
- Inevitability and impossibility $F_{N}(A)=0 不可能事件,F_{N}(A)=1必然事件$
- Addability of frequency $F_{N}(A+B) = F_{N}(A)+F_{N}(B) 对于不同时发生的AB成立$
- The nature of a large number of theorem guarantees $N\longrightarrow \infty~F_{N}(A)\longrightarrow P_{N}(A)$ Probability is the measure of the probability of an event. The frequency is the result of the experiment.
Sample space and events
Here we continue to explain some of the basic concepts of probabilities.
Sample space
Experiments. Trail is a necessary study of random phenomena. Sample point is a possible result of the experiment Sample space is the entire sample point The source of the earthquake. $(x,y,z)$ Sample space, 3D area It doesn't matter the number of sites in the sample space. It's very clear that studying what kind of sample space depends on our problem, and there's some important sample space in probabilistic theory, but when it comes to practical problems, it's mathematical modelling.
Events
Event event is the assembly of some sample points, and at this point, for each event, we can judge whether a sample point is the event. Medium Obviously. $\Omega 整个样本空间 是一个必然事件;\emptyset 是一个不可能事件$
Operation of events
How to calculate an event using the probability of a simple event to study the probability of a complex event is a question of probabilities.
$A\subset B$ All sample points for A are in B.
$A\subsetBandB\subset A$ 2 events equivalent
$\bar{A}$ The opposing event of the A, all of the sample points not in A.
$A\cup B~~A\cap B$ The combination of events means that both occur simultaneously and that one of them is more than the other.
$A\cap B=AB=\emptyset$ AB can't happen at the same time.$A\cup B=A+B$
$A-B$ A, but not B.
Demorgen Law$$(\bigcup_{i=1}^n A_i)^c = \bigcap_{i=1}^n A_i^c~~~(\bigcap_{i=1}^n A_i)^c = \bigcup_{i=1}^n A_i^c$$
Operation laws
$$A\cup B=B\cup A~~~AB=BA~~~(A\cup B)\cup C=A\cup (B\cup C)~~~(A\cup B)\cap C=AC\cup BC$$
It's obvious that this assembly theory is very close, and we'll give him a more detailed explanation in future studies, which is closely linked to the higher form of probabilities, the measurement theory.
Definitions:$P(w_{1})+P(w_{2})+...+P(w_{n})=1$
Definition: the probability of the event is the probability of each sample point
Limited sample space: a limited number of sample sites
Dispersed sample space: sample space in which the number of sample points can be quantified
Classical profile
Let's start with this kind of simpler probabilistic problem.
- The results of the experiment are limited and different. Hand it over.
- Probability of all events These random phenomena are called classical generals, which are exactly the same as the classical generals we studied in high school. The classical definition of probability, the definition of probabilistic theory.$P(A)=m/n其中m是利场合数,n是样本点总数$ In fact, the classical model is because he does use it very widely, and we usually use it to study it, but its application goes far beyond the touch of the ball.
Basic combination analysis formula
Core principles
Multiplication doctrine: serial $n=n_{1}*n_{2}$ Added: Parallel $n=n_{1}+n_{2}$
Arranged questions
Of the n elements, the r elements are selected for ranking, considering not only the elements to be taken, but also the order to be taken.
- There's a playback. $n^{r}$
- No Releases $A_{n}^{r}=n(n-1)(n-2)...(n-r+1)$ It's quite understandable, when... $n=r$When it's full, it's full.$n!$
Cluster issues
Of the n elements, the r elements are selected for grouping, the elements to be considered and the order of removal not considered
- R for n elements $C_{n}^{m}=A_{n}^{r}/r!$ It's also shown in the following form, so it's not an integer, it's also called a binary coefficient. $$\left(\begin{array}{l} n \ v \end{array}\right)$$
- The n elements are divided into k groups, the number of groups has been given, the number divided by $n!/r_{1}!r_{2}!...r_{k}!$ He's also known as the multiple coefficient.$C_{n}^{m}$Accumulated
In fact, most of the grouping problems are deformations that are the most fundamental.
Some supporting formulae and definitions
$$0!=1 $$ $$n>~left (\begin{array}l) I'm sorry. k =0 $$ $$\left(\begin{array}{l} n \ k \end{array}\right)=\left(\begin{array}{l} n \ n-k \end{array}\right)$$
Extension to non-integer form
[Mathematic Analysis 1 Limits and Continuity Theory]
A classic example of classical generalization.
Examples of this section "The Classic Example of the Classic Outline" section
Two distributions and hypergeometric distributions
$a件次品,b件合格品,抽n件,研究k件不合格的概率$ We consider it easier to consider differences between products;
In case of re-entry. General $(a+b)!$A favourable occasion for $C n^k}a^{k}I'm sorry. Probability is $P=C n^k}(a/(a+b))^{k}(b)(a+b))^n-k}$ He's a dichotomy extension, all of which is called the dichotomy.
Without putting it back $$总数\left(\begin{array}{l} a+b \ n \end{array}\right) 合适的样本点数目\left(\begin{array}{l} a \ k \end{array}\right)*\left(\begin{array}{l} b \ n-k \end{array}\right)$$ That's the way it's distributed.
In fact, when the total number is very large but the number of samples is low, the results of the two profiles are close.
An example of mathematical statistics
Examples of this section "Mathematical examples of probabilities" section
Some nature
$$P(A\cup B)=P(A)+P(B)-P(AB)$$ $$P(A)=1-P(\bar{A})$$ $$AB=\emptyset ~~~~P(A_{1}+A_{2}+...+A_{n})=P(A_{1})+P(A_{2})+...+P(A_{n})$$
Geometry
In the geometry section, we have identified the apparent limitation of the limited number of classical overview sample points and have decided to use geometry to address this situation; The core of geometry is the scale.
Geometry tells us that the probability is zero and that the event will not take place at an equal price.
Simple geometry example
Place a circle equal to the diameter in the square; the probability of a random drop point within the square, which falls within the circle, is the ratio of the size of the two.$\frac{\pi}{4}$
Buffoon.
On the plane.$a$A parallel line, one for a long one.$l$The probability that needles, needles and parallel lines will intersect Obviously, if it's used$x$The distance between the midpoint of the needle and the parallel line,$\theta$Angular angle, and our final intersection is $x <l/2*sin(\theta)$,此时的$P=\frac{2\pia} If we do computer simulations, we can calculate the probability.$\pi$The value, then developed into a random simulation of statistics
The Monte Carlo method of calculating points
The ordinary Lehman is essentially a measure, and we can study him in size, while space can be obtained through random simulation experiments, with a plumb injection, which is the method of calculating points in Monte Carlo; Although the method is very inefficient, the error is stable and is still widely applied. We don't discuss this method in detail in our probabilistic research, but rather teach it where it's practice-oriented.
Bertrand.
The core of geometry is the construction of a measure of probability; this gives rise to some questions and subsequent mathematicians' thinking; Bertrandicism is due to a contradictory geometry problem that arises from a probability calibration, which we will study more in the high probability theory. Him.
Probability of conditions
The probability theory of conditions, which is studied in the case of event B, is recorded as the probability of event A occurring at this time.$P(A|B)$He's also an important part of probabilistic theory. We're here to study it. Definitions: Under a given probability space $P(A|B)=\frac {P(AB)}{P(B)}$ That's the probability formula. Obviously, both classical and geometric can easily be used.
Nature of probability of conditions
Inferences :Looks like it's just a migration, but it's important, called a multiplication formula. $$P(A|B)*P(B)=P(AB)$$ $$非负性:P(A|B)\ge 0$$ $$规范性:P(\Omega|B)= 0 $$ $$可列可加性:AB=\emptyset ~~~~P(A_{1}+A_{2}+...+A_{n}|B)=P(A_{1}|B)+P(A_{2}|B)+...+P(A_{n}|B)$$ Inferences: Multiplication formula for promotion $$P(A_{1}A_{2}...A_{n})=P(A_{1})*P(A_{2}|A_{1})*P(A_{3}|A_{1}A_{2})...P(A_{n}|A_{1}...A_{n-1})$$
Full Probability Formula
(a) The full probability formula is the probability of studying another thing by using the probability of something, which is closely linked to the probability of conditions; $$P(B)=P(AB)+P(\bar{A} B)=P(A)P(B|A)+P(\bar{A})P(B|\bar{A})$$ For some of the problems that are very difficult to address, there's an experiment that a previous experiment can refer to, and the full probability formula is a good choice, which is that his core should be, step by step. The full probability formula is freely extended to more dollars. $$P(A)=\sum\limits P(AB_{i})=\sum P(B_{i})P(A|B{i})$$
Bayes Formula
The Bayes formula is about studying the connection between the two events, on which Bayes' decision-making and distinction are based.
If$B$Only incompatible.$A_{i}$ It happens at the same time. $B=\sum P(BA_{i})$
$$P(A_{i}B)=P(A_{i}|B)*P(B)=P(A_{i})*P(B|A_{i})$$
And so...
$$P(A_{i}|B)=P(A_{i})*P(B|A_{i})/P(B)$$
So...
$$P(A_{i}|B)=\frac{P(A_{i})*P(B|A_{i})}{\sum\limits P(A_{i} )~P(B|A_{i})}$$
That's the Bayes formula. $B$ In the event $A_{i}$ Probability of events;
$P(A_{i})$ It's a summary of past experience, which we call probabilistic data, which he learned before the experiment.
$P(A_{i}|B)$ It's what we'd like to study, commonly known as the probability of a posteriori.
The Bayes formula is widely used for disease diagnosis, and at this point B means indicator A means disease; the probability of disease occurring under one indicator can be calculated using the Bayes formula, and the degree of confidence obtained in the calculation can be used by the doctor for reference purposes.
Independence of events
Just as we're looking at probability, we're looking at two things here.
Independence of the two incidents
Definitions:$P(AB)=P(A)P(B)$ Combining the multiplication formula, which means the condition of probability is no longer working.
Inference: Independent events satisfied $P(A|B)=P(A)$
Inference: $AB independent, rule A\bar{B}},{\bar{A}\bar{B}},{AI don't know.
Independence of multiple events
You have to meet the following conditions simultaneously.
$P(AB)=P(A)P(B)P(AC)=P(A)P(C)P(BC)=P(C)P(B)~~P(ABC)=P(A)P(B)P(C)$$
Independence is a matter of high standards, and Bienstein's example shows this.
It's very natural that there's no need to give a clear description of the forms of independence of multiple events.
Independence of the experiment
From events to experiments, experiments are the core of probabilistic science.
If there's an experiment, and they have separate sample space, the total sample space is the calcium of the sample space.
At this point,$P(A^{1}A^{2}...A^{n})=P(A^{1})P(A^{2})...P(A^{n})$ We call these experiments independent.
The independence of the experiment explains what we've learned, but we don't know. Repeated independent experiments will be a very common concept to learn from behind.
Bernuli.
In some cases, we focus only on two types of results, qualified and unqualified, which is the subject of the Benoligue study; Consider repeating.$n$The next Benuli experiment required
- Up to two results per experiment.$A~\bar{A}$ Probability and 1
- $P(A)$ Stabilization
- The experiment is independent of each other.
- Conduct$n$Minor experiments Total number of points for such experiments$2^{n}$In fact, as the number of experiments draws closer, the number of sample points becomes the first infinity. This is a very broad profile, such as the oversale of airline tickets, genetic problems, etc., which is consistent with the repeated 01 results. That's why we're here for a more detailed analysis.
Benuli distribution
Just one experiment. The distribution column is very easy to write.
Two distributions
We have described the two distributions earlier. In fact, in the case of the extraction and return of the defective items, it's a one-one-time test, and that's a Benuli experiment, so we don't want to go on here too much; as for the probability of the two distributions, it's a very good thing. $$b(k;n,p)=C_{n}^{k}p^{k}q^{n-k}$$
Geometric distribution
This is not a supergeometric distribution. The so-called geometric distribution is a study. The first successful experiment is in the first place.$k$Probability of repeats And we can easily calculate the geometric distribution column. $$g(k;p)=q^{k-1}p$$ Geometrically unrememberable It's a very important feature of geometry; it means $P (X)>(m+n)|X>m)=P(X>That's it. He means that no matter how many experiments have been done, it doesn't affect the probability behind it.
Pascal distribution/negative distribution
We're trying to expand the geometrical distribution of boundaries.$k$It's been a success.$r$It's not really hard to calculate. $$f(k;r,p)=C_{r-1}^{k-1}p^{k-1}q^{k-r}p=C_{r-1}^{k-1}p^{k}q^{k-r}$$
Two distributions and porcelain distribution
We're here to study a few examples and a new distribution, which is closely linked to the two previous studies.
A little more.
$$\begin{aligned} &\frac{b(k;n,p)}}b(k+i,r,b)}=1+\frac{(n+1)p-k}{kq}\ I'm sorry. That tells us.$k=(n+1)p=m$, at which time the probability of the two distributions is taken to the maximum; Of course, because of the integer limit, we can only go to a close value. This is what we call the center of two distributions. Unexplained conclusions:$P(m)=(2\pi npq)^{-\frac{1}{2}}$
A simple example.
Assuming a total of 200 machine beds are powered at 1 kw each, with a 60 per cent probability of opening each other independently, the workshop is expected to be stable above 99 per 1,000. It's obviously a probabilistic problem. Calculating \\sum\limitesb(k,200,0.6)>0.999$ 研究找到的$That's all we need. Now the problem is this huge distribution is hard to calculate.
Two distributions approaching.
In a lot of Benuli experiments, if$n$Large$p$They're small.$\lambda$It's more moderate in size and, in this case, Persson has found a more easy form to calculate. $$当np接近\lambda,则在n\to \infty 时有 b(k;n,p)=p(k,\lambda)=\frac{\lambda^{k}}{k!}e^{-\lambda}$$ It's called theorem.
In fact, the porcelain distribution has found many uses.
- Web access (this is all the counting process in unit time)
- Thermal electronic launch and microbiological distribution
- Composition of other random phenomena This has become an important stand-alone distribution rather than a double distribution calculation;
Random variables and their distribution
By now we have concluded a small phase of our research, and from now on, the question of the probabilities of quantitatively uniform research is a very important one, which is random variables, and we will study in detail the dimensions of this chapter.
Random variables and their distribution
We call random variables as variables that represent random phenomena, how they describe the results of random phenomena, and how to further study practical issues using random variables, which is clear from our chapter;
Concept of random variables
Many of the sample points are expressed in a number and because of the randomity of the sample points, they are a random variable;
Definitions: define real-value functions in sample space$X=X(w)$ Called a random variable, usually expressed in capital letters and taken values in lowercase letters;
If the value of a random variable is limited or columnable, it is referred to as an discrete random variable or, if the value is full of space on a number of axes, as a continuous random variable;
Distribution function for random variables
In order to master the statistical regularity of a random variable, we need to know the probability of his going to values, and the probability of random variables is clearly cumulative, so we just need to know.$F(x)=P{X\le x}$ That's it.$F(x)$ It's a definition.$(-\infty,\infty)$ functions;
Definitions: setting X is a random variable for any actual number$x$ Claims$F(x)=P{X\le x}$ It's random.$X$The distribution function called$X$Obey.$F(x)$
Random variables, whether continuous or discrete, have distributed functions
Theorem: Any distribution function$F(X)$ Both have the following characteristics:
- $F(X)$It's a monotony.
- $F(X)$The range of values is 0 to 1 in the closed range, and the limits of both ends are 0 and 1.
- $F(X)$It's a right continuous function. $\lim_{x \to x_{0}+0}F(x)=F(x_{0})$
The function that satisfies these three characteristics must be a distributed function
Distribution column of discrete random variables
For discrete random variables, we often use the following distribution column to express it accurately: It's...
Definitions: Establishment $X$ is a discrete random variable, if$X$ All possible values are $x_{1}, x_{2}, \cdots , x_{n}$ , or$X$ Remove $x_{i}$ Probability$p_{i}=p\left(x_{i}\right)=P\left(X=x_{i}\right), i=1,2, \cdots, n$Yes $X$ the probability distribution column or abbreviation column, as $X \sim\left{p_{i}\right}$
The distribution column can also be expressed as follows: $ \begin{array}ccccc} X & x_{1} & x_{2} & \ldots & x_{n} & \cdots \ \hline P & p\left(x_{1}\right) & p\left(x_{2}\right) & \cdots & p\left(x_{n}\right) & \cdots \end{array}$$ 或者 $$ \left(\begin{array}{ccccc} x_{1} & x_{2} & \cdots & x_{n} & \cdots \ p\left(x_{1}\right) & p\left(x_{2}\right) & \cdots & p\left(x_{n}\right) & \cdots \end{array}\right)$$
The distribution column clearly has two basic properties.
- Non-negative$p(x_{i})\le 0$
- Regularity$\sum p(x_{i})=1$
Probability density function for continuous random variables
The distribution column will certainly not be able to study a range of random variables, but imagine that the distribution column is intended to describe the probability of a single point and corresponds to the continuous, distributed function.$F(X)$And that's how it works, so...
Definitions: set random variables$X$Distribution Functions$F(X)$ There is a non-negative buildup function$p(x)$ Satisfied$\int_{-\infty}^{x} p(t)dt=F(X)$ Claims$p(t)$Yes$X$Probability density function
Probability density functions have two basic properties
- Non-negative$p(x)\ge 0$
- Regularity$\int_{-\infty}^{\infty} p(x_{i})dx=1$ In order to calculate probability from a density function, we have to look at it from the point of view of points that he cannot avoid.$N-L$The formula.
If the density is$0$Use$N-L$The formula must be taken out of this compartment when it takes down points. Because I'm here.$0$It must mean the existence of a non-continuous point.$N-L$I'm sure there's something wrong with the score.
Comparison of two random variables
Probability density functions and distribution columns are basically close, but there are still some differences, as we simply describe here.
- The distribution function of the discrete random variable is right continuous, but the distribution function of the continuous random variable is completely continuous
- The discrete random variable has a zero probability at some points, but the probability of any single point of a continuous random variable is zero.Our probability is zero.
- Because the single point probability of a continuous random variable is zero, you're out. The removal of a few points should not affect; otherwise, the discrete random variable must be measured against each point to ensure that the end result is correct
A mathematical expectation for random variables
Studying the characteristics of random variables, describing the characteristics of random variables as a whole in simple quantities is what we've always wanted to do, and that's the numerical characteristics of random variables, and we present the characteristics of one of the most important random variables, mathematical expectations;
The concept of mathematical expectations
We would like to know where this random variable is going in its entirety, and that's what mathematical expectations or averages want to answer.
arithmetical average: directly calculated average Weighted average: taking into account the impact on overall trends of different numbers, depending on the frequency of occurrence It is clear that in random variables, the use of probability as a power value is a very natural thing;
Definition of mathematical expectations
For discrete random variables, we call$E(X)=\sum p(x_{i})x_{i}$ is the mathematical expectation of discrete random variables, if they are constricted;
It's because it's not the only way to avoid the expectations of non-inclusion, which we mention in the infinity of mathematical analysis; for random variables of a limited number, the expectations must be there; in fact, the expectations of Cauchy distribution are not there, and we make this additional requirement reasonable.
For continuous random variables, we call$E(X)=\int_{-\infty}^{\infty} x_{i}p(x_{i})dx$ It's the mathematical expectation of a continuous random variable.
Nature of mathematical expectations
The mathematical expectations of the functions of random variables
We know the mathematical expectations of random variables.$E(X)$ The only thing that is identified in the distribution, and it is very clear that the random variable function is also a random variable, and in order to solve the mathematical expectations of the random variable function, we need to do the following research. In the most basic context, we can study this further by calculating the distribution column of the new random variable, the probability density function, and in order to facilitate the next calculation, we give the following theory. We have a distribution column for discrete random variables.$p(x_{i})$and Functions$g(X)$ Mathistic Expectations$E{g(X)}=\sum g(x_{i})p(x_{i})$ For continuous random variables, we have a probability density function.$p(x)$and Functions$g(X)$ Mathistic Expectations$E{g(X)}=\int g(x)p(x)dx$ It's natural for theorem to consider it directly intuitively.
Several export properties
- $E(x)=c$
- $E(aX)=aE(X)$
- $E(g(x_{1})+g(x_{2}))=E(g(x_{1}))+E(g(x_{2}))$
- $X,Y$On my own.$E(XY)=E(X)E(Y)$
- $E(X+Y)=E(X)+E(Y)$
The difference of the random variable
The variance of the random variable is presented in order to study the size of the fluctuations of the random variable;
Definition of variance and standard deviation
Of course, it's impossible for any variable to happen to be the average, and there's bound to be a deviation; and...$X-E(X)$It's also a random variable that we choose to erase positive and negative effects and not to use obnoxious absolute values.$(X-E(X))^{2}$ As an image of fluctuations,$E((X-E(X))^{2})$ It'll reflect the overall fluctuations. Definitions If Random Variables$X$The mathematical expectation exists, but it's called$E((X-E(X))^{2})$It's a random variable method.$Var(X)$ For differential calculations, it's easy to see what he's expected to be a random variable function, and it's very easy to calculate. The standard difference is the amount derived from the equation, his schematics are the same as the original random variable, which is what he meant to exist. When mathematical expectations exist, differences don't necessarily exist, but vice versa.
Nature of the difference
- $Var(X)=E(X^{2})-(E(X))^{2}$It's more practical to calculate the difference.
- $Var(c)=0$
- $Var(aX+b)=a^{2}Var(X)$
- $\mathrm{Var}(X)=E(X(X-1))-\mu_X(\mu_X-1)$
- $X,Y$On my own. $Var(X+-Y)=Var(X)+Var(Y)$
Chebby Scheffer's all the same.
A constant of random variables when expectations and differences exist$\varepsilon$ Yes. $$P\left(\left|\xi- E(\varepsilon)\right|\geqslant\varepsilon\right)\leqslant\frac{D\left(\xi\right)}{\varepsilon^{2}}$$ The core of the Chebbyschev heterogeneity is to give a high probability of deviation.
Theorem: Random variable$X$Difference$Var(X)=0$Meaning$X$Almost everywhere equals a constant.$c$
Frequent discrete distribution
Two distributions
$$b(k;n,p)=C_{n}^{k}p^{k}q^{n-k}$$ It's the probability of two distributions, and it's born to study sampling and put back, and we add here about averages and differences.
- $E(X)=np$
- $Var(X)=np(1-p)$
Benuli distribution
He's the two distributions of degradation.
Porsche distribution
$p(k,\lambda)=\frac{\lambda^{k}}{k!}e^{-\lambda}$ The porcelain distribution is exported by the approximation of two distributions.
- $E(X)=\lambda$
- $Var(X)=\lambda$ It's a very amazing feature.
Supergeometric distribution
It was also exported in two distribution experiments. $$总数\left(\begin{array}{l} a+b \ n \end{array}\right) 合适的样本点数目\left(\begin{array}{l} a \ k \end{array}\right)*\left(\begin{array}{l} b \ n-k \end{array}\right)$$ The supergeometric distribution is complicated, and we don't study his averages and differences here.
Geometric distribution
Studying another problem with sampling $$g(k;p)=q^{k-1}p$$
- $E(X)=\frac{1}{p}$
- $Var(X)=\frac{1-p}{p^{2}}$
Pascal distribution
One extension of geometry is called negative dichotomy. $$f(k;r,p)=C_{r-1}^{k-1}p^{k-1}q^{k-r}p=C_{r-1}^{k-1}p^{k}q^{k-r}$$
- $E(X)=\frac{r}{p}$
- $Var(X)=\frac{r(1-p)}{p^{2}}$ The probability is derived from geometry.
Regular and continuous distribution
The continuous distribution of density functions and distribution functions can be exported from one another, but people give more attention to probabilities density functions, which we will also reflect in our subsequent narratives.
Normal distribution
This is the most important continuous distribution in probabilistic and mathematical statistics; we repeat it countless times in a lot of subsequent studies;
Density and distribution functions for normal distribution
If the density function of random variable X is$p(x)=\frac{1}{\sqrt{2\pi}\sigma}e^{-\frac{(x-\mu )^{2}}{2\sigma^{2}}}$ Name$X$Subject to normal distribution $X\sim N(\mu,\sigma)$ His distribution functions are usually expressed directly in the form of points.$\mu$Center called Location Parameters for Normal Distribution $\sigma$It's called a scale parameter that determines the degree of fragmentation.
Standard normal distribution
The normal distribution of position parameters 0 and scale parameters 1 is called the standard normal distribution, i.e.$N(0,1)$ His density function is... $\varphi (x)=\frac{1}{\sqrt{2\pi}}e^-{\frac{\mu^{2}}{2}}$
Since the standard normal distribution does not contain any parameters, we can give them directly. $\Phi(X)$He'll be used extensively to calculate probabilistic values.
- $\Phi(-u)=1-\Phi(u)$
- $P(X>u)=1-\Phi(u)$
- $P(a<X<b)=\Phi(b)-\Phi(a)$
- $P(|X|<c) =2\Phi(c)-1 US$ These formulas are easy to predict.
Standardization of normal distribution
Theorem : If$X\sim N(\mu,\sigma)$ then$U=\frac{X-\mu}{\sigma}\sim N(0,1)$ Based on this theory, we know that when calculating the probability of a non-standard normal distribution, we can choose to translate it directly into a standard form and then proceed with the next calculation. $P(X)<(c) = (\frac{c-mu}{\sigma} This is the common way to calculate the value of any normal distribution. Functions$\Phi(x)$The value can be obtained from a checklist.
Expectations and differences in normal distribution
In fact, we studied in high school for normal distribution.
- $E(X)=\mu$
- $Var(X)=\sigma^{2}$
Normal distribution$3\sigma$ Principles
There are two important explanations for this theory.
- If there's a serious deviation,$3\sigma$ Principle Distribution is not a normal distribution
- If there is a serious deviation from production$3\sigma$ It means production is uncontrolled.
even distribution
We say the distribution of the distribution function meets the following conditions:$U(a,b)$ $$p(x)=\left{\begin{array}{ll} \frac{1}{b-a}, & a<x<b, \ 0, & \text {Other.} {\bord0\shad0\alphaH3D}right.
- $E(X)=\frac{a+b}{2}$
- $Var(X)=\frac{(b-a)^{2}}{12}$
Index distribution
We call the distribution of distribution functions that meet the following conditions as an index distribution.$EXP(\lambda)$ $$p(x)=\left{\begin{array}{ll} \lambda e^{-\lambda x}, & x\ge0, \ 0, & \text {Other.} {\bord0\shad0\alphaH3D}right. He's used to describe life expectancy.
- $E(X)=\frac{1}{\lambda}$
- $Var(X)=\frac{1}{\lambda^{2}}$ The index distribution is immutable, as is geometry.
Gamma distribution
Give Gamma function again $$\Gamma(\alpha)=\int_0^\infty t^{\alpha-1}e^{-t}dt$$ Nature gives a probability density function for Gamma distribution =x(x)=\left{erray}\frac{\a^e\beta}&x>0\0&\text{otherwise}\end{array}\right.$$ $E[X]=\frac\alpha\beta,\quad\text{,}Var(X)=\frac\alpha{\beta^2}$ It's natural.$\alpha = 1$ He's the probability density function for index distribution. If$\alpha=\frac{n}{2};\beta=\frac{1}{2}$ It's freedom.$n$the distribution of the card Gamma distribution is derived from the index distribution; he is the sum of several independent and distributed index distribution variables Trying to prove that this theory can use the theory of rectangular parent function (Moment Generating Fund)
Against Gamma Distribution
It's also called inv Gamma distribution. $$\left.\left{array}\frac{\alpha}{\alpha} (xalpha})1}exp0\ (\frac{\beta}x0\),\geq0\,0\<0\end{array}\right.\right.$$ 计算他的期望和方差有 $$\begin{aligned} &E(x)=\frac{\lambda^{\alpha}}{\Gamma(\alpha)}\int_{0}^{+\infty}x^{-\alpha}e^{-\frac{\lambda}{x}}dx=\frac{\Gamma(\alpha-1)}{\Gamma(\alpha)}\lambda=\frac{\beta}{\alpha-1} \ &E(x^{2})=\frac{\lambda^{\alpha}}{\Gamma(\alpha)}\int_{0}^{+\infty}x^{-\alpha+1}e^{-\frac{\lambda}{x}}dx=\frac{\Gamma(\alpha-2)}{\Gamma(\alpha)}\lambda^{2}=\frac{\beta^{2}}{(\alpha-1)(\alpha-2)} \ &== sync, corrected by elderman == @elder man I'm sorry. It's more common in Bayesian statistics.
Beta distribution
Beta distribution is also a very common continuum. $$f(x;\alpha,\beta)=\frac{1}{B(\alpha,\beta)}x^{\alpha-1}(1-x)^{\beta-1}$$ Note that the parameter range here is $\alpha~beta>$ 0 million Of which$B(\alpha,\beta)$ is the value of the Beta function presented in the mathematical analysis. $$B(p,q)=\int_{0}^{1}x^{p-1}(1-x)^{q-1}dx$$ Their expectations and equations are different. $$E(X)=\frac{\alpha}{\alpha+\beta}$$ $$Var(X)=\frac{\alpha\beta}{(\alpha+\beta+1)(\alpha+\beta)^2}$$
Distribution of random variable functions
From one distribution, look at the other with one function.$g(X)$ It's very central to probabilistic and mathematical statistics.
Distribution of discrete random variable functions
In fact, the function of a discrete random variable must be an discrete random variable, so the problem is very simple. Because of its limited nature, we can calculate which function values are mapped for which new function and add the same probability to it.
Distribution of functions of continuous random variables
Turned out.
If this function leads to a new discrete distribution, we can only calculate it by the way we calculate it, or by the time we calculate the probability of each discrete point.
Strictly monotonous.$g(x)$
We can give a very effective theory to deal with this kind of problem. Theorem Set$X$It's a continuous random variable.$p(x)$ $Y=g(X)$ It's another continuous random variable if$g(x)$It's a strictly monotonous function.$h(y)$Other Organiser$Y$The density function meets $p(y)=\left(begin{array}ll} P x[h(y)]'}(y)|, & a<y<b, \ 0, & \text {Other.} {\bord0\shad0\alphaH3D}right. of which$a=min{g(-\infty),g(\infty)} b=max{g(-\infty),g(\infty)}$ This is a very important theorem in which the non-zero definition of the new density function should be replaced by the definitional domain of the original density function to the transformational function.
Common theorem and nature derived
Set Random Variables$X$Subject to normal distribution$N(\mu,\sigma^2)$ There is.$Y=aX+b\sim N(a\mu+b,a^2\sigma^2)$ It's a very basic nature of normal distribution. Set Random Variables$X$Subject to normal distribution$N(\mu,\sigma^2)$ There is.$Y=e^X$ The probability density function is $$\left.P (y)=\left{begin{aligned}&\frac{1}{\sqrt{2\pi}y\sigma}\exp\left{-\frac{\left(\ln y-\mu\right)^{2}}{2\sigma^{2}}\right},y>0,\&I'm sorry. What we call a logarithmic normal distribution is a very common distribution. Set Random Variables$X$Obey Gamma distribution.$Ga(\alpha,\beta)$ There is.$Y=kX\sim Ga(\alpha,\frac{\beta}{k})$
Other$g(x)$Situation
The basics are that we can use definitions for research. For research.$Y=g(X)$The distribution function we have.$F_{Y}(y)=P(g(X)\le y)$ Turned into about$x$By$y$Probability forms of restriction Different studies are conducted according to different characteristics
Other features of distribution
Rectangular
$k$Paradox:$\mu_k=E(X^{k})$
Central Rectangular
$k$Center Rectangular:$v_k=E(X-E(X))^{k}$
Variable factor
The amount given below is ununited. Defines the variable factor as $$V=\frac{\sigma}{\mu}$$
Bits
The following conditions are met:$x_p$Called$p$It's called the lower.$p$It's not very useful for us to circumvent the concept of the upper fraction. $$F\left(x_{p}\right)=\int_{-\infty}^{x_{p}}p\left(x\right)dx=p$$ $p=0.5$in which case the fraction is called the median
Offset coefficient
If Random Variables$X$The first three steps exist. $$\beta_{S}=\frac{\nu_{3}}{\nu_{2}^{3/2}}=\frac{E\left(X-E\left(X\right)\right)^{3}}{\left[\mathrm{Var}\left(X\right)\right]^{3/2}}$$ For the margin He reflects the degree to which the distribution deviates from symmetry.
Peak coefficient
If Random Variables$X$The first four steps exist. $$\beta_{k}=\frac{\nu_{4}}{\nu_{2}^{2}}-3=\frac{E\left(X-E\left(X\right)\right)^{4}}{\left[Var\left(X\right)\right]^{2}}-3$$ is the peak coefficient He reflects the steepness of the peak or the thickness of the tail.
Random vector and its distribution
In some random phenomena, it's not enough for sample points to use only one random variable to describe it, which leads to the concept of random vectors, and we can't consider these random variables separately.
Random vector and joint distribution
Multi-dimensional Random Variables
Definitions : If$X_{i}(w)$ It's defined in the same sample space.$n$A random variable is called$X(w)=(X_{1}(w),X_{2}(w),X_{3}(w)....)$ It's one.$n$V's random vector
The core must be in the same sample space.
Joint Distribution Functions
Definitions: Yes$n$Could not close temporary folder: %s<x_{1}},{X_{2}<x_{2}},{X_{3}<x_{3}}...$同时发生的概率$F(x_{1},x_{2}...x_{n})=P(X_{1}<x_{1}...X_{n}<x_{n})$ 就是$Distribution function of the n$D random variable In our later study, we're focusing on the combined distribution of two dollars, and the natural extension of more dimensions. Basic properties of multiple joint distribution functions
- Monophonic $F(x,y)$ Both variables are single.
- There's a boundary. $F(x,y)$ The value range is 0-1 and as long as one is close and negative, it's zero.
- Right continuity, right continuity for individual variables.
- Non-negative $P(a)<X<b,c<Y<d) = F(b,d)-F(a,d)-F(b,c)+F(a,c)\ge0 I can also prove that all functions that satisfy this nature are distributed functions.
Joint Distribution Columns
For the discrete binary distribution, we can use the matrix to describe the probability of the event. Icon Joint distribution display matrix. We'll talk about it later. Nature of the joint distribution column
- Non-negative$p(x_{i})\ge 0$
- Regularity$\sum \sum p_{ij}=1$ The core of the joint distribution is research probabilities.
Joint density function
Now handle continuous multiple distribution functions Definitions: If existing$p(x,y)$ Distribution function for random variables of the binary$F(x,y)$Satisfied$$F(x,y)=\iint\limits_{-\infty}^{x,y}p(x,y)dxdy$$It's calculated using a cumulative fraction. Name$p(x,y)$It's a joint density function, and the dilution of the distribution function must be a density function.
- Non-negative$p(x,y)\ge 0$
- Regularity$\iint\limits_{-\infty}^{\infty}p(x,y)dxdy=1$ In the case of multiples, let's stress that when we use points to calculate probability, The range of points is the intersection between the range required by the title and the non-zero zone, which is then calculated as a cumulative fraction. As for the issue of cumulative fractions, it's clear from the mathematical analysis that there's no very difficult cumulative calculation here.
Some common multi-dimensional random variables
Multiple distributions
In high school, there's a certain amount of research, which means we have more than one option. $$P=C_{n}^{k_{1}}C_{n}^{k_{2}}...C_{n}^{k_{n}}p_{1}^{k_{1}}p_{2}^{k_{2}}...p_{n}^{k_{n}}$$ Multiple distributions are a discrete distribution.
Multi-dimensional hypergeometric distribution
Or don't put it back for sampling. $$P(X_{1}=n_{1},X_{2}=n_{2},\cdots,X,=n_{r})=\frac{\binom{N_{1}}{n_{1}}\binom{N_{2}}{n_{2}}\cdots\binom{N_{r}}{n_{r}}}{\binom{N}{n}}$$
Multi-dimensional even distribution
$$p(x)=\left{\begin{array}{ll} \frac{1}{S}, & x\in S, \ 0, & \text {Other.} {\bord0\shad0\alphaH3D}right.
Binary Normal Distribution
The core of the normal distribution is five.$(X,Y)\sim N(\mu_{1},\mu_{2},\sigma_{1},\sigma_{2},p)$ The joint density function is $$f(x_1,x_2)=\frac1{2\pi\sigma_1\sigma_2\sqrt{1-\rho^2}}e^{-\frac1{2(1-\rho^2)}\left(\frac{(x_1-\mu_1)^2}{\sigma_1^2}-2\rho\cdot\frac{x_1-\mu_1}{\sigma_1}\cdot\frac{x_2-\mu_2}{\sigma_2}+\frac{(x_2-\mu_2)^2}{\sigma_2^2}\right)}$$ They have edge density functions. We saw it in the joint density function.$p$He's the relevant coefficient.
Marginal distribution and independence
There's a lot of diversity in distribution that deserves our study.
- Density function for a single variable - marginal density function
- Level of correlation between the two volumes - relevant coefficient
- When a measure is given, another distribution - the condition distribution We'll be working on it later, and we'll be working on it.
Marginal Distribution Functions
It's not a difficult problem. And so... $F(x,y)$ About$x$and$y$The marginal distribution is as follows: $\lim_{y \to \infty} F(x,y)$ $\lim_{x \to \infty} F(x,y)$ Marginal distribution is the distribution of a fraction or parts of a vector It lacks an image of the relationship between weight and weight.
Marginal Distribution Bar
For discrete scenes, the marginal distribution column needs to add up each row and each column.
Marginal density function
We can give this formula. He understands it very well. $p_{X}(x)=\int\limits_{-\infty }^{\infty } p(x,y)dy$ $p_{Y}(y)=\int\limits_{-\infty }^{\infty } p(x,y)dx$ The formula itself is very well understood, or is it an old problem?
Random variable independence
Sometimes the weights between multiple random vectors interact with each other, but sometimes they're independent of each other.
Definitions: If $F(x_{1},x_{2}...,x_{n})=F(x_{1})F(x_{2})...F(x_{n})$ Call these random variables independent of each other.
For separate random variables:
- Disconnection can determine the probability of a big event by building up the probability of every small event.
- The continuous combined probability density is the accumulation of marginal probability density.
In order to determine whether the random variable is independent, the core is to determine whether the combined probability density is the amount of the marginal probability density.
Functions of random vectors
Classical partition scene
It's always our choice.$Y$Once the value is taken out, it's enough to calculate the composition and the final probability.
We'll give a simple example of how his mind is repeated in the back.
Prove the additionality of the porcelain distribution, which is $X\simP (\lambda )Y\sim P(\lambda_{2})\\text{requirements}$ $Z=X+Y\sim P(\lambda 1+\lambda 2) $
It's easy to know. $Z$ Can go to all non-negative integers is a discrete distribution and$Z=k$ Yes.
$$\left{X=i,Y=k-i\right}$$
It's an incompatible set of times, and given the independence,
$$P\left(Z=k\right)=\sum_{i=0}^{k}P\left(X=i\right)P\left(Y=k-i\right).$$
We call itDispersive volume formula So we can calculate.
$$X+Y\sim P(\lambda_1+\lambda_2)$$
The volume is the sum of two random variables.
Maximum distribution
Use definition studies to remove the minimum value mark by the nature of probability as follows: Set$X_1,X_2,...X_n$ It's independent.$n$A random variable.$max{X_1,X_2,...X_n};min{X_1,X_2,...X_n}$ Distribution $ \begin{aligned}F y(y)&=P(\max{X_1,X_2,\cdots,X_n}\leqslant y)=P(X_1\leqslant y,X_2\leqslant y,\cdots,X_n\leqslant y)\&=P(X_1\leqslant y)P(X_2\leqslant y)\cdots P(X_n\leqslant y)=\prod_{i=1}^nF_i(y).\end{aligned}$$ $$\begin{aligned} F_{z}\left(z\right)& =P\left(\min\left{X_{1},X_{2},\cdots,X_{n}\right}\leq z\right) \ &=1-P\left(\min\left{X_{1},X_{2},\cdots,X_{n}\right}>z\right) \ &=1-P\left(X_{1}>z,X_{2}>z,\cdots,X_{n}>z\right) \ &=1-P\left(X_{1}>z\right)P\left(X_{2}>z\right)\cdots P\left(X_{n}>z\right) \ &== sync, corrected by elderman == @elder man I'm sorry. It's a very common treatment, and we'll meet somewhere else.
Continuous volume formula
For continuous$Z=X+Y$ Two random variables are irrelevant. Theoretically:$p_{Z}(z)=\int\limits_{-\infty }^{\infty} P_{X}(z-y)P_{Y}(y)dy=\int\limits_{-\infty }^{\infty} P_{X}(x)P_{Y}(z-x)dx$ He's still easy to remember and understand. The core point is that we have to make sure that neither of the two probability densities is zero, so we need to process an iniquities between the blocks. The volume formula can also be used for two random variables that are not independent.
Variable variant
We studied the function of random variables when they were random. Down Set 2D random variable$(X,Y)$ The joint density function is$p(x,y)$ If there is. $$\begin{cases}u=g_{1}\left(x,y\right)\v=g_{2}\left(x,y\right)\end{cases}$$ There's a continuous deviation and there's a single inverse. $$\left.\left{\begin{matrix}x=x\left(u,v\right)\y=y\left(u,v\right)\end{matrix}\right.\right.$$ Transforming the Yacima one. $$\left.\left.J=\frac(x, y\right)\partial\left(u,v\right)}\left\begin{matrix}\frac(partialx}{\partialu}&\frac{\partial x}{\partial v}\\frac{\partial y}{\partial u}&\frac{\partial y}{\partial v}\end{matrix}\right.\right|=\left(\frac{\partial\left(u,v\right)}{\partial\left(x,y\right)}\right)^{-1}=\left(\left|\begin{matrix}\frac{\partial u}{\partial x}&\frac{\partial u}{\partial y}\\frac{\partial v}{\partial x}&\frac{\partial v}{\partial y}\end{matrix}\right.\right|\right|^{-1}\neq0$$ 如果 $$\left.\left{\begin{matrix}U=g_{1}\left(X,Y\right)\V=g_{2}\left(X,Y\right)\end{matrix}\right.\right.$$ 则$(U,,V)$ 的联合密度为 $$p\left(u,v\right)=p\left(x\left(u,v\right),y\left(u,v\right)\right)|J|$$
Stock and commerce
Set Random Variables$X,Y$Independent$p_X(x),p_Y(y)$ $U=XY$The density function is $$P_{U}\left(u\right)=\int_{-\infty}^{\infty}p_{X}\left(\frac{u}{v}\right)p_{Y}\left(v\right)\frac{1}{\left|v\right|}dv.$$ $U=\frac{X}{Y}$ The density function is $$P_{U}\left(u\right)=\int_{-\infty}^{\infty}p_{X}\left(uv\right)p_{Y}\left(v\right)\mid v\mid dv.$$ The formula is the same as usual.
Multi-dimensional digital characteristics
Multi-dimensional random vector expectations
The formula that we want to be able to give the desired function of a 2-D random variable, which is the new feature of a multi-dimensional situation, is that only the function of a multi-dimensional random variable is a new random variable that can study expectations. Set 2D random variable$(X,Y)$ The distribution is expressed as a joint distribution column$P(X=x,Y=y)$ The joint density function is$p(x,y)$ then$Z=g(X,Y)$ The mathematical expectation is... $$\left.E\left(Z\right)=\left{\begin{matrix}\sum_{i}\sum_{j}g\left(x_{i},y_{j}\right)P\left(X=x_{i},Y=y_{j}\right)\\\int_{-\infty}^{\infty}\int_{-\infty}^{\infty}g\left(x,y\right)p\left(x,y\right)\mathrm{d}x\mathrm{d}y,\end{matrix}\right.\right.$$
Agreement
It's called the center of the equation. If for a 2D random variable$(X,Y)$ Yes.$E\left[\left(X-E\left(X\right)\right)\left(Y-E\left(Y\right)\right)\right]$ Existence It's called an agreement. $$\mathrm{Cov}\left(X,Y\right)=E\left[\left(X-E\left(X\right)\right)\left(Y-E\left(Y\right)\right)\right]$$ Especially. $Cov\left(X,X\right)=Var\left(X\right)$ The difference is a mathematical expectation of two differential multipliers, so the difference is positive and negative, too. When the balance is positive, it's called positive correlation. When the alignment is 0, either two random variables are irrelevant or they have non-linear relationships.
Here are some of the usual characteristics.
- $Cov(X,Y)=E(XY)-E(X)E(Y)$
- It's not like you're gonna be able to be independent.
- $Var(X+Y)=Var(X)+Var(Y)+2Cov(X,Y)$
- $Cov(X,Y)=Cov(Y,X)$
- $Cov(X,c)=0$
- $Cov(aX,bY)=abCov(X,Y)$
- $Cov(X+Y,Z)=Cov(X,Z)+Cov(Z,Y)$
Related coefficient
If for a 2D random variable$(X,Y)$ $Var(X),Var(Y)>$0.00 $$Corr\left(X,Y\right)=\frac{Cov\left(X,Y\right)}{\sqrt{Var\left(X\right)}\sqrt{Var\left(Y\right)}}=\frac{Cov\left(X,Y\right)}{\sigma_{X}\sigma_{Y}}$$ A linear correlation between two random variables is illustrated. There's another explanation for the coefficient.Coordinated random variables Bad
Or give me something.
- $-1\leqslant Corr\left(X,Y\right)\leqslant1,或\left|Corr\left(X,Y\right)\right|\leqslant1$
- $|Corr(X,Y)|=1$ . . . . . .$X,Y$It's almost linear, which is...$P\left(Y=aX+b\right)=1$ The relevant coefficient is$[-1,1]$Here's some evidence of the value taken between. = \ \ \ \ \ = = = = = = = = = = = = = = = = } } } } } } } } } } } } } } } } } = = = \ \ \ \ \ \ \ = Y})=2+2p{xy}$$ $$0\leq Var(\frac x{\sigma _X}-\frac y{\sigma _Y})=\frac{\mathrm{Var}\left[X\right]}{\sigma _X^{2}}+\frac{\mathrm{Var}\left[Y\right]}{\sigma _Y^{2}}-2Cov(\frac x{\sigma _X},\frac y{\sigma Y})=2-2p$ A combination of two variants would prove the original proposition.
Mathematical Expectations Matrix and Convergence Matrix
We use the matrix.$n$The mathematical expectations of the wi-random vectors, the difference between the parties; the difference between the parties, of course, leads to a correlation coefficient. Mathematical Expectations Matrix $$E\left(X\right)=\left(E\left(X_{1}\right),E\left(X_{2}\right),\cdots,E\left(X_{n}\right)\right)^{\prime}$$ Coordinated Matrix (a non-negative matrix) $ = \begin{matrix}\pectorname{Var} (X 1)&\operatorname{Cov}(X_1,X_2)&\cdots&\operatorname{Cov}(X_1,X_n)\\operatorname{Cov}(X_2,X_1)&\operatorname{Var}(X_2)&\cdots&\operatorname{Cov}(X_2,X_n)\\vdots&\vdots&&\vdots\\operatorname{Cov}(X_n,X_1)&\operatorname{Cov}(X_n,X_2)&\cdots&\operatorname{Var}(X_n)\end{pmatrix}$$
Distribution of conditions and expectations
The theory of the overall distribution of conditions is largely consistent with the form of the theory of the distribution of conditions that we have studied in chapter I. It must be a one dollar distribution of the initial distribution of the binary.
Dispersed condition distribution
Or is the probability of simultaneous occurrence divided by the probability of a condition occurring at a time when the probability of a condition occurring is a marginal distribution? First give a distribution column of 2D discrete random variables $$p_{ij}=P\left(X=x_{i},Y=y_{j}\right)$$ So the condition distribution is $$P_{ij}=P\left(X=x_{i}\mid Y=y_{j}\right)=\frac{P\left(X=x_{i},Y=y_{j}\right)}{P\left(Y=y_{j}\right)}=\frac{P_{ij}}{P_{.j}}$$
Continuous conditions distribution
Or is the probability of simultaneous occurrence divided by the probability of a condition occurring at a time when the probability of a condition occurring is a marginal distribution? $$P\left(X\leq x\mid Y=y\right)=\int_{-\infty}^{x}\frac{p\left(u,y\right)}{p_{r}\left(y\right)}du$$
It's a mathematical expectation.
The mathematical expectation of condition is the mathematical expectation of a certain distribution. $$\left.E\left(X\mid Y=y\right)=\left{\begin{matrix}\sum_{i}x_{i}P\left(X=x_{i}\mid Y=y\right)\\int_{-a}^{\infty}xp\left(x\mid y\right)dx,\end{matrix}\right.\right.$$ The mathematical expectation is the expectation, but it's the same.$y$It's a function, so we often give another pattern. $E(X|Y)$
Full expectation formula $$E\left(X\right)=E\left(E\left(X|Y\right)\right)$$
Big-digit theorem and center-critical.
Condensity
Concealed by probability
We've come up with the theory long ago that probability is the stable value of frequency.
- Frequency$v$Right chance.$p$Absolute deviation$|p-v|$Approaching stabilization value
- Because of randomity, we can't rule out the possibility of big deviations, but the probability of big deviations getting smaller.
Now we have a general definition. Definitions: Establishment${X_{n}}$It's a random variable sequence. $X$It's a random variable.$\varepsilon$ Yes. $lim n\info}P (Xn-X|)<\varepsilon) = $1 It's called a random variable with probability.$X_{n}\longrightarrow X(P)$
Weaknesses by distribution
The distribution function is also an important image of the probabilities problem.
We've been able to tell you about the downs and downs of the function in the mathematical analysis, but it's too strong a condition to make the distribution function no longer conform to the distribution function, so let's put it down.
Definition: Set a distribution function column${F_{n}(x)}$ What if...$F(x)$Any point of continuity. $$\lim_{n \to \infty} F_{n}(x)=F(x)$$ is the distribution function column${F_{n}(x)}$Weak harvest$F(x)$ Or a random series of variables.${X_{n}}$Condense X by distribution Recorded$X_{n}\longrightarrow X(L)$ Or... $F_{n}(x)\longrightarrow F(x)(W)$
Theorem: Probability-based roll-out by distribution Theorem: Consistency on the basis of distribution is a sine qua non of probabilization (at the same time)
Almost everywhere. Probability one.
It's almost everywhere. $$P(\lim_{n\to\infty}X_n=X)=1$$ In general, the difference between probability and probability is just changing the position of the limit sign.
In fact, almost everywhere, it's much more mathematically than probabilistic, and he's a version of the probabilistic theory, similar to the concept of a dot-compression function in mathematical analysis.
It's almost everywhere that the random variable sequence is condensed and random at every point. Probability only requires that probability be calculated first, that the probability be limited to one, not every point.
Large-digit laws
In practice, it is recognized that the arithmetical averages of a large number of measurements are also stable.
Benuli's law of big numbers.
He describes the connection between frequency and probability as a retort to what's ahead. Theoretically: Set$S_{n}$Yes.$n$The event at the Hobernuli experiment.$A$Number of incidents$p$For each experiment$A$There's a chance that there's a chance that there's any varepsilon. >Yes. $lim n\info} (\frac{S}n}-p|<\varepsilon) = $1 Benuli's law of big numbers gives us the theoretical basis for using frequency to determine probability.
Chelby Scheffer's Law of Major Numbers.
Theorem: Set${X_{n}}$ It's a series of random variables that are not relevant.$X_{i}$ There's a difference and there's a common upper boundary for any \\varepsilon>0$ $$\lim_{<\to\infty}P\left(\left|\frac{1}{n}\sum_{i=1}^{n}X_{i}-\frac{1}{n}\sum_{i=1}^{n}E\left(X_{i}\right)\right|<\varepsilon\right=$1. We've weakened the distribution requirement.
Markov's Law of Numeracy.
Set${X_{n}}$ Yeah, random variable sequences.$\frac{1}{n^{2}}D(\sum\limits X{i})\longrightarrow0$ For any \\varepsilon>0$ $$\lim_{<\to\infty}P\left(\left|\frac{1}{n}\sum_{i=1}^{n}X_{i}-\frac{1}{n}\sum_{i=1}^{n}E\left(X_{i}\right)\right|<\varepsilon\right=$1. Markov's Law of the Magnificent is a condition for another big law to repeat.
The Law of Sinchin.
Theorem: Set${X_{n}}$ It's a series of random variables that are irrelevant and subject to the same distribution and have mathematical expectations.$\mu$For any \\varepsilon>It's all there. $$\lim_{n \to \infty}P(|\frac{1}{n}\sum\limits X_{i}-\mu|\ge\varepsilon)=0 $$ The Sinchin Big Number Theorem is the basis of the theory of subsequent rectification.
It's very restrictive.
Some kind of deviation is the sum of small errors caused by a large number of small and incidental factors, and the small errors caused by all these different factors are independent of each other, and each of them has little effect on the sum.
It's very difficult to calculate using a volume formula at this point, but when we're drawing the graphics, we find that there's a lot more to be done.$n$The increase and the closeness of the function to the normal distribution, which is the central limit.
The Lindbergh-Levi Center is extremely restrictive.
Theorem: setting random variables$X1, X2,…, Xn$Independently, subject to the same distribution, with the same limited mathematical expectations and differences if Remember $Y }=\frac{X_{1}+X_{2}+\cdots+X_{n}-n\mu}{\sigma\sqrt{n}}$$ 有 $$\lim_{x\to\infty}P\left(Y_{n}^{== sync, corrected by elderman == @elder man The sum of all random variables that are independent and distributed can be approximated by normal distribution.
Under certain conditions, the sum distribution of a large number of stand-alone random variables is normal.
The Mofo-La Plas Center is extremely restrictive.
He's the narrow form of the first theorem in two distributions, considering the two distributions as multiple Bernouli distributions and
Theorem: Set$Y_{n}$, subject to two distributions and with mathematical expectations and differences, random variables $\bar{Y}=\frac{ Y-np}{\sqrt{np(1-p)}}$ The probability density function is$\frac{1}{\sqrt{2\pi}}e^{\frac{-t^{2}}{2}}$
This theorem means that the normal distribution is the limit of the two distributions.
The law of power and power and their strength.
Large-digit law is an important theory in modern probabilistic theory and an important bridge between probabilistic and statistical theory.
Most of the theories in mathematics are named after theorem, which in turn reflects their results through rigorous extrapolation. The law is usually used to describe patterns in nature and is based on observations. The close connection and importance of his application can be seen in the introduction of this law in mathematics.
Basic law of big numbers.
Large-digit laws describe a phenomenon for a random series of variables$\left { X_n\right }$ It's... Front$n$Item average $$A_n=\frac{X_1+X_2+\cdots+X_n}n$$
For meeting certain conditions${X_n}$, when the number of random variables$n$Very large, their averages are highly likely to be valued.$\mu$
$$A_{n}\to\mu $$
This is a value.$\mu$Usually.$X_i$Mathematical expectations. Depending on the mode of consolidation, the law of large numbers is divided into strong and weak law of large numbers.
In essence, we're still looking at the stability of average results in a lot of random phenomena.
The law of powerful numbers.
Its mathematical form is $$P\left{\lim_{n\to\infty}A_n-\mu=0\right}=1$$
MeaningWhen the length of the series of random variables is infinite, their averages necessarily tend to be constant.I don't know. We call this random variable sequence a strong number law.
The law of weakness.
The mathematical form is $$\forall\varepsilon>0:\lim_{n\to\infty}P{|A_n-\mu|<\varepsilon}=1$$
MeaningWhen the sequence of random variables is of infinity length, the probability of their average approaching the fixed value is close to 1.I don't know. Calls this random series of variables consistent with the law of a weak large number.
Distinction and linkage
A strong law of numbers is easy to understand, similar to the contraction of columns. And the law of the big and the weak is relatively difficult to understand and at first glance seems to be no different from the law of the strong. Actually, it's about right.LimitsUnderstand.
The powerful law of numbers is an act of constriction, or he means almost everywhere. And the law of weakness requires only one probability of conceiving.Consistency by probabilityI don't know. That means that if a random series of variables meets the law of strong numbers, then he must also meet the law of weak large numbers.
In particular, if we look back on the section of this paper, "The Lindbergh-Levy Center is extremely restrictive", we can conclude that the form of enrichment in this section is just as distributive.
- The law of powerful numbers: almost everywhere.
- Weaknesses Law: Concealed by probability
- Centers are extremely restrictive: decrease by distribution
Most of the random variable sequences in practical application are also subject to the law of strong numbers, so all the later references to the law of large numbers refer to the law of strong numbers.
Caution: Largest laws are by their very nature established by the innumerable number of tests, and we cannot carry out infinity large experiments, so in a limited number of experiments, any major deviations are not in conflict with the large number laws themselves. The law of big numbers will not affect the independence of the experiment.
- Title: Elementary Probability: Random Events, Probability Models, and Random Variables
- Author: Hyacehila
- Created at : 2023-03-18 13:28:02
- Link: https://hyacehila.github.io//blog/2023/03/18/elementary-probability-notes/
- License: This work is licensed under CC BY-NC-SA 4.0.