Statistics and Truth: How to Use Chance
- In the final analysis, all knowledge is history.
- In abstraction, all science is mathematics.
- All judgments are statistical, based on reason.
The questions in this article can also be addressedAnscom Quartet: Visualized power and statistical illusion、Figures don't lie, but the liars make them up: talk about statistical fraud from Benford's particular law.How the concept of a relatively close read together is developed in different contexts.
Statistics and Truth - How to Use Coincidentally, the issues discussed were:
- How to design experiments to provide the required information
- How to get all the useful information from the results of the experiment
- How to apply this information in practice
Many human efforts are seeking the truth. However, truth is not available in the strict sense of the word.We're only looking for acceptable knowledge.I'm sorry. The blogger adds:Knowledge is not truth, but it should be used as much as possible.
It's not a professional book. It reviews the history of statistics and the many issues that arise in their development, and by way of example, it describes statistical thinking as a science rather than a collection of facts.
Uncertainty, randomity and creation of new knowledge
Just like the movement of basic particles in physics, the insulation of genetic factors and chromosomes in biology, and the behaviour of people in a tense society, in nature,UncertaintyIt's inherent. Instead of following these phenomena,DecisionariesThe rules, rather they are based on the principles of the law.Random theoryUncertainty of the law; it has become the necessary foundation for the development of the natural, biological and social sciences theory.
So how does human beings decide in the face of uncertainty? How can new phenomena be summarized or theory proposed from specific observational data? Is this a process involving art, technology, or science?
It was not until the beginning of the twentieth century that attempts were made to quantify uncertainty in order to answer those questions. This effort cannot be said to have been fully successful, but its results have entered many areas of human activity: They bring new research ideas, promote the development of natural scientific knowledge and change our thinking methods.
The random series show maximum uncertainty (or confusion or entropy) and are the most common method for studying uncertainty. Random columns do not follow any particular pattern and can be generated in a variety of ways, including random and artificially generated columns derived from natural observations.
Many random methods have been developed to address difficult questions that are difficult to obtain and to obtain new information and develop new ideas. For example, Monte Carlo Method、Random sample surveysandExperimental design methodsI'm sorry. Randomness can also be used to address some of the issues in the theory of game and decision-making.
There are interesting examples of randomity:
- Monkey typewriterIf we have a smart monkey that keeps typing, it should be able to hit all Shakespeare's works for a limited but long time. It's about the probability. $10^{-41600}$。
- Gambler MisunderstandingIf there are a significant number of girls born in the last few days, the chances of a couple getting a male baby will be increased. However, simulations or Monte Carlo experiments show that a stable and even system can present some partial imbalances at a certain frequency (i.e., the associated phenomena occur over a short period of time).
- Biological cycleTotal survival of most animal species is approximately three years in one cycle. The prevalence of this phenomenon has led many to believe that a new law of nature may have been found. However, the probability of one number being greater than the other two is one third for the three random numbers. This gives an average time lag of three years between the two peak years of the above-mentioned problems.
Randomness can also be used for sensitive questionnaires. If we ask the question: “Do you smoke marijuana?” I am afraid that we will not get the right answer. In this regard, two questions (one of which is irrelevant) can be listed:
- S: Do you smoke marijuana?
- T: Is your phone number even at the end?
The person questioned was then asked to throw a coin, to answer the face when it appeared, and to answer it when it appeared. At that point, the questioner did not know which question the respondent had answered, so the information could be kept confidential. Based on these answers, the following estimates can be made to extrapolate the proportion of cannabis users.
Randomness has also gradually entered decision-oriented natural science research. For a long time, it has been believed that the phenomenon of nature has clearly been predefined, the most extreme of which can be found in the idea of “magnetic god” in Laplace. The Mathematical Spirit is given unlimited mathematical ability: if at any point all the measurements of the state at that time are known, it predicts all the events that will occur in the future world.The La Plass Demon.)。
It is gradually discovered that this is not possible. Humans have difficulty in knowing the initial state of the system precisely because of measurement errors. In such cases, minor differences in the initial state may result in significant differences in the prediction of the system ' s follow-up status. Three developments followed by random entry into the natural sciences:
- Ketelle (A. Quetel, 1869) Use the concept of probabilistic theory to describe sociological and biological phenomena.
- Mendel, 1870 His genetic laws were framed through simple random structures (such as dice throwing).
- Boltzmann, 1866 One of the basic propositions in theoretical physics, the second law of thermodynamics, was given a statistical interpretation.
The introduction of statistical concepts in physics begins with the need to address astronomical measurement errors. Randomity has now become a basic concept and a technique for expressing quantitative rules.
Uncertainty management - statistical development
The idea of statistics existed in ancient times, but as a discipline, its history was short. The origins of statistics can be traced back to the earliest days of humankind, and it was not until the recent era that they became an important discipline in practical application. This development also raises the question of what statistics are.
The earliest records of statistics can be traced to ancient times. Even before arithmetic emerged, the originals had been scratched on trees to calculate livestock and other property. The need to collect data and record information arises when human beings abandon their own nomadic lives and begin organized social life. The ancient human race must pool the resources at its disposal in order to distribute and use them correctly and to plan for future needs.
Terminology STATISTICS The roots of the words, in Latin, mean the country. STATUSI'm sorry. In the mid-century, the new term, created by the German scholar G. Achenwall, was meant to mean “the collection, processing and use of data by the State”. Countries use statistics to describe the current state of the problem and to guide the way things go.
International cooperation is essential to make statistics a useful research tool. In order to exchange experiences and develop common standards, European countries hosted (about 10) international statistical conferences between 1853 and 1876. In 1885, the fiftieth anniversary of the London Institute for Statistics was marked by a proposal to establish an international statistical institute. After many discussions, the resolution to establish a permanent international organization, the International Statistical Institute, was reached. So, on June 24, 1885,International Statistical Institute, ISI Born.
ISI has been expanding its scope for the past 100 years. Under ISI management, branches of mathematical statistics, probability theory, statistical computation, sample surveys, administrative statistics and statistical education have been formed.
As noted earlier, the statistical roots are meant to be data collection and collation and to be used in public policy formulation. But statistics are not just data per se, but also methods for researching and analysing data containing uncertainties.
Statistical research contains the truth of uncertainty and also addresses such situations in the real world. It is recognized that, although knowledge built up by special to general patterns is uncertain, once the uncertainties can be measured, the knowledge acquired is determined, even if it varies in different types. This structure can be written as follows:
$$ \text{不确定性的知识} + \text{度量不确定性的方法} = \text{可用的知识} $$
This basic equation describes an effective risk management approach and also removes reliance on Mr. Oracle and fortune teller. It places the future within a framework that allows for informed decision-making in the immediate future:
- If we have to make choices without any certainty, mistakes are inevitable.
- Since mistakes are inevitable, it is desirable that the choice (to create new, uncertain knowledge) be made according to a certain pattern, with knowledge of the frequency of errors (i.e. knowledge of uncertainty measures).
- Such knowledge can be used to identify a pattern of decision-making, reduce blindness and minimize the frequency of decision-making mistakes or the loss they cause.
This is precisely the subject of statistical discussion. In the real world, we can only make decisions by drawing conclusions based on incomplete or poor information; the judgement given to the data on the basis of the reasoning lacks precision.
One of the main concepts that leads to conclusions from generalization is thatQuantification of uncertaintyQuantification of uncertainty has been controversial. To that end, statistical institutes have been established to study different methods of measuring uncertainty. The current Bayesian and frequency schools are the two main schools of statistics that have been formed.
Thus, a new discipline of information and extrapolation from data has emerged, and the term “statistics” has also expanded from data itself to interpretation of data. Coincidence is no longer a matter of concern or a manifestation of ignorance; it is a logical way of expressing the knowledge we have.
Statistics are not just a science, but a synthesis:
- Science: It is similar to science and technology that is guided by certain fundamental principles and has broad application. These technologies cannot be applied in fixed models; users must select the technologies to be applied, based on the expertise available, and amend them as necessary.
- Process: It is like the quality control procedures in industrial production processes. The methodological approach to statistics is developed in a management system that ensures that products meet expected quality and maintain stability.
- Arts• Statistical methodological approaches relying on general reasoning are neither fully codified nor uncontroversial. ** Different statisticians may have different conclusions on the analytical treatment of the same data set. ** The data that are usually actually given contain much more information than the information obtained from statistical tools.
The rationale and strategy for data analysis — cross-checking of data
The format of statistical analysis has changed over time, but the “extracting all information from data” or “comprehensive and revealing” as the purpose of statistical analysis has not changed. The methodology of statistical analysis has changed over the years and the development of data analysis is broadly as follows.
The blogger says:Description of statisticsandTheory StatisticsTwo areas of statistically different methodology were considered. The former aims to consolidate the given data set within the meaning of “statistical description” and to use techniques to express the visual and visible characteristics of the data. In theoretical statistics, the synthesis or description of statistics depends on a particular random model. The distribution of these statistics is used to determine the extent of uncertainty in extrapolating certain unknown parameters. So this is calledInfer data analysis。
K. PearsonThe first statisticians to try to communicate with both. He used the results of a descriptive analysis based on the rectangular and histogram to extrapolate the distributional population. It's famous.CalculatorandCarp's on the line.I'm sorry. The government has been making a difference in the past few years.Fisher A series of exceptionally rich statistical ideas emerged. Fishery developed an accurate sample test based on normal assumptions, offering to use standard test tables to assist in testing, which usually give a threshold of 5% and 1%.
In the 1930s, there was also a systematic development of the experimentally designed method of collecting data, pioneered by Freechilles, which enabled analysis of data through a specific method of differential analysis and meaningful interpretation of data: experimental design guides how to analyse data, while data analysis shows the structure of the experimental design.
The development of sampling methods can be seen after the early 20s. This method is used by surveyors to collect a large amount of data based on information obtained from randomly selected individual responses to a set of questions. Ensuring data accuracy and comparability is a common topic in sample surveys.
Common statistical extrapolation methods also require an integrated approach: a proper understanding of the given data and their deficiencies and characteristics, followed by the selection of random probabilities models or model communities suitable for data analysis.Tukey I've been called Explore data analysis EDA (Explorory Data Analysis) The approach, the philosophy of EDA, is to understand the basic characteristics of data and then apply robust processes to adapt data to a possible wider random probabilities model community.
The entire data analysis process can be summarized as data collection, exploratory data analysis, extrapolational data analysis, implemented in an iterative manner, with data collection including historical information (and databases), sample surveys, and experimental design of three acquisition methods.
One principle of data analysis is that no additional assumptions are used that are not proven by current data or past experience; expert advice can be used as a reference.
Data analysis to answer client questions is not the only job of a statistician. To understand the nature of the given data, a broader data analysis is also needed to identify the questions that the data available answer and to base new questions on them and plan further studies. Statisticians are often asked to provide appropriate statistical methods (or software programs) for the processing of a data set, but do not have the opportunity to cross-test the data. If data have certain special features, they must be considered in processing; the entire process must also be monitored continuously to determine whether the original processing needs to be modified.
The objective of the statistical analysis is “to extract all information from the data from the observations”. There are sometimes certain deficiencies in the recorded data, such as errors and anomalies in the records, and sometimes even possible falsifications,One attempt by a statistician should be to examine the data in detail or cross-check them in order to identify possible deficiencies and understand the characteristics of the data. The next step is to present a suitable random probability model for data using a priori information and cross-check techniques. Data extrapolation analysis based on selected models, including estimates of unknown parameters, hypothetical tests, predictions of future observations and decision-making.
Weighted distribution — biased data
Beyond this, a sample survey is not always designed with a suitable sampling structure to ensure that the event has a specified (usually equal) opportunity to become a sample. In fact,Not all natural events produce sampling structures.
For example, some incidents could not be observed and were therefore missing from the record. In such cases, so-called tailing samples, cut-off samples or incomplete samples were produced. Alternatively, an event can be observed only with a certain probability, and its probability size depends on the nature inherent in the event, such as its visibility and the process used for observation, with the result being a sample of varying probabilities. Or the occurrence of an event varies randomly with the time or process of observation, so that what is recorded is actually an amended event. In statistical analysis, such changes or impairments must be modelled appropriately.
Certain events, although they have occurred, may have unobserved components. The distribution observed is thus cut off from a part of the sample space. For example, if we investigate the distribution of the number of eggs laid by an insect, the number of eggs laid is not detectable. At this point, the use of a distribution containing cut-off needs to be considered when conducting statistical probabilities model analysis.
More generally, an event has been recorded (or included in the sample) with a certain probability. That's right.Weighted probability distributionIt also has its own theory of probabilistic models.
An example of the application of weighted distribution can be found in a sample survey conducted using a different probability sampling method or probability ratio P.P.S. Sampling method (probability method to size). Many mega-demographic surveys use these methods.
Statistics and truth
Today, the understanding, research and practical application of statistics has expanded throughout the natural sciences, social sciences, engineering technology, management, economics, art and literature.
The general population uses statistical knowledge (through various data and analyses obtained in newspapers and consumer reports) to make decisions in everyday life or to develop future plans. As with reading and writing skills, one day the statistical thinking approach will become a necessary capacity for efficient citizens.
- For a Government, statistics is a tool for long-term and short-term planning for specific economic and social purposes.
- In scientific research, as I have already mentioned, data collection through effective design experiments, hypothetical tests, estimation of unknown parameters and interpretation of results play an important role in statistics.
- In industrial production, simple statistical techniques are used to improve and maintain product quality to the desired level. The R & D sector conducts experiments to determine the best formulation.
- Not only is statistics used in commerce to predict future demand for commodities; in medicine, the principles of experimental design are used for drug efficacy identification and clinical testing. In literature, statistical methods are used to determine the style of a writer; in court, statistical validation of the probability of an event occurring is used to supplement traditional confessions and other evidence in decisions.
The value of human activities can be enhanced by introducing statistical thinking into the design of the plan and by adopting statistical methods that effectively analyse data, evaluate feedback and control results.
If there were problems to be resolved, reference should be made to statistics rather than to a single expert committee. Statistical and statistical analysis can provide additional clues to the problem than the wisdom of a few experts.
Statistics have no intrinsic object and is a unique subject. Statistics exist and flourish in other areas. L.J. Savage once said:
** Statistics are essentially parasitic: they survive by studying work in other fields. ** This is not a sign of contempt for statistics, as for many host countries, there are no parasites that could die.
- Title: Statistics and Truth: How to Use Chance
- Author: Hyacehila
- Created at : 2026-01-10 04:00:00
- Link: https://hyacehila.github.io//blog/2026/01/10/statistics-and-truth/
- License: This work is licensed under CC BY-NC-SA 4.0.