The Application of Statistics and Statistics of Application

Hyacehila

"Statistics represents the science of learning from data."
Statistical science should be about learning from data.

The questions in this article can also be addressedThe Statistical Crisis in ScienceScientific theory and practical experienceHow the concept of a relatively close read together is developed in different contexts.

Today, when big data and artificial intelligence, especially the generation AI, are sweeping the globe, we hear often the argument that statistics pass the test of the times. Is statistics outdated because of the inability to process big data?

It's a hypocritical proposition. Statistics never fear the size of data, making traditional statistics seem to be a drag in modern applications, not data.VolumeAnd we're dealing with data.AttitudeandParameter

It is intended to discuss why statistics are far removed from application sites on the path to mathematical rigour and how modern data science and machine learning can bring problems back into reality through a “data-driven” approach, focusing on two seemingly entangled concepts: “Application of statistics” and “Statistics of application”.

Lost in the original: from solving problems to pursuing mathematics.

The “leveling” of mathematics and statistics

The original aim of mathematics was to study the real world (e.g. measuring land, calculating astronomicals) but it gradually developed into a clear hierarchy:

  1. Pure mathematician.: Study abstract structures and logic, pursuing the ultimate breakthrough of theory.
  2. Applied mathematician: Although in the name of “application”, it is often the highly abstract structure (e.g. the condensity of the PBDE numerical decomposition) that is studied rather than the immediate reality.
  3. Math ApplicatorResearchers in specific disciplines (physical, engineering, economic) who abstract the real world, form the theory of their own disciplines and use mathematical tools only at a very small point.

Statistics also seem to be unsavory and embark on a similar path.

The purpose of statistical studies, at the beginning of their life (e.g. during Fisher), was to analyse both agricultural experimental data and biogenetic characteristics.Addressing specific application issuesI'm sorry. As disciplines developed, statisticians began to build a more complete theoretical system. To make the system mathematically sound, fine assumptions are introduced and theoretical reasoning is becoming more complex and sophisticated.

Ultimately, statistics also form a tower structure similar to mathematics: the top-level statisticians no longer process specific data but research abstract distributive, progressive and optimal proof.The gradual alienation of applied statistics for statistical applications— That is, to find nails with hammers, to force reality to be fine-tuned to fit the perfect theoretical models.

The application-oriented disciplines gradually distance themselves from the application site; the corresponding student development is also moving away from the application problem and it is ultimately difficult to develop the talent to solve the application problem.

Leo Breiman's Warning

As early as 1994, Leo Breiman, a statistical giant at the University of California in Berkeley, had warned of deafening. In one of his speeches, he mentioned:

"Nor is any discipline as far away from theory and practice as statistics, and most statistical theories and statisticians deal with issues that are far apart, as if they were living in different worlds."

Breiman has a keen point that statistics should not be aimed at being second-rate mathematicians, but should be nurtured.Experts collecting, analysing and drawing conclusionsI'm sorry. If statisticians cannot find pleasure and solve problems in specific applications such as voice processing, astronomy, medicine, etc., then the identity crisis in statistics will never be solved.

Late awakening: statistics at the crossroads

Breiman's voice may have been too radical at the time, but 25 years later, it's become a consensus in the academic world.

In 2019, in a major report funded by the National Science Foundation of the United States, Statistics at a Crossroads, several leading statisticians, including Bin Yu, finally collectively acknowledged this. The report notes that statistics are at risk of being marginalized and that fundamental cultural changes are required.

The most striking point in the report is thatPractice must be re-centre of statistical evaluationI'm sorry. There has long been a false dichotomy in the statistical community: “theory statistics” are considered to be a noble upper-level building, while “application statistics” are considered to be a secondary manual labour. This value leads to the alienation of the discipline: scholars have to prove their intelligence by publishing obscure mathematical theories in order to gain a permanent teaching career (Tenure), and lack interest in solving scientific or social problems.

This late report calls for the evaluation of a statistician standard that should no longer be the complexity of mathematical techniques but rather the solution to actual problems.Impact

The shackles of the big data age: difficult to verify assumptions

When time comes to the twenty-first century, the eruption of big data has exposed the embarrassment of traditional statistics.

Statistics can handle big data, but not data that no longer satisfy the assumptions.

Traditional mathematical statistics are often based on a series of fine assumptions, none of which is more classic than:

  • IID Assumptions: Data are distributed independently and independently.
  • Normality assumptions: The error items are subject to normal distribution.
  • Linear assumptions: The relationship between variables is linear.

In the age of small data, these assumptions are essential to our understanding of the world's crutches. However, in the age of big data, the sources of data became extremely complex, with exponential growth in dimensions, and the correlation between observations became complex.

And now, if we still insist,Let's assume, then extrapolate.The traditional paradigm reveals that data from the real world almost never fully satisfy the perfect assumptions that are derived from mathematics.

  • When we use linear regression to force the integration of a highly non-linear complex system;
  • When we use p to test a data set that does not satisfy the normal distribution and has a large sample mass;

We get one of them all the time.A precise wrong answerI'm sorry. With difficult assumptions to verify, statistical science has lost its promise of application in a truly complex world.

But applied statistics are not in reality. In many cases, statistical models can also describe some of the "causal and consequentials" with relationships, not because the models identify their own causal structures from the data, but because they are not.Causation is embedded in the researchers ' understanding of the problem, in their research design and in their selection of variables.I'm sorry. Often researchers do not “discover causes and consequences” from the data, but first judge what the field knowledge may affect, then use statistical models to estimate strength, comparative direction and quantify uncertainty. Strictly speaking, this is not a causal recognition that is entirely from the data, but it is the reality of many applied statistical studies today: models do not create causal explanations, but they are merely trying to paraphrased and quantified the causal understanding that the researchers have brought into the analysis.

Applied statistics: data science and machine learning relay

Since traditional theoretical statistics have touched the wall of application, who has taken the great thing of applied statistics?

Yes.Data Science or Data MiningandMachine Learning (Machine Learning)

The difference between the two: from Inference to Prevention

People often ask: What is the difference between statistics and machine learning? One of the excellent answers is:They can't all be around the same question: how can we learn from data?

But the difference in their focus explains why machine learning is so eccentric in the age of big data:

  • Statistics: Emphasis on statistical extrapolation (Inference). Focus on confidence interval, hypothetical tests, estimation of parameters. It attempts to explain the model (Explainability), by which it has to make strong assumptions about data distribution (e.g. logical regression).
  • Machine Learning: Emphasis on Forecasting. It treats the data generation mechanism as a black box, not a strong understanding of internal parameters, but a desire to use the data to create a data base. $f(x)$ It's the best way to predict it. $y$。

In high-dimensional, structural realities (e.g. image identification, referral systems), human beings simply cannot predict the right mathematical distribution. Machine learning has abandoned the quest for perfect form and assessed the good and bad of models by empirical means such as cross-checking (Cross-Validation) rather than relying on theoretical progressive normality.

But to date, there are still a number of statisticians and some social scientists who see “interpretation” as a higher and more scientific objective, while demeaning “predictation” as a less academic, engineering-oriented technical activity. This view is in itself untenable.Interpretation and prediction are not a science-non-science opposition, but are two equally important objectives in scientific research. Interpretation helps us understand mechanisms, organizational knowledge, theory; predictions help us to test whether models capture stable structures and help us measure whether a method is useful in unknown samples, future situations and real-life decisions. A model that can only explain but cannot demonstrate stability in new data is hardly truly mastery of the world; and a model that can make accurate predictions on a sustainable basis must not be easily described as “unscientific”.

This is what the first point of the post is:The approach must first solve the problem.

The rise of the fourth paradigm

The data science was born precisely to cope with the explosion and the complexity of the sources of this data dimension. As a scientific researcher,Fourth paradigm (data-intensive scientific discoveries)Data science no longer relies on the a priori knowledge required for theoretical simulation or computational simulation, but is based directly on data and patterns.

It is not simply a confluence of disciplines, but a deep convergence of several types of capabilities:

  • Computer scienceInfrastructure is provided: from core database technology (storage and extraction) to cloud computing and distribution systems, the rapid leap in computing capacity makes the processing of big data possible.
  • StatisticsThe blog provides an evolution of methodology:
    • Computation statistics (e.g., MCMC, EM algorithms) replace some traditional resolution and provide a powerful support for the processing of complex statistical structures.
    • The exploratory data analysis (EDA) is re-energizing in the age of big data, helping us “sniff” information from the data ocean.
    • The combination of compression of perception and thinness processing in high-dimensional statistics provides a mathematical antidote to the “dimensional curse”.
  • Artificial intelligencePowerful tools are provided: from traditional models to evolutionary learning through modern machines, to the outbreak of in-depth learning, providing a powerful capability for automated characterization extraction and non-linear modelling.

This confirms one point:While statistics have many tools to use to complete the system, the fundamental 0-1 breakthrough in statistics must have been the result of addressing major application problems. And neither Fisher nor today's data science was created to solve real problems, not to perfect the mathematical structure.

Statistics and machine learning: The difference between return and return is a repeat of the past.

Kiri Wagstaff is here. Machine Learning That Matters It says:

"Much of current machine learning (ML) research has lost its connection to problems of import to the larger world of science and society."

(Most of the machine learning research is now lost to the scientific and social communities. I'm not sure.

That criticism sounds familiar. If we put that in the first sentence, "Machine Learning" Replace with "Statistics"And it could be seamlessly present in the 1980s, to criticize statisticians who were obsessed with progressive theory and ignored reality data. He seems to be in collusion with Leo Breiman's famous speech at the University of California in Berkeley in 1994.

History seems to be repeating itself at an unprecedented rate: when a field is pursuing simple ** indicatorsSOTA, I'm not.When the problem itself is repeated, it is the same as statistics of the year.

This is also a serious challenge for the AI sector:

  • Engineering overwhelms theory: a great deal of research has focused on “finding” techniques, and the models of End-to-End have become more complex and the black box is becoming more powerful.
  • Theories lag behind practice: while deep learning sweeps the top lists, the mathematics behind them (e.g. interpretation of generalization, rudge-precision paradox) fall far behind. Academics often have difficulty answering the practical questions posed by industry, and even have the embarrassment of “industry is a leader in academia”.

If the crisis in traditional statistics isThe theory is out of line.And the crisis of modern machine learning isReality has abandoned theory.And what's in common is thatResearch is out of touch with real practice

This brings us back to the substance of our discussions in both areas. What's the difference between statistics and machine learning? In short:There is no difference in substance. They all focus on the same question — how can we learn from data?

If one wishes to summarize their main differences at this time:

  • Statistics: Statistical inferences (confidence interval, hypothetical tests, optimal estimates) focused on forms in low-dimensional issues, emphasizingExplanatory
  • Machine learning: Focusing on predictive accuracy in high-dimensional issues, emphasizingBroadening capacity

Although the focus is different, the two areas are increasingly being integrated. The core of data science is not whether you're using t or deep nervous networks, but whether you really use these tools to solve one.Existing scientific or social problems

Direction of application of data science

Then there is a discussion of how data science should be applied, broadly speaking in two directions:

  • Continuing to serve academia: using data science to solve complex problems that are being addressed in the physical, biological, social sciences (this is what many “calculations X” are doing).
  • Turning to industry: addressing the last-end reality world applications. Here, we can further break down into two closely related but different-focused roles:
    • Data Scientist, DS: Focus on extracting insights from data (Insights) to assist enterprises or organizations in scientific decision-making through analysis. They are closer to “consultants” and “discoverers”.
    • Machine Learning Engineer (Machine Learning Engineering, MLE): Focus on construction products. Their first task is to convert algorithms into engineering, landing and practical services.

In theory, MLE is part of a broad DS, but they represent two different forms of value excavated from data: one isTo understand.One of them is...For action.

In either direction, the core is designed to solve that real problem.

Return to the fields: keys to the backyard

Professor Terry Speed of the University of California at Berkeley once had a famous saying:

"Statistics should have been the subject of other disciplines, and I'm so interested in statistics, that it's like putting keys in the backyard of any discipline."

This is perhaps the most important message to be borne in mind by all data workers, whether they call themselves statisticians or data scientists.

Statistics should not be a mathematical game in ivory towers, but rather a tool for solving practical problems. The core value of this, whether it be called applied statistics or data science, is whether we can use the data in our hands to find definitive answers to the problems of biology, economics, medicine and even social sciences.

When we put aside our belief in assumptions and embrace the true complexity of data, we can truly achieve the application of statistics that give new life to this ancient and fascinating discipline in the data age.

  • Title: The Application of Statistics and Statistics of Application
  • Author: Hyacehila
  • Created at : 2026-02-14 04:00:00
  • Link: https://hyacehila.github.io//blog/2026/02/14/application-of-statistics-and-applied-statistics/
  • License: This work is licensed under CC BY-NC-SA 4.0.
Comments