The 60th Year of Data Science
It's a big topic. The main idea of this paper is from David Donoho's 50 Years of Data Science, which also adds something I find interesting. Nearly 10 years after Donoho published this paper, the wave of artificial intelligence generation is transforming society as a whole, and data science is no exception.
The questions in this article can also be addressedStatistics and Truth: How to use the accident (Statistics and Truth)、Anscom Quartet: Visualized power and statistical illusionHow the concept of a relatively close read together is developed in different contexts.
The following elements were taken into account in the preparation of this document, and the rest of the paper will not add any emphasis on the sources of content:
- 50 Years of Data Science (David Donoho)
- A Conversation with John W. Tukey and Elizabeth Tukey
- C. Radhakrishna Rao: A Century in Statistical Science
- What are the most important statistical ideas of the past 50 years? (Andrew Gelman & Aki Vehtari)
- Berkeley Data Science Planning
- Statistics at a Crossroads: Who Is for the Challenge?
Before the text begins
From the beginning of the development of human science,
Science is the tool of humanity to understand and interpret the world. Looking back at the evolution of human science, we can see the evolutionary context from concrete to abstract, and from theory to calculation.
The origins of mathematics in ancient China can be summarized in the nine chapters of the book, which are studying us.How to address some of the specific problems in the real world without addressing the general theory behind them♪ That can be considered ♪First paradigm: experiments and observationsI'm sorry. Natural phenomena are recorded and described, and lessons learned from them.
The widely known scholar in physics, Newton, gave the laws of sport and gravity to transform the laws of movement in the real world.Precision, abstract theory.I'm sorry. At this stage, scientific research focuses on the theoretically abstracting of the universal formula and theorem, which is known asII: Theory evolution。
In the development of modern science, with advances in computer technology,SimulationBe a common method. In the face of theoretical models that are too complex to solve by deciphering, computers are used to simulate (e.g. limited meta-analysis) based on known physical patterns. Relying on simulation research and developing science, or scientific computation, makes up3rd Parameter: Calculator Simulation。
The fourth paradigm of scientific research
Data science defined as after experimental observations, theoretical evolutions, computational simulationsFourth data-driven scientific research paradigmIt's an inspiring idea.
Now we have more problems: we don't know the theoretical mechanisms, we are constrained by physical conditions that do not allow experimental observations, we have too large a calculation or we have no parameters that make it impossible to rely on computational simulations. In this “uninformed” dilemma, the blogger says:Data driverAnother path is provided.
The first three paradigms can be called together.Knowledge paradigmThey are based on a certain a priori knowledge of the problem (experience, theory, equation). Where knowledge is sufficient, they are effective and accurate. However, in the absence of knowledge and incomprehensible mechanisms, the fourth paradigm is based directly on data, discovering patterns, digging linkages and adding to the hard-to-covered aspects of the knowledge paradigm. Data science is concerned with how to build testable and reusable judgements when mechanisms are incomplete.
Definition of data science
As an emerging discipline, data science is difficult to sum up accurately in a few sentences.
Some scholars simply define data science as a discipline used to process data and as a tool. This may be an accurate summary of what we do with it, but whether it's a matter of thinking about where it is.
Based on the above discussion of the scientific paradigm, I thinkFourth ParameterIt is an effective perspective for defining data science:
Data science is the discipline that implements the fourth paradigm of scientific research — the data-driven paradigm of scientific research; it is a discipline that works to make data useful.
This definition emphasizes two points:
- Changes in methodological layersIt adds to the problem that traditional scientific research paradigms are difficult to cover.
- High Applicability: Data science should not remain theoretical,It should be a highly applied discipline.I'm sorry. We have tried to make the data useful, with the goal of bringing it into the process of awareness, decision-making and intervention in the real world.
In this context, discussions could continue on the content of data science, its sources of statistics and computer science, and how it has evolved to this day.
Data science and statistics
A hundred years of statistics.
In the final analysis, all knowledge is history; in abstraction, all science is mathematics; in a rational world, all judgments are statistics. The #C.R. Rao
This statement by Professor C.R. Rao summarizes the place of statistics in the human knowledge system. As a recent statistical conglomerate, Rao experienced statistics from early descriptive analysis (Pearson era) to modern extrapolation theory (Fisher age) to the full 100-year history of today ' s integration with the depth of computing science. His work includes the Cramer-Rao heterogeneity and the Rao-Blackwell Theorem, which also shows how statistics can extract stable information from uncertain data through rigorous mathematical tools.
The core of statistics is “inferment”. It deals not only with data collection and collation, but also with how to predict the total from the sample, how to quantify uncertainty, and how to find signals in noise. In the past 100 years, statistics have provided universal language for natural sciences, social sciences and engineering technologies, allowing for clearer judgements in randomity.
With the advent of the big data age, statistics have also come under new pressure. Traditional statistics tend to rely on strict assumptions about data distribution (e.g., independence and distribution, normality), while large data in the real world tend to be high, heterogeneity and dynamic. As Rao considered in late years, statistics needed to embrace computing, embrace more complex models and re-examine what “inferment” meant in the new data environment.
In the picture of data science, statistics are not outdated because of the emergence of big data.It provides the extrapolation skeleton in data science.I'm sorry. Only by combining statistical inferences with modern computer computations can we understand the patterns behind the data.
Why data science?
“Data scientists” means professionals who use scientific methods to extract and create meaning from raw data.
For statisticians, doesn't that sound like the job of applied statisticians? “Statistics” refers to practice or science that collects and analyses a large amount of data.
For statisticians, this definition of statistics seems to cover most of the definition of data scientists, but it appears to be limited. The DSI programme is therefore incomprehensible: statisticians believe that their daily work throughout their careers is being packaged as new by managers.
Several perspectives on data science and its relationship to statistics are discussed below.
Is big data the core cause?
A widespread voice is that data science has emerged because of Big Data: data are so big that statistics cannot handle them that a new discipline is needed. This view, while popular, cannot be easily refined.
Statistics never fear large-scale data, and their history encompasses theories and practices for processing complex, big data. If only because of the increase in data, we could develop “big data statistics” without having to build another stove. In fact, this view is more derived from misconceptions among non-statistical professionals; the emphasis on the increase in data volumes alone does not explain why a new entity, “data science”, is needed.
Realistic drivers: skills gaps and the “golden fever” of talent
Data science is driven largely by industry successes over the past decade. The use of data by technology giants such as Google and Amazon has generated significant commercial returns and has allowed businesses to begin to systematically compete for talent.
In this heat, businesses have discovered an awkward reality: traditional discipline education does not provide the much needed talent.
- Graduated from traditional statistics Fan.Premise extrapolation and analysis, but often lack the capacity to process large-scale databases, prepare production-level codes and build complex software systems.
- Graduate of Computer ScienceThe sophistication of engineering and systems is often poorly trained in statistical thinking, such as extracting signals from noise, quantifying uncertainties, etc.
Mike Barlow is here. The Culture of Big Data It was noted that this skill gap led to a thirst for “data scientists”. This new title is a yes.Capacity integrationHigh demand: a qualified data scientist must be able to think as closely as a statistician, and to process dirty data and build systems as software engineers. The demand for this mix of skills is a watershed in the job market between “data science” and traditional statistics.
To true science.
However, if data science is only designed to fill the recruitment gap for commercial companies, it is at best a vocational training orientation rather than a science.
The establishment of the “data science” entity should not be limited to commercial recruitment. As mentioned earlier, the fourth paradigm, we need a door.Science on learning from dataI'm sorry. This scientific legacy of rigorous statistical inferences, while embracing modern computing capabilities, is used to address the problems of theoretical evolution and computational simulations that are difficult to grasp.
This vision is also the direction in which many statisticians have been building a foundation for the past 50 years.
And now,Math, statistics, computer science, artificial intelligence (mechanical learning) theory integrationThe evolving data science will be an effective tool. Statistics provide assumptions-based extrapolations, computer science provides high performance query and computing tools, and machine learning techniques provide methods for complex data modelling.
Data science is a highly applied discipline that we use to understand and manipulate the world.
The Future of Data Analysis
1962: The birth of the prophecy
John Tukey, in The Future of Data Analysis (1962), predicted many of today's data science issues more than 50 years ago. Tukey, to be blunt, said he was disturbed by his status as a mathematical statistician; by observing the development of mathematical statistics, he realized that his interest was in the development of a digital system. Data analysis。
The blogger says:Data analysis is a science.And not a branch of math. Mathematics seeks logical consistency and probability, while data analysis has three main elements of science:
- Intellectual Content
- Understandable forms of organization
- Relying on the empirical test as the ultimate criterion for effectiveness
In Tukey ' s vision, the Statistical Form Theory is only a fraction, not all, of this new science. He listed four main factors driving this new scientific development: statistical theory, and the following:Computation capacity, Big Data challenges and quantitative trends across disciplines. This 1962 list, which is in today ' s data science discussion, still has many real problems.
For Tukey, whether it optimizes the trajectory of the Nike anti-aircraft missile at Bell Laboratory or analyses the flow data of U-2 aircraft, he always looks from the practical point of view, looking for answers in an empirical way, rather than the assumptions in superstition textbooks. The blogger says:If a method is not even used in practice, then it's pointless to test whether it's worth it.。
From Tukey to Cleveland
Despite Tukey ' s arms shivers, academic statistics continued to react in a cold manner over the following decades, continuing to be obsessed with purely theoretical evidence. However, Tukey's Bell lab colleagues and a few visionary scholars have taken over the torch and continue to march in the wilderness.
- John ChambersIn 1993, the S Language Developer called for the establishment of “Grey Statistics”, warning of the risk of marginalization if statistics do not embrace the concept of inclusiveness of learning from data.
- Jeff Wu In his inaugural speech in 1997, the speaker directly introduced “Statistics = Data Science” and advocated renaming statistics as Data Science.
- William S. Cleveland The famous Data Science: An Action Plan was published in 2001 and a road map for action was developed for this discipline. He proposed to allocate academic resources to six areas: interdisciplinary research, modelling methods,Data Computation(c) Teaching methods, tools assessment and theory. In addition to theory, five other areas were virtually blank in the traditional statistical faculties of the time.
Counting environmental victories
Industry and practitioners have defined the future by code when definitions are still being debated by the academic community. The evolution of the computing environment changed the game rules from the early SSS/SAS to the S language developed by John Chambers, to the later R language.
Scripts became a new age's paper.I'm sorry. It is an accurate and abstract description of the calculation steps. When the quantitative programming environment like R is popular, data analysis is no longer on paper, but rather becomesRecoverable, shared, verifiable- The practice. People can directly run others ' codes, validate methods on different data, and improve analytical processes through performance measures.
So far, Tukey aboutData analysis is a science.The judgement was finally put into practice through the code and the computational environment.
Forecast-led statistical modelling
Two cultures: Generate vs projections
Leo Breiman published a sensational article in The State of Science in 2001. Statistical Modeling: The Two CulturesI'm sorry. He noted that there were two distinct statistical modelling cultures in the process from data to conclusions:
- Data Modeling Culture: Assumes that the data are generated by a known random process (e.g. linear regression model). The task of statisticians is to extrapolate the parameters of the model. Breiman believes that this represents 98 per cent of the academic statistical community.
- Algorithmic Modelling Culture: Treating the data generation mechanism as an unknown and complex "blackbox" does not attempt to decipher its internal mechanisms, but focuses on finding input through algorithms $x$ and Output $y$ to achieve the preciseest possibleProjections。
The statistical community has long relied heavily on data models, leading to a disconnect between theory and reality. The existing algorithm models, while lacking a critical theoretical underpinning (as in early statistics), have developed rapidly in computer science and industry with the ability to address complex realities. That explains it.Data ScienceWhy does it arise: it contains both traditional statistical assumptions and embraces a predictive-centric algorithm model.
The secret of a culture of prediction: a framework for common tasks (CTF)
If culture of prediction is an important pillar of data science,** the common mission framework (Common Task Framework, CTF)** is one of the mechanisms through which it can make sustainable progress.
The computational linguist Mark Liberman argues that the CTF is an important driver of machine learning and predicting the success of modelling, but is often overlooked by the mainstream statistical community. A typical CTF has three elements:
- Open training data sets: Contains features and labels.
- Competing: Committed to training the best predictive rules.
- The adjudicative system: An objective and automatic assessment of the accuracy of the projections is conducted using the “black box” test set.
From Netflix Challenge to Kagle Competition to modern deep learning revolutions like ImageNet, the CTF paradigm has been repeated. It turns the original blurry of research (such as early machine translation) into a new one.Quantifiable, comparable, re-emergibleThe engineering challenge.
Minimize the predicted error + CTF paradigm = the perfect for the performance of experience.
This model not only filters valid algorithms, but also changes talent needs. In the framework of the CTF,Information technology (IT) skills(Processing data, constructing systems, preparing scripts) becomes more direct than purely mathematical extrapolation. This explains why today's education in data science must include a great deal of computer science: in a new world dominated by prediction, code capability is as important as statistical thinking.
Data science now
As the discipline developed, early controversy over the relationship between data science and statistics gradually subsided. A pragmatic consensus has been reached in practice between academia and industry, first and foremost in the curriculum of higher education.
Consensus in education
The system of data science courses at higher education institutions such as the University of California at Berkeley (UC Berkeley) reveals a deep convergence of statistical and computer science skills. This integration is no longer limited to the discussion of the subject ' s attribution, but focuses on the development of practical competencies:
- Basic status of computing capability• Unlike traditional statistical education, the curriculum for modern data science considers programming to be a core skill. Students need to master languages such as Python or R for production-level code writing, version control (e.g. Git) and large-scale data processing, not just mathematical extrapolation.
- The predictions and extrapolations are weighed equally: “Application Machine Learning” is a key feature of the curriculum. This marks a shift in the focus of education from a single parameter statistical inference to a parallel approach to processing high-dimensional data and achieving accurate predictions.
- Integration of data projects• Data storage, retrieval and management are integrated into the core curriculum system, complementing the knowledge gap in traditional statistics on the principles of databases and distributed systems.
As Donoho has pointed out, there is a consensus that data scientists should be software engineers with statistical thinking and statisticians with expertise in software engineering.
2019: Statistics at the crossroads
This integration is reflected not only in education but also in the reflection of the academic community on its own positioning. In 2019, the National Science Foundation of the United States (NSF) funded a landmark report, Statistics at a Crossroads: Who Is for the Challenge?
This report, written by top statisticians like Xuming He, David Madigan, Bin Yu, echoes Tukey's predictions half a century ago. The report states categorically that statistics are at a crossroads: Without fundamental reform, the discipline risks being marginalized.
The core call of the report coincides with the idea of data science:
- Centre of practiceStatistics must return to the essence of “learning from data”, theoretical research must not exist for the sake of mathematical perfection alone, but for the service of “Better Practice”.
- The Revolution of the Evaluation SystemThe tendency of academia to over-reward the publication of purely theoretical papers for a long time must be changed. The report clearly recommends thatInterdisciplinary cooperation, software code development, data cleansing and collationIt should be seen as an academic contribution of equal importance to the publication of papers.
This marks the official recognition of the mainstream statistical community: In addition to extrapolating formulas, writing codes and processing data are an integral part of scientific research.
Broad Data Science
To define this discipline more comprehensively, Donoho proposes a framework** for a broad data science (Greater Data Science, GDS)**. The framework divides the scope of data science into six dimensions, noting the limitations of the current focus of academic research and the breadth of the needs for practical application.
- GDS1: Data collection, preparation and exploration: Data scientists often devote a significant amount of time to data cleansing and collation in practice. Although this process is often overlooked in traditional academic research, it is the foundation for ensuring data quality and analytical validity.
- GDS2: Data expression and conversion: involves transforming unstructured information (e.g. text, images, audio) and structured data (SQL/NOSQL) into mathematical forms suitable for modelling.
- GDS3: Data-based calculations: Covers programming language proficiency, replicability of analytical stream construction and use of high performance computing resources. In this framework, code scripts that document the full analysis process have the same scientific value as academic papers.
- GDS4: Visualization and presentation: Not limited to the generation of static charts, but also includes the use of modern graphic syntax (e.g. ggpllot2) and interactive tools to discover data patterns and effectively deliver information to audiences.
- GDS5: Data modelling: Traditional, generating models and modern predictive models. This is the area of greatest concentration of current academic research, but it is only part of the broader data science.
- GDS6: Science of data science (Science about Data Science): This is the more forward-looking dimension of the framework. It advocates the use of scientific methods to study the data analysis process itself and to assess the effectiveness and deviation of different analytical methods.
Core challenge: Science on data science
Currently, the GDS6 dimension study is responding to a major challenge for the scientific community -Recoverable CrisisI'm sorry. As the volume of scientific literature has grown, it has become crucial to validate the findings of the study. Data science works by:
- Meta-AnallysisSystematically integrates data from existing literature, assesses the overall effect of a given scientific issue and identifies publication deviations.
- Cross-workstream analysis: Study the impact of different analytical pathways (e.g. choice of data pre-processing methods, differences in model assumptions) on final conclusions, quantify the uncertainty associated with “researcher freedom” and thus seek more robust scientific findings.
This trend indicates that data science is moving from empirical practice to a more rigorous scientific framework, which is dedicated to assessing and enhancing the validity of scientific findings from data analysis.
Reflections: Data-driven limitations
The case of the long-term capital management company (LTCM) provides a lesson on overdependence on data and models. Despite the existence of the best quantitative models, the impact of low-probability extreme events (i.e., the “black swan” event) was ignored by over-reliance on historical statistics, which eventually led to a systemic collapse.
This reveals the potential risk of purely data-driven thinking: it tends to focus on relevance rather than causality. Models based on historical data training may be ineffective once an unknown structural change occurs in the bottom mechanism for data generation.Pure data-driven methods are intrinsically difficult to predict black swan events that never appear in historical data.I'm sorry. This suggests that data science cannot be divorced from the understanding of the knowledge of the field and the mechanisms of causality.
Future of data science
Looking ahead, the development of data science should not be seen as a mere tool for industry to improve efficiency. As an independent discipline, the goal is to createScience of Learning from Data。
Knowledge at the heart
Future data science will go beyond the simple superseding of statistical and computer science to focus on its cognitive issues:How can the extrapolation from the data be ensured is real, reliable and re-emergible?
Developments in this area will be directed in three directions:
- Integration of the two cultures: Breiman proposes a further integration of the “generated model” (explaining mechanism) and “predictional model” (precision prediction). The development of an Explanatory Machine Learning (Ex plainable AI) is the manifestation of this convergence trend, which aims to obtain high predictive precision and model interpretability at the same time.
- Evidence-based methodology: The development of scientific methods for data will be guided by empirical principles. The performance of different algorithms and analytical processes is objectively assessed through the Common Task Framework, CTF and large-scale empirical research, rather than relying solely on theoretical assumptions.
- Scientific effectiveness safeguards• In the data-driven research paradigm, data science will assume responsibility for setting academic standards. To ensure that scientific findings are firmly grounded by developing “sciences on data science” and by establishing rigorous certification systems to remove noise and deviations.
As Donoho said:
“The broad data science is essentially dedicated to understanding and enhancing the validity of research findings and can play a key role in all areas dominated by data analysis modelling.”
In the future, data science will serve as the basic methodological basis for modern scientific research, providing not only computing tools but also a critical thinking framework that will guide researchers from relevance to causality and more reliable conclusions.
- Title: The 60th Year of Data Science
- Author: Hyacehila
- Created at : 2026-02-13 04:00:00
- Link: https://hyacehila.github.io//blog/2026/02/13/60-years-of-data-science/
- License: This work is licensed under CC BY-NC-SA 4.0.