Benford's Law and Statistical Fraud Detection

Hyacehila

Introduction: Why is humans not good at forging random?

The questions in this article can also be addressedStatistics and Truth: How to use the accident (Statistics and Truth)Anscom Quartet: Visualized power and statistical illusionHow the concept of a relatively close read together is developed in different contexts.

If you write the next "random" number right now, like 20 times a coin, you might write something like that. HTHHTHTH... Such a sequence. You probably would have been unconsciously avoiding writing. HHHHHH This is a continuous combination, because the instinct tells you, "It doesn't look random."

But randomly, it doesn't matter if it looks random. In real random processes, it is not only possible but almost inevitable that a long series of repetitions will occur when the sample is sufficiently large.

In order to make the data natural, the fraudster tends to proactively avoid continuous repetition, extreme values and irregular fluctuations, making the results too consistent with the perception of “average” and “random”. These modifications may leave an anomaly in the distribution of numbers. Several common methods of inspection are described below.

The Law of First Laws

Rationale: Why is it "1"?

In our intuition, if a pile of data is randomly distributed, the first number is one to nine, the probability should be the same, each. $\approx 11.1%$。

But Simon Newcomb found an anti-intuitive phenomenon in a logarithmic table that was frequently viewed in the nineteenth century: numbers starting with 1 appear much more frequently than others. Later, the physicist Frank Benford had a more systematic validation of the law.

Benefluents of the Order.It was noted that the first of many data sets (e.g. accounting, demographics, physical constants) that were naturally formed $d$ ($d \in {1, \dots, 9}$) The probability of occurrence follows the logarithmic distribution:

$$ P(d) = \log_{10} \left( 1 + \frac{1}{d} \right) $$

The corresponding probability is that:

  • 1 Probability at beginning:$\approx 30.1%$
  • 2 Probability at the beginning:$\approx 17.6%$
  • ...
  • 9 Probability at beginning: Only $\approx 4.6%$

A visual explanation is derived from the growth process across the orders of magnitude. A country's population has to double from 1 million to 2 million, and from 9 million to 10 million, it needs to grow by about 11 per cent. In this category, values stay longer in the beginning of the first "1" period. However, not all data satisfy the specific law of Benfu, and the way in which the data are generated and the range of values to be taken is still to be checked before they are used.

Application and case studies

This pattern is often used for financial audits and election fraud detection. When the fabricator makes the data, it is often the first number that is evenly distributed for “scrutinizing”, resulting in a low frequency of 1 and a high frequency of 9.

Enron ' s financial fraud cases are often used to discuss such methods. An ex post analysis of each share of the proceeds and other financial data disclosed by it was performed with a Benfo-specific test and a deviation from the first numerical distribution and theoretical values was observed. Such deviations provide a trail for audits, but cannot be independently substantiated for falsification of data, and need to be investigated in conjunction with accounts, transactions and business processes.

Another example that is often discussed is that 2009 Iranian presidential electionI'm sorry. After the election was fraudulently challenged, the statisticians Walter R. Mebane, Jr. conducted a Benfu-specific second-order test of the votes obtained in the open area (2 BL test). The analysis found that there was an anomaly in the distribution of votes and the accumulation of final numbers in some constituencies. These results support further verification of electoral data, but the distribution of figures cannot in itself be a substitute for ballot auditing and evidence of the electoral process.

Statistical test methods: calonian proposed eugenicity test

It's not just the eye, but the use. Carpside Probability Test (Ch-Square Goodness of Fit Test)

  • zero scenario ($H_0$): The first digital distribution of data is consistent with the Benefu-specific law.
  • Alternative scenario ($H_1$): The first numerical distribution of data does not conform to the Benfu law.

Calculate statistics $\chi^2$:

$$ \chi^2 = \sum_{i=1}^{9} \frac{(O_i - E_i)^2}{E_i} $$

of which $O_i$ It's the frequency observed.$E_i$ The expected frequency is calculated on the basis of the Benefig-specific law. Calculate $p$ After value if $p < 0.05 (or more stringent threshold) we have reason to reject the zero assumption and suspect that the data is abnormal.

Late Number Analysis (Last Digit Analysis)

Principle: Human perception of random intuitive deviation

If the first figure is the pattern of “natural growth”, the last figure is the psychology of “man-made intervention”.

In the measurement or counting data, the last number (0-9) should normally beUniform Distribution And the probability of each number is about 10 percent.

Two mistakes the fraudster makes:

  1. Avoidance of duplication: In human subconscious, 889911 Such figures are too false and therefore deliberately avoid them when they are made up. Even in multiple numbers, the adjacent figures are deliberately different.
  2. It's a good idea.: For the sake of economy or psychological comfort, fake data 0 and 5 The frequency of occurrence is often abnormally high (heavy effect).

Application and case studies

Here.Supermarket salesorHeight records.is more common. If the percentage of data ending in 0 or 5 (e.g. 170 cm, 175 cm) is abnormally high in a height record, the record should first be checked for clean-up, estimation or manual entry. The distribution of tails alone does not determine whether the data are false.

In the analysis of the elections in Iran, statisticians also checked the last two figures of the total number of votes cast. The frequency of some of the numerical combinations is higher than random expectations, which may be related to manual filling or other data generation mechanisms, and needs to be continued in conjunction with the electoral process.

Statistical testing methods: evenness tests

This step can also be used for a calibration, but for a balanced distribution.

  • zero scenario ($H_0$): equal probability of the last number (0-9) (i.e. $P = 0.1$)。
  • Statistics: Same calculation $\chi^2 = \sum \frac{(O_i - E_i)^2}{E_i}$But here. $E_i$ It's all the total sample. $1/10$。

It's also possible to introduce Runs Test To check the randomity of the numerical sequence and determine whether there is an anomaly in the serial relevance that is caused by “intentionally avoiding duplication”.

Too Good To Be True

Rationale: Differences and fluctuations

Observation data such as stock prices, temperature and experimental measurements are usually randomly volatile. The size of the fluctuations depends on the object, the method of measurement and the sampling process, and the data cannot be judged only on the basis of the smoothness of the curve.

The forger may interpret “good data” as “pretty data”, erase fluctuations in fabrication or embellishment, smooth the curve too much or over-compatible the theory of the variable relationship. The lack of the desired variance (Lack of Variance) is therefore a signal worth checking, but the determination of “how much difference” must rely on specific models and cross-reference data.

Application and case studies

Madoff Ponzi Scheme. It is a frequent case. Bernard Madoff reported that investment returns remained exceptionally stable over the long term and that there was little significant rebound when markets fluctuated. The Quantification Analyst Harry Markopolos therefore questioned whether his strategy could be realized. The volatility analysis did not single out the fraud, but it pointed to an inexplicable inconsistency between the public performance and the stated trading strategy.

Macroeconomic data can also be cross-checked through relevant indicators. For example, the “gang index” is based on the idea of looking at GDP together with indicators such as industrial electricity, rail freight and bank balance for medium- and long-term loans. If reported GDP growth is long-term deviation from the relevant physical indicators, thisRelevance BreakThis is worth further explanation. However, the relationship between indicators is influenced by industry structures and statistical calibres, and deviations cannot themselves directly prove that data are false.

Statistical testing methods: differential tests and relevance

We're looking at data for such fraud.VolatilityandMulti-dimensional related structures

First of all, we can use it. F-test (F-test) Comparison of target data (%2)$S_1^2$) with baseline data ( )$S_2^2$) The difference. By Calculator $$ F = \frac{S_1^2}{S_2^2} $$If the F value is significantly less than 1, indicate that the target data are subject to a lower rate of volatility than the benchmark. This is an unusual signal to be investigated and whether it is artificially smoothed and judged in conjunction with data sources.

And then...Structure Consistency and Disability Analysis (Structural Consortium) & Residual Analysis)I'm sorry. For time series such as GDP and electricity use, it is possible to use Cointegration Analysis Distinguishing between “false return” and long-term association. Common metaphors are drunks and dogs he's holding: both can move randomly, but distance does not increase indefinitely. In time series, although the two variables are not stable, a linear combination of them may be stable.

In the test, you can. Cannot initialise Evolution's mail component. Check for long-term balanced relationships. If the two historically harmonized indicators begin to deviate, the potential for a widening of the gap and its unstable appearance suggests a change in relationship. Changes may come from statistical calibres, economic structures, external shocks or human intervention, and therefore structural mutations and background information are also needed to determine the causes.

Thrust at the threshold (Bunching / Threshold Effects)

Principle: Avoiding harm by profit

When there is some sort of appraisal indicator, tax threshold or academic publication standard (e.g. $p) < At 0.05.00, data tend to be distorted near the threshold. It's called Clock effect (Bunching)

Application and case studies

This can be seen in academic publications and tax returns. When analysing the distribution of P values for published papers, researchers have observed that values are concentrated in rapid declines before the significant threshold of 0.05. This may be related to selective reports or P-hattering. Similarly, if significant amounts of declared revenue are concentrated below the tax threshold, possible reasons such as disclosure rules, integrity behaviour and tax avoidance incentives need to be examined.

Statistical testing methods: Visualization and McRay density tests (McCray Density Test)

The threshold effect can be identified first.VisualiseStart. A fine histogram shows whether the data are at a certain threshold (e.g., $p=0.05$ The government has also been able to provide information on the situation in the country, and has been able to provide information on the situation. Such graphics are a clue to further testing, not a direct proof of man-made manipulation.

We use it for statistical validation. MacCray Density Test (McCray Density Test)I'm sorry. This is a test commonly used in the Breakpoint Return Design (RDD) in the sense that the probability density function of the variable is continuous at the breakpoint. If the left density of the threshold is significantly higher than the right and the difference is statistically significant (even if random fluctuations are taken into account), we have reason to reject the continuity assumption and to assume that there is artificial manipulation.

Conclusion: anomalies are only the starting point of the investigation

Benfoux-specific laws do not apply to all data, such as heights with limited range of values, lottery numbers and fixed-priced goods. The end figures, differences, related structures and thresholds are also subject to their respective conditions.

The effect of these methods is to identify anomalies that need to be explained and to help auditors decide what to examine next. They cannot be characterized separately from the data generation process, business context and other evidence.

When looking at too smooth a growth curve or an abnormally even digital distribution, one more step can be asked: can the existing data generation process explain this shape?

  • Title: Benford's Law and Statistical Fraud Detection
  • Author: Hyacehila
  • Created at : 2026-02-04 15:00:00
  • Link: https://hyacehila.github.io//blog/2026/02/04/benfords-law-and-statistical-fraud/
  • License: This work is licensed under CC BY-NC-SA 4.0.
Comments