Data analytics mistakes can produce wrong business decisions when inaccurate data, biased samples, weak models, or misleading charts make unreliable results appear trustworthy. A completed pipeline doesn't prove its data is correct: executives reportedly used wrong numbers, and one MRR report overstated monthly recurring revenue by at least 70% [1].
- An MRR report over-reported monthly recurring revenue by at least 70% because it counted trials and conversions incorrectly [1].
- More than percent of companies reportedly rely on stale data for decision-making [2].
- Every chart in one study had at least one design flaw, with an average of 2.47 flaws per chart [3].
- Each reported Excel error took an average of minutes to fix, and percent of surveyed people had seen an Excel error cost employers money [2].
What counts as a data analytics mistake, and how can it lead to a wrong business decision?
A data analytics mistake is a problem in data, interpretation, modelling, or communication that makes a decision less reliable. Data-quality problems include inaccurate, incomplete, inconsistent, outdated, or duplicated data, all of which can undermine decisions based on it [4]. A successful data pipeline run doesn't prove that the underlying data is correct: Jobe reports assuming that a successfully run pipeline meant the data was correct, while executives still made decisions using wrong numbers.
Unstructured data adds separate risks. Erroneous natural-language-processing results can prevent careful data science from producing a good analysis [5], and focusing more on tools than on the problem is identified as a major mistake [5]. NLP analysis should use original text because machine translation can introduce potentially fatal errors [5]. In one MRR report, double-counted trials and conversions over-reported monthly recurring revenue by at least 70% [1]. A Type I error rejects a null hypothesis when no real effect exists, while a Type II error fails to reject it when something is happening [6]. Those errors can waste resources or create missed opportunities [6].


Which data-quality checks can reveal missing, duplicated, outdated, or inaccurate data before analysis?
Data-quality checks should run across data ingestion, data transformation, and data serving rather than being added only after a dashboard is published. Manual dashboard double-checking can indicate a data-engineering failure. A practical control is a “bad data first” pipeline with clearly stated validation rules that reports violations before questionable records are used [2].
During entry or import, use real-time validation, then apply automated checks during ETL to catch invalid data before ingestion [4]. Scheduled data-cleansing jobs can identify duplicates, correct formatting, and standardize values across systems [4]. Compare dashboards with business outcomes: discrepancies are an early warning sign of data-quality problems [4]. More than percent of companies reportedly rely on stale data for decisions [2]. Some organizations use 95% completeness and less than 2% duplication as benchmarks, but these are organizational benchmarks, not universal rules [4]. Where precision matters, exclude double-firing events, affected periods, and test users [1].

How can sampling bias, confounding variables, or mistaking correlation for causation distort an analysis?
Sampling bias can make an analysis appear persuasive while failing to represent the population or the question you need to answer. Two events occurring together don't establish that one caused the other [7]. Before generalizing a result, the sample should be large enough, representative of the target population, and randomized [7].
Survivorship bias is one clear failure mode: an analysis of successful mutual funds that leaves out funds that failed during the same period gives an incomplete sample [7]. A Nevada polling example used of Republican Hispanic respondents and presented that biased sample as if it represented a national result [7]. Confounding variables can also provide alternative explanations, but the ledger doesn't provide a specific confounding-variable test or a numerical sampling-error threshold. Treat model validation cautiously: testing on a different sample is supported, but no single threshold is supplied [7].

What charting mistakes—such as misleading axes, poor labels, or unsuitable chart types—can change how decision-makers interpret results?
Misleading visualizations can change a decision before anyone checks the underlying data. A nonzero Y-axis baseline can distort perceived trends and magnitudes, while a compressed scale can make small variations appear significant [8]. Manipulating the axis can make a 4% increase appear four times larger [7]. Use a consistent scale so viewers can interpret the data correctly [8].
Missing axis labels, unclear legends, excessive colors, and data overload make charts harder to interpret [8]. Chart selection should match both the data and the message [8]; bar charts suit comparisons between discrete categories more than trends over time [8]. Three-dimensional pie-chart effects can distort proportions and exaggerate segments [8]. A chart study found at least one design flaw in every chart, averaging 2.47 flaws per chart; missing units, out-of-context labels, and truncated axes were common [3]. Of incorrect participant insights, arose alongside unusable or less usable charts, visual clutter, or suboptimal design [3].

How should analysts choose and validate metrics, assumptions, and models before using their results to make decisions?
Model choice should match the decision: descriptive models report what is known and the level of certainty, predictive models provide forecasts, and prescriptive models weigh courses of action against objectives [5]. Form a hypothesis before analysis to reduce the risk of data mining [7], test the model on a different sample [7], repeat tests across datasets or periods [6], and use cross-validation to assess whether results generalize beyond one sample [6].
| Model type | Primary purpose | Question it supports |
|---|---|---|
| Descriptive | Report known information and certainty | What is known? |
| Predictive | Provide forecasts | What may happen? |
| Prescriptive | Weigh actions against objectives | What course of action should be considered? |
Metrics should fit the product’s business model and user behavior [1]. A proxy metric should be sensitive, simple, independent of other factors, and directionally related to the target metric [1]. Averaging acquisition costs can hide business-model differences [1], while total and average profit can produce significantly different decisions [3]. Ask a colleague to review conclusions for confirmation bias [7], and remember: “All models are wrong, but some are useful” [7]. The ledger provides no specific evidence for data leakage, overfitting, confidence intervals, sensitivity-analysis thresholds, metric ownership, or data-lineage procedures; those require further evidence. Poor-quality data fed into AI systems can reduce model performance and produce inaccurate predictions [4].
What financial, operational, or compliance costs can result from decisions based on faulty analytics, and how can those risks be measured?
Faulty analytics can surface as mismatched finance reports, customer overcharges or undercharges, and duplicate records that inflate customer or sales KPIs [4]. Type I errors can waste resources, while Type II errors can create missed opportunities [6]. In one government-agency case, reported mean squared error was times greater than the regression coefficient and times the usual acceptable margin of error; the agency couldn't know whether a heavily funded plan would work [5].
Rework has a measurable cost: each reported Excel error took an average of minutes to fix, and percent of people surveyed had seen an Excel error cost employers money [2]. Missing consent records can create compliance risks, especially in regulated industries [4]. Outdated data in healthcare or finance can lead to compliance violations and legal consequences [4]. Reported GDPR fines reached nearly €100 million in the first half of 2022, while the source says Amazon faced an $888 million privacy fine in and regulators can fine up to four percent of company revenue [2]. Measure risk through error rates, reconciliation differences, rework time, affected customers, forecast error, and compliance incidents; these are measurement approaches, not statistics supplied by the ledger.
What does the/20 rule mean in data science, and is the claim that 87% of data science projects fail supported by reliable evidence?
The supplied ledger contains no definition or supporting evidence for what the/20 rule means in data science, so a specific interpretation shouldn't be presented as established fact. It also contains no reliable source supporting or refuting the claim that 87% of data science projects fail. That claim is unverified from the available evidence.
A better conclusion is measurable rather than dramatic. For analytics, BI dashboards, data visualization, data pipelines, and AI adoption, check validation rules before use [2], repeat tests across datasets or periods [6], apply cross-validation [6], review chart quality, and reconcile dashboard results with business outcomes [4]. AI & Data Consulting Desk readers—especially small and mid-size businesses—can use those controls when assessing an analytics project or planning an AI and data roadmap. Document assumptions, check source data, validate outputs, and escalate questionable results. A completed pipeline isn't proof of reliability, and data quality should be addressed during ingestion, transformation, and serving. Users may also be unaware of the assumptions and trade-offs behind analytical and design decisions made by ChatGPT [3].
| Model type | Purpose | Decision question |
|---|---|---|
| Descriptive | Reports what is known and the level of certainty | What is known? |
| Predictive | Provides forecasts | What may happen? |
| Prescriptive | Weighs courses of action against objectives | What action should be considered? |
Key Takeaways
- Validate data during ingestion, transformation, and serving instead of relying on dashboard checks after publication.
- Check whether samples represent the target population and avoid treating correlation as proof of causation.
- Match chart type, scale, labels, and detail to the data and decision.
- Test models on different samples, repeat tests, and use cross-validation before acting.
- Measure errors through reconciliation differences, affected customers, rework, forecast error, and compliance incidents.
Frequently Asked Questions
What is the/20 rule in data science?
The supplied ledger gives no definition or supporting evidence for a specific/20 rule in data science. Treat any particular interpretation as unverified from the available evidence.
What are some common mistakes people make when doing data analysis?
Common mistakes include using inaccurate or duplicated data, confusing correlation with causation, relying on biased samples, choosing unsuitable charts, and accepting a completed pipeline as proof that its data is correct.
What are some common problems faced in data analysis?
Common data-analysis problems include incomplete, inconsistent, outdated, or duplicated data; sampling bias; weak validation; misleading visualizations; unsuitable metrics; and unclear assumptions.
Do 87% of data science projects fail?
The supplied ledger contains no reliable source supporting or refuting the claim that 87% of data science projects fail. The claim should therefore be treated as unverified from the available evidence.
Sources
- KPIs Done Wrong: Fixing Common Reporting Mistakes (2024-10-23)
- How to navigate your (avoidable) data errors (2022-11-07)
- Vibe Visualizing: How Visualization Novices Try (and Fail) to Generate and Interpret Visualizations with Conversational AI
- 9 Common Data Quality Problems and How to Fix Them in 2026 (2025-11-25)
- The Top NLP Mistake Made by Data Scientists (2020-11-24)
- medium.com
- Statistics Done Wrong: How to Avoid Common Stats Errors w/ Dr. Debbie Berebichez @Debbiebere (Episode 10)#DataTalk (2017-09-15)
- Are You Making These Bad Data Visualization Errors? (2024-05-24)


