15 Data Mistakes You Should Avoid
Pitfalls and Obstacles in Business Intelligence Analyses

No matter how large your company is or what industry it operates in, data analysis is a key element for gaining valuable insights and achieving sustainable success. Unfortunately, however, the results of such analyses are often sobering, as there are many factors that can lead to flawed insights from the data. We’ll explain the pitfalls of data analysis you should be aware of and how you can avoid them to unlock the full potential of your data. Join us as we explore data errors such as the “anchor effect,” the “Simpson’s paradox,” and the infamous “gambler’s fallacy” to gain a deeper understanding of the most common misconceptions in business intelligence.
Data errors and how to avoid them
1. Cherry Picking:
Cherry-picking is a common data trap in which only certain data points or pieces of information are selectively chosen to support a thesis, while other relevant data is intentionally ignored. This misleading approach can completely distort the results of a data analysis, as it skews the overall picture of the data. For example, cherry-picking can make a situation appear significantly better or worse than it actually is. Imagine, for instance, that your marketing department wants to analyze the effectiveness of a product. If only positive customer reviews or success stories are used for this purpose, you can assume that the analysis will present a distorted picture of reality. In this specific case, the analysis will paint a very positive picture of the product’s efficiency. However, if there are actually many negative reviews or critical voices that are not taken into account in the analysis, it could be that your product is not particularly efficient in reality. Therefore, to optimize or further develop your product, you should definitely include these negative voices in the analysis.
Solution: To prevent cherry-picking, it is crucial to conduct a systematic and transparent data analysis. Therefore, all available data should be collected, objective statistical methods should be applied, and all data should be published. Take special care to include data that does not support the hypothesis you are trying to prove. Peer reviews and external evaluations by independent experts can also help identify and correct such biases. By presenting the data honestly and completely, you ensure that analyses are based on a solid foundation and are not influenced by cherry-picking.
2. Survivorship Bias:
Survivorship bias is a distortion that occurs when only the successful or surviving cases are included in an analysis, while the unsuccessful or non-surviving cases are omitted. This leads to an unrealistic portrayal of the prospects for success, as important data on failures or setbacks are missing. This data error can thus lead to incorrect conclusions, since the data that is not taken into account may constitute an important part of the overall picture. Survivorship bias is often found, for example, in studies of successful companies or famous individuals. Thus, the stories of successful companies or people are often examined, while failed companies or unknown individuals are not taken into account. This leads to a distorted assessment of the factors contributing to success. A particularly frequently cited example of survivorship bias is the study of aircraft during World War II. To determine where armor should be reinforced, researchers first examined returning aircraft with bullet holes. Based on this, the parts with the most bullet holes were to be reinforced. What sounds logical at first, however, contained a crucial flaw: all aircraft that had crashed due to a bullet hit were missing from the data set under examination. It was subsequently discovered that the parts of the aircraft that showed the fewest bullet holes in the study were the ones that absolutely needed to be reinforced—since it was hits in these areas that caused most of the aircraft to crash.
Solution: Use a comprehensive dataset that includes both successful and failed cases to prevent this phenomenon. Since, as in the example above, it is not always certain that all data is available, you should critically examine the data before analyzing it to avoid drawing incorrect conclusions from the analysis. Therefore, always be aware that data may be missing, and actively look for such cases to minimize or prevent distortion caused by survivorship bias.
3. Cobra effect / perverse incentive:
The cobra effect refers to a situation in which a proposed solution to a problem has unintended side effects that exacerbate the problem or create new ones. It is, in other words, a perverse incentive. The term originates from an anecdote from the colonial era in India: At that time, many people in India were dying from cobra bites. To rid the population of cobras, the British colonial authorities offered a reward for every cobra captured. Unfortunately, they had not considered that this might create a perverse incentive. In response, the locals began breeding cobras to exchange them for the reward. After the government ended the initiative, these captive-bred cobras were often released into the wild, leading to a dramatic increase in the cobra population rather than a decrease.
We can often observe the cobra effect in the economy as well: For example, if a government tries to lower inflation by drastically reducing the money supply, this can lead to a deterioration in economic conditions. This is because people then have less money to invest and spend. This, in turn, can lead to a decline in economic activity.
Solution: To avoid the cobra effect, it is crucial to carefully consider the long-term implications of any proposed solution to ensure that unintended side effects are avoided. Consulting with experts and stakeholders can help take different perspectives into account and identify unforeseen consequences before a solution is implemented. Continuous monitoring and adjustment of the measures are also important to ensure that the cobra effect and similar unintended consequences are avoided.
4. False causality:
False causality is a fallacy that occurs when it is assumed that a cause-and-effect relationship exists between two events, even though they are merely coincidentally correlated or other hidden variables explain the relationship. A classic example is the link between the rise in ice cream sales and the increase in swimming pool accidents during the summer. A quick glance at such an analysis might lead one to assume that the swimming pool accidents are caused by increased ice cream consumption. In fact, however, both events are caused by the warm season.
Solution: Be sure to carefully distinguish between correlation and causation to avoid this mistake. A correlation measures the statistical relationship between two variables. In contrast, causal relationships provide information about cause and effect. A correlation can therefore indicate a causal relationship, but this is not necessarily the case. Statistical methods such as experiments and control groups can help identify actual cause-and-effect relationships. Therefore, analyze all available data and examine alternative explanations for observed correlations. In addition, in-depth knowledge of the specific field can help you better understand relevant relationships and avoid unfounded assumptions. Conscious, critical analysis and an open-minded attitude toward different possible interpretations are crucial for preventing erroneous conclusions regarding false causality.
5. Data Fishing:
Data fishing, also known as p-hacking or data grabbing, refers to the practice of sifting through large amounts of data in search of statistically significant results or patterns without testing a specific hypothesis. This can lead to misleading results, since statistically significant results are expected if enough tests are conducted, even if there is no actual effect. For example, researchers might test hundreds of variables against a specific outcome and then present only the results that appear statistically significant. If, for instance, a drug trial tests the effects of different dosages of the drug on a variety of symptoms, researchers should consider all the results. However, if data fishing is used to select only the dosage that shows a statistically significant effect on one symptom—without considering the other tests—this can lead to a distorted representation of the results.
Solution: To prevent data fishing, it is important to formulate a clear hypothesis before data collection and to plan the analysis methods in advance. If multiple tests are conducted, a correction method such as the Bonferroni test should be applied to reduce the risk of false-positive results. Transparency and openness are also crucial. You should document all tests conducted and their results, even if they are not statistically significant. This enables a comprehensive evaluation and prevents the selective reporting of results that could be skewed by data fishing.
6. Confirmation Bias:
Confirmation bias, also known as the confirmation error, is the tendency to seek out information or data that confirms existing beliefs or hypotheses, while ignoring or rejecting contradictory information. This is because people subconsciously seek confirmation of what they already believe, rather than objectively evaluating all available information. This can lead to a one-sided interpretation of data. A practical example would be an investor who tends to pay attention only to news and analyses that support his positive assessment of a stock, while ignoring negative reports or warnings about potential risks.
Solution: To prevent confirmation bias, it is important to be aware of this tendency and actively work to counteract it. A first step is to foster an open and critical mindset. In science, methods such as double-blind studies and peer reviews help ensure objective evaluations. In your company, you can seek out opinions and feedback from people with different perspectives and experiences to challenge and broaden your own point of view. It’s also helpful to regularly check whether you’re remaining objective when evaluating information or unconsciously seeking confirmation. Through conscious self-reflection and the use of different perspectives, the influence of confirmation bias can be minimized.
7. Regression to the Mean:
Regression to the mean describes the phenomenon in which extremely high or low values in a measurement tend to return to less extreme values when the measurement is repeated. This occurs independently of any intervention or change and is due to random fluctuations in the data. One example of this is academic performance. For instance, it is likely that students who perform exceptionally well on a test will not achieve quite as outstanding results when they retake the test later. This is due to normal fluctuations, such as those caused by the students’ performance on a given day.
Solution: To prevent regression to the mean, it is important to understand that extreme values can often occur randomly and do not necessarily indicate a cause-and-effect relationship. When evaluating performance or results, one should therefore not overreact to outliers, as they tend to revert to less extreme values when the measurement is repeated. It is advisable to use statistical methods to identify the random nature of extreme values and to always consider the context when interpreting data. Regular reviews and critical analysis can help draw reliable conclusions without being influenced by random fluctuations.
8. Anchoring effect:
Anchoring heuristic, also known as anchoring bias or the anchoring effect, refers to the tendency to be strongly influenced by an initial value or piece of information when making decisions. Even if this anchor is irrelevant or based on a false assumption, people tend to rely heavily on it. For example, the first price mentioned during a price negotiation serves as an anchor that has been shown to strongly influence the outcome of the negotiation. If a seller sets a very high price, for instance, buyers will tend to base their own offers closer to that high price.
Solution: Be aware of how anchors can influence our decisions. To counter this, actively distance yourself from the value mentioned first and use objective evaluation criteria. It can be helpful to consider alternative anchor values based on objective data and use them as the basis for decisions. For example, during negotiations, it may be wise to focus on relevant facts and comparative prices to minimize the influence of an arbitrarily set starting point. Conscious decision-making based on sound data and analysis can help minimize the effects of the anchoring heuristic. Conversely, this also applies, for example, when you want to collect data. For example, when designing a survey, you should be aware that respondents may be influenced by the anchoring effect, which in turn can compromise the survey’s validity. Therefore, in such cases, choose anchor values very carefully or avoid using them altogether if possible.
9. Simpson’s Paradox:
The Simpson paradox describes a statistical illusion in which a trend in the overall data is opposite to the trend in the individual groups. This means that an observation that appears in an overall analysis may be reversed when the data is divided into different subgroups. A real-world example could be a study comparing the treatment outcomes of two different hospitals. In the overall analysis, one hospital might show a higher survival rate. However, if the data is broken down by disease severity, it might turn out that the other hospital has a better survival rate across all severity levels.
Solution: To avoid the Simpson’s paradox, it is important to carefully consider potential interactions between variables during statistical analyses. When significant differences are observed in the overall data, it is advisable to examine more closely whether these differences are consistent across the subgroups. A more in-depth analysis that takes various variables into account and examines possible interactions among them can help identify and understand the paradox. When dealing with highly complex data, it is often advisable to collaborate with experienced statisticians or data analysts to ensure an accurate and reliable interpretation of the results.
10. Ecological Fallacy:
The ecological fallacy refers to the erroneous inference about individual characteristics based on aggregated, group-level data. This bias occurs when statistical correlations at the group level are extrapolated to individuals without taking individual differences into account. Consider, for example, a wealthy city where the average income of residents is very high; you might conclude from this that all residents of the city are wealthy. In reality, however, it is more likely that even in such a city there are significant income disparities among individual residents, meaning that some residents might be very wealthy, while others might be very poor.
Solution: To avoid the ecological fallacy, it is important to distinguish between aggregated and individual characteristics when interpreting data. Data analyses should therefore be conducted not only at the group level but also at the individual level to gain a more accurate understanding of the actual differences. Be aware that aggregated data cannot necessarily be generalized to individual experiences or characteristics; pay attention to the context of the data, and rely on appropriate data sources and analytical methods to avoid drawing false conclusions.
11. Goodhart's Law:
Goodhart’s Law, named after the British economist Charles Goodhart, states that any observed statistical relationship that is turned into a rule loses its predictive power as soon as it is used for decision-making. Simply put, this means that when a specific key figure or metric is made the basis for rewards or penalties, people or organizations develop strategies to optimize that metric. This often leads to unintended side effects. For example, if a company uses a product’s sales figures as a performance indicator for its sales staff and as the basis for a bonus, those employees might be tempted to use short-term sales strategies to earn the bonus. So while you may have sold a lot of products, this strategy could still have a detrimental effect on your company in the long run.
Solution: To prevent Goodhart’s Law, it is important to develop a holistic and balanced performance evaluation. This can be achieved by using multiple performance metrics to assess various aspects of performance. It is advisable to consider different perspectives when evaluating the overall performance of an individual or an organization. Furthermore, it is important to regularly review and adjust the metrics and indicators to ensure that they continue to provide relevant and meaningful information without creating incentives for undesirable behavior. A critical review of the performance metrics used and their potential impact on behavior can help minimize the negative effects of Goodhart’s Law.
12.Gambler's Fallacy:
The Gambler's Fallacy is a cognitive bias in which people believe that random events are influenced by their previous outcomes or frequencies. They mistakenly assume that a certain series of events—such as a long losing streak while gambling—must lead to a future positive outcome in order to restore balance. A simple example is the assumption that when flipping a coin, a tails flip is more likely after a series of heads flips. Statistically speaking, it is indeed likely that the number of heads and tails flips will eventually even out to 50 percent each. Nevertheless, each individual toss is independent of the previous one and therefore also has a 50 percent probability for every possible outcome. The situation is similar in sales, for example. So you shouldn’t assume that a salesperson’s chances of selling your product during the next customer meeting will increase just because they weren’t successful in recent meetings. Rather, statistically speaking, the salesperson has the same probability of making a sale in every single meeting.
Solution: To avoid the Gambler's Fallacy, it is important to realize that random events are not influenced by previous outcomes. Statistically speaking, the odds do not change based on past results. An understanding of the basic principles of probability can help you develop realistic expectations and overcome the Gambler’s Fallacy.
13. Regression Bias:
Regression bias occurs when not all relevant variables are taken into account during data analysis, leading to a false relationship between the variables. This can result in inaccurate predictions or incorrect conclusions. For example, a study examining the relationship between chocolate consumption and life expectancy without taking into account factors such as diet, exercise, or genetic predisposition would not be meaningful. If only chocolate consumption and life expectancy are analyzed without considering the other influencing factors, a distorted picture of reality emerges.
Solution: To prevent regression bias, it is important to consider all relevant variables that could influence the relationship between the variables under study when analyzing data. This requires thorough preliminary research and a deep understanding of the subject matter to identify potential confounding factors. The use of statistical techniques such as multivariate regression can help analyze multiple variables simultaneously and isolate their individual effects. For complex analyses, it is also helpful to consult experts and specialists in the relevant field to ensure that all relevant variables are taken into account. Careful and comprehensive data analysis that takes all influencing factors into account is crucial for minimizing the risk of regression bias and obtaining accurate results.
14. Data Mining Bias:
Data mining bias refers to distortions in the results of data analyses that can arise from the improper selection or interpretation of data. It can occur, for example, when analyses are performed on large datasets to identify patterns, correlations, or trends, and in the process, certain groups are unintentionally favored or disadvantaged. A real-world example would be a job application screening algorithm that uses historical data but, due to existing gender or racial biases, unintentionally favors or disadvantages candidates from certain groups.
Solution: To prevent data mining bias, it is important to exercise caution when selecting and interpreting data. A thorough data analysis should ensure that all relevant factors and groups are adequately represented. Regular reviews of the analyses can help identify and correct biases early on. Transparent and ethical guidelines for data use should be developed to ensure that data analyses are conducted fairly and impartially. Training and awareness-raising initiatives for data analysts and decision-makers can help raise awareness of data mining bias and ensure that analyses are objective and fair. Finally, it is important to critically examine the results and seek alternative explanations for the observed patterns in order to identify and correct potential biases.
15. Disposition Effect:
The attribution bias is a cognitive distortion in which people tend to attribute positive outcomes to their own abilities and wise decisions, while attributing negative outcomes to external circumstances or bad luck. This leads to an imbalance in self-perception and can result in irrational decisions. This bias can often be observed, for example, in the stock market: Many investors view a profit as the result of their own sound analysis, while losses are attributed to unpredictable market fluctuations.
Solution: As with most other data errors, the first step in avoiding the disposition fallacy is to recognize that it exists and can occur in many situations. Self-reflection on decisions and a willingness to view failures as learning opportunities can help mitigate the disposition bias. It is also helpful to seek external perspectives, whether through peer reviews, feedback from colleagues, or expert advice. An objective analysis of successes and failures, taking all relevant factors into account, can help develop a more realistic self-perception and prevent irrational decisions. Regular reflection and awareness of one’s own thought patterns are crucial for recognizing the disposition bias and actively countering it.
Conclusion
True to the saying, “Never trust a statistic you haven’t faked yourself,” it’s important to be aware that there can be many pitfalls and obstacles when analyzing data with business intelligence. Once you’re familiar with the various types of data errors—from cherry-picking to disposition errors—you can approach the results of analyses critically and thus ensure that the right decisions are made. Data errors can be avoided by adopting a transparent approach and taking into account different perspectives, analytical methods, and techniques.
The myPARM BI business intelligence softwareact offers an optimal solution for overcoming these challenges. With its advanced data analysis capabilities, transparent reporting options, and built-in mechanisms for validating your data, myPARM BIact enables precise and reliable data analysis, thereby providing a solid foundation for data-driven decision-making. In addition, with myPARM BIact you can immediately translate the decisions you’ve made into concrete actions.
Weitere Informationen über die Business Intelligence Software myPARM BIact:
Möchten Sie myPARM BIact in einer Demonstration kennenlernen? Dann vereinbaren Sie gleich einen Termin!



