Why Your 'Data-Driven' Decisions Aren't Actually Data-Driven

Why Your 'Data-Driven' Decisions Aren't Actually Data-Driven

Why Your ‘Data-Driven’ Decisions Aren’t Actually Data-Driven

Dr. Elena Vasquez, Ph.D. in Artificial Intelligence


We live in an era that is obsessed with metrics. Open any modern dashboard, and you are greeted by a cascade of sparklines, heat maps, and KPIs that pulse with the quiet authority of numbers. We have been sold a beautiful narrative: if you just look at the data long enough, the right decision will reveal itself. If the algorithm says to launch the product, you launch. If the model predicts the customer will churn, you send the retention email. We call this data-driven decision making, and it has become the badge of honor for every modern knowledge worker, product manager, and executive.


But after two decades of building and studying artificial intelligence systems, I can tell you a secret that the industry is slow to admit. Most of the time, your decisions are not data-driven at all. They are narrative-driven, dressed up in the language of data. The numbers are not driving the car. They are just sitting in the passenger seat, nodding along with the driver.


This is not an argument against data. Data is essential. But it is not a substitute for understanding, context, or judgment. And when we mistake correlation for causation, when we treat a model’s output as a prophecy rather than a probabilistic estimate, we end up making decisions that feel objective but are actually just as subjective as the gut feeling they were supposed to replace.


Let’s pull back the curtain.

The Illusion of Objectivity

The first trap is the most seductive: the belief that numbers are neutral. They are not. Every number is the product of a choice. Someone decided what to measure. Someone decided how to measure it. Someone decided what counts as a success and what counts as a failure.


Consider a classic example from the early days of machine learning. A hospital in the United States built a model to identify patients who needed additional medical care. The model was trained on historical data: which patients had received extra care in the past, and what their clinical records looked like. The model learned to predict which patients were most likely to have received extra care. It did not learn which patients needed extra care.


When the hospital looked at the model’s output, they noticed something strange. Black patients were being deprioritized. The model was scoring them lower than white patients with similar clinical needs. The doctors were frustrated. “How can this be?” they asked. “We’ve removed the race variable from the model. It can’t be biased.”


And they were right. The model was not using race directly. But the model had learned a proxy. In the historical data, white patients had, on average, received more extra care than black patients — partly because of historical inequities in access to care, partly because of subtle differences in how symptoms were documented, partly because of a host of social and economic factors that were never recorded in the clinical notes. The model had absorbed all of those invisible signals and encoded them into its predictions. The race was not in the data. The race was in the data, woven through it like a thread in fabric.


The decision to trust the model was data-driven. The decision to trust the model was also a narrative-driven choice to believe that the data was a fair and complete record of reality.


This is the first lesson: data is not neutral. It is a record of past decisions, past biases, and past assumptions. When you make a decision based on data, you are making a decision based on the choices that produced that data. And those choices were made by humans, with all their blind spots and interests intact.

The Correlation-Causation Confusion

The second trap is the one that trips up even the most numerate decision makers. We see two things move together in the data, and we assume one causes the other.


A retail chain notices that sales of ice cream and sunglasses move in lockstep. They decide to bundle the two products. A software company notices that customers who use the new feature for more than three days are 40% more likely to renew their subscription. They push the new feature harder. A marketing team notices that customers who watched the explainer video converted at a higher rate. They decide to lead with the video in every campaign.


In each case, the data shows a correlation. But correlation is not causation. Ice cream sales and sunglasses sales are both caused by sunny weather. Customers who use the new feature more are likely the ones who were already more engaged and more likely to renew, regardless of the feature. Customers who watched the video may be the ones who were already further along in the purchase journey.


The data tells you what happened. It does not tell you why it happened. And the “why” is what matters for decision making.


A truly data-driven decision would require not just a correlation, but a causal model. That means designing experiments, running A/B tests, using instrumental variables, or building structural causal models that separate the effect of the intervention from the effects of everything else that happened at the same time. Most organizations do not do this. They look at the dashboard, see a pattern, and decide. The decision feels data-driven. It is actually pattern-driven, and patterns are fragile. They break when the context changes.


This is the second lesson: data tells you what is. It does not tell you what would happen if you did X.

The Problem of Selection and Framing

The third trap is the most subtle, and the most common. We do not look at all the data. We look at the data that confirms the decision we are already leaning toward.


This is not a criticism of bad faith. It is a description of how human cognition works. We are pattern-matching, hypothesis-testing creatures. We form a hypothesis, then we look for evidence that supports it. We are remarkably good at finding that evidence in a dataset. We are remarkably bad at looking for evidence that contradicts it.


A product manager suspects that adding a new feature will increase retention. She looks at the data. She finds a subset of users who used the feature and saw a retention lift. She presents this to the team. The decision is made. The feature is built. It ships. Retention doesn’t move.


The data did not drive the decision. The narrative drove the decision, and the data was recruited to support it. The data that would have told a different story — the users who used the feature and still churned, the users who did not use the feature and still stayed — was there in the database. It just wasn’t looked at.


This is the third lesson: data is a library, not a verdict. You can find almost any answer in a large enough dataset. The question is which data you choose to look at, and why.

The Overfitting of Our Own Assumptions

As an AI researcher, I see a fourth trap that is unique to the age of machine learning. We build models, we train them on data, and we treat the model’s output as if it is a clean, objective signal. But the model is a compressed representation of the data, filtered through the choices we made in building it. The choice of features. The choice of the loss function. The choice of the architecture. The choice of the hyperparameters. Each of these is a decision, and each one shapes what the model can and cannot see.


A model that is trained to predict revenue will optimize for revenue. A model that is trained to predict user satisfaction will optimize for satisfaction. These are not the same thing. A customer can be satisfied with a product and still not renew. A customer can renew out of habit or necessity and still be unhappy. The model will tell you what you asked it to predict. It will not tell you what you should have asked it to predict.


And here is the key: you only know what you should have asked it to predict after you see the results. By then, the model has already made its decision.


This is the fourth lesson: a model is a lens, not a window. It shows you a particular slice of reality, shaped by the choices that built it. And those choices are made by humans, with all their assumptions and blind spots intact.

The Context Gap

The fifth trap is perhaps the most practical. Data is collected in a context. Decisions are made in a context. And most of the time, the two contexts are not the same.


A model is trained on data from last year. The decision is made this year. A model is trained on data from the East Coast. The decision is made for a West Coast market. A model is trained on data from the last economic cycle. The decision is made in a new economic cycle.


The data is real. The model is real. The prediction is real. But the context has shifted, and the data does not know that. The dashboard does not know that. The decision maker, however, does know. And that knowledge — the knowledge of what has changed, what is different, what the data does not capture — is not in the data. It is in the head of the person making the decision.


A truly data-driven decision would require not just the data, but the context that the data cannot provide. The new competitor that just launched. The regulation that is about to change. The shift in customer sentiment that is visible in the support tickets but not yet in the metrics. The macroeconomic signal that is in the news but not in the database.


This is the fifth lesson: data is a map, not the territory. And the territory is always moving.

So What Should We Do?

None of this is an argument for abandoning data. It is an argument for using data well. And using data well means treating it as one input among many, not the sole input.


It means asking: what are we measuring, and why? Whose choices are encoded in this data? What is not in this data? What context is missing? What assumptions are baked into the model? What would it look like if we ran the experiment? What would it look like if we looked at the data that contradicts our hypothesis?


It means building a culture where data is respected but not worshipped. Where a number is a starting point for conversation, not an ending point for decision. Where the decision maker is a judge, not a juror. The data presents the evidence. The decision maker weighs it, contextualizes it, and decides.


It means being honest about the limits of data. We are not gods who can see all the variables. We are humans with a dataset, a model, and a hunch. And the best decisions are the ones that integrate all three, with a healthy respect for the uncertainty that runs through all of them.

The Quiet Power of Judgment

In a world that has elevated data to a near-religious status, it is almost radical to say that judgment still matters. That understanding still matters. That context still matters. That the human in the loop is not a bug in the system but the feature.


A doctor does not diagnose from the lab results alone. A judge does not sentence from the statistics alone. A general does not deploy troops from the map alone. They use the data. They also use the data. They also use the experience, the intuition, the knowledge of the particular situation that no dataset can fully capture.


Your decisions are not data-driven because you looked at the data. Your decisions are data-driven when the data has been interrogated, contextualized, and woven into a richer understanding of the situation. When the numbers are in service of a judgment, not a substitute for it.


So the next time you open your dashboard, take a moment before you make the decision. Ask yourself: what is this data telling me? What is it not telling me? What assumptions are baked into it? What context is missing? What would the experiment say?


And then make the decision. Not because the data told you to. But because you understood the data, and the context, and the situation, and you judged that this is the right call.


That is what data-driven really means. Not that the data made the decision. That you made the decision, and the data was one of the tools you used to make it well.