Your Data Warehouse is Lying to You

Most organisations have no idea if their data is accurate. 

Many organisations have built massive infrastructure to answer questions, but there’s no systematic way to verify if the answers are correct.

The Trust Erosion Cycle

A business user finds an error in a report. They stop trusting the data team. They build their own “corrected” version in Excel. Now there are multiple “sources of truth.” The data team doesn’t know their data is being bypassed. The cycle repeats and compounds.

Why This Keeps Happening?

The industry has become obsessed with modern data stacks, real-time pipelines, fancy visualisations, and scalable architectures. Most data teams spend 80% of their time building pipelines and 20% validating outputs.

However, there’s a fundamental question that gets lost in all the technical sophistication: How do we know this number is right?

Consider the organisation that discovered their revenue reporting was off by 15% for 18 months. Not because of a technical failure. The error was in the business logic. A join condition that looked correct but was silently dropping transactions. Hundreds of decisions were made on wrong numbers. Nobody noticed because the numbers looked reasonable.

This isn’t an isolated incident. It’s a pattern that repeats across industries because the focus has been on building faster, more sophisticated systems rather than more trustworthy ones.

What Actually Works

The most successful data organisations are shifting to what might be called ‘validation-first data engineering’. This means fundamentally rethinking priorities.

Before building anything, define what right looks like. 

  • What should this number be? 
  • How can it be verified? 
  • What are the acceptable tolerances? 

 

These questions should be answered before a single line of code is written, not after the dashboard is already in production.

Validation needs to be built into the pipeline, not tacked on afterward. This means checking whether row counts match source systems, verifying that key business metrics fall within expected ranges, maintaining referential integrity, and ensuring historical trends remain consistent. These checks shouldn’t be manual processes that someone remembers to run occasionally. They should be automated gates that prevent bad data from propagating downstream. 

Data quality needs to be visible and measurable. There should be clear dashboards showing validation test results, automated alerts when checks fail, and transparent data lineage showing exactly where numbers originate. When something goes wrong, the system should make it obvious immediately, not three months later when someone notices their analysis doesn’t make sense. 

Creating feedback loops is essential. This means regular reconciliation with source systems, making it easy for business users to flag suspect numbers, and conducting root cause analysis for every discrepancy. The goal isn’t to blame people when errors occur – it is to understand why the system allowed the error to reach production in the first place.

Perhaps most importantly, organisations need to accept that 100% accuracy is impossible. The question isn’t “How do we eliminate all errors?” but rather “How do we define acceptable error rates, prioritise validation for high impact metrics, and maintain transparency about data limitations?”

The Real Issue:

Data quality has been treated as a technical problem when it’s actually a trust problem. Trust isn’t built by promising accuracy, it is built by: 

  • being transparent about limitations, 
  • catching errors before users do, 
  • having systematic validation processes, and 
  • fixing issues quickly when they’re found.

The shift needs to be from “we built you a data warehouse” to “we built you a trustworthy data warehouse.” Because having data at your fingertips is meaningless if you can’t trust what it’s telling you.

How does your organisation validate data accuracy? Or is there an unspoken assumption that the technology itself guarantees correctness? 

Picture of Cate Bernard

Cate Bernard

Graduate Data Consultant

Get the latest on data management in your inbox

We are an established data consultancy, working on some of Australia’s
biggest data management projects across seven capital cities.

"*" indicates required fields

Getting Your Critical
Data Sorted

Don’t Ruin Your Organisation by
Tolerating Poor Data Quality

Tuesday 27 February, 11-11.45am AEDT

Tim Goswell Practice lead

Tim Goswell

James Bell

James Bell

Tim Goswell Practice lead

Connect with Tim

Todd Heather

Connect with Todd

James Bell

Connect with James

Lloyd Robinson Director

Connect with Lloyd

How can we help