Catching a Broken Schema Before It Reaches a Client's Dashboard
A data analytics consultancy's pipeline has a job most software pipelines don't: it has to catch a broken result, not just broken code. A dbt model can run without error and still produce a dashboard number that's quietly wrong because an upstream source changed its schema.
Here's a worked example of the kind of failure that CI/CD for a data pipeline needs to catch, and how GitHub Actions and GitLab CI each support the testing that catches it.
Vendors Covered in this Article
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Worked example: a client's source system renames a column
Say a client's CRM export quietly renames a customer_id column to account_id during a routine vendor update. A code-only test suite that checks whether your transformation code runs without throwing an error will pass, because nothing in the SQL itself is broken, it just silently starts joining on a column that no longer means what your model assumes.
A data quality test, checking row counts, null rates, and whether key columns still exist with the expected name, catches this before it reaches a client's dashboard. Tools like dbt's built-in testing framework or a lightweight custom check running as a pipeline step both work; what matters is that the check runs on every pipeline execution, not just when someone remembers to look.
Testing data, not just code
A software pipeline's tests ask whether the code behaves correctly against known inputs. A data pipeline needs a second layer: does the data itself still look like what the code assumes it looks like. Build both into your CI setup, code tests that run against fixture data on every pull request, and data quality checks that run against the actual current source data on a schedule.
The second kind of test is the one clients actually notice when it's missing, since a wrong number quietly appearing on a client's dashboard is a worse outcome than a build failure they can see and understand.
Where GitHub Actions and GitLab CI diverge for data pipelines
Both platforms run scheduled pipeline triggers well, which matters for data work since a lot of your pipeline runs are date-driven rather than triggered by a code change. GitHub Actions' cron-based scheduled workflows and GitLab CI's scheduled pipelines both cover this cleanly.
The real difference shows up in how each surfaces a failure. GitLab CI's pipeline dashboard groups scheduled and merge-triggered runs together in one view, which is convenient when you're triaging whether last night's data refresh failed for a code reason or a data reason. GitHub Actions requires a bit more setup, usually a dedicated Slack or email notification step, to get the same visibility without someone manually checking the Actions tab each morning.
Treat compiled documentation as a build artifact
A dbt project's compiled documentation, showing lineage from source to dashboard, is genuinely useful to a client during a handoff or an incident, and it's easy to generate as a pipeline step rather than something someone runs locally and forgets to share. Publish it as a build artifact on every successful pipeline run so it's never more than one deploy out of date.
On deployment frequency: teams with pipelines mature enough to deploy small model changes routinely tend to sit in the higher-performing DORA clusters rather than batching changes into large, infrequent releases1. For a data pipeline specifically, that usually looks like shipping one model change at a time instead of a monthly bundle of unrelated fixes.
What to check before a data pipeline change ships
Before merging a change to a client-facing data model, confirm:
- Does the change include a test for the specific data quality issue it's meant to fix, not just a code-level assertion?
- Will the change break a downstream dashboard or report that isn't obviously connected to the model being changed?
- Is there a rollback plan if the change produces wrong numbers that reach a client before anyone notices?
- Does the pipeline log which source data version fed a specific run, so a wrong number can be traced back to its cause?
A data pipeline's failure mode is usually silent, not loud, which is exactly why these checks need to be automated rather than left to someone noticing a number looks off.
Explaining a bad number to a client without sounding defensive
Eventually a wrong number reaches a client despite your checks, and how you explain it matters almost as much as how fast you fix it. Because your pipeline logs which source data version and which model version produced a given run, you can point to the specific upstream change that caused it rather than offering a vague apology with no explanation attached.
Clients tolerate an occasional data issue far better when the explanation is specific and the fix includes a new automated check preventing the same failure mode next time, rather than a promise to be more careful.
What Good Looks Like
Good looks like a pipeline that tests both code correctness and data quality on every run, publishes lineage documentation automatically, and can trace a wrong number on a dashboard back to the specific run that produced it.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
For pipelines feeding a client's warehouse on AWS, scoping the pipeline's access to exactly the tables it needs to write limits the blast radius of a bad model run.
Google Cloud's BigQuery pairs naturally with a pipeline that schedules transformations and quality checks together on the same trigger.
For clients who need evidence that access to their data warehouse is controlled and monitored, Vanta can track that alongside your pipeline's own access scoping.
Frequently Asked Questions
How do we test for a data quality issue we haven't seen yet?
You mostly can't, and that's fine; build tests for the failure modes you've actually seen, source schema changes, unexpected nulls, duplicate keys, and add new tests each time a new kind of issue reaches a client. The test suite should grow every time something slips through.
Should data quality checks run on every pipeline execution or on a schedule?
Both, for different reasons. Run lightweight checks on every execution to catch obvious issues immediately, and run a fuller data quality sweep on a schedule to catch slower drift that a single run might not reveal.
Is a failed data quality check the same severity as a failed code test?
Usually more severe for a client-facing pipeline, since a code test failure blocks a deploy before anyone sees it, while a data quality failure often means something already looks wrong on a live dashboard. Treat data quality alerts with at least the same urgency as a production incident.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Deployment frequency by DORA performance cluster (max days between deploys). DORA Accelerate State of DevOps 2024 (Google Cloud), cluster table via Octopus Deploy analysis, 2024.
Related Guides
GitHub Actions vs GitLab CI vs CircleCI: Continuous Integration Comparison
Compare GitHub Actions, GitLab CI, and CircleCI: build speeds, runner pricing, matrix testing, Docker orchestration, secret management, and DORA metrics.
Managing CI/CD Across Client Networks Without Losing Track of Access
A checklist for IT consulting firms and MSPs choosing between GitHub Actions and GitLab CI across many client environments, with access pitfalls to avoid.
Securing a Client's Data Pipeline: A Worked Example
A worked example of scanning an Airflow and dbt pipeline for a business intelligence and data engineering consultancy, Snyk versus GitHub Advanced Security.
Cursor vs GitHub Copilot for Data and Analytics Consultants
The unit of work for a data consultant is a query, a DAG node, or a notebook cell. How that changes the Cursor vs GitHub Copilot decision for client warehouses.
SOC 2 for BI and Data Engineering Consultancies
SOC 2 for business intelligence and data engineering firms building pipelines across client warehouses, and how Vanta, Drata and Secureframe compare.
Choosing Endpoint Security for a BI and Data Consultancy
Exported CSVs and cached query results sit on analytics consultants' laptops long after the work ends. How CrowdStrike and SentinelOne fit that gap.