You've just taken over a data stack. The previous vendor is gone, the documentation is a README from two years ago, and someone in finance has already told you the revenue number is wrong.
The instinct is to rebuild. New repo, clean models, a fresh start. It feels decisive, and it is usually the most expensive way to find out what the old stack was actually doing.
Why rebuilding first goes wrong
An inherited stack is not just code. It is a set of answers that people already rely on, some of them right, some of them quietly wrong, and nobody can tell you which is which.
Rebuild before you know, and you carry three risks into the new system. You rebuild things nobody uses. You drop logic that someone depends on but nobody wrote down. And you reproduce the wrong number with cleaner code, which makes it harder to question.
An audit costs a week. It tells you what to keep, what to fix and what to switch off. Sometimes the decision is the deliverable.
The audit, in seven steps
-
Inventory what runs. Every scheduled thing: dbt jobs, Cloud Scheduler triggers, Airflow DAGs, Fivetran or Airbyte connectors, scripts on someone's laptop. For each one, write down what triggers it, what it writes, and who gets told when it fails. "Nobody" is a common answer.
-
Find what nobody reads. Most stacks carry tables that are refreshed every night and opened by no one. In BigQuery, the job history tells you exactly which tables people query. The query below is the first thing we run.
-
Trace the three numbers that matter. Ask leadership which three figures they look at every week. Trace each one from the dashboard back to the source system, and count how many definitions you find along the way.
-
Check whether the tests can fail. A green test suite proves nothing if the assertions are too loose to ever trip. We wrote about a migration where every task ran green while loading duplicate rows.
-
Check who holds the keys. Service accounts, personal credentials, deploy keys and anyone with write access. Inherited stacks often run on a key that belongs to someone who left.
-
Price what's left. Bytes billed per table and per model, from the same job history. The expensive model and the important model are rarely the same one.
-
Write the decision down. For every component: keep, fix or retire, with one line of reasoning. That document is what you rebuild from, if you rebuild at all.
Step 2 in practice: tables nobody has read in 90 days
-- BigQuery: tables with no reads in the last 90 days.
-- JOBS history covers 180 days. Exclude your pipeline's own
-- service account, or every table it rebuilds will look "read".
WITH reads AS (
SELECT
ref.dataset_id,
ref.table_id,
MAX(creation_time) AS last_read
FROM `region-us`.INFORMATION_SCHEMA.JOBS_BY_PROJECT,
UNNEST(referenced_tables) AS ref
WHERE creation_time > TIMESTAMP_SUB(CURRENT_TIMESTAMP(), INTERVAL 180 DAY)
AND job_type = 'QUERY'
AND state = 'DONE'
AND user_email NOT LIKE '%gserviceaccount.com'
GROUP BY 1, 2
)
SELECT
t.table_schema AS dataset_id,
t.table_name,
r.last_read
FROM `region-us`.INFORMATION_SCHEMA.TABLES AS t
LEFT JOIN reads AS r
ON r.dataset_id = t.table_schema
AND r.table_id = t.table_name
WHERE r.last_read IS NULL
OR r.last_read < TIMESTAMP_SUB(CURRENT_TIMESTAMP(), INTERVAL 90 DAY)
ORDER BY r.last_read NULLS FIRST;Excluding service accounts matters. If your BI tool queries through one, add it back in by name, otherwise you will flag your most-used dashboard tables as dead.
On the dbt side, compare dbt ls --resource-type model against what your dashboards actually reference. Models with no downstream consumer are candidates for retirement, not migration. In most Talend audits we've run, 20 to 30 percent of jobs turned out to be dead weight. dbt projects are not immune.
What the audit usually finds
At a B2B security SaaS, step 3 found three different definitions of "converted lead", one in the CRM, one in a spreadsheet and one in a dashboard filter. Each was defensible. None of them agreed. Settling a single definition in dbt did more for trust in the numbers than any new dashboard would have.
That is the typical shape. The problem is rarely the tool. It is metrics that don't match, jobs nobody owns, and a few tables everyone depends on without knowing it.
What to do next
Before you approve a rebuild, ask for the keep, fix or retire list. If nobody can produce one, that's the first piece of work.
If you've inherited a stack and want a second pair of eyes on it, book a 30-minute call. We'll tell you honestly whether it needs a rebuild or a week of cleanup.
Rebuilding is a decision. The audit is how you earn the right to make it.