Small pipelines fail quietly
A big data platform fails loudly. Alerts fire and someone gets paged. The small pipelines an analytics team runs on its own, the scheduled scripts and nightly syncs, fail in a different way. They stop, and the dashboard keeps showing yesterday’s numbers with total confidence.
After enough of those, I wrote a monitor for my own jobs. This is what I learned building it.
A job that didn’t run sends no error
Most alerting watches for failures. The common failure is absence: the machine was asleep, the schedule was disabled, a token expired before the first line ran. Check for the last success, not the last error.
Check the data, not the job
A green run tells you the script exited cleanly. It doesn’t tell you rows landed. For every pipeline I check two things at the destination: when the table last changed, and how many rows it has. Stale or shrinking tables get flagged even when the job reports success.
Make “unknown” a status
If the monitor can’t reach a source, it should say so. A check that silently passes when it can’t see is worse than no check.
Reconcile on a schedule
Incremental syncs drift. Records get deleted or archived upstream and nothing tells you. A periodic full comparison against the source catches what the increments miss.
Put freshness on the dashboard
The cheapest fix of all is a “data as of” line where people can see it. It turns a silent failure into something a reader notices on their own.
None of this needs a platform. Mine is one script and a status page. The work is deciding, for each pipeline, what healthy means and writing that down as a check.
