Reading a data engineering CV like an engineer, not a recruiter

Reading a data engineering CV like an engineer, not a recruiter

A typical data engineering CV lists a stack: Spark, Airflow, dbt, Snowflake, Kafka, pick your combination. Screening on that list alone tells you almost nothing, because tool familiarity is the easiest thing to learn on the job and the easiest thing to pad on a CV. What it doesn't tell you is whether someone can design a pipeline that survives contact with messy, late, or duplicated real-world data — which is the actual job.

So instead of starting with tools, we start with scale and consequence: what broke when their pipeline failed, and who found out first? A candidate who's run production data infrastructure has a specific, slightly uncomfortable failure story ready — a backfill that took down a downstream dashboard the CFO was using, a schema change that silently corrupted a week of records before anyone noticed. If a candidate can't produce one of these without prompting, we treat that as a real signal, not a technicality.

The second thing we look for is ownership of data quality, not just data movement. Plenty of engineers can move data from A to B. Fewer can describe how they'd catch a subtle problem — a currency field silently switching units, a join key that started producing duplicates after an upstream API change — before it reaches a dashboard a VP is looking at. That's the difference between someone who builds pipelines and someone who's accountable for what comes out the other end.

None of this means tooling doesn't matter; if a role is built entirely around a specific stack, obviously that experience helps. But we weight it after we've confirmed the more important thing, which is whether this person has actually been on call for data infrastructure that mattered to the business, not just built it in a sandbox. That's what separates a strong data engineering hire from a CV that reads well.

Hiring, or looking for your next role?

Start a conversation