
Half the Data Engineer CVs I read in 2026 look like a Data Analyst CV with Airflow tacked on, or a DevOps CV with dbt sprinkled in. Recruiters cannot tell who you actually are. A Data Engineer owns the pipes other people depend on. If your bullets do not show that ownership, you blur into the dozen analysts or platform engineers in the same inbox.
Stop sounding like a Data Analyst or DevOps
The bullet "built dashboards in Looker for the finance team" is an analyst sentence, not yours. The bullet "set up Terraform for the EKS cluster" is a DevOps sentence. Your bullets need to live in the middle: you built the thing that feeds the dashboard, and you set up the infra that runs the thing. Concretely: pipeline ownership, schema design, orchestration, freshness SLA. The recruiter is scanning for that exact territory.
And the Analytics Engineer overlap in 2026 is the trap nobody warns about. If every bullet is dbt models and tests, you read as AE, not DE. Mix in ingestion, streaming, infra cost work. That contrast is what tells a hiring manager which side of the data team you sit on.
The bullet shape that proves ownership
"Owned the events pipeline (25M events/day) for the product analytics team, freshness SLA 99.4% over 14 months, on-call rotation of one." Three load-bearing things: scale, reliability metric with a window, the fact you were the person paged. Compare with "worked on the events pipeline". Same job. Completely different read. The first sentence tells the recruiter you carry a beeper. The second tells them you did some tickets.
If you cannot share volume because of NDA, share architecture: "12 TB warehouse, 40+ DAGs, 600M events/month". Those numbers almost never violate NDA, but they are read as proof of scale.
Modern stack realism in 2026
Listing every warehouse vendor on Earth makes you look junior. In 2026 the actual market split is roughly: Snowflake and Databricks for enterprise and ML-heavy shops, BigQuery for GCP-native and product analytics, Redshift mostly as legacy you migrate off of. Pick the one you really shipped in. List the second one as "working knowledge". Drop the rest.
Same logic for orchestration and streaming. dbt is the transformation default; Airflow is still the orchestration default but Dagster and Prefect are no longer exotic. For streaming, the realistic split is Kafka if you are on-prem or multi-cloud, Kinesis if you live in AWS, Pub/Sub if you live in GCP. If you list all five with no context, you signal you have not run any of them in production.
Spark / Scala vs Python in 2026
The honest answer in 2026: PySpark won the day-to-day, Scala stayed alive in legacy enterprise and a few performance-critical teams. If your CV leads with "Scala expert" and you are applying outside of finance, telecom, or specific Databricks shops, you are filtering yourself out. List PySpark first, mention Scala only if you actually shipped Scala in the last two years. Pure Python with pandas is fine for under-100GB jobs - say that out loud, do not pretend you would Spark every batch.
Run your CV through the CV Analyzer against a real Data Engineer job and watch which keywords miss. The gap between your bullets and the JD is usually freshness SLA, cost per query, or data contracts. Those three carry the senior signal.
The senior bar nobody writes about

Mid-level Data Engineer CVs talk about pipelines they built. Senior CVs talk about systems other teams depend on, and the contracts that hold them together. Four phrases that signal senior in 2026:
- Data Contracts ("introduced contracts between backend and data, breaking schema changes in prod dropped to zero for the quarter")
- Lineage ("OpenLineage rolled out across 200+ models, MTTR for upstream incidents cut from 4h to 35min")
- Cost optimization ("reduced monthly Snowflake bill by $14k via clustering and warehouse autosuspend tuning")
- Governance ("led PII tagging rollout for GDPR audit, 18 sources, signed off by legal in one round")
If at least two of those four are missing from your CV and you are applying senior, hiring managers re-grade you to mid before the call. Not out of malice, just pattern-matching on what they read all day.
Three numbers worth more than your stack
- Uptime / freshness SLA with a window ("99.4% over 14 months")
- Cost per query or cost per TB processed, before and after
- Number of downstream consumers ("8 product squads", "40+ dashboards", "3 ML models")
ATS gotcha for Data Engineers
Data engineers love writing "ETL/ELT" with a slash. Some ATS index that as "ETL" only, or worse, "ETL" and "ELT" lose each other in tokenization. Same trick as analysts: spell it out once, "ETL and ELT pipelines", then shorten. Also: "PySpark" and "Apache Spark" are different keywords in many ATS - put both in skills if you have shipped both.
When you apply, log it in the Job Tracker with the JD attached. After 8-10 applications you will see which keywords keep appearing - that is your real keyword set, not the one you guessed when writing the CV.
The interview answer that pushes you up a band
"Tell me about a pipeline that broke in production." Mid engineers describe the fix. Senior engineers describe the blast radius, what they communicated, and what they changed in the system so it cannot happen the same way again. "Schema change upstream killed our event ingestion for 3 hours. I posted in the incident channel within 10 minutes with affected dashboards, rolled back the consumer, then wrote a data contract test that fails CI on the producer side. We have not had a similar incident since." That second half is the senior signal.
Organise your job search with Trackr
Track applications, analyse your CV with AI, and prepare for interviews - free.
Get started free

