Data engineer job description template: pipeline scope, an honest stack and screening questions
On this page
A data engineer job description often reads like a list of every logo in the modern data stack. That attracts keyword-heavy resumes and tells strong engineers nothing about the actual work: how much data, how fresh it must be, what breaks, and who gets paged. The title also covers several jobs, from modeling warehouse tables for analysts to running streaming systems for a product. Below is a copy-ready data engineer job description template, the on-call, data access, pay and EEO lines, how each must-have becomes a screening question and scorecard row, and the mistakes that make data engineering postings indistinguishable. The general method is in how to write a job description that screens.
Name the kind of data engineering first
| Kind | What fills the week | Must-have that separates it |
|---|---|---|
| Analytics engineering | Transforming raw data into tested models analysts query; documentation | Has built data models others rely on, with tests that caught real problems |
| Batch pipeline engineering | Ingesting from sources and APIs, scheduling, backfills, failure handling | Has owned pipelines in production and fixed them when they broke |
| Streaming or real-time | Event pipelines feeding products or alerts, latency and ordering problems | Has run a streaming pipeline with a latency requirement |
| Data platform | Warehouse or lake infrastructure, cost, access control, tooling for other engineers | Has run shared data infrastructure, including cost and permissions |
| ML data engineering | Feature pipelines, training data, model inputs in production | Has built data pipelines for a model that runs in production |
Then settle the facts that define the job: rough data volume and freshness (daily batches or seconds), the number of sources, who consumes the data, the size of the data team, and whether there is an on-call rotation. Those facts tell a candidate more than the tool list, and they let you level the role by scope rather than years.
The data engineer job description template
[Senior] Data Engineer, [team: Analytics / Platform / Product data]
[Company], [city] — [On-site / Hybrid: days / Remote within: states
or time zones]
Pay: $[min]–$[max] per year, plus [bonus / equity, if any].
[Benefits summary, if required.]
SUMMARY
You will build and run the pipelines that bring data from [N sources:
e.g. our product database, payments provider and CRM] into [warehouse
or lake], and model it for [analysts / product features / machine
learning]. Data arrives [daily / hourly / in real time], and [teams]
rely on it by [time]. You will join a team of [N] data engineers and
report to [title].
WHAT YOU WILL DO
- Build and maintain pipelines from [sources] with [orchestration
tool], including backfills and schema changes.
- Model data for [use], with tests and documentation others can rely
on.
- Monitor data freshness and quality, and fix failures at the cause.
- Work with [analysts / product engineers / data scientists] to agree
definitions and requirements.
- Keep costs and access under control in [warehouse / cloud
platform].
- [Streaming: run event pipelines with a latency target of N.]
YOU MUST HAVE
- Built and maintained data pipelines in production, and been
responsible when they failed.
- Strong SQL, including modeling data for other people to query.
- [Python / Scala / Java] used to build pipelines or tooling.
- Handled a data quality problem: found it, fixed it and stopped it
recurring.
- [Platform: managed shared data infrastructure, including access
and cost.]
NICE TO HAVE
- Experience with [warehouse], [orchestration], [transformation
tool], [streaming platform], [cloud provider].
- Knowledge of [domain: payments, healthcare, advertising data].
If you meet the must-haves and none of these, please apply.
FIXED CONDITIONS
- On-call: [rotation of N engineers; one week in N; how it is
compensated].
- Data access: the role works with [customer / financial / health]
data and requires [privacy or security training, background check
as permitted by law].
HOW WE HIRE
[Recruiter screen; hiring manager conversation; a [60-minute live
pipeline design discussion / take-home of at most N hours]; a
conversation with the team.]
[EEO STATEMENT]
If you need an adjustment to apply or interview, email [address].
Pay range and EEO statement
Pay. Leave the range bracketed until it is approved on the requisition. Remote data engineering roles open in several states can trigger several pay transparency laws at once; see pay transparency laws by state. For an outside reference, be careful with BLS data: the closest Occupational Outlook Handbook profile, database administrators and architects (last modified August 27, 2026), describes designing and managing databases and does not mention data pipelines. Treat any BLS figure, such as those for Database Architects (SOC 15-1243) in its OEWS program, as a loose comparison; this page does not quote figures.
EEO statement. The EEOC lists the federal protected characteristics, including age (40 or older), and gives "recent college graduates" as an example of ad wording that may discourage older applicants (EEOC). "Digital native" and "young team" carry the same risk. Close with your standard EEO statement and an accommodation contact.
From must-haves to screening questions and scorecard rows
A recruiter can screen these without deep technical knowledge: each question asks for a specific system the candidate worked on, and depth shows in the detail.
| Must-have | Screening question | Scorecard competency | Evidence of a strong answer |
|---|---|---|---|
| Production pipelines | "Describe a pipeline you owned: where the data came from, where it went, how often, and who used it." | Pipeline ownership | Concrete sources, schedule, volume and consumers; says "I" where they did the work |
| Failure handling | "Tell me about the last time a pipeline broke. How did you find out, and what did you change afterward?" | Reliability and debugging | Monitoring or alert, root cause, and a fix that prevents a repeat |
| SQL and modeling | "How did you model a table that analysts used every day? What did you get wrong at first?" | Data modeling | Grain, keys and definitions; a real correction they made |
| Data quality | "When did bad data reach a report or product? How was it caught?" | Data quality judgment | Tests or checks added, communication with users, source fixed |
| Cost and access (platform) | "What did you do the last time warehouse costs jumped?" | Platform stewardship | Found the cause (query, job, table) and changed it; knows who has access to what |
The full first-round screen is in data engineer phone screen questions. If the role is closer to analysis, compare it with the data analyst job description template; if it is mostly backend services, the software engineer job description template may fit better. The interview guide for software engineers covers the technical rounds that follow.
Common mistakes in data engineer job descriptions
The logo wall
Listing every warehouse, orchestrator, streaming platform and cloud as required describes no real stack and rewards keyword stuffing. Require the languages used daily; name the rest as the environment.
No scale or freshness
"Big data" means nothing on its own. A line such as "about 200 million events a day, available to analysts within an hour" (an invented example; use your own numbers) tells a candidate whether this is their kind of problem.
Hidden on-call
If the pipelines feed a morning executive report or a product, someone is on call. Say how often and how it is handled.
An analyst job with an engineer title
If most of the week is ad hoc queries and dashboards, engineers will leave and analysts will not apply. Title the role by the work.
Unbounded exercises
A take-home that builds a complete pipeline over a weekend filters for free time. Cap the time, state it, and use the same rubric for every candidate.
Before you publish
- The kind of data engineering, sources, volume, freshness and consumers are in the summary.
- Daily languages are required; other tools are described, not demanded.
- On-call, data access conditions and exercise length are stated.
- An approved pay range, the EEO statement and an accommodation contact close the posting.
If you paste the finished posting into Interview Signal, it builds the screen's question guide from the must-haves, and the scorecard quotes what the candidate said about the pipelines they owned and the failures they fixed.
Questions people ask
What is the difference between a data engineer and a data analyst posting?
A data engineer builds and runs the systems that move, store and model data: pipelines, warehouses and the jobs that keep them current. A data analyst uses that data to answer business questions. If the person will mostly write analysis queries and dashboards, title it analyst; if they will own pipelines and their failures, title it engineer.
Which tools should a data engineer job description require?
Require the languages used daily, usually SQL and often Python or another general-purpose language, and describe the rest of the stack as the environment. Engineers move between warehouses and orchestration tools fairly quickly; the harder skills are designing reliable pipelines, handling bad data and debugging failures.
Should a data engineer posting mention on-call?
Yes. If pipelines feed reports or products that people rely on each morning, someone will be paged when they fail. State the rotation, how often each person is on call and how it is compensated, if at all. Candidates ask about it, and leaving it out costs trust later.
How should a data engineer technical exercise be described in the posting?
Say what kind of exercise it is and how long it takes, for example a 60-minute live session to design and discuss a pipeline, or a take-home capped at a few hours with a provided dataset. Give everyone the same exercise and rubric, and discuss the candidate's choices with them rather than just grading the output.