Sourcing by role
How to find a data engineer.
The title is barely a decade old. The people with fifteen years of this experience spent most of it called something else — and most searches quietly exclude them.
Data engineering has a vocabulary problem that hides its most experienced practitioners. Someone who has built data infrastructure since 2010 spent the first half of that career as an ETL Developer or Data Warehouse Developer, and a search for 'Data Engineer' never reaches them.
The second problem is more expensive. A great many advertised data science roles are data engineering roles in disguise: the organisation wants models but has no reliable pipelines, and the first year is infrastructure work. Naming that honestly attracts someone who wants the job as it actually exists.
Job titles worth searching
Grouped by what the person actually does, because searching all43 at once produces a result set you cannot triage. Decide which group you need first — that decision does more for the search than any string below.
Core titles
Data Engineer became the standard term only in the last decade. Anyone with fifteen years in this work called themselves an ETL Developer or Data Warehouse Developer for most of it, and those people are frequently excluded by searches that only use the current title.
- Data Engineer
- Senior Data Engineer
- Staff Data Engineer
- Principal Data Engineer
- Big Data Engineer
- Data Platform Engineer
- Data Infrastructure Engineer
Legacy and warehouse terms
The largest overlooked pool. ETL developers and warehouse specialists have deep modelling discipline — dimensional design, slowly changing dimensions, data quality — that many self-taught modern-stack engineers lack. Whether they transfer depends on their appetite for code over GUI tools.
- ETL Developer
- ETL Engineer
- Data Warehouse Developer
- Data Warehouse Engineer
- BI Developer
- Informatica Developer
- SSIS Developer
- Datastage Developer
- Database Developer
Modern stack
Where most current demand sits. This group works in code rather than drag-and-drop tools, treats pipelines as software, and expects version control and testing. The tooling names below are more reliable than any title for identifying them.
- Analytics Engineer
- dbt Developer
- Data Pipeline Engineer
- Snowflake Engineer
- Databricks Engineer
- Lakehouse Engineer
- DataOps Engineer
- Data Reliability Engineer
Streaming and real-time
A genuinely distinct specialisation. Batch and streaming involve different failure modes, different reasoning about correctness, and different tooling. A strong batch engineer does not automatically handle exactly-once semantics or late-arriving data well.
- Streaming Data Engineer
- Real-Time Data Engineer
- Kafka Engineer
- Event Streaming Engineer
- Flink Developer
- Spark Streaming Engineer
Platform and adjacent engineering
Where data engineering blurs into infrastructure. In larger organisations these people build the platform other data engineers use, and they need genuine software and systems engineering depth rather than pipeline authoring alone.
- Data Platform Engineer
- Data Infrastructure Engineer
- Backend Engineer (Data)
- Distributed Systems Engineer
- Software Engineer (Data)
- Cloud Data Engineer
- Data Architect
Governance and quality
Increasingly separated out as data volumes and regulation grow. These roles suit people who care about lineage, contracts, and correctness rather than throughput, and they are a good landing spot for former warehouse specialists.
- Data Quality Engineer
- Data Governance Engineer
- Master Data Management Specialist
- Data Steward
- Metadata Engineer
- Data Observability Engineer
Where data engineers actually are
GitHub reveals working style as much as capability here. Published dbt models, Airflow DAGs, and pipeline code show an engineer who treats data work as software — version controlled, tested, reviewed — which is precisely what modern data teams expect and what many legacy-tool practitioners have never done.
Open source contribution is the strongest available signal. The core tools in this field are open — Airflow, dbt, Spark, Delta Lake, Kafka — and contributors to the stack your team depends on are both highly capable and rarely approached. As with all maintainer outreach, a message that shows no familiarity with their work will be ignored.
The conference circuit is unusually tractable. Data Council, Coalesce, Kafka Summit, and Data + AI Summit are the main gatherings, and their speaker lists are small enough to function as a regional shortlist of senior practitioners. For the legacy pool, dimensional modelling vocabulary — Kimball, star schema, slowly changing dimensions — finds experienced people that tool-name searches miss entirely.
Boolean search strings
Written to be pasted as-is. Each one is built around an intent rather than a platform, since the useful question is what you are trying to find, not which site you happen to be on.
LinkedIn profiles, direct X-ray
Google (LinkedIn)site:linkedin.com/in/ ("data engineer" OR "analytics engineer" OR "ETL developer") (Airflow OR dbt OR Snowflake OR Spark) "{city}"Reasonably productive by the standards of this cluster, since data engineers list tooling in headlines and LinkedIn still indexes some headline text — but titles and locations are no longer exposed to crawlers, so pair the tool names with the title rather than relying on the title alone.
Modern stack practitioners by tooling
Google (GitHub)site:github.com ("dbt" OR "airflow" OR "dagster" OR "prefect") ("models" OR "dags" OR "pipeline") -awesome -tutorialThe most reliable filter available. Engineers who publish dbt models or Airflow DAGs treat pipelines as software — version controlled, tested, reviewed — which is exactly the working style modern data teams expect.
Legacy warehouse specialists
Google("ETL developer" OR "data warehouse developer" OR Informatica OR SSIS OR Datastage) ("dimensional" OR "star schema" OR "Kimball") -jobsThe Kimball and dimensional modelling vocabulary identifies people with genuine warehouse discipline. This population is large, experienced, and consistently overlooked by searches using only current terminology.
Streaming specialists
Google("Kafka" OR "Flink" OR "Kinesis" OR "Pulsar") ("exactly-once" OR "event-driven" OR "stream processing") -jobs -courseThe correctness vocabulary is the filter. Someone discussing exactly-once semantics or watermarking has genuinely worked in streaming rather than having run a tutorial.
Cloud platform specialists
Google("BigQuery" OR "Redshift" OR "Snowflake" OR "Databricks") ("cost optimization" OR "partitioning" OR "clustering") -jobs -salesCost and performance tuning vocabulary separates people who have run these platforms at scale from those who have only queried them. Warehouse spend is a real operational concern, and engineers who manage it are valuable.
Conference speakers and community contributors
Google("speaker" OR "talk") ("Data Council" OR "Coalesce" OR "Data + AI Summit" OR "Kafka Summit") 2024..2026 -jobsThe data engineering conference circuit is small enough that speaker lists are a workable shortlist of senior practitioners in a region.
Open source contributors to data tooling
Googlesite:github.com ("contributor" OR "maintainer") ("apache/airflow" OR "dbt-labs" OR "delta-io" OR "apache/spark")Contributors to the core tools your stack depends on are the highest-signal candidates available. Small population, rarely approached, and they expect outreach that shows familiarity with their work.
Engineers writing about pipeline architecture
Google("how we" OR "lessons" OR "migrating") ("data pipeline" OR "data platform" OR "warehouse") ("scale" OR "cost" OR "reliability") -courseWriting about a migration or a platform rebuild reveals architectural judgement that no CV conveys. The cost and reliability framing filters toward practitioners rather than vendors.
Mistakes that cost the most time
Advertising a data engineering role as data science
This is the most common miscast in the whole data field. An organisation without reliable pipelines advertises for a data scientist, hires a modeller, and the actual first-year work turns out to be building infrastructure. The data scientist leaves. If the pipelines do not exist, the honest advert is for a data engineer, and it will attract someone who genuinely wants that work.
Excluding ETL developers by vocabulary
Data Engineer is a recent title. Practitioners with fifteen years in this work spent most of it as ETL Developers or Data Warehouse Developers, and many have modelling discipline that self-taught modern-stack engineers lack. Searching only current terminology quietly removes the most experienced part of the market.
Treating batch and streaming as one skill
Streaming introduces problems batch does not have — exactly-once processing, late-arriving data, watermarking, and state management under failure. A capable batch engineer needs real ramp time to reason about these correctly. If the role is genuinely real-time, the streaming experience is a requirement rather than a preference.
Screening on tool names alone
Tooling in this field turns over quickly, and a competent engineer moves between orchestrators without difficulty. What does not transfer easily is the modelling discipline and the systems thinking. Screening for an exact tool match selects for keyword-matched CVs rather than for the people who will still be effective when the stack changes.
Confusing analytics engineers with data engineers
Analytics engineers model data inside the warehouse, usually in SQL and dbt, and sit close to the business. Data engineers build the systems that get data there and keep them running. The roles are complementary and often confused, and hiring one when you need the other leaves a visible gap that neither can cover.
Underestimating the on-call reality
Data pipelines fail at night and someone has to fix them. Candidates ask about on-call rotation, alerting quality, and how much of the job is firefighting versus building, because these determine daily experience far more than the tech stack does. Adverts that avoid the subject are read as concealing a bad answer.
Common questions
- What job titles should I search for when hiring a data engineer?
- Search both current and legacy terminology, because the title changed recently. Current terms include Data Engineer, Analytics Engineer, Data Platform Engineer, and DataOps Engineer. Legacy terms cover a large and experienced population: ETL Developer, Data Warehouse Developer, BI Developer, and tool-specific titles such as Informatica or SSIS Developer. Streaming specialists use Kafka Engineer, Flink Developer, and Real-Time Data Engineer. The most reliable identifiers are not titles at all but tooling names — Airflow, dbt, Spark, Snowflake, Databricks.
- What is the difference between a data engineer and an analytics engineer?
- A data engineer builds and operates the systems that move data — ingestion, orchestration, storage, and reliability — and needs genuine software engineering ability. An analytics engineer works inside the warehouse, modelling data into usable form with SQL and dbt, and sits closer to the business and its definitions. They are complementary rather than interchangeable: hiring an analytics engineer when the pipelines do not exist leaves the ingestion problem unsolved, and hiring a data engineer when the need is business modelling leaves the warehouse full of raw tables nobody can use.
- Can an ETL developer transition to modern data engineering?
- Frequently yes, and this is the most overlooked pool in the market. ETL developers often have stronger data modelling discipline than self-taught modern-stack engineers, including dimensional design, slowly changing dimensions, and data quality practice that predates current tooling. The genuine question is appetite rather than capability: the modern stack is code-first, version controlled, and tested, which is a real shift from GUI-driven tools. Candidates who have already started working in SQL and Python rather than drag-and-drop interfaces usually transition well.
- Why do so many data science hires turn out to need data engineering?
- Because organisations underestimate how much infrastructure precedes analysis. A company decides it needs machine learning, advertises for a data scientist, and hires someone strong at modelling — who then discovers there are no reliable pipelines, no consistent definitions, and no accessible historical data. The first year becomes data engineering work the person did not want and may not be good at, and they leave. If the foundational data infrastructure is not in place, hiring a data engineer first produces better outcomes for everyone.
- Where can I find data engineers outside LinkedIn?
- GitHub is the strongest source, because published dbt models, Airflow DAGs, and pipeline code demonstrate both capability and working style — engineers who version control and test their pipelines are visibly different from those who do not. Contributors to core tools such as Airflow, dbt, Spark, and Delta Lake are the highest-signal candidates available. The conference circuit is small enough to be useful: Data Council, Coalesce, Kafka Summit, and Data + AI Summit publish speaker lists that amount to a shortlist of senior practitioners.
The method behind the strings
Sourcing, in full.
Full Stack Recruiter devotes its first seven chapters to search: Boolean fundamentals, search engines beyond Google, research sources, contact discovery, and responsible public-source research. The titles change by role; the method under them does not.