Sourcing by role
How to find a data scientist.
One title, three genuinely different jobs. Until you know whether the role produces an experiment readout, a deployed model, or a paper, no search string will help you.
Data Scientist is the least informative title in technology, and that is saying something. The same two words describe a product analyst running experiments at a large company, an engineer shipping recommendation models at a startup, and a PhD publishing papers at a research lab. All three are correctly titled. None of them can do each other's jobs well.
Company size inverts the meaning further: at a large technology firm data scientists often do analytics while ML engineers handle production, and at a small one the same title covers the entire pipeline. The fix is not a better string. It is establishing at intake what the person will actually produce, and only then deciding where to look.
Job titles worth searching
Grouped by what the person actually does, because searching all47 at once produces a result set you cannot triage. Decide which group you need first — that decision does more for the search than any string below.
Core titles
One title covering at least three different jobs. At a large technology company a Data Scientist is often a product analyst who runs experiments; at a startup the same title means the person who builds and ships the models. Never source this role from the title alone — establish which of the three jobs it is first.
- Data Scientist
- Senior Data Scientist
- Staff Data Scientist
- Principal Data Scientist
- Lead Data Scientist
- Applied Scientist
- Research Scientist
- Decision Scientist
Analytics-leaning
Experimentation, causal inference, and product metrics rather than model deployment. These people are frequently stronger statisticians than the ML-leaning group and weaker engineers. If the role is about measuring whether something worked, this is the group you want.
- Product Data Scientist
- Product Analyst
- Quantitative Analyst
- Analytics Engineer
- Business Scientist
- Experimentation Scientist
- Growth Data Scientist
- Marketing Data Scientist
Machine learning leaning
Where the title blurs into ML engineering. The practical distinction is whether the person ships models into production or hands them to someone else. A candidate who has only ever worked in notebooks will struggle in a role that expects deployment, and the CV rarely makes this clear.
- Machine Learning Scientist
- Applied Machine Learning Scientist
- ML Research Engineer
- Deep Learning Scientist
- NLP Scientist
- Computer Vision Scientist
- Recommendation Systems Scientist
- Forecasting Scientist
Research and academic
PhD-heavy, publication-driven, and evaluated on papers rather than shipped products. Genuine research roles are rare, and candidates from this group are often mismatched into applied positions where the expectations are entirely different. Check what the role actually needs before sourcing here.
- Research Scientist
- Senior Research Scientist
- Research Engineer
- Postdoctoral Researcher
- Scientist
- Member of Technical Staff
- AI Researcher
Domain-specific
Where domain knowledge gates the role as firmly as technical skill. A clinical biostatistician and a quantitative researcher at a hedge fund are both data scientists and are not remotely interchangeable. In these fields the domain term is more discriminating than the data term.
- Biostatistician
- Computational Biologist
- Bioinformatics Scientist
- Quantitative Researcher
- Risk Data Scientist
- Fraud Data Scientist
- Clinical Data Scientist
- Geospatial Data Scientist
- Actuarial Data Scientist
Adjacent and easily confused
Frequently conflated with data science in job adverts, and genuinely different work. A Data Engineer builds pipelines, an Analytics Engineer models warehouse data, and an ML Engineer productionises models. Posting a data science role that actually needs one of these is a common and expensive miscast.
- Data Engineer
- Machine Learning Engineer
- MLOps Engineer
- Data Analyst
- Business Intelligence Analyst
- Statistician
- Operations Research Analyst
Where data scientists actually are
This role leaves a strong published trail, and it is the most reliable way to read someone's actual depth. arXiv and Google Scholar carry author affiliations and show what a person genuinely works on. For applied modelling ability, Kaggle rank is a real signal — with the caveat that it measures modelling in isolation and says nothing about production engineering or stakeholder communication.
GitHub separates the notebook population from the shipping population, which is the distinction most job descriptions fail to test. A data scientist whose repositories touch MLflow, Airflow, dbt, or containerisation has worked beyond exploratory analysis. That single filter resolves a large share of applied-role mismatches before the first conversation.
The most underused pipeline is academic. PhDs and postdocs leaving research are technically strong and comfortable with ambiguous problems, but their CVs are written for academic audiences and consistently undersell them against industry keyword screening. They are also far less contested than candidates already working in technology.
Boolean search strings
Written to be pasted as-is. Each one is built around an intent rather than a platform, since the useful question is what you are trying to find, not which site you happen to be on.
Published researchers by topic
Google Scholar / arXiv(site:arxiv.org OR site:scholar.google.com) "{topic}" ("{city}" OR "{company}") -jobsFor research and applied science roles, publications are the strongest available signal. arXiv listings carry author affiliations, and the papers show exactly what someone works on rather than what they claim.
LinkedIn profiles, direct X-ray
Google (LinkedIn)site:linkedin.com/in/ ("data scientist" OR "applied scientist" OR "research scientist") ("{technique}" OR "PhD") "{city}"Runs into the same indexing limits as every other X-ray search, and into a role-specific problem on top: the title is so ambiguous that even a working filter would not tell you whether the person does analytics, applied ML, or research. Publications and repositories answer that question and a LinkedIn headline does not, which is why the sources above matter more here than the profile search.
Kaggle competitors and notebook authors
Googlesite:kaggle.com ("competitions" OR "notebooks") ("master" OR "grandmaster" OR "expert") "{technique}"Kaggle rank is a genuine skill signal for modelling ability, though it says nothing about production engineering or stakeholder work. Useful as a filter for the ML-leaning group specifically.
Data scientists who ship, not just model
Google (GitHub)site:github.com ("mlflow" OR "airflow" OR "dbt" OR "docker") ("data scientist" OR "machine learning") -tutorial -awesomeThe infrastructure tools are the tell. A data scientist whose repositories include orchestration or deployment tooling has worked beyond the notebook, which is the distinction most job descriptions fail to test for.
Analytics and experimentation specialists
Google("A/B testing" OR "causal inference" OR "experimentation" OR "difference-in-differences") ("data scientist" OR "product analyst") -jobs -courseCausal inference vocabulary separates genuine experimentation practitioners from people who have run a t-test. Excluding 'course' strips the enormous volume of training content on these terms.
Conference speakers and workshop presenters
Google("speaker" OR "talk" OR "workshop") ("PyData" OR "NeurIPS" OR "ODSC" OR "Strata" OR "useR") 2024..2026 -jobsData science conferences publish speaker lists with affiliations. This population is small, senior, and demonstrably able to explain their work — which is much of what the interview is trying to establish.
Technical writers on data science
Google(site:towardsdatascience.com OR site:medium.com OR site:substack.com) ("how we" OR "lessons" OR "at scale") "{technique}" -course -tutorialWriting about applied work reveals depth and judgement. The 'how we' and 'at scale' phrasings filter toward practitioners describing real systems rather than explainers of textbook methods.
Domain specialists in regulated fields
Google("biostatistician" OR "clinical data scientist") ("SAS" OR "CDISC" OR "clinical trial" OR "FDA submission") -jobsIn pharma and clinical research the regulatory vocabulary is the filter. Someone who has worked on an FDA submission has experience a general data scientist cannot substitute for.
PhD holders leaving academia
Google("PhD" OR "postdoc" OR "doctoral") ("machine learning" OR "statistics" OR "computational") ("industry" OR "transitioning" OR "leaving academia") -jobsA large, underexploited pipeline. Academics moving to industry are strong technically, undersold on CVs written for academic audiences, and far less contested than candidates already in tech.
Mistakes that cost the most time
Not establishing which of the three jobs it is
Data Scientist covers product analytics, applied machine learning, and research science. These need different people, and a candidate excellent at one may be weak at another. The single most useful question at intake is what the person will produce — an experiment readout, a deployed model, or a paper. Everything else follows from the answer.
Reading the title without reading the company
Company size inverts the meaning. At a large technology company a Data Scientist often does product analytics while ML work sits with dedicated ML engineers; at a small company the same title covers everything from pipeline to production model. A candidate's title tells you almost nothing until you know where they held it.
Requiring a PhD by default
A PhD is genuinely necessary for research science roles and largely irrelevant for product analytics and most applied ML work. Requiring one by habit removes a large share of the strongest applied practitioners while attracting researchers who will find the role unsatisfying. Ask whether the work involves novel methods; if not, drop it.
Screening on tool lists instead of problems solved
Tool requirements in data science job descriptions are frequently arbitrary — a strong practitioner moves between pandas, R, and Spark without difficulty. Screening on named libraries filters for CV keyword matching rather than ability. What does discriminate is whether the person has solved a comparable class of problem.
Confusing data science with data engineering
A great many advertised data science roles are actually data engineering roles: the organisation has no reliable pipelines, and the first year will be spent building them. Data scientists hired into these roles leave. If the data infrastructure does not exist yet, the honest advert is for a data engineer.
Ignoring the academic pipeline
PhDs and postdocs leaving academia are technically strong, accustomed to ambiguous problems, and dramatically less recruited than candidates already in industry. Their CVs are written for academic audiences and undersell them against industry norms, which means keyword-based screening rejects them systematically.
Common questions
- What job titles should I search for when hiring a data scientist?
- The title itself is unreliable, so search by what the person actually does. For experimentation and product work, look for Product Data Scientist, Decision Scientist, Quantitative Analyst, and Experimentation Scientist. For model building, search Machine Learning Scientist, Applied Scientist, and specialisation terms like NLP Scientist or Computer Vision Scientist. For research roles, Research Scientist and Member of Technical Staff are common. In regulated domains the domain term is more discriminating than the data term — Biostatistician, Quantitative Researcher, or Clinical Data Scientist.
- What is the difference between a data scientist, data engineer, and ML engineer?
- A data engineer builds and maintains the pipelines and storage that make data usable. A data scientist analyses that data to answer questions or build models. A machine learning engineer takes models and makes them run reliably in production. The distinction matters because many advertised data science roles are actually data engineering roles in disguise — if an organisation has no reliable pipelines, whoever is hired will spend the first year building them, and a data scientist hired into that situation will leave.
- Does a data scientist need a PhD?
- For genuine research science roles, usually yes — the work involves novel methods and the ability to read and extend current literature. For product analytics and most applied machine learning, no, and requiring one by default is actively counterproductive. It removes many of the strongest applied practitioners, who came through engineering or analytics routes, while attracting researchers who will find the role unsatisfying and leave. The useful test is whether the work requires developing new methods or applying established ones well.
- Where can I find data scientists outside LinkedIn?
- The published work trail is strong for this role. arXiv and Google Scholar carry author affiliations and show precisely what someone works on, which matters for research and applied science roles. Kaggle rank is a real signal of modelling ability, though it says nothing about production engineering. GitHub reveals whether someone has worked beyond notebooks — repositories touching MLflow, Airflow, or dbt indicate production experience. Conference speaker lists from PyData, NeurIPS, ODSC, and similar events identify senior practitioners who can explain their work.
- Why do data science hires fail so often?
- Usually because of a mismatch established before the candidate was ever sourced. The most common pattern is hiring a modeller into an organisation with no usable data infrastructure, where the actual first-year work is data engineering. The second is title mismatch: hiring someone whose experience is product analytics into a role that expects production model deployment, or the reverse. Both are intake failures rather than sourcing failures, which is why establishing what the person will produce — an experiment readout, a deployed model, or a paper — is the highest-value part of the process.
The method behind the strings
Sourcing, in full.
Full Stack Recruiter devotes its first seven chapters to search: Boolean fundamentals, search engines beyond Google, research sources, contact discovery, and responsible public-source research. The titles change by role; the method under them does not.