AI, data and technology

Sourcing by role

How to find an AI specialist.

Prompt engineering as a standalone job is already dissolving. What remains scarce is evaluation — knowing whether a change actually made the output better.

The prompt engineer title is being absorbed into broader roles faster than almost any job title in recent memory, because prompting turned out to be a component of building AI systems rather than a discipline in itself. Many people write a good prompt. Far fewer can tell you whether a change improved anything.

That is where the actual scarcity sits: systematic evaluation, retrieval that grounds answers in real data, and understanding failure modes in production. The field also attracts an unusually high volume of shallow claims, which makes public artefacts more valuable here than in almost any other role.

Job titles worth searching

Grouped by what the person actually does, because searching all47 at once produces a result set you cannot triage. Decide which group you need first — that decision does more for the search than any string below.

Core AI specialist titles

A rapidly shifting set of labels. 'Prompt Engineer' as a standalone title is already being absorbed into broader roles, because prompting alone turned out to be a component of the work rather than the job itself.

  • Prompt Engineer
  • AI Specialist
  • AI Engineer
  • Applied AI Engineer
  • GenAI Specialist
  • LLM Engineer
  • AI Solutions Specialist
  • Conversational AI Designer

Evaluation and quality

Where the discipline is actually maturing. Evaluating model outputs systematically — building test sets, defining quality criteria, and measuring regression — is harder and more valuable than writing prompts, and far fewer people can do it.

  • AI Evaluation Engineer
  • Model Evaluation Specialist
  • AI Quality Engineer
  • LLM Evaluation Lead
  • Red Team Specialist (AI)
  • AI Test Engineer
  • Benchmark Engineer
  • Output Quality Analyst

System and pipeline building

Where prompting meets software engineering. Retrieval systems, agent orchestration, and tool integration require genuine engineering ability, and this is where most production AI work actually sits.

  • RAG Engineer
  • AI Agent Engineer
  • AI Integration Engineer
  • Retrieval Engineer
  • Applied Scientist
  • AI Product Engineer
  • Forward Deployed Engineer
  • Solutions Engineer (AI)

Content and linguistic

Where language expertise matters more than code. Conversation design and AI content roles draw from linguistics, technical writing, and UX writing backgrounds, and these people are often better at prompting than engineers are.

  • Conversation Designer
  • AI Content Designer
  • Computational Linguist
  • AI UX Writer
  • Dialogue Designer
  • Localisation AI Specialist
  • Voice Interaction Designer
  • Taxonomy Specialist

Domain-embedded AI roles

The most valuable and least contested group. Domain experts who became AI-fluent — lawyers, clinicians, analysts — evaluate AI outputs in ways generalists cannot, because they know when an answer is subtly wrong.

  • Legal AI Specialist
  • Clinical AI Specialist
  • Financial AI Analyst
  • AI Research Assistant
  • Domain AI Lead
  • AI Enablement Specialist
  • AI Trainer
  • Subject Matter Expert (AI)

Enablement and adoption

Where the work is organisational rather than technical. AI enablement roles help teams adopt tools effectively, which requires teaching ability and change management more than deep technical skill.

  • AI Enablement Manager
  • AI Adoption Lead
  • AI Programme Manager
  • AI Training Specialist
  • Internal AI Consultant
  • AI Champion
  • Productivity Specialist

Where AI practitioners actually are

Published work separates building from describing, which matters more in this field than any other because claims so far outpace capability. Repositories containing retrieval pipelines, agent orchestration, or evaluation harnesses demonstrate real work; a compelling description of AI experience demonstrates nothing.

Honest writing about failure is an unusually strong filter. Anyone can describe a successful demo, but explaining why a chunking strategy failed or why an evaluation set gave misleading results reveals someone who has operated a system in production rather than built one for a presentation.

Two pools are consistently overlooked. Conversation designers from linguistics and UX writing backgrounds are frequently better at prompt precision than engineers, since language exactness is their existing craft. And domain experts who became AI-fluent — lawyers, clinicians, analysts — can recognise when an output is confidently and subtly wrong, which is the judgement that determines whether an application is safe to deploy.

Boolean search strings

Written to be pasted as-is. Each one is built around an intent rather than a platform, since the useful question is what you are trying to find, not which site you happen to be on.

LinkedIn profiles, direct X-ray

Google (LinkedIn)
site:linkedin.com/in/ ("AI engineer" OR "prompt engineer" OR "applied AI") ("RAG" OR "evaluation" OR "fine-tuning" OR "LLM") "{city}"

Use with caution here. This title attracts a very high volume of shallow claims, and LinkedIn no longer indexes titles and locations anyway. The technical terms do the real filtering, and even then the published work below is a far better signal than any profile.

Practitioners with published AI systems

Google (GitHub)
site:github.com ("langchain" OR "llamaindex" OR "dspy" OR "instructor" OR "eval") ("retrieval" OR "agent" OR "pipeline") -tutorial -awesome

The strongest available filter. Published AI systems demonstrate whether someone builds working pipelines or writes prompts in a chat window, which the title cannot distinguish.

Evaluation and measurement depth

Google
("eval set" OR "evaluation harness" OR "LLM as judge" OR "regression testing" OR "golden dataset") ("LLM" OR "model") -course -vendor

Evaluation vocabulary is the clearest maturity signal in this field. Anyone can write a prompt; building systematic evaluation that catches regressions is genuinely difficult and rare.

Retrieval and grounding specialists

Google
("RAG" OR "retrieval augmented" OR "chunking strategy" OR "embedding model" OR "reranking") ("production" OR "latency" OR "accuracy") -course -vendor

Retrieval quality determines whether a system is useful or confidently wrong. Chunking and reranking vocabulary indicates someone who has debugged real retrieval failures.

Domain experts who became AI-fluent

Google
("lawyer" OR "clinician" OR "analyst" OR "researcher") ("using AI" OR "LLM" OR "built a tool" OR "automated") ("in my practice" OR "in our team") -jobs -course

The most valuable and least contested pool. Domain experts who can evaluate AI output know when an answer is subtly wrong, which generalist AI specialists frequently cannot detect.

Writers and demonstrable practitioners

Google
(site:substack.com OR site:medium.com OR site:github.io) ("we built" OR "lessons from" OR "what failed") ("LLM" OR "AI agent" OR "RAG") -course

The 'what failed' framing is particularly useful in a field full of confident claims. People writing honestly about what did not work have shipped something real.

Conversation and content designers

Google
("conversation design" OR "conversational AI" OR "dialogue design") ("chatbot" OR "voice" OR "assistant") -jobs -vendor -course

Conversation designers come from linguistics and UX writing backgrounds and are frequently better at prompt construction than engineers, because language precision is their existing craft.

Community and competition participants

Google
("AI hackathon" OR "LLM hackathon" OR "prompt competition" OR "AI engineer summit") ("winner" OR "speaker" OR "built") 2024..2026 -jobs

The applied AI community runs frequent hackathons and events. Participants are demonstrably building rather than describing, which matters in a field where claims outpace capability.

Mistakes that cost the most time

  1. Hiring for prompting when the job needs evaluation

    Writing a good prompt is a skill that many people can develop quickly. Building systematic evaluation — test sets, quality criteria, regression detection, and knowing whether a change improved anything — is genuinely difficult and rare. Most production AI problems are evaluation problems, and hiring for prompting alone leaves them unsolved.

  2. Underestimating the volume of shallow claims

    This title attracts an unusually high proportion of candidates whose experience is using consumer AI tools competently. That is not nothing, but it is not building production systems. Public artefacts — repositories, published systems, honest writing about failures — separate the two quickly where interviews frequently do not.

  3. Overlooking domain experts who became AI-fluent

    A lawyer or clinician who can evaluate AI output knows when an answer is confidently and subtly wrong, which a generalist AI specialist often cannot detect. For domain-specific applications this judgement matters more than technical AI depth, and these candidates are far less contested.

  4. Treating the title as stable

    Prompt engineering as a standalone job is already being absorbed into broader roles, because prompting turned out to be one component of building AI systems rather than the whole discipline. Hiring against the title rather than the actual work risks a role that looks outdated within a year.

  5. Ignoring conversation designers and linguists

    People from linguistics, technical writing, and UX writing backgrounds are frequently better at constructing precise prompts than engineers, because language precision is their existing craft. They are rarely considered for AI roles, which keeps them available.

  6. Screening on framework familiarity

    The AI tooling landscape changes faster than any other in technology, and today's framework may be legacy within a year. What transfers is understanding the underlying problems — grounding, evaluation, latency, cost, failure modes — since frameworks are learned in days by anyone who grasps those.

Common questions

What job titles should I search for when hiring an AI specialist?
Search the work rather than the title, since this vocabulary is shifting rapidly. For system building, RAG Engineer, AI Agent Engineer, AI Integration Engineer, and Applied AI Engineer. For the maturing evaluation discipline, AI Evaluation Engineer and Model Evaluation Specialist. For language-focused work, Conversation Designer and AI Content Designer, roles often filled from linguistics and UX writing. Domain-embedded titles such as Legal AI Specialist and Clinical AI Specialist identify the most valuable and least contested group.
Is prompt engineering still a real job?
As a standalone title it is already being absorbed into broader roles, because prompting proved to be one component of building AI systems rather than a discipline in itself. What remains genuinely valuable and scarce is everything around it: designing evaluation that detects whether a change actually improved outputs, building retrieval systems that ground responses in real data, orchestrating multi-step agent behaviour, and managing cost and latency in production. Hiring for prompting alone typically leaves the harder problems unaddressed.
How do I tell genuine AI capability from confident claims?
Look at artefacts rather than descriptions, because this field attracts a high volume of shallow claims. Published repositories showing retrieval pipelines, agent systems, or evaluation harnesses demonstrate building rather than describing. Writing about what failed is particularly informative — anyone can describe a successful demo, but explaining why a retrieval strategy did not work reveals genuine experience. In interview, questions about evaluation methodology and failure modes separate practitioners quickly, because those are the problems that only appear in production.
Should an AI specialist come from a technical or domain background?
It depends on the application, and domain background is more often the right answer than recruiters assume. For building AI infrastructure and pipelines, engineering ability is essential. But for applying AI within a specific field — legal, clinical, financial — a domain expert who became AI-fluent can evaluate outputs in ways a generalist cannot, because they recognise when an answer is confidently and subtly wrong. That judgement is frequently the binding constraint on whether an AI application is safe to deploy.
Where can I find AI practitioners outside LinkedIn?
GitHub is the strongest source, since published retrieval systems, agent frameworks, and evaluation harnesses demonstrate real building. Technical writing that discusses failures rather than successes identifies people who have shipped something. The applied AI community runs frequent hackathons and events whose participants are demonstrably building. Two under-searched pools are worth attention: conversation designers from linguistics backgrounds, who are often better at prompt precision than engineers, and domain experts who became AI-fluent within their own field.

The method behind the strings

Sourcing, in full.

Full Stack Recruiter devotes its first seven chapters to search: Boolean fundamentals, search engines beyond Google, research sources, contact discovery, and responsible public-source research. The titles change by role; the method under them does not.