AI QA engineer is one title for two different jobs: testing AI systems, and testing ordinary software with AI tools. The first checks whether a model, a prompt or an agent gives answers good enough to ship; in the second, AI tools do part of the testing work.
Both get posted as “AI QA engineer”. They need different skills, sit in different teams and, judging by live postings, pay very differently. And a third thing hides under nearby search terms like “AI tester” or “AI test engineer”, which isn’t a QA job at all.
Which of the two jobs is in the ad
The title rarely tells you. The verbs in the responsibilities do.
Testing AI postings talk about evals, golden datasets, rubrics, LLM judges, red-teaming and regressions between model versions. Glean’s evals engineer posting describes systems that “decide whether new models, prompts, retrieval strategies, and agent workflows are ready to ship”. These roles are usually staffed as ML or infrastructure engineers, and the pay follows: Anthropic’s post-training evaluations role lists $500,000 to $850,000, at the far end of the market.
Testing with AI postings talk about generating test cases from requirements, self-healing scripts, AI-assisted failure triage and coding assistants. DevRev’s quality engineer posting wants to “reduce reliance on manual test case creation”, and lists experience validating AI systems only under nice-to-have. Perplexity’s senior QA role, $140,000 to $160,000, asks testers to use tools like Cursor or Claude Code to fix small bugs themselves.

Neither family is fake. The trouble starts when an ad borrows the vocabulary of the first and pays for the second, or when you prepare for one and interview for the other.
Deterministic tests vs evals
If you come from classic QA, the second family will feel familiar. The first one breaks a few habits you probably rely on:
| Deterministic test | Eval | |
|---|---|---|
| Pass criteria | Output equals the expected value | Output scores above a threshold on qualities you defined (correct, grounded, safe, on tone) |
| Who knows the right answer | The spec, or the developer | Often nobody up front; a domain expert labels a sample first |
| Same input run twice | Same result | Can differ, so you sample and look at rates |
| A regression is | A test that went red | A score that dropped on a fixed dataset after a model, prompt or retrieval change |
| Grading at scale | Assertions in code | Code checks where possible, a model as judge where not, humans on a sample |
| Typical tooling | Playwright, pytest, JUnit, CI | Eval frameworks, labelled datasets, tracing, judge prompts |
The row that causes the most pain is the second one. A model acting as judge is only useful if it agrees with a person who knows the domain, so the judge needs its own test. Hamel Husain’s guide to LLM judges recommends labelling a set of real outputs with a domain expert, measuring how often the judge agrees on both passes and failures separately, and preferring a binary pass/fail over a 1 to 5 scale, because the gap between a 3 and a 4 is hard to define. Anthropic’s own evaluation docs push in the same direction from the other side: more questions with automated grading usually beat a few hand-graded ones, and the judging model should be a different one than the model being tested. In practice, a lot of the early work in a real AI QA job goes into writing labels and arguing about what “correct” means, before any test exists. A Firecrawl evals posting says the metric itself “has to be invented”.
In Europe there’s also a regulatory reason for this work. Article 15 of the EU AI Act requires high-risk AI systems to reach an appropriate level of accuracy and robustness and to declare their accuracy metrics. Somebody has to produce those numbers and defend them.
AI tester jobs: a real AI QA role or an annotation gig?
Search “AI tester” and Google mostly shows AI-text detectors. Search “AI tester jobs” and the results mix QA postings with something else: Google’s own suggestions next to that search include “Outlier AI”, “AI training jobs”, “What is an AI task reviewer job?” and “Does AI task pay real money?”.
That something else is data annotation. Platforms such as DataAnnotation and Outlier pay contractors to rate, correct and write model answers. It’s real work and some of it pays decently: DataAnnotation’s own FAQ advertises general tasks starting at $25 to $50+ an hour and coding tasks from $40 to $150+. But the conditions are those of gig work. TIME reported workers who passed the assessment and were then offered no tasks, and one account deactivated with $2,869 of work unpaid. A Business Insider report on Scale AI’s contractors quoted a tasker who spent close to 40 hours in one month in unpaid onboarding without landing any work.

How to tell the two apart before you spend an evening on an assessment:
| Signal | AI QA or evals role | Annotation gig |
|---|---|---|
| Employer | A named company, on its own careers page or ATS | A platform; the client is often anonymous |
| Pay unit | Annual salary, often with equity | Per hour or per task, often paid weekly |
| Hiring step | Interviews with the team you’d join | A screening assessment, frequently unpaid |
| What you produce | Eval sets, grading scripts, reports for release decisions | Ratings and rewritten answers to someone else’s rubric |
| Continuity | Employment contract and notice period | Projects open and close; access can be removed |
| Who you report to | An engineering or AI lead | A queue |

Annotation isn’t worthless for a career. Writing rubrics and grading model output is the raw skill behind evals, and Harvey, the legal AI company, hires people to run human evaluation quality at a salaried level. The step from one to the other is building something, an eval set or a grading script, and showing it.
Real AI QA roles carry a quieter risk. Of the ten live AI QA and evals postings reviewed for this page, none gave the role explicit authority to block a release. You may own “the quality bar” and still only inform the decision. Ask in the interview who decides when an eval fails a week before launch, and what happened the last time.
If you’re weighing a move from classic QA into evals, or two offers that both say “AI” and mean different things, that’s a decision worth talking through. It’s the kind of career question I work on in 1:1 career mentoring.



