AI QA Engineer: Two Jobs Behind One Title

· by Marian Kamenistak

Marian Kamenistak speaking at an engineering leadership event.

AI QA engineer is one title for two different jobs: testing AI systems, and testing ordinary software with AI tools. The first checks whether a model, a prompt or an agent gives answers good enough to ship; in the second, AI tools do part of the testing work.

Both get posted as “AI QA engineer”. They need different skills, sit in different teams and, judging by live postings, pay very differently. And a third thing hides under nearby search terms like “AI tester” or “AI test engineer”, which isn’t a QA job at all.

Which of the two jobs is in the ad

The title rarely tells you. The verbs in the responsibilities do.

Testing AI postings talk about evals, golden datasets, rubrics, LLM judges, red-teaming and regressions between model versions. Glean’s evals engineer posting describes systems that “decide whether new models, prompts, retrieval strategies, and agent workflows are ready to ship”. These roles are usually staffed as ML or infrastructure engineers, and the pay follows: Anthropic’s post-training evaluations role lists $500,000 to $850,000, at the far end of the market.

Testing with AI postings talk about generating test cases from requirements, self-healing scripts, AI-assisted failure triage and coding assistants. DevRev’s quality engineer posting wants to “reduce reliance on manual test case creation”, and lists experience validating AI systems only under nice-to-have. Perplexity’s senior QA role, $140,000 to $160,000, asks testers to use tools like Cursor or Claude Code to fix small bugs themselves.

Marian Kamenistak mid-talk on a dark conference stage, lines of code projected on the screen behind him.
Same title on the ad, different code on the screen: an eval pipeline or a test suite with an AI assistant attached.

Neither family is fake. The trouble starts when an ad borrows the vocabulary of the first and pays for the second, or when you prepare for one and interview for the other.

Deterministic tests vs evals

If you come from classic QA, the second family will feel familiar. The first one breaks a few habits you probably rely on:

Deterministic testEval
Pass criteriaOutput equals the expected valueOutput scores above a threshold on qualities you defined (correct, grounded, safe, on tone)
Who knows the right answerThe spec, or the developerOften nobody up front; a domain expert labels a sample first
Same input run twiceSame resultCan differ, so you sample and look at rates
A regression isA test that went redA score that dropped on a fixed dataset after a model, prompt or retrieval change
Grading at scaleAssertions in codeCode checks where possible, a model as judge where not, humans on a sample
Typical toolingPlaywright, pytest, JUnit, CIEval frameworks, labelled datasets, tracing, judge prompts

The row that causes the most pain is the second one. A model acting as judge is only useful if it agrees with a person who knows the domain, so the judge needs its own test. Hamel Husain’s guide to LLM judges recommends labelling a set of real outputs with a domain expert, measuring how often the judge agrees on both passes and failures separately, and preferring a binary pass/fail over a 1 to 5 scale, because the gap between a 3 and a 4 is hard to define. Anthropic’s own evaluation docs push in the same direction from the other side: more questions with automated grading usually beat a few hand-graded ones, and the judging model should be a different one than the model being tested. In practice, a lot of the early work in a real AI QA job goes into writing labels and arguing about what “correct” means, before any test exists. A Firecrawl evals posting says the metric itself “has to be invented”.

In Europe there’s also a regulatory reason for this work. Article 15 of the EU AI Act requires high-risk AI systems to reach an appropriate level of accuracy and robustness and to declare their accuracy metrics. Somebody has to produce those numbers and defend them.

AI tester jobs: a real AI QA role or an annotation gig?

Search “AI tester” and Google mostly shows AI-text detectors. Search “AI tester jobs” and the results mix QA postings with something else: Google’s own suggestions next to that search include “Outlier AI”, “AI training jobs”, “What is an AI task reviewer job?” and “Does AI task pay real money?”.

That something else is data annotation. Platforms such as DataAnnotation and Outlier pay contractors to rate, correct and write model answers. It’s real work and some of it pays decently: DataAnnotation’s own FAQ advertises general tasks starting at $25 to $50+ an hour and coding tasks from $40 to $150+. But the conditions are those of gig work. TIME reported workers who passed the assessment and were then offered no tasks, and one account deactivated with $2,869 of work unpaid. A Business Insider report on Scale AI’s contractors quoted a tasker who spent close to 40 hours in one month in unpaid onboarding without landing any work.

Marian Kamenistak with hand on forehead, looking at his laptop, contemplative.
Read the ad twice. The pay unit and the employer usually tell you more than the title.

How to tell the two apart before you spend an evening on an assessment:

SignalAI QA or evals roleAnnotation gig
EmployerA named company, on its own careers page or ATSA platform; the client is often anonymous
Pay unitAnnual salary, often with equityPer hour or per task, often paid weekly
Hiring stepInterviews with the team you’d joinA screening assessment, frequently unpaid
What you produceEval sets, grading scripts, reports for release decisionsRatings and rewritten answers to someone else’s rubric
ContinuityEmployment contract and notice periodProjects open and close; access can be removed
Who you report toAn engineering or AI leadA queue

Hand-drawn 2x2: Testing with AI and Testing AI on the horizontal axis, annual salary and per-task pay on the vertical one; both salaried AI QA jobs fill the top half, the annotation gig sits below in per-task pay.

Annotation isn’t worthless for a career. Writing rubrics and grading model output is the raw skill behind evals, and Harvey, the legal AI company, hires people to run human evaluation quality at a salaried level. The step from one to the other is building something, an eval set or a grading script, and showing it.

Real AI QA roles carry a quieter risk. Of the ten live AI QA and evals postings reviewed for this page, none gave the role explicit authority to block a release. You may own “the quality bar” and still only inform the decision. Ask in the interview who decides when an eval fails a week before launch, and what happened the last time.

If you’re weighing a move from classic QA into evals, or two offers that both say “AI” and mean different things, that’s a decision worth talking through. It’s the kind of career question I work on in 1:1 career mentoring.

Other AI roles on the map

Frequently asked

How can I be an AI tester?+
To become an AI tester, first decide which of three different jobs you mean. Testing AI systems needs Python, statistics and practice building eval sets with clear pass criteria. Testing software with AI tools needs classic QA automation plus fluency with AI coding assistants. Annotation platforms also call their work testing, but that's contract task work, not a QA career.
What is an AI task reviewer job?+
An AI task reviewer rates or corrects answers that an AI model produced, usually on a contract platform. You follow a rubric, score responses and sometimes rewrite them, and the data trains or evaluates the model. Pay is per hour or per task, and access to projects can stop without notice.
Does AI task work pay real money?+
AI task work does pay, but the hourly rate is only part of the picture. DataAnnotation advertises $25 to $50+ an hour for general tasks and more for coding or specialist knowledge. The catch is unpaid assessments, projects that dry up, and accounts that can be closed, so treat it as income, not a career step.
Which AI course is best for QA engineers?+
No single course closes the gap, because the hard skill is judgment about what good output looks like. Pick a real feature, write fifty test inputs with expected qualities, grade them, and then automate the grading with a model. That small eval project shows employers more than a certificate does.

New posts, straight to your inbox.