---
title: "AI QA Engineer: Two Jobs Behind One Title"
description: "AI QA engineer can mean testing AI systems with evals or testing software with AI tools. How to tell them apart, what each pays, and how to spot a gig."
url: "https://www.marian.coach/ai-qa-engineer-role/"
lang: "en"
image: "https://www.marian.coach/_astro/talk-11.BJeQ21aF_Z20DEx6.jpg"
---

**AI QA engineer** is one title for two different jobs: testing AI systems, and testing ordinary software with AI tools. The first checks whether a model, a prompt or an agent gives answers good enough to ship; in the second, AI tools do part of the testing work.

Both get posted as “AI QA engineer”. They need different skills, sit in different teams and, judging by live postings, pay very differently. And a third thing hides under nearby search terms like “AI tester” or “AI test engineer”, which isn’t a QA job at all.

## Which of the two jobs is in the ad

The title rarely tells you. The verbs in the responsibilities do.

**Testing AI** postings talk about evals, golden datasets, rubrics, LLM judges, red-teaming and regressions between model versions. [Glean’s evals engineer posting](https://job-boards.greenhouse.io/gleanwork/jobs/4712438005) describes systems that “decide whether new models, prompts, retrieval strategies, and agent workflows are ready to ship”. These roles are usually staffed as ML or infrastructure engineers, and the pay follows: Anthropic’s [post-training evaluations role](https://job-boards.greenhouse.io/anthropic/jobs/5198255008) lists $500,000 to $850,000, at the far end of the market.

**Testing with AI** postings talk about generating test cases from requirements, self-healing scripts, AI-assisted failure triage and coding assistants. [DevRev’s quality engineer posting](https://job-boards.greenhouse.io/devrev/jobs/6114345004) wants to “reduce reliance on manual test case creation”, and lists experience validating AI systems only under nice-to-have. Perplexity’s [senior QA role](https://jobs.ashbyhq.com/perplexity/4b5e02a8-6a94-494d-9d49-037f2cfc144d), $140,000 to $160,000, asks testers to use tools like Cursor or Claude Code to fix small bugs themselves.

![Marian Kamenistak mid-talk on a dark conference stage, lines of code projected on the screen behind him.](https://www.marian.coach/_astro/talk-37.CpcLkiqr_Z148QtM.webp)

Same title on the ad, different code on the screen: an eval pipeline or a test suite with an AI assistant attached.

Neither family is fake. The trouble starts when an ad borrows the vocabulary of the first and pays for the second, or when you prepare for one and interview for the other.

## Deterministic tests vs evals

If you come from classic QA, the second family will feel familiar. The first one breaks a few habits you probably rely on:

|  | Deterministic test | Eval |
| --- | --- | --- |
| Pass criteria | Output equals the expected value | Output scores above a threshold on qualities you defined (correct, grounded, safe, on tone) |
| Who knows the right answer | The spec, or the developer | Often nobody up front; a domain expert labels a sample first |
| Same input run twice | Same result | Can differ, so you sample and look at rates |
| A regression is | A test that went red | A score that dropped on a fixed dataset after a model, prompt or retrieval change |
| Grading at scale | Assertions in code | Code checks where possible, a model as judge where not, humans on a sample |
| Typical tooling | Playwright, pytest, JUnit, CI | Eval frameworks, labelled datasets, tracing, judge prompts |

The row that causes the most pain is the second one. A model acting as judge is only useful if it agrees with a person who knows the domain, so the judge needs its own test. Hamel Husain’s [guide to LLM judges](https://hamel.dev/blog/posts/llm-judge/) recommends labelling a set of real outputs with a domain expert, measuring how often the judge agrees on both passes and failures separately, and preferring a binary pass/fail over a 1 to 5 scale, because the gap between a 3 and a 4 is hard to define. Anthropic’s own [evaluation docs](https://docs.claude.com/en/docs/test-and-evaluate/develop-tests) push in the same direction from the other side: more questions with automated grading usually beat a few hand-graded ones, and the judging model should be a different one than the model being tested. In practice, a lot of the early work in a real AI QA job goes into writing labels and arguing about what “correct” means, before any test exists. A [Firecrawl evals posting](https://jobs.ashbyhq.com/firecrawl/25092c0e-9a32-4191-af79-050738213704) says the metric itself “has to be invented”.

In Europe there’s also a regulatory reason for this work. [Article 15 of the EU AI Act](https://artificialintelligenceact.eu/article/15/) requires high-risk AI systems to reach an appropriate level of accuracy and robustness and to declare their accuracy metrics. Somebody has to produce those numbers and defend them.

## AI tester jobs: a real AI QA role or an annotation gig?

Search “AI tester” and Google mostly shows AI-text detectors. Search “AI tester jobs” and the results mix QA postings with something else: Google’s own suggestions next to that search include “Outlier AI”, “AI training jobs”, “What is an AI task reviewer job?” and “Does AI task pay real money?”.

That something else is data annotation. Platforms such as DataAnnotation and Outlier pay contractors to rate, correct and write model answers. It’s real work and some of it pays decently: [DataAnnotation’s own FAQ](https://www.dataannotation.tech/faqs) advertises general tasks starting at $25 to $50+ an hour and coding tasks from $40 to $150+. But the conditions are those of gig work. [TIME reported](https://time.com/6962608/data-annotation-legit-tech-jobs-ai/) workers who passed the assessment and were then offered no tasks, and one account deactivated with $2,869 of work unpaid. A [Business Insider report](https://www.aol.com/articles/pay-cuts-poaching-pivoting-inside-091101909.html) on Scale AI’s contractors quoted a tasker who spent close to 40 hours in one month in unpaid onboarding without landing any work.

![Marian Kamenistak with hand on forehead, looking at his laptop, contemplative.](https://www.marian.coach/_astro/portrait-thinking.BI-2Z8YX_lAoLb.webp)

Read the ad twice. The pay unit and the employer usually tell you more than the title.

How to tell the two apart before you spend an evening on an assessment:

| Signal | AI QA or evals role | Annotation gig |
| --- | --- | --- |
| Employer | A named company, on its own careers page or ATS | A platform; the client is often anonymous |
| Pay unit | Annual salary, often with equity | Per hour or per task, often paid weekly |
| Hiring step | Interviews with the team you’d join | A screening assessment, frequently unpaid |
| What you produce | Eval sets, grading scripts, reports for release decisions | Ratings and rewritten answers to someone else’s rubric |
| Continuity | Employment contract and notice period | Projects open and close; access can be removed |
| Who you report to | An engineering or AI lead | A queue |

![Hand-drawn 2x2: Testing with AI and Testing AI on the horizontal axis, annual salary and per-task pay on the vertical one; both salaried AI QA jobs fill the top half, the annotation gig sits below in per-task pay.](https://www.marian.coach/images/blog/ai-qa-engineer-role/infographic.webp)

Annotation isn’t worthless for a career. Writing rubrics and grading model output is the raw skill behind evals, and Harvey, the legal AI company, hires people to [run human evaluation quality](https://jobs.ashbyhq.com/harvey/f0e4e24a-2a7a-4fc4-8872-9a8a762fecf5) at a salaried level. The step from one to the other is building something, an eval set or a grading script, and showing it.

Real AI QA roles carry a quieter risk. Of the ten live AI QA and evals postings reviewed for this page, none gave the role explicit authority to block a release. You may own “the quality bar” and still only inform the decision. Ask in the interview who decides when an eval fails a week before launch, and what happened the last time.

If you’re weighing a move from classic QA into evals, or two offers that both say “AI” and mean different things, that’s a decision worth talking through. It’s the kind of career question I work on in [1:1 career mentoring](https://www.marian.coach/engineering-career-coach/).

## Other AI roles on the map

-   [QA engineer: the classic role, its ladder and its risks](https://www.marian.coach/qa-engineer-role/)
-   [AI platform engineer: the layer evals run on](https://www.marian.coach/ai-platform-engineer-role/)
-   [AI program manager: shipping AI with moving targets](https://www.marian.coach/ai-program-manager-role/)
-   [Director of AI: three versions of the seat](https://www.marian.coach/director-of-ai-role/)
-   [Forward deployed engineer, the customer-side engineering role](https://www.marian.coach/forward-deployed-engineer-role/)

## Frequently asked

How can I be an AI tester?+

To become an AI tester, first decide which of three different jobs you mean. Testing AI systems needs Python, statistics and practice building eval sets with clear pass criteria. Testing software with AI tools needs classic QA automation plus fluency with AI coding assistants. Annotation platforms also call their work testing, but that's contract task work, not a QA career.

What is an AI task reviewer job?+

An AI task reviewer rates or corrects answers that an AI model produced, usually on a contract platform. You follow a rubric, score responses and sometimes rewrite them, and the data trains or evaluates the model. Pay is per hour or per task, and access to projects can stop without notice.

Does AI task work pay real money?+

AI task work does pay, but the hourly rate is only part of the picture. DataAnnotation advertises $25 to $50+ an hour for general tasks and more for coding or specialist knowledge. The catch is unpaid assessments, projects that dry up, and accounts that can be closed, so treat it as income, not a career step.

Which AI course is best for QA engineers?+

No single course closes the gap, because the hard skill is judgment about what good output looks like. Pick a real feature, write fifty test inputs with expected qualities, grade them, and then automate the grading with a model. That small eval project shows employers more than a certificate does.

[1:1 mentoring](https://www.marian.coach/engineering-manager-mentor/)

## Ready to move this from advice to action?

1:1 mentoring for engineering leaders. 317 mentees since 2019. Rate a session under 7/10 and you don't pay for it. First intro is free.

[Book a free 30-min intro →](https://www.marian.coach/meet)

## Read next

-   [![Marian Kamenistak speaking at an engineering leadership event.](https://www.marian.coach/_astro/talk-21.BdOplLV4_mJDuE.webp)
    
    ### AI Platform Engineer: What You Build and What You Own
    
    AI platform engineer: how the role differs from ML platform and MLOps, the five layers you run, who owns what, and a plan for your first 90 days.
    
    9 October 2026 · By Marian
    
    Read →
    
    
    
    ](https://www.marian.coach/ai-platform-engineer-role/)
-   [![Marian Kamenistak smiling at his laptop in an outdoor courtyard.](https://www.marian.coach/_astro/portrait-laptop-smile.D82UYdFF_Z2gghCm.webp)
    
    ### AI Program Manager: Four Targets That Keep Moving
    
    An AI program manager gets AI from pilot to production while the model, data, evals and EU AI Act keep moving. A risk register and what to negotiate.
    
    9 October 2026 · By Marian
    
    Read →
    
    
    
    ](https://www.marian.coach/ai-program-manager-role/)
-   [![Marian Kamenistak smiling at the camera in a brown blazer, large monstera leaves behind him.](https://www.marian.coach/_astro/portrait-plants-hero.5FpoUpQC_1LGXlk.webp)
    
    ### Applied AI Engineer: The Role, and the Ladder Above It
    
    What an applied AI engineer owns, how the job differs from AI engineer, staff and principal AI engineer, and what a promotion case above it must prove.
    
    9 October 2026 · By Marian
    
    Read →
    
    
    
    ](https://www.marian.coach/applied-ai-engineer-role/)

### New posts, straight to your inbox.

Your emailSubscribe

You're on it.

Nothing hits your inbox today. First send goes out once the list is live, and you're already on it.

[Connect on LinkedIn →](https://www.linkedin.com/in/mariankamenistak/)[See the free tools →](https://www.marian.coach/ai-coaching-tools/)

## Structured data

```json
{"@context":"https://schema.org","@type":"ProfessionalService","name":"Marian Kamenistak | Engineering Leadership Mentoring","url":"https://www.marian.coach/ai-qa-engineer-role/","image":"https://www.marian.coach/_astro/talk-11.BJeQ21aF_Z20DEx6.jpg","description":"AI QA engineer can mean testing AI systems with evals or testing software with AI tools. How to tell them apart, what each pays, and how to spot a gig.","serviceType":["Engineering leadership mentoring","Engineering manager coaching","CTO coaching","Fractional CTO","Leadership development for software engineers"],"hasOfferCatalog":{"@type":"OfferCatalog","name":"Engineering leadership mentoring and advisory","itemListElement":[{"@type":"Offer","itemOffered":{"@type":"Service","name":"1:1 engineering leadership mentoring","description":"One-to-one mentoring for engineering leaders from Staff Engineer to CTO. 60-minute sessions, weekly or bi-weekly.","url":"https://www.marian.coach/pricing/"}},{"@type":"Offer","itemOffered":{"@type":"Service","name":"Engineering manager coaching","description":"Coaching for first-time and experienced engineering managers: delegation, performance, hiring, team scaling.","url":"https://www.marian.coach/engineering-manager-mentor/"}},{"@type":"Offer","itemOffered":{"@type":"Service","name":"CTO and VP Engineering mentoring","description":"Mentoring for CTOs, VPs of Engineering and Heads of Engineering running 30-200 person organisations.","url":"https://www.marian.coach/cto-mentor/"}},{"@type":"Offer","itemOffered":{"@type":"Service","name":"Mentor in Residence","description":"A dedicated onsite mentor for your engineering organisation, booked by the quarter.","url":"https://www.marian.coach/mentor-in-residence/"}},{"@type":"Offer","itemOffered":{"@type":"Service","name":"Fractional CTO","description":"Part-time CTO engagements for companies without a full-time technology leader.","url":"https://www.marian.coach/fractional-cto/"}}]},"telephone":"+420 736 519 879","email":"marian@marian.coach","sameAs":["https://www.linkedin.com/in/mariankamenistak/","https://www.youtube.com/@engineeringleaders","https://github.com/marian-kamenistak","https://www.engineeringleaders.io/","https://www.elc-conference.io/","https://www.marianslist.com/","https://www.mentoringhub.io/"],"address":{"@type":"PostalAddress","streetAddress":"Varšavská 715/36, Vinohrady","addressLocality":"Praha 2","postalCode":"12000","addressCountry":"CZ"},"geo":{"@type":"GeoCoordinates","latitude":50.0734418,"longitude":14.4385757},"openingHoursSpecification":[{"@type":"OpeningHoursSpecification","dayOfWeek":["Monday","Tuesday","Wednesday","Thursday","Friday"],"opens":"09:00","closes":"17:00"}],"priceRange":"296–395 EUR","areaServed":[{"@type":"Place","name":"Europe"},{"@type":"Country","name":"Czechia"},{"@type":"Country","name":"Slovakia"},{"@type":"Country","name":"Poland"},{"@type":"Place","name":"North America"}]}
{"@context":"https://schema.org","@type":"Person","@id":"https://www.marian.coach/#person","name":"Marian Kamenistak","url":"https://www.marian.coach/","image":"https://www.marian.coach/_astro/talk-11.BJeQ21aF_Z20DEx6.jpg","jobTitle":["Engineering Leadership Mentor","Fractional CTO & Chief AI Officer","Founder, Engineering Leaders Community","Engineering Manager Coach","CTO Coach","Keynote Speaker"],"hasOccupation":[{"@type":"Occupation","name":"Engineering Leadership Mentor"},{"@type":"Occupation","name":"Fractional CTO & Chief AI Officer"}],"worksFor":{"@type":"Organization","name":"Engineering Leaders Community","url":"https://www.engineeringleaders.io/"},"alumniOf":[{"@type":"Organization","name":"Mews"},{"@type":"Organization","name":"Databricks"}],"knowsLanguage":["English","Czech","Slovak"],"knowsAbout":["Engineering leadership","Engineering management","Scaling engineering teams","CTO mentoring","Engineering org design","VP of Engineering role","Engineering director role","CTPO role","Team lead development","Staff engineer career path","Fractional CTO engagements","AI-assisted engineering leadership","Conference keynotes on engineering leadership","Panel moderation","Podcast guest — engineering leadership"],"homeLocation":{"@type":"Place","name":"Prague, Czechia"},"sameAs":["https://www.linkedin.com/in/mariankamenistak/","https://www.youtube.com/@engineeringleaders","https://github.com/marian-kamenistak","https://www.engineeringleaders.io/","https://www.elc-conference.io/","https://www.marianslist.com/","https://www.mentoringhub.io/","https://www.crunchbase.com/person/marian-kamenistak","https://www.wikidata.org/wiki/Q140488098","https://developers.mews.com/author/marian/","https://startupdisrupt.com/speakers/marian-kamenistak/","https://theorg.com/org/engineering-leaders-community/org-chart/marian-kamenistak"]}
{"@context":"https://schema.org","@type":"WebPage","url":"https://www.marian.coach/ai-qa-engineer-role/","name":"AI QA Engineer: Two Jobs Behind One Title","inLanguage":"en","dateModified":"2026-10-09T00:00:00.000Z"}
{"@context":"https://schema.org","@type":"FAQPage","inLanguage":"en","mainEntity":[{"@type":"Question","name":"How can I be an AI tester?","acceptedAnswer":{"@type":"Answer","text":"To become an AI tester, first decide which of three different jobs you mean. Testing AI systems needs Python, statistics and practice building eval sets with clear pass criteria. Testing software with AI tools needs classic QA automation plus fluency with AI coding assistants. Annotation platforms also call their work testing, but that's contract task work, not a QA career."}},{"@type":"Question","name":"What is an AI task reviewer job?","acceptedAnswer":{"@type":"Answer","text":"An AI task reviewer rates or corrects answers that an AI model produced, usually on a contract platform. You follow a rubric, score responses and sometimes rewrite them, and the data trains or evaluates the model. Pay is per hour or per task, and access to projects can stop without notice."}},{"@type":"Question","name":"Does AI task work pay real money?","acceptedAnswer":{"@type":"Answer","text":"AI task work does pay, but the hourly rate is only part of the picture. DataAnnotation advertises $25 to $50+ an hour for general tasks and more for coding or specialist knowledge. The catch is unpaid assessments, projects that dry up, and accounts that can be closed, so treat it as income, not a career step."}},{"@type":"Question","name":"Which AI course is best for QA engineers?","acceptedAnswer":{"@type":"Answer","text":"No single course closes the gap, because the hard skill is judgment about what good output looks like. Pick a real feature, write fifty test inputs with expected qualities, grade them, and then automate the grading with a model. That small eval project shows employers more than a certificate does."}}]}
{"@context":"https://schema.org","@type":"Article","headline":"AI QA Engineer: Two Jobs Behind One Title","description":"AI QA engineer can mean testing AI systems with evals or testing software with AI tools. How to tell them apart, what each pays, and how to spot a gig.","inLanguage":"en","articleSection":"career","datePublished":"2026-10-09T00:00:00.000Z","dateModified":"2026-10-09T00:00:00.000Z","author":{"@id":"https://www.marian.coach/#person"},"publisher":{"@type":"Organization","name":"Marian Kamenistak","url":"https://www.marian.coach/","logo":{"@type":"ImageObject","url":"https://www.marian.coach/favicon-192x192.webp","width":192,"height":192}},"image":{"@type":"ImageObject","url":"https://www.marian.coach/_astro/talk-11.BJeQ21aF_Z20DEx6.jpg","width":1200,"height":630},"mainEntityOfPage":{"@type":"WebPage","@id":"https://www.marian.coach/ai-qa-engineer-role/"},"keywords":"career, ai, quality","wordCount":1254,"timeRequired":"PT6M","isAccessibleForFree":true}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://www.marian.coach/"},{"@type":"ListItem","position":2,"name":"Resources","item":"https://www.marian.coach/resources/"},{"@type":"ListItem","position":3,"name":"AI QA Engineer: Two Jobs Behind One Title","item":"https://www.marian.coach/ai-qa-engineer-role/"}]}
```
