An AI platform engineer is the engineer who builds the shared layer product teams use to put large language models into production. That layer is the gateway model calls pass through, the evals that decide whether a change ships, the traces, and the limits on cost and data.
The title is new enough that postings disagree on what it means. Eleven of them, read in October 2026 across the US and Europe, fall into two families. One builds a control plane for models the company buys: access, routing, quotas, cost, guardrails. The other runs compute for models the company hosts: GPUs, scheduling, inference servers. Same title, and the pay, the on-call and the career path differ between them.
AI platform engineer vs ML platform engineer vs MLOps: what’s the difference?
An AI platform engineer extends classic platform work to model access, GPU compute, cost governance, guardrails and compliance. MLOps automates the lifecycle of models a company trains itself, from training to deployment to retraining, and ML platform engineering builds the infrastructure under that.
Classic platform engineering paves the road for shipping any code: pipelines, runtime, observability. TrueFoundry’s guide to AI platform engineering describes the AI version as the same mandate extended to model access, GPU compute, cost governance, guardrails and compliance.
MLOps is older and narrower. Google Cloud defines it as automating the lifecycle of models you build yourself, from training to deployment to retraining. ML platform engineering builds the infrastructure under that. Mistral’s ML platform posting reads like a Kubernetes scheduler job for researchers, with queueing, quotas, preemption and an explicit on-call rotation across GPU infrastructure.
The AI platform engineer mostly serves models the company didn’t train. Scalable Capital in Berlin has the cleanest version: model access through LiteLLM, tracing in Langfuse, rate limits and strict API cost management, for a team of internal AI software engineers.
Even the people defining the field split it two ways. platformengineering.org uses “AI platform engineering” on its topic hub for developer platforms that use AI, and its own guide separates those from platforms built to run AI workloads. When an ad says “AI platform”, check which of the two it means before the first call.
Pay follows the family more than the title. Code Metal, which runs its own inference servers and gateway, offers a staff base of $205,000 to $250,000. Sysdig’s staff role, a control plane for business teams on AWS Bedrock, runs $141,000 to $194,700, so its top end sits below Code Metal’s floor. In Prague, DEK offers 90,000 to 120,000 CZK a month.
Who owns which decision
Most of the friction in this role comes from three titles touching the same feature. The applied AI engineer column follows Latent Space’s essay on the AI engineer, which puts product data and evals on the AI engineer and foundation-model work on ML engineers.
| Decision or asset | ML engineer | AI platform engineer | Applied AI engineer |
|---|---|---|---|
| Which model a feature uses | Trains or fine-tunes candidates | Makes approved models available, sets defaults | Picks from the approved list |
| Prompts, retrieval, tool calls | Not involved | Provides the building blocks: vector store, tool registry | Owns them |
| Evals | Model-level benchmarks | Runs the eval suite and the release gate in CI | Writes the test cases for the feature |
| Gateway, keys, quotas | Uses it | Owns it | Uses it |
| GPU and inference capacity | Asks for it | Schedules and serves it, when models are self-hosted | Rarely sees it |
| Token spend | Training budget | Attributes it to teams, alerts on spikes | Answers for the feature’s bill |
| PII and guardrails | Training data hygiene | Enforces filters at the gateway | Designs the feature to need less data |
| Provider outage at 2am | Not paged | Paged | Paged if the feature has no fallback |
The table looks obvious on a page. Inside a company it usually exists only in people’s heads, which is why the 90-day plan below starts with writing it down.
The five layers you end up running

- Compute and serving. When models are self-hosted, this is GPUs, inference servers and the question of who gets priority overnight. Trezor’s Prague posting describes one multi-GPU server running vLLM, LiteLLM and Open WebUI, owned end to end by one person. When models come from providers, the layer shrinks to contracts, regions and rate limits.
- The gateway. One door for every model call. Code Metal’s posting lists what it has to do: “consistent auth, routing, failover, quotas, and cost attribution.” When I test a candidate for a Head of AI seat, one of my questions is what they run as an AI gateway (the full list is in my AI transformation playbook).
- Evals. An eval runner, golden datasets per feature, and a gate that blocks a release when quality drops. TrueFoundry’s guide, for all its detail on gateways and cost, doesn’t list evaluation as a domain. My rule on it is in the same playbook: “AI is a probability model, not binary. If there are no validations, I can guarantee you it does not save you time.”
- Observability. Traces per request, latency, token use per team, answer quality over time. Sysdig pipes Bedrock usage into Snowflake, matched to identity, with threshold alerts on spend.
- Cost and guardrails. Budgets per team, PII filters, and sandboxes for unproven models (Sysdig’s wording is “keep unproven models sandboxed”). In Europe this layer also produces audit evidence. Under Article 12 of the AI Act, a high-risk system has to record events automatically over its lifetime, and Article 26 has deployers keep those logs for at least six months. The Digital Omnibus regulation moved the high-risk obligations for Annex III systems, such as hiring and credit scoring, to 2 December 2027 (timeline). A gateway with identity-matched logs is where that evidence will come from.

Your first 90 days
The risks in this seat show up in the postings before you sign: scope that’s bigger or smaller than the title, accountability for things other teams control, and a new team whose main users are pilots. A plan for the first three months should defuse them in that order.
Days 1 to 30: find every model call. List the providers, the API keys, who pays each invoice, and which features are live versus pilots. Assume more usage than anyone admits. MIT NANDA’s State of AI in Business 2025 found workers at over 90% of surveyed companies using personal AI tools, while only 40% of companies had bought an official LLM subscription. This month also tells you which job you took. DEK’s “AI platform” posting has 13 duties, and 11 of them are taking over vendor apps, CI/CD and Docker, with AI experience optional. If your week looks like that, renegotiate the scope or the title while you’re still new.

Days 31 to 60: put the gateway in front, and get the ownership table signed. Route the most-used features through one gateway first, then fill in the table above with the engineering leaders whose teams build those features. Sysdig’s posting says why openly. The gateway is “co-owned”, and the job is to get other teams building against a written standard “without owning their roadmaps”. Without a signed table, a bad answer in someone else’s feature becomes your incident.
Days 61 to 90: show the money and one save. Bring a cost report per team per month, and one release your eval gate stopped. Gartner expects over 40% of agentic AI projects to be cancelled by the end of 2027, and names escalating costs, unclear business value and weak risk controls as the reasons. Two of those three are line items on your platform. None of the postings mentions budget risk, but a platform built to serve pilots is exposed when the pilots stop, unless it can show what it saved.
If you’re the first AI platform hire, or the engineering manager that team reports to, this is the kind of plan I work through in AI engineering manager mentoring. The first 30 minutes are on me.



