AI Platform Engineer: What You Build and What You Own

· by Marian Kamenistak

Marian Kamenistak speaking at an engineering leadership event.

An AI platform engineer is the engineer who builds the shared layer product teams use to put large language models into production. That layer is the gateway model calls pass through, the evals that decide whether a change ships, the traces, and the limits on cost and data.

The title is new enough that postings disagree on what it means. Eleven of them, read in October 2026 across the US and Europe, fall into two families. One builds a control plane for models the company buys: access, routing, quotas, cost, guardrails. The other runs compute for models the company hosts: GPUs, scheduling, inference servers. Same title, and the pay, the on-call and the career path differ between them.

AI platform engineer vs ML platform engineer vs MLOps: what’s the difference?

An AI platform engineer extends classic platform work to model access, GPU compute, cost governance, guardrails and compliance. MLOps automates the lifecycle of models a company trains itself, from training to deployment to retraining, and ML platform engineering builds the infrastructure under that.

Classic platform engineering paves the road for shipping any code: pipelines, runtime, observability. TrueFoundry’s guide to AI platform engineering describes the AI version as the same mandate extended to model access, GPU compute, cost governance, guardrails and compliance.

MLOps is older and narrower. Google Cloud defines it as automating the lifecycle of models you build yourself, from training to deployment to retraining. ML platform engineering builds the infrastructure under that. Mistral’s ML platform posting reads like a Kubernetes scheduler job for researchers, with queueing, quotas, preemption and an explicit on-call rotation across GPU infrastructure.

The AI platform engineer mostly serves models the company didn’t train. Scalable Capital in Berlin has the cleanest version: model access through LiteLLM, tracing in Langfuse, rate limits and strict API cost management, for a team of internal AI software engineers.

Even the people defining the field split it two ways. platformengineering.org uses “AI platform engineering” on its topic hub for developer platforms that use AI, and its own guide separates those from platforms built to run AI workloads. When an ad says “AI platform”, check which of the two it means before the first call.

Pay follows the family more than the title. Code Metal, which runs its own inference servers and gateway, offers a staff base of $205,000 to $250,000. Sysdig’s staff role, a control plane for business teams on AWS Bedrock, runs $141,000 to $194,700, so its top end sits below Code Metal’s floor. In Prague, DEK offers 90,000 to 120,000 CZK a month.

Who owns which decision

Most of the friction in this role comes from three titles touching the same feature. The applied AI engineer column follows Latent Space’s essay on the AI engineer, which puts product data and evals on the AI engineer and foundation-model work on ML engineers.

Decision or assetML engineerAI platform engineerApplied AI engineer
Which model a feature usesTrains or fine-tunes candidatesMakes approved models available, sets defaultsPicks from the approved list
Prompts, retrieval, tool callsNot involvedProvides the building blocks: vector store, tool registryOwns them
EvalsModel-level benchmarksRuns the eval suite and the release gate in CIWrites the test cases for the feature
Gateway, keys, quotasUses itOwns itUses it
GPU and inference capacityAsks for itSchedules and serves it, when models are self-hostedRarely sees it
Token spendTraining budgetAttributes it to teams, alerts on spikesAnswers for the feature’s bill
PII and guardrailsTraining data hygieneEnforces filters at the gatewayDesigns the feature to need less data
Provider outage at 2amNot pagedPagedPaged if the feature has no fallback

The table looks obvious on a page. Inside a company it usually exists only in people’s heads, which is why the 90-day plan below starts with writing it down.

The five layers you end up running

Marian Kamenistak pointing at a slide scoring AI-assisted engineering skills — agentic coding, AI code review, MCP integration — against a mid-level developer salary.
MCP integration and AI code review now show up as skills on engineering salary slides, and as layers on the platform.
  1. Compute and serving. When models are self-hosted, this is GPUs, inference servers and the question of who gets priority overnight. Trezor’s Prague posting describes one multi-GPU server running vLLM, LiteLLM and Open WebUI, owned end to end by one person. When models come from providers, the layer shrinks to contracts, regions and rate limits.
  2. The gateway. One door for every model call. Code Metal’s posting lists what it has to do: “consistent auth, routing, failover, quotas, and cost attribution.” When I test a candidate for a Head of AI seat, one of my questions is what they run as an AI gateway (the full list is in my AI transformation playbook).
  3. Evals. An eval runner, golden datasets per feature, and a gate that blocks a release when quality drops. TrueFoundry’s guide, for all its detail on gateways and cost, doesn’t list evaluation as a domain. My rule on it is in the same playbook: “AI is a probability model, not binary. If there are no validations, I can guarantee you it does not save you time.”
  4. Observability. Traces per request, latency, token use per team, answer quality over time. Sysdig pipes Bedrock usage into Snowflake, matched to identity, with threshold alerts on spend.
  5. Cost and guardrails. Budgets per team, PII filters, and sandboxes for unproven models (Sysdig’s wording is “keep unproven models sandboxed”). In Europe this layer also produces audit evidence. Under Article 12 of the AI Act, a high-risk system has to record events automatically over its lifetime, and Article 26 has deployers keep those logs for at least six months. The Digital Omnibus regulation moved the high-risk obligations for Annex III systems, such as hiring and credit scoring, to 2 December 2027 (timeline). A gateway with identity-matched logs is where that evidence will come from.

Hand-drawn diagram: five pillars, Compute and serving, The gateway, Evals, Observability, and Cost and guardrails, holding up one beam labelled AI platform.

Your first 90 days

The risks in this seat show up in the postings before you sign: scope that’s bigger or smaller than the title, accountability for things other teams control, and a new team whose main users are pilots. A plan for the first three months should defuse them in that order.

Days 1 to 30: find every model call. List the providers, the API keys, who pays each invoice, and which features are live versus pilots. Assume more usage than anyone admits. MIT NANDA’s State of AI in Business 2025 found workers at over 90% of surveyed companies using personal AI tools, while only 40% of companies had bought an official LLM subscription. This month also tells you which job you took. DEK’s “AI platform” posting has 13 duties, and 11 of them are taking over vendor apps, CI/CD and Docker, with AI experience optional. If your week looks like that, renegotiate the scope or the title while you’re still new.

Marian Kamenistak leading a two-day engineering leadership workshop for a small group of managers, presenting at a screen.
The first month is mostly listening: who uses which model, and who believes they own it.

Days 31 to 60: put the gateway in front, and get the ownership table signed. Route the most-used features through one gateway first, then fill in the table above with the engineering leaders whose teams build those features. Sysdig’s posting says why openly. The gateway is “co-owned”, and the job is to get other teams building against a written standard “without owning their roadmaps”. Without a signed table, a bad answer in someone else’s feature becomes your incident.

Days 61 to 90: show the money and one save. Bring a cost report per team per month, and one release your eval gate stopped. Gartner expects over 40% of agentic AI projects to be cancelled by the end of 2027, and names escalating costs, unclear business value and weak risk controls as the reasons. Two of those three are line items on your platform. None of the postings mentions budget risk, but a platform built to serve pilots is exposed when the pilots stop, unless it can show what it saved.

If you’re the first AI platform hire, or the engineering manager that team reports to, this is the kind of plan I work through in AI engineering manager mentoring. The first 30 minutes are on me.

Roles around the AI platform

Frequently asked

What is an ML platform engineer?+
An ML platform engineer builds the infrastructure a company uses to train, evaluate and serve its own models. That means GPU scheduling, training pipelines, feature stores and model registries. The AI platform engineer is the newer cousin. Most models now come from providers through an API, so the work shifts to gateways, evals, cost limits and guardrails.
How do AI platforms work?+
An internal AI platform puts one shared layer between product teams and the models they call. Requests pass a gateway that checks access, routes to a provider or a self-hosted model, applies quotas and logs the call. Evals test changes before release, traces show what happened in production, and cost reports split the bill by team.
What is the pay for an AI platform engineer?+
Published US staff-level base bands in October 2026 ran from about $141,000 to $250,000. The company and whether you run GPUs move it most. Salary sites have little data yet because the title is new. Central Europe pays far less; one Prague posting offered 90,000 to 120,000 CZK a month.
Is AI platform engineering a stressful job?+
It can be, mostly because of accountability without authority. You often answer for AI spend, data leaving the company and answer quality on features other teams build. The pressure drops once ownership is written down: who picks the model, who owns the prompt, who pays the bill, and who gets paged when a provider goes down.

New posts, straight to your inbox.