AI Transformation Playbook: The Money Isn't in R&D

· by Marian Kamenistak

AI Transformation Playbook: The Money Isn't in R&D

In August 2026 I found out that every failure alert from my own scheduled AI jobs had been going nowhere. The script I built to tell me when a job broke was breaking itself, silently.

That’s the Head of AI Transformation job in small: moving a whole company up the AI adoption ladder, not only yourself, and proving in finance’s numbers that it paid off.

Here is what happened. I run part of my business on scheduled jobs and AI agents on my Mac. When one fails, a small helper posts an alert to Slack. The Python on that machine had never had its certificates installed, so every secure web call from it died. The helper used exactly that Python, and it swallowed its own errors.

So a job that ran and failed told me nothing. A job that never started still got caught, because that watchdog runs in the cloud. I fixed it on 21 August, and the rule I wrote down that day is in this whole post: check before you run, and fail loudly. Never fall back quietly.

I run my own scheduled agents, which is L3 on the ladder below, and I still had automation that looked like it worked and wasn’t validated. Now scale that to 80 developers and 130 people in the business. That’s the seat. Below is how I check who can sit in it, the five levels with two KPIs each, why the money isn’t in R&D, and a six-month plan worked through on a made-up company.

Using AI yourself is a way different discipline

Many people claim the background for this seat. Usually that background is them using AI on their own.

Expanding it to your team, your departments, and let’s say 100 people in the company, is a different discipline. You can tell quickly how far someone got, if you ask them to show instead of tell. This individual ladder is how I check.

LevelWhat the person doesAsk them to show you
L0Uses AI as a chatbotNothing to show
L0.5Asks AI well toward a goal, has connectors and a few pluginsThe tools connected to their AI today
L1Builds AI skills, so they stop copy-pasting the same contentTheir skills, and which ones others use
L2Assigns skills to professional roles, which turns them into agentsOne agent and the role it plays
L3Runs agents on a scheduleLast week’s run log
L4Has automated processes that are audited, logged, integrated, and save real timeThe hours saved, with evidence
L5Runs complex, multi-step workflows that are scheduled, tested and validatedThe validation step that stopped a run, and why

Not many people get to L4, because L4 means proving it saves you time. I see plenty of people bragging about how they use AI, and it doesn’t save them any time. They sit in front of the screen all day answering the questions their agents ask.

People with a technical background usually climb faster. They use the right terminology, closer to the technical definition, so AI understands better what they want and which edge cases matter.

How I test a Head of AI

Job ads for this seat ask for the same six things. I read 13 recent ones for Head of AI Transformation, AI Transformation Lead, AI Adoption Lead and Head of AI.

The ad asks forHow one posting puts it
Kill and rank use cases by valueChip: rank by capacity released, cost avoided, revenue enabled
Measure outcomes, not activityDeepgram: “measurable productivity and quality, not activity”
Build with your own handsN-iX: “approximately half of this role is hands-on engineering”
Drive adoption nobody asked forDeepgram: “environments that didn’t start out asking for it”
Put guardrails into the platformDeepgram: “embedded into platforms rather than enforced through gates”
Redesign rolesRelay: “including which roles don’t need to exist”

An ad can’t check any of that. These are the questions I use.

Start with you, then your teams

Tell me how you use AI on your own. Then tell me how you made people around you use it, a team or teams, and how successful it was. You’ve got to prove it, on the technical level, the org chart level and the individual level.

The org chart

Have you changed roles from backend and frontend to full stack or to product builders? Have you flattened your structure? How? What are your dirty tips and tricks, such as code pairing?

The knowledge base

How did you build the guidelines and the knowledge base for engineers, marketing, sales, HR and product managers? That’s one of the most important things, by the way.

The tokens

How do you measure that you spend your tokens wisely? What do you run as an AI gateway? How fast can you take a simple app from prototype to production, with all non-functional requirements?

The validation

The best question I use: how do you build the full workflow, technically, and the validation for each and every step of the AI software development lifecycle?

AI is a probability model, not binary. If there are no validations, I can guarantee you it does not save you time.

My silent alerts are the cheap version of that lesson. So ask for one run their own validation stopped, and why.

The money

Prove to me it was really worth it, in money. And show me you didn’t stop at optimizing developers, that you moved on to the people closer to the customer.

Marian Kamenistak mid-talk on a dark conference stage, lines of code projected on the screen behind him.
On stage at an engineering leadership event.

Sit next to a developer

The other trick I use very often: I sit next to a developer to see how far they are with AI. When new functionality breaks, most people fix the specification. But if it’s the process that is broken, they should fix the full process, so the same failure doesn’t come back next time. That’s a big difference.

Keep the AI freaks

The best Heads of AI and AI engineering managers I see these days are AI freaks with their own startups or pet projects. That’s how they learned how AI works in a business, and I would not go against it at all. It’s up to the company to stay attractive enough, or help them develop their ideas, so they stay loyal through the transformation.

Sign the contract before you write the strategy

A Chief AI Officer needs a proper contract: what they will achieve in 3, 6 and 12 months, and what they need for it. That’s how you ask for a mandate, and it’s pretty tricky.

Sometimes I see CTOs or Chief AI Officers who are solo, with no direct reports. That’s a path down to hell. They write strategies with no people to execute them and end up talking about how nice it’s going to be one day. IBM’s study of 600+ Chief AI Officers found the average CAIO team has five people, and smaller teams are less successful.

The contract one-pager is the first file in the download. Fill it in before your first strategy slide.

The money isn’t in R&D

Usually I see companies start AI adoption on the technical level and stay there. I think that’s a huge mistake, because the money is usually not in tech or in R&D.

The money is in what comes after the software is written: getting functionality to market fast enough, marketing, sales, activation, adoption, customer support. Inside engineering, coding is about 16% of a developer’s day, and DX measured a median PR throughput gain of 7.76% across 400+ companies.

Outside it, the numbers are bigger. McKinsey’s 2026 survey found revenue gains from AI are most often attributed to marketing and sales.

MIT’s NANDA report puts the best ROI in back-office automation instead. That’s outside R&D too.

So when I advise companies, I build the first AI-native team in tech, and after three months I distribute its members into marketing, sales, HR, support, you name it. If you help the company sell its value fast, that hugely helps to prove the AI transformation is worth it.

What should you expect there? Published studies give honest ranges, not miracles: 14% higher productivity across 5,179 support agents, and 34% for novices, and in online retail anything from no detectable effect to 16.3% more sales, mostly through conversion. Months 4 and 5 of the playbook below walk through the move with illustrative numbers from the made-up company, not from a client.

A prototype that works is not a product

What business people and top managers don’t understand well is the huge gap between a prototype that looks nice and an app in production. Production needs security, high availability, backups, multi-tenancy, multi-region access, single sign-on and incident management. Without them the product won’t live for long.

Explaining that is your job. Talk cost of delay and risk, in their numbers, one requirement at a time.

Missing in the prototypeWhat breaksHow to say it to the CEO
Single sign-onEnterprise deals stall in security review”Two deals wait for SSO. Every week we don’t have it, their revenue waits too.”
BackupsOne bad deploy loses customer data”Three weeks of work now, or one night that costs us the customer.”
Incident managementNobody owns the 2 a.m. failure”Without an owner, our first outage becomes their first outage.”

That’s also why the month-3 hackathon in the plan counts only what reached production.

Not everyone will make it

This part is what I see, not what a survey proves. Not so many people understand AI to the extent you would want, which is natural for a new way of working. In some mid-sized and larger companies it’s very hard to compose even one team of five people who already have a taste for AI.

And there’s a secret nobody talks about: not every person can make it to be AI compatible. You’ll need to filter certain people out of the game. The data only shows how uneven it is: 38% of developers have no plans to adopt AI agents, per Stack Overflow’s 2025 survey.

How to handle it: at the month-3 scorecard, list who didn’t move a single level on the individual ladder. Pair each of them with a champion for a month. Then decide with their manager, openly, what happens next.

Marian Kamenistak leading a two-day engineering leadership workshop for a small group of managers, presenting at a screen.
A two-day workshop with a small group of managers.

When companies can’t build the AI team from inside, they bring in consultants. Be careful what you demand from them. You don’t want the type that will suck your blood and money. You want people who shadow your engineers and pair with them while coding, almost invisible, because they should help your people shine, not themselves.

Five AI maturity levels, two KPIs each

Six of the ten KPIs below live in engineering, and that’s on purpose. The first AI team starts in tech because that’s where the people who can build it sit. The money gets proven outside it, at L3, before anyone climbs further. The playbook further down has each level in detail.

LevelWhat it looks likeKPI 1KPI 2
L1Spending money on licences like a foolAI spend per weekly active userPilots with a ship or kill decision in 90 days
L2Connectors, a knowledge base and a skills marketplace by roleSkills reused by 2+ teamsAgent-mode share in the lowest team
L3Automation, the first AI team, tokens spent wiselyVerified hours saved per person per weekNet AI return per person per month
L4Token telemetry, AI-native and spec-driven deliveryAI cost per shipped changeSpec-to-production lead time, guarded by change failure rate
L5Self-running workflows, tested, validated, auditedUnattended run rateAgent-delivered work share

Measure only the two KPIs of your level. And stop reporting these four as success on their own:

Stop reportingWhy it lies
Seat utilisation90% of DORA’s 2025 respondents already use AI at work
Suggestion acceptance rateDX’s Laura Tacho calls it “the new lines of code measurement”
Share of AI-written codeDX measured 52.7% while the innovation ratio stayed flat
Self-reported speedMETR’s developers felt 20% faster and were 19% slower

The playbook, month by month

Each section below has the guidance on the left and a filled-in example on the right, for a made-up company called Ferrymark: 80 engineers, 130 people in the business, and a new Head of AI Transformation named Filip. You can follow his $11,200 monthly AI bill all the way to a support team answering customers in 48 minutes instead of four hours.

Want blank pages instead? Download the templates (5 KB zip): contract one-pager, KPI scorecard with formulas, pilot register, six-month checklist, month-6 memo.

Part 1

The company ladder. Five levels, two KPIs each.

Level 0 is a company that doesn't use AI at all. Few engineering companies are still there, so the ladder starts at 1. Measure only the two numbers of the level you're on. The KPIs of the level above are noise until you get there.

Level 1Company ladder · 1 of 5

Licences, bought like a fool

The company pays for AI and calls it adoption. Seats for everybody, a hackathon or two, demos that never shipped, and nobody can say what last month's bill bought.

What to do

  • Put every tool, seat, API key and pilot on one sheet, with cost, owner and the one number each should move.
  • Give each pilot 90 days. Then it ships or it dies. MIT's NANDA study found the top performers took about 90 days from pilot to full implementation, against nine months or more for large enterprises.

The two KPIs

1AI spend per weekly active user

monthly AI spend (seats + tokens) ÷ people who used AI on at least one day that week

The waste meter. It falls when you cut dead seats.

2Pilot decision rate

pilots older than 90 days with a written ship or kill decision ÷ all pilots older than 90 days

Gartner: at least 50% of GenAI projects were abandoned after proof of concept. A kill on purpose is the KPI working.

The trap

Treating seat utilisation as success. Once 90% of people have a seat it tells you nothing, and 90% of DORA's 2025 respondents already use AI at work.

Ask your CTO"What did last month's AI bill buy us? One sentence, with a number."

Example page · fictional company
Engineering / AI transformation

🧾 AI spend and pilots, the day Filip started

FNOwned by @Filip, Head of AI Transformation · 6 Jan 2027

!

145 paid seats across 5 tools. 41 people used any of them at least once last week.

ToolSeatsWeekly active$/monthOwner
Coding assistant A80291,520nobody
Chat assistant B3591,050Marketing
AI editor C183720nobody
Chat assistant D120360Sales ops
API keys, 7 of themn/an/a7,5504 teams
Total1454111,200

The two numbers

  • Spend per weekly active user: $11,200 ÷ 41 = $273
  • Pilot decision rate: 7 pilots from two hackathons, oldest 7 months, decided 0 of 7

Level diagnosis: L1 The API keys are the biggest line, and nobody had looked at them since the hackathon in June.

Level 2Company ladder · 2 of 5

Connectors, knowledge, skills

AI gets access to how the company works. Connectors into the tools people already use, one knowledge base it can read, and a marketplace of skills that people in different roles reuse instead of copy-pasting the same prompt again and again.

What to do

  • Connect the four places work lives: the code host, the ticket tracker, the wiki and the support desk. One page of data rules, written as a paved road, not a list of prohibitions.
  • Start the skills marketplace with ten skills that each save a named person time this week. A skill for writing a spec from a ticket. A skill for release notes. A skill for answering a support question from the docs.
  • One champion per team, with two hours a week for it, on paper.

The two KPIs

1Reused skills

shared skills used by at least two teams in the last 30 days

A skill only its author uses is a personal prompt.

2Agent-mode share, lowest team

people who used agent or chat mode (not only autocomplete) in a week ÷ people with a seat, per team; report the lowest team

At Sabre, adoption hit 74% while only 25% of users used agent mode (DORA 2025). The average hides the team that stayed behind.

The trap

Mandating usage. DX warns against top-down mandates and against AI metrics in individual performance reviews, because both invite people to game the number.

Ask your champions"Which skill did someone from another team use this week, and what did it save them?"

Example page · fictional company
Engineering / AI transformation / Skills marketplace

🧰 Skills marketplace, top 6 by reuse

MWCurated by @Mei, champion, Booking stream · 28 Mar 2027

SkillAuthor's teamTeams using itRuns, 30 days
Ticket to spec draftBooking9412
Release notes from merged PRsPlatform11 + Marketing96
Answer from the operator docsSupport2 + Support1,840
Timetable import checkerOperations3220
Incident timeline from logsPlatform631
Test cases from acceptance criteriaPayments5188

The two numbers

  • Reused skills: 0 → 14 of 31 published
  • Agent-mode share: 22% → 47% org-wide. Lowest team: Payments, 4% → 21%

Connected: code host, ticket tracker, wiki, support desk. Data rules: one page, linked from every skill.

Level 3Company ladder · 3 of 5 · the brake

Automation and the first AI team

Skills turn into scheduled agents and automations that run without somebody sitting next to them. A small AI team owns them. People start spending tokens wisely, because somebody now asks what each run is worth.

What to do

  • Build one AI team of 4 to 6 people. Only a handful of people in a company are usually ready for it, so expect to hire one person from outside who has already been through this change.
  • Move the best skills into scheduled runs: a nightly dependency upgrade, a morning incident digest, a weekly churn-risk list for customer success. Every run is logged and has an owner.
  • Stop climbing here until finance can see the money. The rest of the ladder costs more, and it only pays off if this level already does.

The two KPIs

1Verified hours saved per person per week

hours from scheduled runs (runs × minutes each replaces), cross-checked with a short survey; never the survey alone

Never the survey alone: perception is a bad meter (see METR above).

2Net AI return per person per month

half the value of the verified hours, minus AI spend per person

Only half, because saved time leaks: Atlassian found developers save about 10 hours a week with AI and lose about 10 to friction. The full formula is in the scorecard download.

The trap

People who sit in front of the screen all day answering the questions their agents ask. They use AI a lot and save nothing. If a run needs a human every ten minutes, it isn't automation yet.

Ask your CFO"Which of these numbers would you sign off in a budget review, and which would you laugh at?"

Example page · fictional company
Engineering / AI transformation / AI team

🤖 Scheduled runs owned by the AI team

KSOwned by @Kryštof, AI team · 30 Mar 2027

RunWhenReplacesHuman touch
Dependency upgrade PRsnightly40 min / team / weekreview only
Incident digest for on-call07:30 daily25 min / daynone
Flaky test triageafter each main build3 h / week1 in 6 runs
Spec draft from new ticketshourly30 min / ticketPM edits
Release notes + changelogon release2 h / releasereview only

The two numbers, R&D only

  • Verified hours saved: 1.2 → 2.6 h per engineer per week. The survey says 4.1 h.
  • Net return: (2.6 × 4.33 × $60 × 0.5) − $112 = +$226 per engineer per month
i

Ines, CFO: "Positive, but it's 80 people saving a bit each. Show me where it reaches a customer." That question is month 4.

Level 4Company ladder · 4 of 5

Token telemetry, AI-native delivery

You can see who spends tokens on what, per person and per team. And the company changes how it builds software: spec-driven development, where the spec becomes the main artefact and AI writes, reviews, tests and ships against it. At about 100 developers this takes three to six months.

What to do

  • Turn on usage telemetry for every tool and API key, by team.
  • Pick two teams. Rebuild their flow around five places where the time and money sit: specifications, code review, testing, automatic deployment and fixing incidents after release.
  • Write specs with product, engineering and business together. A spec written by an engineer alone is a ticket with more words.

The two KPIs

1AI cost per shipped change

AI spend of a team (seats + tokens) ÷ changes it deployed to production that month

Cost per change up while lead time stays flat means burned tokens.

2Spec-to-production lead time, guarded by change failure rate

median days from approved spec to running in production; counts only if change failure rate didn't rise

Faros measured median time in PR review up 441.5% in AI-heavy teams. Code got faster and review became the queue. DX's median change failure rate for tech companies under 100 engineers is 4.02%.

The trap

Speeding up coding, the part that wasn't slow. DX estimates coding is about 16% of a developer's day. The rest is specs, review, waiting and fixing.

Ask the two teams"Where did last month's slowest change wait, and for whom?"

Example page · fictional company
Engineering / AI transformation / AI-native pilot

🚢 AI-native pilot: Booking Core + Payments

DVSponsor @Dana, CTO · started 5 Jul 2027 · review 4 Oct 2027

MetricBeforeWeek 8
Spec-to-production, median9.5 days6.0 days
Time in review, median19 h7 h
Change failure rate4.8%4.6%
AI cost per shipped change$41$58

What changed in the flow

  • Specs written by PM + engineer + support lead together, with acceptance tests inside.
  • AI does the first review pass. A named human owns the merge.
  • PRs over 400 lines go back to be split.

Cost per change went up. Lead time went down by 3.5 days. Ines accepted the trade for the two pilot teams, not yet for all eleven.

Level 5Company ladder · 5 of 5

Self-running workflows

Complex workflows with several steps, scheduled, that the business relies on without watching them. Each one is tested, validates its own output, logs what it did and can be audited, across business and tech.

What to do

  • Promote only workflows that ran at L3 for a quarter with a stable owner.
  • Give each one a validation step that can stop it, and an audit log someone reads weekly.

The two KPIs

1Unattended run rate

scheduled runs that finished and passed validation with no human touch ÷ all scheduled runs

Automation, as opposed to a person babysitting an agent.

2Agent-delivered work share

merged PRs or completed tasks executed by agents under a team's ownership ÷ all of them, same review and failure-rate rules as L4

Faros found under 1% of PRs opened by AI agents as of March 2026. Most companies are a long way from here.

The trap

Agent washing. Gartner estimates only about 130 of the thousands of agentic AI vendors are real. A cron job with a prompt in it isn't L5.

Ask the workflow owner"When did this last stop itself, and who read why?"

Example page · fictional company
Operations / Workflows

⚙️ Timetable change workflow, month-12 target

BROwner @Bára, Operations · target for Jan 2028

  1. Operator emails a new timetable (PDF or spreadsheet).
  2. Agent parses it, diffs it against the live schedule, flags conflicts.
  3. Validation: every route, stop and time checked against the operator's rules. Fail = stop and page Operations.
  4. Pass: change scheduled, affected bookings found, customers notified in their language.
  5. Audit log to the operator account; weekly sample read by Bára.
MetricToday (L3)Month-12 target
Unattended run rate0% (manual)70%
Time from email to live3 days4 hours

Not part of the first 6 months. Filip named it in the month-6 memo as the first L5 bet.

Part 2

The plan. Six months, from L1 to the L3 brake.

Months 1 to 3 happen inside tech: the contract, the clean-up and the first AI team. Months 4 to 6 move that team into the business, where the money is, and end with a decision about L4.

First 3 monthsMonth 1 of 6

Sign the contract, take the baseline

Before you change anything, agree in writing what you'll achieve at 3, 6 and 12 months and what you get to do it. A title with no people and no budget is a strategy you'll present for a year.

What to do

  • Weeks 1-2: write the one-page contract with your CTO. Goals per horizon, people (named, with % of time), budget, access to finance data, the day it gets reviewed.
  • Weeks 2-3: inventory spend and pilots (L1 sheet). Meet every engineering manager and the heads of sales, marketing and support.
  • Weeks 3-4: run the individual ladder check for everyone who wants in. Ask people to show their skills, agents and scheduled runs, not to describe them.
  • Take the baseline of the L1 and L2 KPIs. Don't announce a strategy yet.

The trap

Announcing a strategy in week two. You don't know the numbers yet, and you'll rewrite it by month three.

Ask your CTO"If I do everything on this page, what would make you call it a failure in month 6?"

Example page · fictional company
Engineering / AI transformation

📝 AI transformation contract, v1

FN@Filip and @Dana, CTO · signed 24 Jan 2027 · review 30 Jun 2027

ByWe will have
Month 3Every pilot decided. AI spend per weekly user under $120. An AI team of 5 with scheduled runs in production.
Month 6AI team members inside sales, marketing, support and operations. A net return finance accepts. A go or no-go on AI-native delivery.
Month 12Two teams delivering AI-native. The first self-running workflow in operations.

What Filip gets

  • 4 engineers at 50% from month 2, 100% from month 3, plus 1 external hire
  • AI budget: current $11,200/month, capped at $15,000
  • Monthly 30 minutes with Ines (CFO) and read access to the cost data
  • Decision rights: tools, data rules, pilot kills
i

Baseline, 31 Jan: spend per weekly user $273, pilot decisions 0/7, agent-mode share 22%, reused skills 0.

First 3 monthsMonth 2 of 6

Clean up the licence mess

Get out of L1. Decide every pilot, cut the tools nobody uses, connect the rest to where work happens and start the skills marketplace.

What to do

  • One 30-minute decision per pilot, owner present: ship, kill or park with a date.
  • Keep the tools people chose with their feet. API keys on one billing account, labelled by team.
  • Connect the four tools, publish the data rules, name the champions, publish ten skills.

The trap

Keeping a pilot alive because its author is senior. Every zombie pilot tells the rest of the company that decisions here are political.

Ask each pilot owner"What number did this move, and would you pay for it from your own team's budget?"

Example page · fictional company
Engineering / AI transformation / Pilot register

🗂️ Pilot register, decided

FNOwned by @Filip · updated 27 Feb 2027

PilotNumber it should moveDecision
Support answer botfirst response timeShip
Timetable import checkerimport errors per operatorShip
Route price predictornone definedKill
Internal docs chatbotnone definedKill
Code review bot v1review timeKill rebuilt in M3
Slack meeting summarisernone definedKill
SQL helper for opsops tickets to engineeringPark until 1 May
  • Pilot decision rate: 0/7 → 7/7
  • Spend per weekly user: $273 → $93 ($8,900 ÷ 96). Two tools cut, 52 dead seats removed.

First 3 monthsMonth 3 of 6

Build the first AI team

Pick the people who are already far up the individual ladder and give them one job: turn the best skills into scheduled runs that save real hours. Then have the hard talk about who can't follow.

What to do

  • Choose 4 people from inside, usually the ones with AI pet projects, plus one hire who has done this before. Consultants only for pairing.
  • Give everybody else a learning path: a two-day workshop, pairing with a champion, weekly office hours.
  • Run a hackathon with one rule: a project counts only if it's in production by the end of the month, with an owner, logs and security review. A prototype isn't a result.
  • End of month: first scorecard to CTO and CFO. Show the spread of people across the individual ladder, team by team.

The trap

Pretending everybody will make it. Some people won't become AI-compatible in six months, and some never want to. Name it early, help them try, and plan what happens if they don't.

Ask the AI team"Which run would you trust to work if you were all on holiday for a week?"

Example page · fictional company
Engineering / AI transformation / AI team

👥 AI team and the month-3 scorecard

FNPresented by @Filip to Dana and Ines · 31 Mar 2027

MemberFromIndividual level
KryštofPlatform, staff engineerL4 (runs agents for his own side project)
MeiBooking, senior engineerL3
OndraPayments, engineerL3
BáraOperations, QA leadL3
SvenExternal hireL5 (led this change at his last company)

Engineers by individual level, n = 80

L0-0.5: 21 · L1: 30 · L2: 18 · L3: 8 · L4+: 3

!

9 engineers didn't move a level in 3 months and 4 of them said they don't want to. Plan with their managers by 30 April.

Next 3 monthsMonth 4 of 6

Send the team into the business

The money is rarely in R&D. It's in getting the product to market faster and in marketing, sales, activation, adoption and support. So after three months the AI team splits up and its members sit with the business teams.

What to do

  • One AI team member per business department, sitting with them three days a week. Champions in R&D keep the engineering runs going.
  • Each member brings one automation that touches a customer within four weeks: a support triage, account research for sales, launch content from the release notes, an onboarding nudge for new users.
  • Turn on token telemetry by team now. Business teams spend differently, and you want to see it before the bill does.

The trap

Keeping the AI team in engineering because it's comfortable there. Engineering hours saved are real, but they reach a customer slowly, and that's the gap finance will ask about.

Ask each business head"What does your team do every week that a customer is waiting for?"

Example page · fictional company
Company / AI transformation

🧭 Where the AI team sits, April to June

FNOwned by @Filip · 7 Apr 2027

WhoSits withFirst automationNumber it moves
MeiSupport, 22 peopleTicket triage + draft answer from operator docsfirst response time
OndraSales, 18 peopleOperator account brief before every first callprep hours per deal
KryštofMarketing, 9 peopleLaunch pack from release notes in 3 languagesdays from release to announcement
BáraOperations, 15 peopleTimetable import checker, scheduledimport errors per operator
SvenStays in R&DOwns the engineering runs and telemetryAI cost per shipped change

Telemetry live 21 Apr: tokens by team, by tool, by run. Support is the biggest spender within two weeks.

Next 3 monthsMonth 5 of 6

Prove the money

This is the L3 brake. Eight weeks of the two L3 KPIs, measured the same way as the baseline, across engineering and the business teams. Then a review where the CFO decides if the number is real.

What to do

  • Lock the window and the team roster before it starts.
  • Translate each business automation into money with the department head, not for them.
  • One page for the review: spend, verified hours, net return per department, what you'd stop.

The trap

Pricing every saved hour at full salary. McKinsey's 2026 survey found 80% of companies report individual productivity gains and 37% see any EBIT impact. Your CFO knows that gap.

Ask your CFO"What would you need to see to approve the next step out of this budget?"

Example page · fictional company
Company / AI transformation / Reviews

💰 Finance review, 8 weeks to 30 May

IMReviewed by @Ines, CFO · 6 Jun 2027

TeamPeopleVerified h / person / weekNet return / person / month
Engineering803.1+$283
Support225.4+$398
Sales183.8+$311
Marketing92.2+$134
Operations151.9+$97

What reached a customer

  • Support first response: 4 h 10 min → 48 min
  • Release to announcement: 12 days → 2 days
✓

Ines: "Support and the launch pack I can defend to the board. Engineering I believe, but I can't see it yet. Go to L4 with two teams, and show me lead time."

Next 3 monthsMonth 6 of 6

Decide L4, sign the next contract

Month 6 ends with a written decision and a new contract. If finance saw the money, start AI-native delivery in two teams. If it didn't, stay at L3 and fix what the review found. Both are honest results.

What to do

  • Write the month-6 memo: where you are on the ladder, the KPIs against the contract, what you killed, what you'd do differently.
  • Hand the L1-L3 numbers to the managers who own the teams. They go into the normal team health review, not into a separate AI report.
  • Sign the 12-month contract: L4 in two teams, telemetry for all, the first L5 bet.
  • Give your best AI people room for their own ideas inside the company.

The trap

Starting L4 because it's on the plan, not because L3 paid. AI-native delivery takes 3 to 6 months at this size. Without the L3 proof, the second half of that gets cancelled in the next budget round.

Ask yourself"Which of my numbers would survive if I left tomorrow?"

Example page · fictional company
Company / AI transformation

📌 Month-6 memo

FN@Filip to @Dana and @Ines · 28 Jun 2027

KPIBaselineMonth 3Month 6
L1 · spend per weekly user$273$93$104
L1 · pilot decision rate0/77/712/12
L2 · reused skills01427
L2 · agent mode, lowest team4%21%35%
L3 · verified h / engineer / week1.22.63.1
L3 · net return / engineer / monthn/a+$226+$283

Decisions

  • Go L4 in Booking Core and Payments from 5 July.
  • Next First L5 bet: the timetable change workflow, January 2028.
  • People Kryštof gets 20% time for his routing side project, as an internal product bet.
  • Me 12-month contract signed. AI KPIs move into each EM's team health review.

Start with the contract and two numbers

Once I could see my broken alerts, the fix was one command per Python install. Finding them was the hard part. Look for the same thing in your company: work that looks fine because nothing checks it.

So tomorrow, do two things. Fill in the contract one-pager with your CTO, and take the baseline of the two KPIs for your level. If you’re stuck in the middle with a blurry assignment and a solo seat, change the strategy, or change the contract about what you plan to accomplish and what you need for it.

If you want somebody next to you for those six months, that’s what my Head of AI mentoring is for.

Which level is your company on today, honestly?

[Writing time: a couple of evenings]

Frequently asked

What should a new Head of AI Transformation do in the first 90 days?+
Get a written contract with the CTO or CEO first: goals at 3, 6 and 12 months, named people, a budget and access to finance data. Then put every AI tool and pilot on one sheet, decide each pilot within 90 days, connect AI to the tools where work happens and build one small AI team that runs scheduled automations. Strategy decks come after that, not before.
What KPIs should a Chief AI Officer track for AI adoption?+
Two per level of adoption, and only for the level the company is on. At the licence stage, AI spend per weekly active user and the share of pilots with a ship or kill decision. At the automation stage, verified hours saved per person and net AI return per person per month. Seat counts, suggestion acceptance rates and the share of AI-written code are poor KPIs on their own.
What are the levels of AI adoption in a company?+
Marian Kamenistak's company ladder has five: licences bought without a plan, then connectors, a knowledge base and shared skills, then automation run by a first AI team, then token telemetry with AI-native, spec-driven software delivery, and finally self-running workflows that are tested, validated and audited across business and tech.
Does an 80-developer company need a Chief AI Officer?+
It needs one person who owns AI adoption with a clear mandate, a small team and a budget. The title matters less than the contract behind it. A Chief AI Officer with no direct reports and no budget at this size usually ends up writing strategies nobody executes.
How do you interview a Head of AI Transformation candidate?+
Ask for proof at three levels: how they use AI themselves, how they got teams and departments to use it, and what it changed in roles, structure and money. The strongest question is how they validate every step of an AI-driven delivery workflow. Without validation, AI output is a probability, and the time saved goes into checking it.

New posts, straight to your inbox.