What 35+ AI implementations taught me.
Since February 2025, in my prior role, I ran more than 35 AI implementations at companies of 50 to 500 people. Insurers, manufacturers, wholesalers, an exam institute, a logistics group. Not every project went smoothly. A few ended with what we called internally a "light-green tick": it works, it's usable, but it wasn't the home run we wanted to deliver. One project ran 25 percent over in hours. One platform turned out, halfway through, to be too small for what we were building. One tool ended in an invoice dispute over how you measure.
This memo is not theory. It is the fourteen patterns I kept running into, with the numbers attached. What took a pilot to production, and what left it stranded as a nice demo. Ordered in three parts: before you sign, while building, and after go-live. At the bottom is a checklist of ten questions to ask any agency before you say yes to anything. Including me.
Part 1. Before you sign.
The most expensive mistakes are made before a single line of code exists. In the intake, in the scope, in the platform choice, in what nobody says out loud about what "done" means, and in the measurement nobody took.
Lesson 1The problem is rarely where the client says it is.
First conversations start with a solution. "We want a chatbot." Or: "We want to make our documents searchable with AI." Rarely with the problem itself.
The first step in every project: wind that solution back to the question. Which problem does this solve, for whom, how much time does it save, and is this the biggest problem there is?
At an exam institute they asked for automation of their knowledge base. After probing, the biggest pain sat in candidate registration in an external portal. One hour of manual work per ten candidates, error-prone. That is what we tackled first. The knowledge base came later. And certificate generation, which took thirty minutes each, now takes seconds.
At an IT services firm with 1,500 employees the ask was "more AI in sales". The baseline said something else: of fifteen salespeople, three used the AI features already built into their CRM. The problem was not a lack of tooling. It was that nobody got the tooling into their day. That is a different project from "build another tool".
The question worth your time is not "can we build it", but "is this the problem we need to solve".
Lesson 2Price on assumptions, pay in hours.
The second phase at that same exam institute was about the knowledge base. The proposal was based on 152 documents across fourteen subject areas. Reality: about a thousand documents across twenty. More than six times as many, discovered once the scope was already on paper. One item, a print-ready PDF, turned out to be delivered by another system already. Priced twice.
At a timber processor the pattern was the same, with harder numbers. An AutoCAD component was estimated at 50 hours and took 600. A module we had agreed "not to do" got built anyway: 160 hours. The whole phase ran from 1,140 planned hours to 1,421 worked hours, 25 percent over. The end result was good, thirty hours a week saved. But nobody had drawn the road there like that.
Now I do one thing differently: first an audit of the material, then a price. How many documents are there really. How many exceptions sit in the process. Who touches it. One day of looking prevents a quarter of overrun.
Lesson 3The platform sets the ceiling. The client's own platform too.
A kickstart at a large energy supplier on Copilot Studio. Halfway through it turned out: the platform blocks external tool integrations, offers only GPT without reasoning, and behaves inconsistently between the Teams and the Copilot interface.
Those are fundamental limits you don't route around with a better prompt. The client put it sharpest, about the final agent: "When I open Copilot in SharePoint it knows more context and gives better answers." Fair. The platform's limits had not been put on the table clearly enough up front. That was our mistake, not the client's, and we said so. The deliverable was redefined, together, into two smaller MVPs.
The ceiling is far from always in the AI platform. More often it sits in what the client already has. At a software company we saw at intake that the team limit on their AI subscription could block the build day: the user was already at a quarter of his weekly quota, so an extra seat had to be arranged up front. At a training provider the client asked on day two for central reuse of media across courses. We rebuilt the design, and five days later the learning platform's vendor confirmed that one shared library across content libraries is not possible. That question should have been asked before the rebuild. At a theatre company a European sovereignty requirement set the entire stack before a single feature was discussed. And at the energy supplier, an integration that could technically be done in an afternoon was stopped by the works council of the parent company, which is critical of AI.
Since then: platform assessment before the kickoff, not after. The AI platform, the client's systems, the licences, and the governance. Sometimes the answer is: pick another platform. Sometimes: build this use case somewhere else. You only see the ceiling if you ask.
Lesson 4Expectations are asymmetric. Write down what "done" is, and how you measure it.
Clients hear "PoC" and think "working product". They hear "four weeks" and think "everything done".
That is not bad faith. It is the result of marketing that presents AI as a simple solution, and of consultants (myself included, sometimes) who don't set the bottom of the expectation bar explicitly enough. A client in timber processing summed it up when a delivery was technically correct but, in his eyes, not finished: "When I buy a pen, I want to write with it. Not just get the ink cartridge."
The approach since: in week one, write a definition of done. Not vague ("a working chatbot") but concrete: which questions does it answer, at what error rate, in which system, judged by whom? That document is now a fixed part of project start, with a formal acceptance moment.
And one rule I only learned later: write down the measurement method too, not just the target. At a web shop two KPIs were on paper, logo recognition and email matching, each with a floor. Logo recognition we hit: 93.8 percent. Email matching we did not: our logs said 73.2 percent, the client measured 53, the agreement was 80. Two problems in one. A missed KPI, and twenty points between two measurements of the same thing, because nobody had written down how it would be measured. Invoice disputed, app paused. The targets were on paper. The way of measuring was not.
That conversation is uncomfortable. It is also the most useful conversation you can have early in a project.
Lesson 5Without a baseline, every result is an opinion.
Most AI projects end with a feeling. "It really saves time." "People are enthusiastic." That is not a result, that is a mood. A result is a delta: this is how it was, this is how it is now, this is the difference.
So the baseline lives in the rhythm, not in good intentions. In week two I measure how the process runs today: minutes per task, errors per batch, how many people actually use the existing tooling. In week three I measure the delta. Without a delta it is an opinion. With a delta it is evidence, and evidence is what you need when someone asks in month four why this cost money.
The baseline also changes the scope. At that IT services firm from lesson 1, "three out of fifteen" was not a number in a report. It was the project. Without that measurement we would have built a tool for twelve people who didn't use the previous one either.
Lesson 6Billing by the hour rewards overrun.
Look again at the timber processor from lesson 2. Those 281 extra hours were not the result of bad work. They were the result of a contract without tight phasing, without interim evaluation moments and without a financial brake. Every sub-flow dug itself into extra hours, and nobody had a reason to stop.
That is why I now work on a fixed price and a fixed scope. Not because hours are unfair, but because the incentive points the wrong way. Whoever bills by the hour earns from scope creep. Whoever agrees a fixed price earns from sharp scoping. The latter is what you want as a client.
There is one more reason, and it is new. An analysis that used to take two days now takes two hours with AI. Code that used to take three days is done in one. Pass that advantage through an hourly rate and the client earns from your speed while you gain nothing. Then you stop getting faster as an agency. A fixed price on a fixed result rewards the agency that gets faster, and the client who knows what they get.
The mirror image belongs here too: don't deliver before the signature. At a building-materials wholesaler the first sessions were already running before the contract was closed, because the relationship felt good. It worked out. It could also not have. The commercial clock and the operational clock rarely run in sync; let the first one lead.
Part 2. While building.
A good start is no guarantee. The next four lessons are about what goes wrong between kickoff and delivery, and how to see it early.
Lesson 7Small scope, big trust. And a checkpoint on day five.
The projects that turned out best were not the most ambitious. They were the most tightly bounded. One problem, one team, four weeks. Then measure. Then decide whether to go on.
Clients who want to start with "a complete AI platform for the whole organisation" I slow down. Not because it can't be done, but because it is rarely a good starting point. One working tool that twenty people use daily says more than a hundred-page roadmap. At a tile merchant the scope was one calculation: pulling material quantities out of a working drawing. That was usable within weeks. At the timber processor the scope was "the planning and the knowledge base and the drawings", and you have read what that cost in hours.
The honest lesson behind that small scope: our own four-week kickstart was, for a long time, too ambitious for the price. Training, analysis and an automation live in four weeks only works for simple cases. On more complex projects we only discovered in week two or three that it was harder than it looked, and that is where scope creep starts. So the decision point now sits on day five, with three outcomes. Simple: go ahead and build in weeks three and four. Medium: the client chooses, a simpler automation now or a larger engagement. Complex: several things at once, so we drop the kickstart format and move straight to a longer engagement. One decision point, early, with the client at the table. That is cheaper than a surprise in week three.
Lesson 8Internal green is not proof. Client data is the truth, and the test should belong to the client.
The tool for that tile merchant worked on our own test set. Green, all of it. When we ran it against a real workbook from a client of the client, the truth came out: pieces 0.0 percent deviation, wall area 8.1 percent, skirting 6.8 percent. Not bad. But not what the internal test suggested either, and the calibration that made the difference could only start once we had real client data.
At the exam institute the truth sat in one cell. The import into the exam system failed for an entire batch as soon as one field in one row was wrong. The client's data was the bottleneck, not the model. That is true more often than not.
Lesson: build against your own test set, but only celebrate once it holds up against the client's real data. Report the deviation honestly, ugly numbers included. A client who hears 8.1 percent trusts you more than a client who hears "it works" and later finds the deviation themselves.
And go one step further: give the client the test. I publish the evaluation set I build with, so it runs on the client's machine. Your machine, my test, your verdict. A claim you cannot recompute is not a claim but marketing. That goes for "95 percent accurate" as much as for "35+ implementations".
Lesson 9The security blocker is often a phantom blocker.
At a wholesaler, an integration with the inventory system sat still for months on a security clearance. When we worked out which data we really needed, most of it turned out to be reachable somewhere else already. The integration wasn't needed. Neither was the blocker.
The reverse happens too: legal and IT taking months over an AI policy, while employees use personal ChatGPT for work documents in the meantime. Without guardrails it is not your worst people who drop out, but your most careful: the people who wait for permission.
Two lessons in one. At every blocker, first ask: what data do we minimally need, and does it already sit somewhere we are allowed to go? And make sure guardrails exist before the tool, not after. A one-page policy on day one beats a legally watertight document in month six.
Lesson 10The champion is the key. And the bus factor.
Every successful project had a champion: someone inside who genuinely wanted to understand it. Not the IT manager, not the director. An employee who saw what AI meant for their own work and took others along.
Projects without a clear champion stall, even when the tool is good. The tool goes unused, the feedback stays away, and after three months nobody remembers what was built.
But one champion is also a risk. At a company where an email-bot pilot saved forty minutes a day, only one person knew the technology. When the follow-up came up, it was parked until after the summer. Not because it didn't work. Because the knowledge sat with one person.
The counterexample is an insurer. We trained 300 people there. After six weeks, ten of them were driving it themselves. After three months, sales processing was sixty times faster. After six months they no longer needed us. That last part is the goal. An agency you still need every week after half a year hasn't built champions, it has built dependency.
Identify the champion in week one. Give them access, time, input. And by week four, make sure there are two.
Part 3. After go-live.
Live is not done. The four lessons that separate a tool that gets used from a tool that sits idle in month three.
Lesson 11Adoption doesn't start after launch.
The most common mistake: planning adoption as a separate step after implementation. "We build the tool, then we run a training, then people will use it."
That doesn't work. Adoption starts at problem analysis. The people who will use the tool must be involved in defining the problem. Not as a checkbox, but because their input makes the tool better and their involvement drives adoption.
At a logistics group, a manager asked whether his team's training could be paid directly by the company, not expensed. That sounds like an administrative question. It was an adoption signal. He wanted to put his team behind this. That willingness is gold, and it does not start on launch day.
Lesson 12The silent majority decides adoption, not the believers.
In every organisation you see three groups. The accelerators, who are experimenting on day one. The holdouts, fifteen to twenty percent, sceptical or simply too busy, who only move once they see proof. And in between, the middle group: sixty to seventy percent of people, who think it's fine but don't start on their own.
The mistake I saw most often: all the attention goes to the accelerators. They are enthusiastic, they give feedback, they show up to the sessions. But they don't carry the middle group along by themselves. A tool that twenty enthusiasts use and two hundred others don't is not adoption. It's a hobby club.
What does work: the middle group gets proof from their own work, not from a demo. One colleague next to you showing that it saves her an hour on Tuesday does more than a day of training. That is why the insurer from lesson 10 runs on ten internal champions and not ten external consultants. Peer-to-peer beats top-down, every time.
Lesson 13One mistake weighs more than 99 good answers.
AI agents now run without errors 95 to 99 percent of the time. Still, people remember that one mistake. If they get no explanation with it, distrust grows, and a tool that is distrusted gets bypassed.
The fix is not a higher percentage. It is visibility. Let the tool say what it doesn't know. Let it show its source. Build in checkpoints where the agent shows what it did and why, before it moves on. Give a human the button to correct, and show that the correction does something. A client in timber processing said it about a tool we built for him: "Doesn't matter how long it takes, as long as it's right." That is the bar. Not speed. Being right, and explaining when it isn't.
Lesson 14What you leave behind is not a system but an agent with an owner.
The two examples in lesson 10 show the same thing: ownership determines what happens after you leave. The insurer runs on ten internal champions; with the email bot that saved forty minutes a day, only one person knew the technology. Dan Shipper of Every, which runs its own consulting practice through an agent in Slack, puts it sharply: "every agent needs a human". Gartner expects over 40 percent of agent projects to be cancelled by the end of 2027 because of escalating costs, unclear business value or inadequate risk controls.
That is why, since August 2026, my handover is a fixed list of five. One name that owns it. The evaluation set from lesson 8, running on the client's machine and not on mine. A one-page runbook: what it does, what it must not do, where to look when it stalls. An agreement on who steps in when it errs, with the button from lesson 13. And a path for the next model version, because the model that works today can be replaced within six months, and then a retest has to run.
Honestly: that list is younger than most of the projects in this memo. The first engagement where all five points are ticked at the client still has to be completed. That is why it sits here as a lesson and not as proof. Ask any agency, me included: what do you leave behind, and who owns it then?
The thread.
35+ implementations are not proof that I do everything right. They are proof that I've done enough wrong to know what works.
The thread through the successful projects is always the same: measure before you build, and be honest about what you measure. A problem you checked first. A scope priced on real numbers. A platform whose limits you know, the client's included. A definition of done with a measurement method. A baseline in week two. A fixed price. A decision point on day five. Validation on client data, with a test the client can run themselves. A champion plus backup. Adoption that starts on day one and aims at the middle group, not the fans. And a tool that explains what it does.
That sounds simple. Most companies still skip at least one of these steps. Usually because the agency doesn't ask.
Ten questions before you sign.
Ask them of every agency. Including me.
- Which problem are we solving, for whom, and how much time or money does it save per week? If the answer is "a chatbot", it isn't an answer.
- Did the agency see the material before naming a price? How many documents, how many exceptions, how many systems?
- What can the chosen platform not do, and what can our own systems and licences not do? Ask for three concrete limits. "None" is a red flag. And who at the works council, security or legal still has to say yes?
- What is the definition of done, on paper, with error rate, measurement method and reviewer?
- What is the baseline, when is it taken, and who measures the delta in week three?
- Is the price fixed or hourly? Who pays for overrun?
- When is the first moment we can stop without losing face? If that is only after four weeks, it's too late.
- Which data is it tested on: the agency's or ours? And can we run the test ourselves?
- Who is the internal owner, and who is the second?
- What happens in week one with the people who will use the tool, and what is the plan for the sixty percent who won't start on their own? If the answer is "training after delivery", read lessons 11 and 12 again.
Which one are you skipping?
One AI engagement, honestly reviewed?
Book 30 minutes. Tell me what you're building or considering. I'll tell you what I'd change, and what I'd kill.
Book a working session