Why most AI projects never show up on the P&L
Nick Dreyfus explains the 98% vs 95% paradox, the one-sentence tripwire, and how to turn ‘saved hours’ into real revenue.
Nick Dreyfus opens with two numbers that should make every CFO sit up: EY reports 98% of senior leaders believe AI delivered value, while MIT finds roughly 95% of generative AI pilots show no measurable return. Both are true. The space between them is where money quietly disappears.
The real problem: a name sells a destination, not a product
Nick uses Tesla’s naming as a lens: a feature called “full self‑driving” sits next to a firm disclaimer that the vehicle is not autonomous. The name promises a destination; the product is a prediction engine today. People act on the name. They assume the technology thinks when it actually guesses the next most likely word.
That difference matters because a failed database tells you nothing; a failed prediction engine gives you a confident, polished, and often wrong answer. Hallucinations are not bugs — they’re how these models were trained and rewarded. For leaders that means the biggest risk isn’t a technical fault; it’s a budgeting and expectation fault.
"You cannot have AI thinking. You can have AI flagging." — Nick Dreyfus
Saved hours are an option, not a return
Someone will tell you the new tool saves 10 hours a week. Stop there and ask: per person or total, and per week or per month? Ten hours across 100 people is six minutes each — a rounding error. Ten hours per salesperson per month is a different business: that could generate 3–5 quality opportunities and real revenue.
Nick’s framing is simple and ruthless: saved hours create the option to collect value only if you redeploy them to revenue‑producing or revenue‑protecting work. Handing a team a tool and saying “figure it out” is the fastest way to see those hours evaporate into morale benefits, earlier Fridays, or unpaid improvements that never touch the P&L.
Concrete fixes: align incentives and aim.
Incentives: share upside tied to the outcomes you want — revenue, retention, cross‑sell — but document it correctly. Non‑discretionary bonuses, overtime rules, and verbal revenue shares can create legal and payroll problems if you don’t structure them.
Aim: point freed time at your worst number. If attrition is 25% a year, redeploy time to client success to stop churn. Retention is the cheapest revenue because you already paid to acquire those customers.
Not all freed hours are equal; pick functions that convert
Free up accounting time and you usually get more accounting. Free up sales or client success time and you get pipeline, cross‑sales, or lower attrition. Nick’s rule before automating anything: answer, "If we give this department back time, what will they do with it that moves the needle?" If you can’t answer, don’t start.
A reasonable early win is invoice processing — treat it as training, not ROI. But if your strategic goal is revenue, prioritize sales, client success, or marketing campaigns that can be measured against a single business metric.
Decide how the project dies before you start
Nick prescribes a tripwire in the approval document: if the agent is not operating in production completing real work within 60–90 days, stop. Operating on its own means completing narrow tasks without constant human intervention; it does not mean no one owns the outcome. A human must still review exceptions and own the result.
Why projects stall: teams try to build giant multi‑purpose systems, they’re learning a new discipline while building, and the target keeps moving as APIs and standards change. Those are solvable problems — but they require honesty about skills and realistic timelines.
If your internal team has built agents before and can show them running, let them lead. If they haven’t, partner with a vendor who can deliver and support outages quickly. Outages will happen; you want someone who has seen them and fixed them faster.
Two vendor questions that save you from slide decks
Before you sign: ask for references for this exact capability in production. And ask for a 30‑day trial (or a low‑cost proof). If they can’t prove against your data, walk. Smaller teams often do this well; newer vendors that won’t demonstrate or trial are selling hope, not outcomes.
A few numbers to put in your pocket
98% (EY): leaders saying AI produced value. 95% (MIT): pilots with no measurable return. 61% (Beam.ai): projects approved where ROI was never measured after launch.
Tripwire: 60–90 days to see production work, then stop.
The Monday action that costs nothing
Talk to each department individually. Ask: what repetitive, clear, black‑and‑white tasks could a junior $15/hour person do? Those are the flagging jobs AI does best. Then ask the second, harder question: if we give you that time back, what specifically will you do that moves the needle? Start where they have a strong answer.
One sentence you can use in every approval
You cannot have AI thinking. You can have AI flagging. Make the machine do black‑and‑white work, make a human own outcomes, and measure one business metric at 90 days.
If you want a proven way to check an AI you’ve already bought, Nick’s team will run a verticalized review, show a case study from your industry, and connect you to references. No pressure — the value is in the verification, not the pitch.
Takeaway to keep: align incentives, pick the right department, set a 60–90 day tripwire, and ask vendors to prove it in 30 days. That sequence turns saved hours into measurable revenue instead of mysterious optimism.