The single biggest cause of “we can’t tell if AI worked” is that no one measured what things cost before AI arrived.

You want to prove a vendor’s tool cut your customer service response time by 40%? You need to know what your response time was. You want to show AI-assisted development shipped features faster? You need cycle time from six months before adoption. You want to demonstrate that automated report generation saved 200 hours per quarter? You need to know how many hours the manual version took.

Most SMBs skip this step. They adopt AI in the excitement of the moment, and six months later they’re trying to reconstruct what “normal” looked like from memory, which is not a defensible answer to a CFO. The result is that AI ROI conversations become anecdote-versus-anecdote — the enthusiasts remember the wins, the skeptics remember the misses, and no one has data.

The fix is unglamorous: measure for six to eight weeks before touching anything. Cycle time, defect rate, hours per process, customer satisfaction, whatever the AI initiative is supposed to move. Even rough measurement beats no measurement — the goal is a baseline your team agrees on, not a statistically pristine study.

Here’s the discipline that separates AI programs that produce defensible business cases from ones that produce PowerPoint theater:

Write the baseline down before you deploy. Send it to the finance team. Attach a specific number to the “success” criteria. Six months later, when you measure again, you have an actual comparison instead of a narrative.

I worked with an operations team that skipped this once and never again. They adopted an AI-powered scheduling tool that felt like it was helping. But when the CEO asked for the ROI at review time, they couldn’t produce it — nobody had recorded scheduling error rates from before. They kept the tool, but the executive team’s confidence in the next AI initiative dropped substantially, because it looked like they’d bought something without a business case.

The lesson: don’t let vendor case studies do your measurement for you. A “30% productivity improvement” that some other company reported is not evidence about your business. Your baseline is the only thing that translates a change into a decision.

If your engineering team, operations team, or customer-facing team can’t tell you the baseline numbers for their most important processes, that’s the pre-work. Baseline first. Adopt second. Measure honestly. Repeat.

Question 5 of the diagnostic tests exactly this.