AI evaluation
Find what slows delivery before buying another AI tool.
I evaluate the path from a useful AI output to a shipped decision: who owns it, where it stalls, what requires human judgment, and what proof is missing.
The read
Boring evals stop at a grade. Useful ones produce a decision.
You leave with one constraint, one first move, the person who owns it, and the proof that will show whether it worked.
Staged scorecard
Leave with a decision your team can run.
Four common failure modes, rewritten as clear ownership, reusable systems, production proof, and review gates.
| Area | What stalls | What good looks like |
|---|---|---|
| Strategy | A broad AI roadmap with no named owner | One valuable problem, owned end to end |
| System | Prompts and process scattered across teams | One reusable workflow with review built in |
| Shipping | Demos accumulate while production stays the same | A working change ships behind your existing checks |
| Governance | Every output needs another manual inspection | Review and escalation match the actual risk |
The method
Three steps to one clear decision.
01
Map the drag
Trace where time, money, risk, and attention leave the workflow.
02
Find the constraint
Separate a tooling problem from an ownership, process, or review problem.
03
Define the first proof
Name what should change first, who owns it, and how you will know it worked.
The deliverables
Four artifacts turn the diagnosis into action.
If the answer is not a retainer, I will say so. If the first move is deletion, I will name it.
Start with the constraint