Using AI prototypes to find the point of leverage
Two AI prototypes appeared successful in isolation. Testing them inside real workflows revealed what had to change before the capability could matter.
Live testing revealed that the real question was not whether AI could give useful guidance. It was whether that guidance arrived at the moment it could actually help.
Can AI do this?
Where can it change the outcome?
The prototypes did not fail because AI was incapable. They worked well enough to reveal that capability was the wrong test.
Each prototype exposed the condition that determined whether the capability could create value: shared criteria for Anne-bot, shared awareness for Capture Caddy.
AI creates pressure to build before anyone finds the leverage point
AI concepts can look viable early. The technology performs the task in isolation, so of course it looks ready. But capability alone doesn't show where a product should intervene, what the workflow can absorb, or what has to change first.
I built two lightweight LLM assistants: Anne-bot for enterprise PRD coaching, Capture Caddy for a live creative workflow. Both were plausible product ideas, and both could generate useful output.
The more important question appeared only after use. Each prototype exposed a smaller, non-AI condition that determined whether the capability could create value, or just accelerate the wrong system.
AI could extend coaching, but couldn't resolve what "approved" meant
On-demand AI coaching could shorten PRD approval cycles. Ford Digital Cabin's PRD reviews were rigorous and Socratic, often stretching across several rounds. Coaching access was concentrated in live review sessions, which meant PMs often had to wait for the next review to get meaningful feedback.
Anne-bot tested whether an assistant could reproduce that Socratic coaching style before review. Instead of giving answers, it asked clarifying questions, challenged vague framing, and helped PMs strengthen the PRD before putting it in front of leadership.
Coaching became more useful when the assistant paused for missing context
The first version gave feedback on the entire PRD at once. It could identify vague claims and ask useful follow-up questions, but without access to context that lived outside the document, roughly half of the feedback missed the mark.
Through PM feedback and iteration, Anne-bot changed from a full-document critic into a section-by-section coach. It paused to gather missing context directly from the PM, then supported one part of the draft at a time.
A third approval path existed, but it was undocumented
At first, I treated the missed feedback as a context problem inside the assistant. Anne-bot needed to ask better questions, gather more information, and avoid commenting on parts of the PRD it could not see clearly.
Observation of review meetings showed a different source of leverage. The official process had two visible outcomes: a PRD was approved in live review and moved to delivery, or it waited for the next live review.
But research surfaced a third path. When the problem statement was approved in the room, the director was sometimes willing to review the remaining PRD asynchronously.
The prototype pointed away from more coaching and toward the approval dependency controlling cycle time. The fastest route was getting one part of the PRD strong enough to unlock async approval.
Until the approval standard existed, Anne-bot could help PMs draft, but it could not resolve the agreement needed for approval. Further AI investment should wait until the evaluation criteria are explicit, because automating around a missing standard would not solve the approval problem. It would only make the gap harder to see.
AI could surface pose plans, but hidden guidance broke the moment it was meant to support
Hands-free access to a pose plan, delivered privately through an earbud, would be less disruptive during a live session than checking a phone or printed notes. Capture Caddy tested whether AI could help a photographer stay oriented during a shoot without visibly breaking connection with the subject.
The same guidance felt different when it was hidden versus shared
The first test used earbuds with a familiar subject. It did not go well. Keeping the guidance private meant splitting my attention between the voice in my ear and the person in front of me. Environmental noise and unreliable transcription made the interaction harder, and I still had to pull out my phone to speak clearly, which defeated the hands-free premise.
My instinct after that test was to stop. The prototype seemed too awkward for the moment it was meant to support.
A second test changed one variable: the guidance played on speaker instead of through an earbud. The test happened at home with the same familiar subject, in a lower-stakes setting where the interaction could be messier. Because I was not trying to hide the assistant, many of the technical and social pressures dropped away.
That surprised me. The assistant felt less like a private whisper competing for my attention and more like having another voice in the room. Both of us could hear the guidance, respond to it, and treat it as part of the interaction.
Two tests, one variable changed. Earbud delivery outdoors, then speaker delivery at home. Same subject, same tool, different result.
The guidance helped only when it became part of the shared interaction
The prototype did not just expose an attention problem. It exposed a social one.
A hidden earbud, however reliable, asked the subject to trust guidance she could not hear. Speaker delivery made the guidance more inclusive: less like managing the subject invisibly, more like a shared prompt both people could respond to.
The comparison also changed how I understood the product's role. Capture Caddy was less convincing as live-session assistance, but more promising as a lower-stakes practice tool that could build confidence, rehearsal, and muscle memory before a shoot.
Through an earbud
- Attention splits between voice and subject
- The subject cannot hear or respond to the assistant
- Latency and noise compound the disruption
Played on speaker
- Both people can hear it
- The subject can act on guidance directly
- The assistant supports the interaction instead of interrupting it
Both prototypes revealed that the surrounding workflow determined whether AI could create value.
Anne-bot and Capture Caddy encountered different limitations. Anne-bot could not resolve evaluation criteria that the organization had never made explicit. Capture Caddy could provide useful guidance, but private delivery divided attention and introduced a hidden layer into a shared interaction.
In both cases, improving the assistant would not have resolved the condition controlling the outcome. Anne-bot needed shared criteria before coaching could scale. Capture Caddy needed shared context before guidance could support practice or a live moment.
The prototypes therefore changed the product question. Instead of treating useful AI output as evidence that the product should be built, I began evaluating whether the surrounding workflow could absorb the intervention without weakening judgment, attention, trust, or coordination.
- Evaluation criteria were not shared
- Private guidance divided a shared interaction
Each workflow needed a human condition to be established before AI assistance could improve the outcome.
Clarify or change that condition first, then decide whether and where AI belongs.
What has to stay human?
I used to ask whether AI could perform the task. Now I ask what has to stay human for the workflow to still work.
Both prototypes looked promising in isolation. The more important evidence appeared once each one entered real use, with a real person on the other side of it.
Anne-bot's evidence pointed to a recommendation, not a launch: pause further investment until the team had shared criteria. Capture Caddy's evidence came from a single test that changed one variable, not a testing program. Both are honest about how much they prove, how much they don't, and where human judgment and accountability still have to hold the work.
That changed how I evaluate early AI product ideas. A working assistant isn't the same as a viable product. The prototype has to reveal more: what capability ceiling it hits, what has to stay human because of that ceiling, and how much the evidence actually supports before I call it a finding.