When I am brought in to lead the experience of a product or feature, I like to start by inviting everyone to share in the vision I have outlined, which is often an accumulation of their many individual ideas, aspirations, and strategies to achieve particular business outcomes. This is where we dream of the ideal state without immediate technology limitations attached, so everyone in the room can see how the experience will enable its users, build their trust, and deliver unique value.
This part frequently goes well. People can imagine themselves and others using the product, and we often end with everyone in alignment. We break and leave excited.
Then the real work starts, and all of us who agreed on that vision must meet to negotiate what P0 will include. What ships first on the way to our vision? How does each milestone inform and enable the next?
In my nearly 20 years of product design, this conversation almost always focuses on features and functionality that are hoped to drive adoption. What should this product do? What jobs will it enable? What use cases does it cover? This is the wrong lens. Instead, the conversation must start from “what will bring people back to us, and trust us to enable them over any competing product?”
It’s understandable that these conversations go so predictably. Debating whether a product or feature has enough reason for being is an exercise in validating its existence at all, and making a business case for attention and resourcing. The people asking it aren’t being short-sighted on purpose. They’re looking out for the interests of the business, and it’s a good question. It’s also focused only on the immediate impact.
More often than not, these conversations end here. At an articulation of value delivery. Not trust building. Not adoption. Those discussions come later, and are often backed into by flexing feature definitions, or worse, by hoping an “amazing experience” will be differentiating enough to win customers.
This simply doesn’t work.

By used I mean handed work that matters, frequently, and without constant oversight or rework.
Value delivery wins for a reason that has nothing to do with conviction. It has a number attached to it. Trust doesn’t. So the trust work is never debated, but instead omitted entirely.
Building for retention is simple product thinking, with roots long before agents. Building a well-functioning product is not enough. You need users to want to use your product, and to trust it with high-value workflows.
Agents and agentic platforms have made the consequences worse. The entire conversation around them is a capability conversation. What use cases do they solve? What workflows do they support? What data do they have access to? How do they interact with other agents and tools? Very little of it is about how we build trust. Almost all of it is about demonstrating value quickly enough that people adopt the thing at all.
A release that delivers value can create adoption. It won’t create retention, and it certainly won’t create advocacy.
I have watched this happen to me. Most of my experience as a user is with tools that claim to take a designer from design to code. I’ve used Vercel, Figma Make, Builder.io, and of course Claude Code with Anthropic’s models. Every one of them is capable, and I still won’t let any of them work unsupervised.
The test I give them is deliberately easy to grade. I start from a tangible piece of design and ask the tool to build it verbatim. Design output is usually judged subjectively, but when I hand over a finished frame and ask for a faithful copy, the judgment is objective. It matches or it doesn’t.
They all get close. None gets close enough that I’d delegate more than a single task. Stacking tasks is out of the question.
That last part isn’t me being cautious. It’s arithmetic. If I hand an agent a stack of tasks and the outputs come back substandard across all of them, I’ve created more work rather than less. Whatever efficiency I was trying to gain gets drained through review.
I wrote about this in my framework of the AI Flywheel, when I argued that multi-prompt chains compound the potential for a negative experience, because a single ineffective step poisons everything that follows. What I didn’t say then is what the compounding costs the person on the other end of it.
The design-to-code tools are trained on design systems and methodologies that are one-size-fits-all, and they have trouble interpreting a unique point of view. So for now, all of them have only earned the trust to receive one supervised task at a time, which is a long way from what they were sold to me as. My work with them stays highly iterative and demanding of my time. They are helpful in making me feel like work is progressing, but far from a full replacement of my attention, skill, and craft. I keep using them, but none of them has yet earned the right to deliver work I haven’t reviewed and edited.
I don’t think I’m unusual in my thinking or approach to using AI to aid my process. Thales’ Digital Trust Index this year found 93% of IT leaders already using, deploying or planning AI initiatives, while 77% of consumers are still concerned about AI agents acting on their behalf. Plenty of companies are shipping capabilities. User trust hasn’t kept up.
A release that delivers value can create adoption. It won’t create retention, and it certainly won’t create advocacy.
I have watched this happen to me. Most of my experience as a user is with tools that claim to take a designer from design to code. I’ve used Vercel, Figma Make, Builder.io, and of course Claude Code with Anthropic’s models. Every one of them is capable, and I still won’t let any of them work unsupervised.
The test I give them is deliberately easy to grade. I start from a tangible piece of design and ask the tool to build it verbatim. Design output is usually judged subjectively, but when I hand over a finished frame and ask for a faithful copy, the judgment is objective. It matches or it doesn’t.
They all get close. None gets close enough that I’d delegate more than a single task. Stacking tasks is out of the question.
That last part isn’t me being cautious. It’s arithmetic. If I hand an agent a stack of tasks and the outputs come back substandard across all of them, I’ve created more work rather than less. Whatever efficiency I was trying to gain gets drained through review.
I wrote about this in my framework of the AI Flywheel, when I argued that multi-prompt chains compound the potential for a negative experience, because a single ineffective step poisons everything that follows. What I didn’t say then is what the compounding costs the person on the other end of it.
The design-to-code tools are trained on design systems and methodologies that are one-size-fits-all, and they have trouble interpreting a unique point of view. So for now, all of them have only earned the trust to receive one supervised task at a time, which is a long way from what they were sold to me as. My work with them stays highly iterative and demanding of my time. They are helpful in making me feel like work is progressing, but far from a full replacement of my attention, skill, and craft. I keep using them, but none of them has yet earned the right to deliver work I haven’t reviewed and edited.
I don’t think I’m unusual in my thinking or approach to using AI to aid my process. Thales’ Digital Trust Index this year found 93% of IT leaders already using, deploying or planning AI initiatives, while 77% of consumers are still concerned about AI agents acting on their behalf. Plenty of companies are shipping capabilities. User trust hasn’t kept up.
The resistance to this gets stronger the higher in the organization it goes. Trust is not an easily tracked metric, certainly not through a direct quantitative one. At the executive level, many find it hard to understand how a specific resource investment in trust actually drives revenue, especially in the near term.
They’re not wrong to ask.
Despite strongly advocating for investment in user trust, I don’t have a direct measure for it. Instead I have proxies.
That last one is the closest I get to a direct reading, and also the softest, because agency granted is somewhat qualitative.
The issue is that those all lag. Every one of them moves after the investment, and usually after the quarter that funded it. So an executive who says they can’t see the return inside the window they’re accountable for is describing the situation accurately.
But one cost doesn’t lag. Success is time spent on the work rather than on the system. Review is what it looks like when that fails. The work comes back, somebody checks it, and the minutes you were trying to save get spent anyway.
This is where my argument runs out. I can name the cost. I haven’t measured it. Everything I’ve said about review burden comes from my own desk, not a dashboard, and until someone instruments it properly, an executive is right to treat it as a story and not a number.
You’re not deciding whether to pay for trust. You’re deciding when to pay. Do it up front, in the first release, or pay for it later, in release after release spent rebuilding trust with the users who stuck around. And if you wait too long, in the ones you burn trying to win back the users who already left.
