When is an AI product ready to trust?

Erad Fridman
Co-founder & CEO
August 12, 2026
Share article
When is an AI product ready to trust?
Contents
Summarize this article

Copy the prompt, open the AI tool, then paste and send.

Open AI tool ↗

“We built 80% of it in a few weeks. Can you help with the last 20%?”

We hear some version of this regularly. A team has built a convincing prototype, the main flow works, and the finish line looks close. 

As the team prepares for production, a different set of questions come up. Is the output good enough? What can happen automatically? When should someone step in?

The decisions influence everything from the interface to the architecture. They’re also why a demo can hide how much of the work is left.

Defining what good means

A team recently showed us a tool for architects. Typing “make the kitchen bigger” redrew a floor plan in seconds. Walls moved, rooms resized, and at first glance, everything looked right.

Outside the demo, we found plausible-looking plans with rooms too small to use and staircases that led nowhere. The model had followed the instruction, but the building no longer made sense.

That’s where the domain expertise comes in. An awkward room layout might be fine while exploring an idea, but one that violates building codes isn’t. The bar is different for an early-concept tool than for one producing construction plans.

Reliability isn’t a universal threshold. It depends on the job the product is doing and the stakes involved.

Learning from real examples

Once that standard is clear, you can turn it into tests. The best evaluation sets come from real examples: unusual requests, vague instructions and examples that previously produced unusable plans. Some can be checked automatically. Others need an architect’s review.

These evaluations are how a team finds out whether the product is good enough for its intended use. A convincing demo shows that it can work. Evals help establish how often it works, where it fails and whether those failures are acceptable.

Run those examples again whenever the model, prompt or workflow changes. An improvement in one area can create a problem somewhere else. Without a consistent set of evaluations, it’s hard to know whether a change made the product better or just made a few examples look better. Keep adding to the set after launch. Every new issue is another case the product should learn to handle.

Keeping the software simple

Some familiar engineering practices become even more useful when part of the system behaves unpredictably. Break the problem into smaller problems. Keep each step focused. Test the parts as well as the whole.

“Make the kitchen bigger” might involve clarifying what can change, proposing a new layout and checking that the result still meets the constraints. Keep each step focused. Test the parts as well as the whole, using software tests for predictable behaviour and evals to assess the model’s output.

Understanding what happened

A workflow can complete every step successfully and still produce a staircase that leads nowhere. To understand why, teams need the relevant inputs, the context the model saw, the tools it used and the checks applied afterwards. Was an important constraint missing? Did the request need clarification? Did a check miss something? Each points to a different fix.

For an architect, getting this right means more than avoiding bad plans. It means being able to explore more layouts, compare ideas earlier and bring better options to a client.

That’s why domain expertise matters so much when building AI products. The people who know the work help define what good looks like, where automation makes sense and where judgment still matters. Evals turn that knowledge into evidence the team can use. Good software engineering practices make it easier to act on what they find.

The demo starts with “make the kitchen bigger.” The useful product lets an architect explore ten versions, discard nine and find the one they might never have had time to consider.

Erad has 15+ years of leadership experience. At Google, Erad led product & design teams in gTech, Ads & Finance. Erad started coding aged 6. By 19 he had completed dual CS & Math degrees.