AI

What building EVIE taught us about AI agents

We set out to save product managers an hour a day. The useful half of that turned out to have very little to do with the model, and almost everything to do with finding things.

EVIE, a purple owl in glasses holding a magnifying glass
5 minute read 14 September 2026 Veris Labs
EVIE, a purple owl in glasses holding a magnifying glass
EVIE says

This one is about me, which is a little strange to introduce. The honest lesson: the clever part was never the hard part. Finding the right four messages out of four hundred was.

EVIE started with a specific annoyance. Opening a Jira ticket tells you what needs doing, and almost nothing about what has already been decided. That lives in a Teams thread, an Outlook chain, a Confluence page nobody linked, and a conversation somebody had in a meeting.

Reconstructing it takes a few minutes per ticket. Across a day of ticket triage, it adds up to somewhere between thirty and sixty minutes for the product managers we watched. That was the problem we set out to remove.

Here is what we got wrong and right along the way.

The hard part was finding things, not generating them

We expected the difficult work to be in the reasoning: producing a good summary, spotting risk, suggesting a next action. It was not. Modern models do that well enough that it stopped being the constraint quite early.

The hard part was retrieval. Working out which of four hundred Teams messages relate to this ticket is a genuinely difficult problem, and getting it wrong poisons everything downstream. A beautifully written summary of the wrong three sources is worse than no summary, because it is confidently wrong and reads as authoritative.

Most of the engineering went into deciding what to look at. Almost none went into what to say about it.

What eventually worked was unglamorous: per-ticket custom search terms, so a user can teach it that an incident is discussed by name rather than by key, plus heavy weighting towards recency and towards people already on the ticket. Not a clever architecture. Just a lot of tuning against real cases.

Where an assistant makes things worse

Our first version of notifications was, in hindsight, obviously wrong. It alerted on every change to a watched ticket. Status, priority, assignee, due date, every comment.

That is not an assistant. That is a second inbox, and people turned it off within a week.

The fix was to stop treating every change as equal. A status moving to Blocked matters. A due date slipping matters. Someone adding a comment saying "thanks" does not. Once notifications became selective, usage recovered and stayed.

The general lesson we took: an assistant that surfaces everything has moved the work rather than removed it. The value is entirely in what it decides not to tell you, and that is a much harder product problem than generating the notification.

Trust is asymmetric and it breaks fast

The thing we underestimated most is how differently people treat a wrong answer from a missing one.

If EVIE cannot find something, users shrug and go and look themselves. They lose thirty seconds and their opinion of the tool barely moves. If EVIE confidently states something that turns out to be wrong, they stop trusting the whole feature, including the parts that work. One bad summary costs more than ten missing ones.

That pushed us towards being visibly conservative. Show the sources next to the summary. Say when confidence is low. Prefer "I could not find a decision on this" over a plausible guess. It makes the product feel less impressive in a demo and considerably more useful on a Tuesday afternoon.

Design for the failure, not the demo

The question that mattered was never how good the best output is. It was what happens when it is wrong, how quickly the user notices, and how much that costs them. Getting that right is what earns the second week of use.

Narrow beat general, repeatedly

Every time we broadened what EVIE tried to do, it got less useful. A general assistant that can answer anything about your work sounds better than a tool that investigates a ticket. In practice the general version gave vaguer answers, because it could not make assumptions about what you wanted.

The features people actually use are the narrow ones. Investigate this ticket. Draft a reply with the full history in mind. Tell me what changed on the things I care about. Each has a defined input, a defined output and an obvious way to tell whether it worked.

This matches what we now tell clients: pick one frequent, repetitive, checkable task and do that properly. It is also the core of our guide on where AI earns its place, which is largely this lesson written down for people outside software.

Measuring it honestly was harder than expected

We claimed thirty to sixty minutes saved per product manager per day. We are reasonably confident in that range, and getting to it was more work than building several of the features.

Self-reported time savings are close to useless. People are generous about tools they like and harsh about tools they do not, and neither has much to do with minutes. What we ended up doing was timing the specific task, ticket investigation, before and after, with the same people on comparable tickets.

That produced a smaller number than the enthusiastic self-reports and a much more defensible one. It also told us which features to cut, because two of them saved no measurable time at all despite being well liked in interviews.

What we would tell anyone building one

  • Assume retrieval is the project. If your assistant works over existing information, finding the right information is most of the work and most of the risk.
  • Be selective or be ignored. Surfacing everything is not assistance. The value is in the filtering.
  • Optimise for being trusted, not for being impressive. Show sources, admit uncertainty, and prefer silence to a confident guess.
  • Keep the scope embarrassingly narrow until it is genuinely good, then widen.
  • Time the task before you start. Without a baseline you will be arguing about impressions six months later.

EVIE is still in development. You can see where it has got to on the EVIE page, and if you are weighing up something similar in your own business, the AI readiness scorecard covers the questions worth answering before you start.

Keep reading

More from the blog

Vee, the Veris Labs robot, holding a spanner Built by Veris Labs

We build the systems that run growing businesses

These calculators are a small, public version of what we do. The Veris Labs suite covers marketing and delivery, and where nothing off the shelf fits, we build it around your business instead.