Skip to content
Updates

Gemini Enterprise now reports how slow your AI agents are and how often they fail

Google added latency and error-rate views to Gemini Enterprise, so admins can see agent performance instead of guessing at it.

Google added two views to the Observability tab in Gemini Enterprise on 3 September. One reports latency, the other reports error rates. Both describe how the AI agents you have deployed are actually behaving.

The latency view breaks response times into Time to First Token, Time to First Answer and Time to Last Token, and reports each at p50 and p95. The p50 is the middle case. The p95 is close to the worst case, which is the number that tells you what a slow day looks like for the people using the agent. The error rate view groups requests into response classes: OK, client errors and server errors.

This is dull on the surface and useful underneath. Most businesses that deploy an internal agent find out it is failing when someone finally mentions it in a meeting, weeks later. A chart that separates a client error from a server error tells you whether the problem is how your team is using the thing or whether the service itself is falling over, and those need very different responses.

Worth knowing what it does not tell you. Latency and error rates measure whether the agent replied, not whether the reply was any good. An agent can answer fast, return a clean OK, and still be wrong. Keep a human checking output quality on anything that touches a customer or a number, and treat these dashboards as the operational half of the picture.

If you have agents running and no one has looked at how they perform, the Observability tab is the first place to check.

Wondering what this means for your business?

That's exactly what the discovery call is for. Free, useful, and no pressure.