Back to insights
Infrastructure ·September 20, 2026 ·7 min read

Monitoring AI services: which signals actually matter?

An AI service can be technically "up" while users get slow, expensive or low-quality answers. Which signals beyond uptime — latency, cost, drift, hallucinations — your monitoring actually needs to catch.

For an AI service, knowing that the server is reachable is not enough. The signals that actually matter include latency, error rates, cost per request, token usage, rate-limit events, availability of underlying models and providers, and indicators of response quality. An AI system can be technically "up" while users are receiving slow, expensive or poor-quality answers.

Effective monitoring therefore covers three layers at the same time: technical availability, operational performance and output quality. The third layer is what makes AI monitoring fundamentally different from traditional application monitoring. An HTTP 200 response tells you nothing about whether the model used the correct source, started hallucinating or suddenly began producing systematically different answers.

Why uptime alone is not enough

For a conventional web application, uptime is often a useful first indicator. If the application responds, the database is available and error rates remain low, the system is usually functioning as intended.

AI systems are different. A request can complete successfully from a technical perspective while the actual answer is unusable. A model may ignore relevant RAG context, misinterpret a source or produce an answer that looks correct structurally but is wrong in substance.

A dashboard showing only "API online", CPU utilisation and memory usage therefore misses a substantial part of the picture. Monitoring needs to show not only that the infrastructure is running, but also that the AI service is behaving within expected boundaries.

Latency: look beyond average response time

Latency is one of the most visible performance signals for users, but an average response time often hides important problems. A service may appear fast on average while a smaller but significant percentage of requests becomes exceptionally slow.

It is therefore useful to monitor response-time distributions, for example through percentiles. You should also separate the different parts of the request chain: retrieval time from a vector or search index, waiting time at the model provider, time to first generated token, total generation time, and execution time for tools or external APIs used by an agent.

For streaming interfaces, time to first token is particularly important. Users often perceive a system that starts answering quickly as more responsive than one that returns the entire answer only after several seconds.

Error rates: technical success is not functional success

HTTP errors, time-outs and failed tool calls obviously need to be monitored, but AI systems require a broader definition of failure.

Examples include an agent that cannot invoke a required tool, a RAG pipeline that finds no relevant documents, a model that produces output that does not match the required JSON schema, a workflow that repeatedly retries a step, or a fallback model being used far more often than expected.

Recording errors by component makes it easier to see where the actual problem occurs. A single overall "failed requests" percentage is usually too coarse.

Token usage and cost per request

Token consumption is both a technical and a financial metric. More tokens generally mean more processing time, higher costs and potentially more irrelevant context being sent to the model.

For that reason, monitor more than total monthly spending. Track usage per request, workflow, model and task type. A sudden increase may point to a prompt change, a retrieval problem or an agent executing unnecessary steps.

For agents, the full workflow matters. A single user request may trigger dozens of model calls, searches and tool executions. Measuring only the cost of the final model response gives an incomplete picture.

Rate limits and quotas

Many AI services depend on external model providers and APIs that impose limits on requests, token throughput or concurrent processing. Rate-limit events are therefore an important operational signal.

An occasional HTTP 429 response is not necessarily a problem if retries and queues are configured correctly. A sustained increase is more significant. It may indicate rising traffic, missing batching, overly aggressive parallelisation by an agent or insufficient provider capacity.

Retry counts and waiting time after throttling should also be monitored. Otherwise, the service may appear technically successful while spending an increasing amount of time waiting.

Provider and model availability

If your AI service relies on an external provider, that provider's availability becomes part of your own operational risk.

Provider-side failures should therefore be distinguished from failures in your own application. Relevant events include time-outs, overload, authentication failures, retired model versions and endpoint changes.

In systems that use multiple models or providers, it should also be visible when a fallback is activated. A fallback may improve availability while producing different latency, cost or output quality.

Drift in response quality

One of the hardest forms of degradation to detect is quality drift. The service continues to operate technically, but its answers gradually or suddenly become worse or simply different.

There are several possible causes. A provider may update a model, the knowledge base may change, retrieval performance may deteriorate or a new prompt version may have unintended effects.

Quality should therefore be checked periodically against a representative evaluation set. Useful checks include whether the correct source document is retrieved, whether required information is present, whether answers remain consistent with known reference answers, whether citations actually support the statements being made, and whether the system answers outside its permitted scope. These checks do not need to be entirely automated — manual sampling remains valuable for important workflows.

Detecting hallucinations and abnormal behaviour

Hallucinations are difficult to represent as one reliable metric. In practice, it is more useful to detect concrete forms of deviation.

For a RAG system, you can check whether factual claims are supported by retrieved sources. You can also flag cases in which an answer cites documents that were never retrieved, or where the model provides a confident answer even though retrieval produced no relevant context.

For agents, other deviations may matter more: unexpected tool calls, unusually high step counts, repeated actions, abnormal execution order or attempts to operate outside predefined boundaries. The objective is not to invent a universal "hallucination score", but to define measurable conditions under which system behaviour differs from what the application is expected to do.

Monitor the entire chain, not just the model

Production problems do not necessarily originate in the language model. An AI service may include authentication, application logic, a vector database, document storage, embedding models, model providers, external tools and sometimes multiple agents.

A useful monitoring setup must therefore be able to follow an individual request across the full chain. Tracing makes it possible to see which components were called, how long each stage took and where failures occurred. For agents this is even more important: without an audit trail, reconstructing why an agent performed a particular action can be extremely difficult.

When should an alert fire?

Not every deviation requires an immediate notification. Excessive alerting makes genuine incidents easier to miss.

Alerts are most useful for conditions that may require direct action: rapidly increasing latency, high request failure rates, sustained rate-limit events, provider outages, unexpected cost increases or clear quality degradation.

Other metrics are better suited to trend analysis. Token consumption, retrieval quality or the average number of agent steps can be monitored over days or weeks to detect gradual changes. Effective AI monitoring combines real-time operational alerts with periodic quality evaluation.

Frequently asked questions

Which metrics are most important for an AI service? Latency, error rates, token usage, cost per request, rate-limit events, provider availability and response-quality indicators together form the core monitoring set.

How can you monitor hallucinations in an AI model? Usually indirectly, by checking whether claims are supported by sources, whether citations are valid and whether responses remain within predefined quality and behavioural rules.

Why is ordinary uptime monitoring not sufficient for AI? Because an AI service can remain technically available while producing incorrect, slow, expensive or abnormal responses. AI monitoring therefore needs to track behaviour and output quality as well as infrastructure status.

Conclusion

The most important monitoring question for an AI service is not simply "is the system running?" but also "is it still behaving as intended?" Infrastructure metrics therefore need to be combined with signals for latency, failures, token usage, costs, rate limits, provider availability and output quality. For RAG systems and agents in particular, tracing, source validation and deviation detection are essential because they expose problems that conventional uptime monitoring cannot see.

Visibility into your AI services

Do you know what your AI system did last night?

We build monitoring and tracing that looks beyond uptime — latency, cost, drift and abnormal behaviour visible, with alerts that actually matter.