Back to Insights
Agents ·14 September 2026 ·6 min read

How do you see what your agent did last night?

An agent that runs autonomously overnight should be traceable the next morning. What to log, how tracing with something like OpenTelemetry helps, and how to build a simple audit trail and dashboard.

If an AI agent runs autonomously overnight, you should be able to reconstruct its activity the next morning: what task it received, which steps it took, which tools it called, what data it read or changed, which errors occurred and what the final result was. The practical way to achieve this is through structured logging, tracing and a persistent audit trail.

A final message saying "task completed" is not enough. An autonomous agent chooses intermediate actions on its own. You therefore need visibility not only into the outcome, but also into the path that produced it. When something fails, you need to know where it failed. When everything appears to have worked, you still need to verify that the agent did not perform unintended actions along the way.

What should you record for every agent run?

Start by assigning every scheduled execution a unique run ID. Every event generated during that execution carries the same identifier. That gives you a reliable way to reconstruct one complete overnight job later.

At minimum, record:

  • when the run started and finished;
  • which agent and configuration version were used;
  • which task or trigger started the run;
  • which tools or APIs the agent called;
  • relevant inputs and outputs for those calls;
  • which files, records or systems were read or modified;
  • errors that occurred;
  • retries or fallback routes;
  • token or compute usage;
  • API or model costs where applicable;
  • the final result and run status.

This does not mean storing every piece of raw data indefinitely. If the agent processes personal data, customer information or confidential documents, excessive logging can create an additional privacy problem — the same data-minimisation principles that apply under GDPR apply to your logs too. Capture enough information to reconstruct the execution, while masking passwords, credentials, personal information and other sensitive values whenever possible.

Why are ordinary application logs not enough?

Traditional software usually follows a relatively predictable execution path. An AI agent may choose different actions each time it receives the same type of task. It might query a database, open a document, perform a search and then decide to generate or modify content.

A log entry such as "Agent completed task" tells you almost nothing.

A useful log contains context. For example, you may want to know that run 2026-09-14-001 performed a document search at 02:13, received three matches, opened one document and subsequently executed a write operation.

This is where structured logging becomes useful. Instead of relying entirely on free-form log messages, events are stored using consistent fields such as: run_id, timestamp, agent, action, tool, status, duration, error and cost. That makes it far easier to search, aggregate and display the data later.

How do you see the route the agent actually took?

For more complex agents, tracing provides more insight than logging alone. A trace represents one complete execution, while individual spans represent the steps inside that execution.

An overnight agent run might contain:

  • receive task;
  • create plan;
  • query database;
  • retrieve document;
  • call model;
  • validate result;
  • update file;
  • store final status.

Tracing shows not only that those actions occurred, but also their order and duration.

OpenTelemetry is a widely used open standard for this type of observability. It was not designed specifically for AI agents, but its core model works well for agent systems: logs, metrics and traces can share identifiers, allowing an operator to move from an error message to the complete execution that produced it.

Should you log the agent's decisions as well?

Yes, but that does not mean storing unrestricted internal model reasoning. For operational monitoring, explicit application-level decisions are usually more useful. Examples include:

  • "no relevant documents found";
  • "primary API unavailable, fallback selected";
  • "write operation blocked because permission was missing";
  • "output failed validation, retry started";
  • "confidence below threshold, human review required".

This makes the reason for a branch in the workflow visible without relying on long, unstructured model output. It also gives you data you can measure over time: how often the agent uses a fallback, how many tasks require manual review, or which validation step fails most often.

How do you build a simple audit trail?

For a small or medium-sized organisation, agent observability does not necessarily require a large monitoring platform. A central database can already provide a useful foundation: one table for runs and a second table for events.

The run table stores properties such as start time, end time, agent, status and total resource usage. The event table stores every action associated with the same run ID.

A simple dashboard on top of that data could show:

  • successful versus failed runs;
  • stalled or unfinished tasks;
  • execution duration;
  • failures by tool;
  • retry counts;
  • token and API usage;
  • write or delete operations;
  • runs requiring human review.

The most important feature is not the visual design of the dashboard. It is the ability to open one abnormal run and inspect the events underneath it.

Which events should trigger an alert?

Nobody should have to read hundreds of log entries every morning. Monitoring should surface exceptions. Useful alerts may include:

  • a scheduled agent did not start;
  • a run takes unusually long;
  • the same tool repeatedly fails;
  • retry counts suddenly increase;
  • token or API usage rises unexpectedly;
  • the agent attempts to operate outside its permitted scope;
  • a destructive action fails or is blocked;
  • a run finishes without valid output.

Dashboards are useful for investigation and trends. Alerts are for situations that require attention.

What should you inspect after an incident?

Imagine that someone discovers in the morning that an agent incorrectly updated a report. Do not start by guessing what the language model might have been thinking. Start with the run — identify which execution changed the file and follow its trace. Check:

  • What task did the agent receive?
  • Which data did it read?
  • Which tool calls did it perform?
  • What did those tools return?
  • Which decisions followed?
  • Were there errors or retries?
  • Which write operation changed the file?
  • Which agent, configuration and prompt version were active?

If that information exists, an incident becomes a technical investigation. Without it, the investigation quickly turns into speculation.

How long should agent logs be retained?

Retention depends on why the data is being stored and what it contains. Operational troubleshooting logs may require a different retention policy from a formal audit trail. It is therefore useful to distinguish between debugging data, operational monitoring and audit records: detailed technical error information may not need to be retained for as long as evidence that an agent modified a financially or legally relevant document.

In a privacy-first architecture, logging should follow the same principles of data minimisation and access control as the AI system itself. There is little value in running a protected on-prem RAG environment if sensitive prompts and document contents are then copied indefinitely into an unsecured logging system.

Conclusion

An autonomous AI agent should not only be capable of doing work; its actions should also be reconstructable and auditable afterwards. Give every run a unique identifier, record tool calls, relevant inputs and outputs, errors, retries, resource usage and important decisions, and connect those events through tracing. Then, when you arrive in the morning, you do not have to guess what your agent did last night — you can see it.

Control over autonomous agents

Do you know what your AI agent did last night?

We build logging, tracing and audit trails into the design of an agent — not as an afterthought. Curious what that means for your situation?