When a coding agent runs unattended in CI, on a schedule, or as part of an autonomous pipeline, its final summary may be the only explanation of what happened. That creates a problem if the report is incomplete or simply wrong, since nobody is watching the transcript in real time.
For example, a subagent once had tests genuinely fail during a run, but the primary agent still reported that all tests passed. In an unattended workflow, that failure could easily go unnoticed unless someone manually inspected the logs afterward.
How are you handling this in practice? Do you use deterministic post-run checks, independent log or artifact validation, approval gates, or some other mechanism? Do you treat the agent's report as useful evidence, or assume it may be unreliable and verify everything separately?
This concern also led me to build a small open-source tool that records what actually happened during an agent session independently of the agent's own summary, but I'm mainly interested in the general approaches people are using.
3 Answers
Nothing should reach production based solely on an agent's summary. Keep production changes deterministic and put them behind the usual tests, artifact checks, and human approval gates. For prototypes or non-production work, the results still need to be validated before promotion.
I would be very cautious about giving an autonomous agent control over CI/CD decisions. An agent can help write code or workflows, but a competent reviewer and deterministic pipeline should decide what runs and whether it is allowed to move forward. Independent logs are useful, but they do not replace gates and reproducible checks.
The safest model is to treat the agent's report as a convenience for humans, not as an authoritative record. Have the pipeline determine success from exit codes, test results, generated artifacts, checksums, deployment state, and other machine-readable evidence. If those signals disagree with the summary, the run should fail or require review.

That matches my thinking: use the agent to produce deterministic code or configuration, then verify the resulting behavior independently rather than trusting the agent to repeat a process consistently.