Green telemetry can hide silent data loss when workloads shift tools
A developer’s weekly AI report stayed healthy while missing most new data because the monitoring only checked the old tool, not the new one.
Bài này hiện chỉ có bản tiếng Anh.
A software engineer discovered that their automated telemetry system reported normal health for weeks while silently failing to capture the majority of their AI agent activity. The incident occurred in October 2026 after the user shifted their workflow from Claude Code to Codex, a change that moved data generation outside the scope of their existing monitoring pipeline.
What happened
The engineer maintains a weekly stratified comparison report that analyzes agent sessions, tracking metrics like tool error rates and output tokens. This report is generated by a cron job every Monday at 09:30, which processes transcripts archived from a specific directory used by Claude Code. For several weeks, the report continued to publish with a status of "normal," showing zero data quality issues and consistent statistical distributions.
However, the underlying data had changed dramatically. On September 6, 2026, the user configured their system to route unattended scheduled runs and headless queries to Codex by default, restricting Claude Code to just two sessions per day. Later, on October 5, they moved Claude Code back to the main orchestrator role, but unattended runs remained on Codex. The ingestion pipeline was never updated to read from the Codex session directory, meaning it only captured a tiny fraction of the actual work being performed.
Despite this massive blind spot, the health checks passed every week. The system verified that the copy command exited with code 0, that the archive file count did not decrease, and that internal data consistency rules were met. Because the cron job itself ran successfully and processed the few remaining Claude Code files without error, the monitoring system had no reason to flag an incident. The report accurately described the small slice of data it received, creating a false sense of security.
Key details
- The telemetry report showed identical medians and interquartile ranges for weeks, differing by only 12 lines related to metadata and minor counts.
- Between September 13 and September 20, the system recorded zero Claude Code transcript files but 192 Codex rollout files.
- From September 27 to October 4, the ingest captured only 3 Claude Code files while 1,058 Codex files were generated.
- The newest session in the database ended on September 11, yet the September 21 report still marked the system as normal.
- Health checks validated process execution and data consistency but did not verify if the volume of ingested data matched actual activity.
- A separate staleness check passed because the database file was rewritten weekly by the build process, regardless of whether new data was added.
Background
Telemetry systems often rely on "heartbeat" or exit-code monitoring to determine health. If a script runs to completion without crashing, it is considered healthy. This approach works well for detecting crashes but fails to detect "silent failures" where the script runs but processes no meaningful data. In this case, the monitoring was coupled tightly to a specific tool path. When the workload shifted to a different tool with a different file structure, the monitor continued to watch the empty old path.
This illustrates a common pitfall in self-hosted observability: monitoring the mechanism rather than the outcome. The system confirmed that the data pipeline was operational, but it did not confirm that the pipeline was receiving the expected input. Without an external denominator—a count of total work performed across all tools—the system could not recognize that its view of the world had shrunk.
Why it matters
For teams running their own software, this scenario highlights the risk of static monitoring in dynamic environments. As infrastructure and workflows evolve, hardcoded paths and assumptions become liabilities. If your backup script runs successfully but backs up an empty directory because the data source moved, your monitoring will likely stay green until you attempt a restore. The absence of errors is not evidence of success.
This incident also demonstrates the danger of over-relying on internal consistency checks. Metrics like "zero orphan tool results" are valuable for data integrity but useless for coverage. A system can be perfectly consistent while being completely irrelevant. Engineers must ensure that their health checks include validity constraints on data volume and freshness relative to external reality, not just internal process states.
Furthermore, the silence of the failure made it harder to detect than a crash. A broken script generates alerts; a working script processing stale data generates confidence. Teams need to design monitors that fail loudly when data stops arriving, rather than quietly accepting whatever is available. This requires decoupling the definition of "healthy" from "executed without error."
What you can do
- Implement heartbeat monitoring for critical jobs that requires an explicit signal from within the process, not just a successful exit code.
- Add external denominator checks that compare ingested data volume against known sources of truth, such as file counts in all relevant directories.
- Configure staleness alerts based on the timestamp of the newest data record, not the last modification time of the database file.
- Review routing policies and tool changes to ensure monitoring scopes are updated whenever data sources shift or expand.
- Test monitoring logic by simulating data starvation to verify that low-volume scenarios trigger appropriate warnings.
- Decouple alerting from the primary tool usage by sending notifications to channels independent of the monitored workflow, ensuring visibility even when the primary tool is idle.



