DevOps & monitoring

Silent tracing gaps and disk bloat in self-hosted Langfuse

A developer found their self-hosted Langfuse instance was only tracing 7% of AI agent traffic due to configuration gaps, while ClickHouse system logs consumed excessive disk space.

Illustration of partial monitoring on a server rack
Illustration created for this article

A developer discovered that their self-hosted Langfuse observability platform was capturing only a fraction of their AI agent activity, leading to misleading performance data. The investigation revealed that configuration oversights caused silent monitoring failures across most agent profiles, while the underlying database accumulated gigabytes of unnecessary internal logs.

What happened

The issue came to light when an AI agent took twenty-five minutes to respond to a simple command to post a reply on a forum. Instead of executing the task, the model generated a fabricated summary of the target article. The developer consulted their Langfuse traces to diagnose the delay, expecting to see network latency or tool execution bottlenecks. The trace data showed that web tools executed in under five seconds, while the local language model spent over twenty-five minutes processing. The root cause was not performance but context: the agent session had started before the posting skill was available, so the model improvised rather than using a missing tool.

This incident prompted a deeper audit of the observability setup. The developer compared the number of sessions recorded in their agent’s internal database against the traces stored in Langfuse over a fourteen-day period. The agent had executed 1,387 model calls across 188 sessions, yet Langfuse contained only 618 events from 79 traces. Further analysis showed that tracing was enabled on only one of ten agent profiles, meaning the system was monitoring roughly 7% of total traffic. The most active profiles, including those for coding and security tasks, were completely invisible to the observability platform.

The second major finding involved storage consumption. Langfuse version 3 uses ClickHouse as its primary trace store, alongside PostgreSQL and Redis. While the actual trace data occupied just 2.2 MiB, the ClickHouse instance was using over 6 GiB of disk space. The majority of this space was consumed by ClickHouse’s own diagnostic system tables, such as trace logs and metric logs, which were writing continuously. This excessive logging created significant input-output overhead on the host machine, contributing to periodic performance storms on the hard drive.

Key details

  • Tracing was active on only one of ten agent profiles, capturing approximately 7% of all model calls.
  • The untraced profiles included high-volume agents for coding, security, and IT administration tasks.
  • ClickHouse system logs consumed 6.03 GiB of disk space, while actual trace data used only 2.2 MiB.
  • The tracing plugin fails silently when API keys are missing, providing no error messages or warnings in logs.
  • One-shot jobs running in safe mode bypass plugins entirely, requiring separate SDK integration for tracing.
  • Removing ClickHouse system log tables via configuration stops new writes but does not delete existing data automatically.

Background

Langfuse is an open-source observability platform designed for large language model applications. It helps developers track costs, latency, and user feedback by recording traces of interactions between agents and models. Self-hosting Langfuse gives teams control over their data but requires managing the underlying infrastructure, including the database layer. In version 3, Langfuse relies on ClickHouse, a column-oriented database management system optimized for online analytical processing. ClickHouse is powerful for handling large volumes of time-series data but includes extensive internal logging features that monitor its own performance and operations.

Fail-open design patterns are common in software plugins to ensure that a missing dependency does not crash the main application. In this context, if the tracing plugin cannot find valid API credentials, it simply disables itself rather than throwing an error. While this prevents application crashes, it creates a blind spot where monitoring stops without alerting the operator. Understanding the distinction between application-level errors and silent configuration gaps is critical for maintaining reliable observability in distributed systems.

Why it matters

For teams running their own AI infrastructure, incomplete observability can lead to incorrect conclusions about system performance and cost. In this case, the developer initially suspected network issues due to the long response time, but the trace revealed the problem was model reasoning within a stale session. If the tracing had been active on the coding agent, which accounted for the majority of calls, the team would have had visibility into how often local models hit latency tails compared to cloud alternatives. Without comprehensive tracing, resource allocation decisions are based on habits rather than data, potentially leading to inefficient use of expensive GPU resources or cloud APIs.

Storage management is another practical concern for self-hosted services. Diagnostic logs are useful for troubleshooting database performance, but they can quickly overwhelm small-scale deployments. When internal logs consume thousands of times more space than the actual application data, they degrade disk performance and increase backup times. For operators managing homelabs or small server clusters, unchecked log growth can cause service interruptions or require frequent manual cleanup. Configuring the database to retain only relevant operational data ensures that the observability platform remains lightweight and sustainable.

What you can do

  • Compare the number of requests recorded in your application logs with the number of traces in your observability platform to identify coverage gaps.
  • Verify that API keys and tracing plugins are configured for every agent profile, worker, and service, not just the default one.
  • Send a test request from each distinct agent profile and confirm that a tagged trace appears in the dashboard.
  • Inspect ClickHouse system tables to determine if diagnostic logs are consuming disproportionate disk space relative to your data.
  • Apply a custom configuration file to disable unnecessary ClickHouse system logs and set appropriate retention periods for query logs.
  • Schedule regular audits of tracing configurations, especially after infrastructure changes such as server address updates or credential rotations.

More news

All news