DevOps & monitoring

Why atomic file writes fail to protect retry budgets in concurrent jobs

A developer found that atomic JSON writes did not prevent duplicate API calls in a Python pipeline, requiring a shift to pre-call reservation and locking.

Illustration showing that atomic writes do not secure the entire retry process chain.
Illustration created for this article

A software engineer discovered that using atomic file writes in a Python comment processing pipeline failed to prevent duplicate external API calls during concurrent execution. Published on October 11, 2026, the analysis details how crash recovery and race conditions bypassed local safety measures, leading to wasted retry budgets. The author subsequently implemented a locking mechanism and state reservation system to enforce strict attempt limits across multiple entrypoints.

What happened

The engineer maintains a pipeline that attaches AI model verdicts to comments on dev.to. When the model returns malformed JSON or the provider is overloaded, the system retries the request using a cached prompt. To control costs and load, each comment has a strict budget of three attempts. The pipeline runs via two entrypoints, one for Claude and one for Codex, which can execute the same canonical source code simultaneously. Initially, the developer ensured data integrity by using atomic writes: saving to a temporary file and then replacing the original. This technique guarantees that readers never see a partially written or corrupted JSON document.

Despite these safeguards, the retry budget was frequently exceeded. Regression tests using real subprocesses and a fake model revealed two specific failure modes. In the first scenario, the model successfully generated a response, but the subsequent step to save the verdict to disk failed. In the second, the process handling the model call was terminated unexpectedly. In both cases, a queued recheck command would read the disk, find the state unchanged from before the failed attempt, and invoke the model again. The log showed two calls for what should have been a single retry, proving that atomic writes protected file structure but not logical state consistency.

Key details

  • Atomic writes ensure a file is either complete or absent, but they do not track whether an external action, such as an API call, actually occurred.
  • The pipeline involves a sequence of reading state, calling an external model, and writing results to three separate files, which cannot be treated as a single transaction.
  • Failures occurring between the model call and the final save erase all evidence that the attempt took place, causing subsequent processes to repeat the work.
  • The fix introduces a shared pipeline lock with a five-second timeout, ensuring only one process handles a specific comment at a time.
  • Attempt counts are now reserved atomically in an inbox file before the model is contacted, so crashed attempts are still counted against the budget.
  • Validation included 38 focused concurrency tests and a full suite of 223 tests, confirming that recovered failures no longer trigger duplicate calls.

Background

Atomic writes are a common strategy in systems programming to prevent data corruption. By writing to a temporary location and then renaming or moving the file to its final destination, developers ensure that other processes never read a half-written file. This is crucial for configuration files or state records where partial data could cause crashes. However, atomicity applies only to the file operation itself. It does not provide transactional guarantees across multiple steps or external services. In distributed systems or concurrent applications, maintaining consistency often requires coordinating state changes across multiple resources, which simple file replacements cannot achieve alone.

When a process interacts with external services like AI APIs, the cost is incurred at the moment of the request, not when the result is saved. If a system crashes after the request but before the result is persisted, the external service has already charged for the call. Without a mechanism to record the intent to call the service before the call happens, the system loses track of its spending. This gap between internal state persistence and external side effects is a frequent source of bugs in batch processing and job scheduling systems.

Why it matters

For teams running self-hosted software or managing CI pipelines, understanding the limits of atomic writes is critical for reliability and cost control. Many automated tasks rely on retry logic to handle transient failures. If these retries are not properly synchronized, systems can enter infinite loops or exhaust API quotas unnecessarily. This is particularly relevant for small and mid-sized companies using paid third-party services, where every duplicate call directly impacts the bottom line. Ensuring that retry budgets are respected requires more than just safe file I/O; it demands careful orchestration of state and external interactions.

Furthermore, as more development workflows incorporate AI models, the complexity of managing these interactions increases. Models can be slow, expensive, and prone to rate limiting. A pipeline that inadvertently doubles its API usage due to poor concurrency handling can quickly hit rate limits or incur unexpected charges. Engineers must design systems that assume failure at any point in the process. By reserving resources before use and using locks to prevent concurrent access, teams can build more resilient automations that behave predictably even under stress or partial failure conditions.

What you can do

  • Audit your retry logic to ensure that attempt counts are incremented before external calls, not after successful completion.
  • Implement cooperative locking mechanisms for shared resources to prevent multiple processes from acting on the same stale state simultaneously.
  • Use short timeouts for lock acquisitions to avoid deadlocks, treating a timeout as a hard error rather than silently skipping the task.
  • Design state updates to be idempotent where possible, allowing later runs to reconcile differences without repeating expensive operations.
  • Test concurrency scenarios explicitly by simulating process crashes and parallel executions to verify that budgets and limits are enforced.
  • Separate the reservation of resources from the execution of tasks, ensuring that a failed task still consumes its allocated quota to prevent runaway retries.

More news

All news