DevOps & monitoring

Designing resilient edge observability with AWS and Prometheus

A detailed architecture guide explains how to monitor thousands of edge devices using heartbeats, serverless ingestion, and VictoriaMetrics for scalable time-series storage.

Illustration of edge devices sending data to a central monitoring dashboard
Illustration created for this article

Viren Patel published a technical guide on October 6, 2026, detailing how to build a production-grade observability pipeline for edge devices. The article outlines an architecture that combines AWS serverless components with Prometheus-compatible metrics to monitor the health and liveness of distributed hardware at scale.

What happened

The guide addresses the challenge of monitoring thousands of edge devices deployed across various customer locations. These devices run local software dependent on network connectivity, hardware health, and correct configuration. Without active monitoring, a device can silently go offline, lose local network access, run low on storage, or fall out of compliance with operating system requirements. The proposed solution answers two critical questions for operators: is the device alive, and is it healthy?

The architecture uses two distinct ingestion paths that converge into a single time-series backend. The first path involves direct heartbeats, where each edge device pushes a small payload on a predictable cadence. The second path uses a scheduled Lambda function to poll an external device management API. This pull-based approach enriches the data with hardware and compliance telemetry that devices do not push themselves, such as CPU temperature or battery health. Both streams are normalized into Prometheus-format metrics before storage.

The system relies heavily on AWS serverless components to handle scale without managing always-on infrastructure. API Gateway and Lambda authorizers handle authentication and routing, ensuring that device serial numbers are validated against the caller’s identity. Simple Queue Service (SQS) buffers incoming heartbeats, protecting the system from bursts and allowing for retry logic. A dead-letter queue captures malformed or failed messages for investigation, preventing data loss. Finally, a dequeue Lambda converts these payloads into metrics and writes them to the backend.

Key details

  • Metrics Backend: The system uses VictoriaMetrics instead of standard Prometheus because it supports the remote-write protocol while offering lower costs and better performance for high-cardinality data.
  • Heartbeat Payload: Devices send minimal JSON payloads containing a serial number and a counter, optionally including network connectivity status.
  • Data Validation: The schema is strict, rejecting unexpected fields to prevent silent errors, and serial numbers are validated against authenticated identities.
  • Audit Trail: Every metrics payload is backed up to Amazon S3 with lifecycle expiry, allowing teams to replay or inspect raw data if needed.
  • Dual Alerting: Device health is monitored via Grafana alerts on VictoriaMetrics data, while pipeline health (queue age, error rates) is monitored via CloudWatch alarms.
  • Batching Strategy: The scheduled telemetry sync chunks output into sub-1.5MB batches with exponential backoff to stay within payload limits during full-fleet updates.

Background

Observability in distributed systems often relies on time-series databases, which store data points indexed by time. Prometheus is a popular open-source tool for this purpose, using a query language called PromQL. However, Prometheus can struggle with high cardinality, which occurs when metrics have many unique label combinations, such as thousands of unique device serial numbers. VictoriaMetrics is a compatible alternative designed to handle this scale more efficiently.

Serverless architectures, such as those built on AWS Lambda, allow code to run in response to events without provisioning servers. This model is ideal for ingestion pipelines that experience variable traffic, as the infrastructure scales automatically. Using queues like SQS decouples the ingestion of data from its processing, ensuring that temporary spikes in device reports do not overwhelm the metrics database.

Why it matters

For teams running their own software, especially those managing IoT or edge deployments, visibility is critical. A device that appears online but has lost its local network connection represents a different failure mode than one that is completely powered off. By separating liveness checks from deep telemetry, operators can diagnose issues faster. The architecture described ensures that these distinct signals are unified in a single dashboard, reducing the cognitive load on support engineers.

Reliability of the monitoring pipeline itself is often overlooked. If the ingestion system fails, dashboards may show stale data, leading teams to believe devices are healthy when they are not. By monitoring the pipeline’s internal health metrics, such as queue depth and Lambda error rates, teams can distinguish between a fleet-wide outage and a monitoring failure. This separation prevents false confidence and speeds up incident response.

The use of strict validation and dead-letter queues also protects data integrity. In large fleets, misconfigured devices or malicious actors could attempt to inject bad data. Validating serial numbers against authenticated identities and rejecting unknown fields ensures that the metrics remain trustworthy. The S3 backup provides an additional safety net, allowing teams to recover from processing errors without losing historical data.

What you can do

  • Implement strict schema validation for incoming telemetry to reject malformed data early in the pipeline.
  • Use a dead-letter queue to isolate failed messages for later analysis rather than dropping them silently.
  • Separate monitoring of the ingestion pipeline from monitoring of the devices to detect tooling failures independently.
  • Store raw payloads in object storage like S3 to create an audit trail that can be used for debugging or replay.
  • Normalize different data sources into a common metrics format, such as Prometheus gauges, to simplify querying and alerting.
  • Validate device identities against authenticated callers to prevent one device from spoofing another’s metrics.

More news

All news