Back to overview
Resolved

US datacenter: endpoint sensors unable to connect (2026-09-19)

Sep 19, 2026 at 9:50pm UTC
Affected services
Endpoint Agent

Resolved
Sep 19, 2026 at 11:16pm UTC

Resolved. Sensor connectivity to the US datacenter was restored at 23:13 UTC and traffic was back to normal levels by 23:16 UTC. Some sensors took until approximately 23:25 UTC to reconnect.

What happened (all times UTC, 2026-09-19)

  • 20:21 - A scheduled automatic upgrade began replacing the servers that receive sensor connections in our US datacenter, one at a time.
  • ~21:50 - The network load balancer in front of those servers was not updated with the replacement servers and kept sending connections to ones that had been removed. Sensors lost their connection progressively as the upgrade advanced.
  • 22:17 - Our sensor connectivity alert paged the on-call team.
  • 23:13 - We forced the load balancer to resynchronize. Traffic returned to normal by 23:16.

Impact

US datacenter only, approximately 80 minutes. Sensors kept running on their hosts and buffered telemetry locally while disconnected. That local buffer is bounded, so hosts with higher event volume may show gaps in telemetry for this window. Detection & Response rules did not run on telemetry that was not delivered. The web application, APIs, and non-sensor data sources (adapters, cloud integrations) were not affected.

Root cause

A load balancer configuration setting (a 60-minute connection drain timeout) caused the component that keeps the load balancer's server list up to date to fall far behind, so the replacement servers were never registered.

What we are doing

  • The configuration that caused the delay has been corrected in all datacenters, including the US datacenter (completed 2026-09-22 04:00 UTC).
  • We are moving the sensor load balancer to a newer design that tracks server membership directly.
  • We are adding monitoring that tests the sensor connection path end to end, and tightening existing connectivity alerts so they fire sooner and stay open until service is fully restored.
  • We are connecting that end-to-end check to this status page so the Endpoint Agent component reflects sensor connectivity.

We apologize for the disruption.

Created
Sep 19, 2026 at 9:50pm UTC

This incident is being reported retroactively.

Starting at approximately 21:50 UTC on 2026-09-19, endpoint sensors connecting to our US datacenter progressively lost their connection to the LimaCharlie cloud. By 22:08 UTC nearly all sensor traffic to the US datacenter was affected. Other datacenters (Canada, Europe, UK, India, Australia) were not impacted.

Our connectivity alerting paged the on-call team at 22:17 UTC and engineers began investigating.