Welcome to our status page. If you are looking for help, please check our documentation guides or contact us on our community forum. All products listed below have a target availability of 99.9%.
Uptime over the past 90 days. View historical uptime.
Between September 2, 2026 at 7:59 pm UTC and September 4, 2026 at 4:30 pm UTC, the Insights dashboard did not display Maestro process runs created after the start of the affected window. The underlying data was never lost. It was held safely in our staging layer and became visible once processing resumed.
The United States, Singapore, India, Canada, Australia, Japan, and United Kingdom regions were affected, with the exact duration varying by region from approximately 24.6 hours to 44.5 hours. The European Union region hit the same failure but recovered on its own within about 45 minutes, as described under Response below. Element-level Maestro data continued to update normally throughout, so the dashboard remained usable and did not appear broken; what was affected was the count and listing of recent process runs.
A data processing task that moves Maestro process-run records from our staging layer into the database serving the Insights dashboard stopped running.
The cause was a deployment sequencing problem. Two related changes are delivered by two separate deployment pipelines: one adds a new field to the staging layer, the other updates the processing task that reads it. In production the second pipeline ran approximately two hours before the first. The updated task therefore referenced a field that did not exist yet, and it failed on each scheduled run. Our data platform automatically suspends a task after a number of consecutive failures, which it did. The missing field was added shortly afterwards, but a suspended task does not resume on its own, so processing stayed stopped until we resumed it manually.
Our pre-production environments received the same two changes in the correct order, so this condition never appeared there.
Our monitoring did detect the task failures automatically, within 2 to 10 minutes of the first failure in each region, and notified our on-call team.
That notification then cleared on its own. The alert measures how often the task is failing, and once the platform suspended the task it stopped running altogether and therefore stopped reporting failures. The signal returned to normal while the data was still not moving. As a result the condition was not escalated at the time, and the full extent of the delay was identified after a customer reported the missing data on the morning of September 4. A formal incident was declared later that day.
After the customer report, our engineering team traced the delay to the suspended processing task within about an hour and resumed it for the first affected region immediately. We then reviewed each production region that received this deployment, one at a time, resumed the task in each affected one, and verified recovery in each case by confirming three things: the task was running again, the dashboard database had caught up to the staging layer, and no delayed records remained.
All 123,090 delayed records were replayed automatically on resume. No manual data reconstruction or backfill was required, because the task resumes from where it stopped rather than from the current time. Recovery was complete across all affected regions by 4:30 pm UTC on September 4.
The European Union region experienced the same failure but recovered on its own, because the missing field was added before that region's task reached the suspension threshold. Its data was delayed by up to 45 minutes.
We published a status page notice the same day, after recovery had completed. For an incident of this shape, where data is delayed rather than unavailable, that notice came later than it should have. We have added a follow-up item to communicate during such incidents rather than after them.