Welcome to our status page. If you are looking for help, please check our documentation guides or contact us on our community forum. All products listed below have a target availability of 99.9%.
Uptime over the past 90 days. View historical uptime.
On , between approximately 15:45 UTC and 18:45 UTC, customers hosted in the Europe, Australia and Japan regions were unable to see recent robot logs in Orchestrator. Logs written during the impact window did not appear in the Jobs and Logs views or through the Robot Logs API, and logs older than approximately 03:00 UTC that day were available only intermittently. The impact window differed per region:
Job execution, queues, triggers and all other Orchestrator functionality were not affected. No robot logs were lost. Logs produced during the impact window continued to be collected and became visible again once the issue was mitigated.
We were performing a scheduled upgrade of the Elasticsearch service that stores Orchestrator robot logs. The upgrade uses a blue/green approach: a new cluster is created alongside the existing one, the data is restored into it from a backup, the pipeline that ingests new logs is started on it, and only then is traffic switched over.
In the affected regions, the traffic switch was approved by the engineer running the upgrade before the data restore and the ingestion catch-up had completed, which is a deviation from our upgrade procedure. This was a human error. As a result, Orchestrator was reading robot logs from a cluster that did not yet contain recent data, while new logs continued to be collected but were not yet indexed on the cluster serving requests. The error was not caught because our deployment tooling relies on the operator to verify data completeness and does not enforce the order between the data restore and the traffic switch.
Our synthetic monitoring, which continuously runs test jobs and verifies that their logs are returned, detected missing logs in the affected regions and raised alerts. Our engineering team correlated the alerts with the in-progress upgrade, declared an incident and published a status page notification.
Our engineering team reverted the traffic to the previous Elasticsearch clusters in all three regions, which immediately restored access to all robot logs, including those produced during the impact window. We confirmed recovery through synthetic monitoring in every affected region and continued monitoring before marking the incident resolved at approximately 18:45 UTC.
The upgrade of the new clusters continues in the background. Traffic will be moved to them only after the data restore is complete and ingestion is confirmed to be up to date.
To prevent this from happening again, we are: