About This Site

Welcome to our status page. If you are looking for help, please check our documentation guides or contact us on our community forum. All products listed below have a target availability of 99.9%.

Current UiPath Status

All Systems Operational

Maestro workflow intermittent failures occuring

Incident Report for UiPath

Postmortem

Customer Impact

Between 14:03 and 21:26 UTC on June 18, 2026, and between 17:02 and 23:59 UTC on June 23, 2026, a subset of customers in the U.S. Region, using Maestro and Agent executions—including Conversational Agents and some previously deployed Autonomous Agents—experienced performance degradation and workflows getting stuck in pending state. The two events, a week apart, shared the same underlying root cause. In each case, new workflow operations were stabilized before all pending workflows were fully recovered. During these windows, affected workflows could experience higher-than-normal latency, fail intermittently, or remain stuck in a pending state until recovery processing completed.

Root Cause

The incident was caused by our workflow execution backend exhausting its configured memory capacity and restarting repeatedly. Those restarts disrupted workflow state processing, causing some workflow executions to experience latency, fail intermittently, or remain pending.

The primary driver of memory growth was an experimental multi-cluster (active/passive) replication feature for disaster recovery purpose. This feature accumulated large, long-lived workflow-history objects in memory faster than they could be released. Each time a history-processing component exceeded its memory limit and restarted, it had to reload its in-flight work—including the replication workload — which drove memory back over the limit and triggered another restart. This self-reinforcing cycle continued until responders intervened, and it was not fully resolved by the initial mitigation applied on June 18, which is why the condition recurred on June 23.

Detection

Both events were detected quickly through automated monitoring, including synthetic tests that continuously exercise the platform and alerts for pending/stuck workflow tasks. At the time of the incidents, there was no dedicated alerting on the memory saturation and restart behavior of the affected components, which lengthened the time required to identify the true root cause. Public status updates began after customer impact was confirmed.

Response and Recovery

A cross-team engineering bridge was established to investigate and mitigate the incidents. Early mitigation attempts focused on scaling backend infrastructure and restarting the affected history-processing components. These provided limited relief because they did not address the underlying memory-growth mechanism.

Once the memory-retention path was pinpointed through diagnostic heap analysis, the incidents were mitigated by two actions: increasing the memory capacity allocated to the affected history-processing components, and disabling the multi-clusters replication feature that was the main contributor to the memory growth. After these changes, new workflow operations returned to normal and the region's capacity to process workflows was restored.

Workflows that had become stuck in a pending state were recovered by the engineering team where possible. A small number of workflows that customers had already canceled were not replayed, and a limited set of Agent executions that had failed during the incident had to be retried manually by customers. Both incidents were marked resolved after all recoverable workflows had been processed successfully.

Follow-Up

To prevent recurrence and improve detection, we are implementing the following:

  1. Right-size the memory capacity and limits for the affected components, and add automatic memory-cap safeguards so the components stay within safe bounds under load.
  2. Redesign and re-validate the cross-region replication feature—including load testing with replication enabled—before it is re-enabled, so it can support disaster-recovery needs without causing memory pressure.
  3. Add repeatable memory-diagnostic collection capability and operational runbooks for the service, so memory analysis and mitigation can be performed quickly during future incidents.
Posted Jul 09, 2026 - 23:10 UTC

Resolved

The issue has been resolved and Maestro workflow performance has returned to expected levels after degraded performance impacted Maestro workflow in US.
Impact: No ongoing user impact. Agent executions that failed will have to be retried manually.
Posted Jun 18, 2026 - 21:25 UTC

Update

We have resolved the situation that was causing the degraded performance for Maestro workflows.
Impact: Users may still experience degraded performance, but this should continue to improve. We are monitoring closely to ensure stability.
Posted Jun 18, 2026 - 21:03 UTC

Update

We have resolved the situation that was causing the degraded performance for Maestro workflows.
Impact: Users may still experience degraded performance, but this should continue to improve.
We are monitoring closely to ensure stability.
Posted Jun 18, 2026 - 19:13 UTC

Update

Mitigation has been applied and performance is improving for the issue that impacted [impacted functionality / user symptom] in Maestro across US.
Impact: Users may still experience degraded perfornance but this should continue to improve.
We are monitoring closely to ensure stability.
Posted Jun 18, 2026 - 17:48 UTC

Update

We have resolved the situation that was causing the degraded performance for Maestro workflows.

We had to scale up the capacity for the underlying services. We are actively monitoring the backlog of requests for latencies and will take the action as needed.
Posted Jun 18, 2026 - 17:46 UTC

Monitoring

We have resolved the situation that was causing the degraded performance for Maestro workflows.

We had to scale up the capacity for the underlying services. We are actively monitoring the backlog of requests for latencies and will take the action as needed.
Posted Jun 18, 2026 - 16:39 UTC

Investigating

We are investigating reports of degraded performance impacting workflows for Maestro in east US.
Impact: Users may experience intermittent failures or stuck workflows.
Our teams are working to identify the cause and will share more details as the investigation progresses.
Posted Jun 18, 2026 - 15:00 UTC
This incident affected: United States (Context Grounding, Agentic Orchestration, Agents).