About This Site

Welcome to our status page. If you are looking for help, please check our documentation guides or contact us on our community forum. All products listed below have a target availability of 99.9%.

Current UiPath Status

All Systems Operational

Component Status Overview

Operational
Degraded Performance
Partial Outage
Major Outage
Maintenance
— Not available
Loading component status...

Uptime over the past 90 days. View historical uptime.

degraded-performance in MLS Observer thruough DU - US

Incident Report for UiPath

Postmortem

Customer impact

Between August 31, 2026 at 3:40 pm UTC and September 1, 2026 at 12:00 am UTC, a subset of customers using Document Understanding custom-model extraction and classification in the US regions experienced intermittent request failures (HTTP 503 "Server overloaded") and slow processing.

The impact was limited to document extraction requests served by custom models. It manifested in five waves of roughly twenty to forty minutes each, the heaviest between 5:10 pm and 5:50 pm UTC. During a wave the majority of requests to the affected models were rejected; between waves most requests completed normally. Document Understanding retried failed requests automatically, so most documents completed after a delay. A small share of documents submitted during the waves did not complete within the incident window and had to be resubmitted. No data was lost or exposed, and no incorrect results were returned. Extractions that appeared to return empty results, or remained in Running, were client-side timeouts inside the incident window

Root cause

Two independent defects combined under an ordinary volume of traffic.

The first was in our request de-duplication mechanism. A small number of very large documents occupied shared GPU capacity and slowed all processing on the affected pool, so ordinary documents began exceeding their processing deadline. When a request's deadline expired, the de-duplication layer treated the work as abandoned even though the worker was still processing the document. The automatic retry of that request was then admitted to a worker, where it correctly avoided reprocessing the document but waited for the original result while holding a unit of worker capacity. With several retries waiting per hot document, the workers ran out of capacity and rejected unrelated requests. Each rejection told clients to retry after five seconds, and Document Understanding's internal services also retried, which kept the workers saturated.

The second was in the shared result cache used to serve repeated requests without recomputation. A configuration defect had left the cache running at its storage limit since early August; in that state it acknowledged some writes and then discarded them. Documents that had already been processed successfully were therefore recomputed when retried, which added real processing load and restarted the cycle after each wave subsided.

Detection

Automated monitoring detected the failures at 4:10 pm and 4:14 pm UTC on August 31, approximately 34 minutes after the first rejections. Engineers began investigating at 4:32 pm UTC, a customer-impacting incident was declared at 4:47 pm UTC, and a public status page notice was posted at 4:55 pm UTC.

Response

At 4:38 pm UTC, GPU capacity was freed by scaling down idle workloads, which reduced the failure ratio from approximately 78% to approximately 50%. Automatic scaling of the affected pool had reached its configured maximum by 5:15 pm UTC. Additional capacity was freed at 5:21 pm UTC when a second wave began.

At 5:26 pm UTC the investigation established that total request volume was rising while the number of distinct documents was flat, indicating retry amplification rather than new customer load. At 5:44 pm UTC an internal retry layer was reduced to a single attempt, and by 5:59 pm UTC request volume had converged with the number of distinct documents. The same change was applied to the delayed US environment by 7:12 pm UTC. The status page was moved to monitoring at 6:24 pm UTC and to resolved at 7:39 pm UTC.

Failures recurred in three further, smaller waves at approximately 6:30 pm, 8:20 pm and 10:50 pm UTC as new large documents arrived, and cleared by 12:00 am UTC on September 1 once those documents completed and evening load declined.

A detailed analysis on September 2 identified the result cache defect. The cache configuration was corrected at 3:15 pm UTC on September 2, which stopped the loss of completed results, and further tuning on September 9 removed a remaining write stall; no discarded writes have been observed since.

Follow up

  1. Improve the request de-duplication mechanism so duplicate requests are rejected before they consume worker capacity. The new implementation is complete and its rollout to production is in progress.
  2. Result cache: capacity sizing for the US working set, explicit error reporting when a write cannot be stored, and monitoring and alerting on cache health. The configuration fix is deployed.
  3. Introduce per-tenant GPU usage limits so that a single workload of very large documents cannot saturate shared capacity for other customers.
  4. Return accurate retry guidance on overload responses and propagate it through Document Understanding's internal layers, and make the reduced internal retry policy permanent in configuration.
  5. Add leading-indicator alerting on request duplication and worker saturation, and an operational runbook for this overload pattern, so it is recognised before customers are affected.

Customers processing very large documents are encouraged to use asynchronous submission where available and to honour the Retry-After header on 503 responses.

Posted Sep 17, 2026 - 21:29 UTC

Resolved

The issue has been mitigated, and the service is currently operational. We will continue to closely monitor the service health and take further action if needed.
Posted Aug 31, 2026 - 19:39 UTC

Monitoring

The issue has been mitigated, and the service is currently operational. We will continue to closely monitor the service health and take further action if needed.
Posted Aug 31, 2026 - 18:24 UTC

Identified

We have identified the issue and implemented the necessary fix
Posted Aug 31, 2026 - 17:07 UTC

Investigating

We are currently investigation the issue.
Posted Aug 31, 2026 - 16:55 UTC
This incident affected: United States (Document Understanding) and Delayed US (Document Understanding).