Welcome to our status page. If you are looking for help, please check our documentation guides or contact us on our community forum. All products listed below have a target availability of 99.9%.
Uptime over the past 90 days. View historical uptime.
Between July 22 and July 29, 2026, customers using Document Understanding in the Europe region experienced intermittent delays and, in a limited number of cases, failed document operations. The issue affected document extraction and splitting/classification operations using Helix 2.0-based models— other model types were not affected.
The impact occurred in three windows, each detected by our automated monitoring:
During these windows, operations that normally complete in under a minute could take several minutes, and a small number of operations failed after all retry attempts. Between the first two windows, a low rate of intermittent delays persisted. The issue was isolated to the Europe region and did not affect other geographies.
The root cause was a concurrency issue in the inference service that runs these models. Two stages of processing—preparing incoming requests and generating results—shared an internal component that only one operation can use at a time. Under high-load, operations queued for this shared component, slowing result generation even though compute capacity was available. Slow operations kept their processing slots occupied, so new requests queued behind them and some were rejected, and a small number of operations ultimately failed after exhausting all retry attempts. Large, image-heavy documents increased use of the shared component and made the slowdown worse.
Retried requests for the same document are automatically deduplicated, so retries did not multiply the processing work; however, rejected requests did not tell clients when to retry, which added some pressure during the constrained periods.
The July 29 recurrence happened because a protective configuration change applied after the earlier occurrences was unintentionally reverted during a routine deployment. It was re-applied the same day and has now been made permanent.
Automated alerts detected the initial issue at 5:14 pm UTC on July 22, approximately four minutes after failures began to rise. Monitoring showed elevated failed requests and longer processing times.
After the first window was resolved, our engineering team continued to track a low rate of intermittent delays through end-to-end tests and telemetry reviews while the investigation continued. A new automated alert detected the July 24 recurrence, and the July 29 recurrence was detected by automated alerts within minutes of onset.
On July 22, automated capacity scaling restored throughput and the failure rate returned to near-baseline levels. The status page was updated at 5:45 pm UTC, and after continued monitoring showed no further customer-visible failures, the incident was closed at 7:39 pm UTC. After closure, our engineering team continued investigating the underlying cause—that investigation was still in progress when the issue recurred on July 24.
On July 24, following the recurrence, our engineering team restarted the affected inference workers and applied a protective configuration change that reduces concurrent processing per worker and increases minimum capacity. The status page was updated at 9:22 am UTC, and the incident was resolved at 3:01 pm UTC after extended monitoring.
On July 29, engineers identified that the protective configuration had been unintentionally reverted, re-applied it, and restarted the affected workers. The status page was updated at 11:54 am UTC. Service was confirmed stable—including by affected customers—and the incident was resolved at 2:05 pm UTC.
We are taking the following steps to prevent recurrence and improve resilience: