Welcome to our status page. If you are looking for help, please check our documentation guides or contact us on our community forum. All products listed below have a target availability of 99.9%.
Uptime over the past 90 days. View historical uptime.
Between August 4, 2026, at 3:25 AM UTC and August 5, 2026, at 3:39 AM UTC, a subset of customers in the US region experienced degraded performance affecting Agent and Maestro workflow runs. This issue was isolated to workflow steps executing Large Language Model (LLM) calls in tandem with PII masking. The event lasted approximately 24 hours. Affected users experienced delayed or failed executions, characterized by extended runtimes, workflows exceeding the five-minute threshold, and timeout errors.
An upstream, third-party service powering our PII detection capabilities experienced a localized configuration anomaly within its US region. This restriction prevented the service from scaling dynamically to accommodate transaction volumes. The upstream provider verified the issue and applied a manual correction. The infrastructure constraint had caused requests involving PII detection and related LLM workflow steps to experience elevated latency or time out.
Initial support tickets regarding timeouts were received on August 4, 2026, at 3:25 AM UTC. Because failures were intermittent and a portion of requests continued to process successfully, an incident was not declared at that time. Following subsequent reports from multiple customers, a formal incident was declared at 6:27 PM UTC. Automated latency alerts did not trigger as intended, preventing proactive internal detection prior to customer escalation.
At 7:14 PM UTC on August 4, 2026, we published a public status update acknowledging the degraded Agent and Maestro workflow performance in the US region. By 7:19 PM UTC, teams identified a probable cause and initiated remediation efforts. An initial mitigation patch was deployed at 8:31 PM UTC, followed by close performance monitoring; however, customer feedback at 8:41 PM UTC and 8:50 PM UTC indicated persistent latency.
Continued investigation at 11:04 PM UTC confirmed ongoing elevated response times and timeout errors originating from the third-party service, prompting deeper escalation and joint investigation with the upstream cloud provider. At 2:06 AM UTC on August 5, 2026, the provider confirmed their scaling configuration issue was resolved. Telemetry indicated a downward trend in workflow durations by 2:22 AM UTC, with performance steadily normalizing over the subsequent three hours. The incident was officially closed at 3:39 AM UTC after Agent and Maestro metrics fully returned to baseline operational levels.