CircleCI Platform Status

Real-time health of CircleCI's own infrastructure and services.

Elevated wait times for machine jobs

Incident Report for CircleCI

Postmortem

Summary

Since October 1, CircleCI customers have experienced multiple incidents which have caused delays and failures in customer pipeline execution. There are two (unrelated) causes, both of which the CircleCI engineering team is actively mitigating:

  • Available Cloud Computing Capacity
  • Internal Infrastructure

Available Cloud Computing Capacity

Internal Infrastructure

  • Problem: We assign customer workflows over multiple compute providers via a workflow orchestration service. This service is currently gated by the write throughput of a single database.
  • Incidents:

  • Mitigations and Resolutions:

    • By October 23, 2026, we plan to implement multiple redundant workflow orchestration service instances, each with its own database, which will enable us split the orchestration load between them.

Our Commitment

For avoidance of doubt,

  • These particular incidents are not related to ongoing outages at Github
  • We are not currently migrating our infrastructure or services
  • We are not currently under cyberattack

Customer experience is our top priority, and we commit to continually improving the reliability of our systems to match the trust that our customers place in us. We thank our customers for their patience while our team worked to resolve these incidents and works to prevent future incidents.

Please reach out to our support team with any questions or concerns.

Posted Oct 08, 2026 - 17:54 UTC

Resolved

Between 12:40 UTC and 16:00 UTC on October 8, customers using Linux machine jobs and remote Docker experienced elevated wait times. The issue has been resolved and wait times have returned to normal. We thank you for your patience while our team worked on implementing a fix.
Posted Oct 08, 2026 - 16:35 UTC

Monitoring

Wait times for customers using Linux machine jobs and remote Docker have returned to normal. We are monitoring to confirm wait times remain stable while we continue to add capacity. We will provide another update by 16:30 UTC.
Posted Oct 08, 2026 - 16:09 UTC

Update

Wait times continue to decrease, but customers using Linux machine jobs and remote Docker are still experiencing delays. Wait times currently average about 90 seconds, with the longest waits up to about 6 minutes on some resource classes. We are working to add capacity as quickly as possible. We will provide another update by 16:30 UTC.
Posted Oct 08, 2026 - 15:59 UTC

Update

Wait times continue to decrease, but customers using Linux machine jobs and remote Docker are still experiencing delays. Wait times currently average about 6 minutes. The longest waits, up to about 20 minutes, are on the 2xlarge, arm.2xlarge and gpu.nvidia.small resource classes. We are working to add capacity as quickly as possible. We will provide another update by 16:00 UTC.
Posted Oct 08, 2026 - 15:31 UTC

Update

A fix has been deployed and wait times are decreasing, but customers using Linux machine jobs and Remote Docker are still experiencing delays. Wait times currently average about 11 minutes, with the longest waits exceeding 35 minutes on some resource classes. We are working to add capacity as quickly as possible. We will provide another update by 15:30 UTC.
Posted Oct 08, 2026 - 15:01 UTC

Update

Customers using Linux machine jobs and remote Docker are experiencing elevated wait times. Wait times have started to decrease and now average about 20 minutes, with the longest waits exceeding 40 minutes on some resource classes. Our engineers have identified the issue and are working on a fix. We will provide another update by 15:00 UTC.
Posted Oct 08, 2026 - 14:37 UTC

Update

Customers using Linux machine jobs are experiencing elevated wait times, averaging about 40 minutes, with the longest waits exceeding 50 minutes. Most Linux machine resource classes are affected, including medium, large, xlarge, 2xlarge and their Arm equivalents. Our engineers have identified the issue and are working on a fix. We will provide another update by 14:30 UTC.
Posted Oct 08, 2026 - 14:04 UTC

Update

Customers using Linux machine jobs are experiencing elevated wait times, averaging about 11 minutes, with the longest waits exceeding 30 minutes on the medium, arm.medium and arm.large resource classes. Our engineers have identified the issue and are working on a fix. We will provide another update by 14:00 UTC.
Posted Oct 08, 2026 - 13:34 UTC

Identified

Customers may be experiencing elevated wait times for machine jobs. We are working to resolve this.
Posted Oct 08, 2026 - 13:00 UTC
This incident affected: Machine Jobs.