Some systems are experiencing issues.
It has been a while since our last update, and so far, the situation appears to be stable. The underlying issue has been fixed during the maintenance, as explained in our last update.
There is currently no impact on users, and the situation is not considered critical. We have taken action to bring the cluster back into service and will continue monitoring it closely over the next week to ensure it remains stable.
As we cannot yet fully guarantee reliable service, we will avoid placing critical services on this cluster until the situation has been fully validated.
The incident will remain in "Watching" status, and all potentially affected services will be marked as "Under Maintenance", including those that are currently running on the cluster.
Once we have confirmed that everything is stable and the issue has been fully resolved, we will mark the incident as "Resolved" and return the cluster to "Active" status. This validation process may take up to one week.
The API issues were caused by some DNS records still pointing to the affected as well as not affected clusters and not being updated accordingly.
We have manually overridden the affected DNS entries to reduce the error rate back to normal levels. We will address the underlying issue during the next maintenance window, scheduled in a few days.
As a result, bots may have experienced issues starting during the period between the initial outage and the DNS fix. External requests to our public API may also have failed during this time.
We will leave this incident marked as "Identified" until Cluster 1 is back in service. However, no services should be affected while Cluster 1 remains unavailable.
We have received an alert indicating that Cluster 1 is partially unreachable. In addition, we are seeing an increase in API errors. We are investigating the issue.
Incident UUID 91decedd-caaa-4e9e-bfce-44552caee374