All affected models have been scaling correctly for more than 30 minutes at this point, and we see no residual prediction queues. Thank you for your patience!
We have been monitoring for a while and services seem to be back. Thank you for your patience
We have restarted the affected component and service should be back to normal. Thank you for your patience
We're looking stable now, and with caching back on. Thanks again for your patience!
This issue is resolved
We have not seen elevated rates of HuggingFace/model setup errors for a couple hours now, so we believe this incident is cleared
H100 capacity has returned to normal levels
We are back under max capacity for H100 hardware. Thank you for your patience!
The H100 capacity is back in good health. Thanks for your patience!
System is back to operating normally
This issue has been resolved and queue times are back to normal
Message flows are healthy.
H100 hardware contention has resolved. Thank you for your patience!
·