AI

Grok Down for Three and a Half Hours, SpaceXAI Apologizes

SpaceXAI apologizes for 3.5-hour Grok outage due to Memphis failure. Claude and ChatGPT also had issues at the same time.

6 min read Reviewed & edited by the SINGULISM Editorial Team

Grok Down for Three and a Half Hours, SpaceXAI Apologizes
Photo by Taylor Vick on Unsplash

SpaceXAI said on September 3, 2026, that its conversational AI Grok was down for about three and a half hours due to a failure at its Memphis compute facility. The company apologized to Grok users and to users of its compute resources. It said all systems have been restored and are operating normally. Reporting by staff@engadget.com (Karissa Bell) of Engadget noted that multiple AI providers experienced issues at the same time.

We are sorry for the issues you may have experienced with Grok following an outage at our Memphis compute center this morning. We’d also like to apologize to our impacted compute partners. All systems have now been restored and are functioning nominally.

The above is the official statement from SpaceXAI. It referred to the morning outage and recovery, and to future corrective action. It did not address details of the cause itself.

Memphis Compute Facility Outage and Recovery

The outage began at 6:30 a.m. U.S. Pacific Time. It was recorded as a “models outage” on the status page. The downtime lasted three and a half hours. The impact extended to Grok on X as well as the Android and iOS apps. Models remained unresponsive. Users were unable to use assistance features for posts and search.

The Memphis compute facility is a core site for SpaceXAI. The company said the failure at the site was the direct trigger. It has not disclosed the specific point of failure or the sequence of events. It is also unclear whether the issue involved power, cooling, or networking systems. It said normal operation was confirmed after recovery. Elon Musk said corrective measures will be taken to prevent a recurrence.

How outage notifications are issued affects trust in providers. Update times and details on status pages serve as primary sources. In this case as well, the start time and the declaration of recovery were clear. However, with no cause stated, verification is difficult. As shown in Microsoft Defender Privilege Escalation Vulnerability “RoguePlanet” Disclosed, clearly stating the scope of impact and countermeasures is important in disclosing vulnerabilities and outages. The same level of disclosure is required for AI infrastructure.

Claude and ChatGPT Also Hit at Same Time

On the morning of September 3, Anthropic’s Claude also experienced issues. It began at around 6:30 a.m. Pacific Time. The status page reported elevated error rates across multiple models. The affected models were Claude Mythos 5.1, Claude Fable 5.1, and Claude Opus 5. It was shown as resolved by 9:16 a.m. It is unclear whether the cause is related to SpaceXAI’s facility.

Anthropic signed a contract this year to borrow compute resources from Musk’s AI company. The scale of the contract and which parts are used have not been disclosed. There is also no identification of the users involved this time. Therefore, a causal link between the two outages cannot be confirmed. Anthropic has not responded to press inquiries.

OpenAI also had an outage that same morning. Elevated error rates occurred for ChatGPT and Codex. User reports of issues increased from around 7:30 a.m. This is based on records from an external outage-tracking site. The company said it was resolved by 9:55 a.m. It has not provided details on the cause. OpenAI has also not responded to press inquiries.

The outage times for the three companies are close together. The start times for Grok and Claude match. The increase in ChatGPT reports also falls within an hour later. The simultaneity suggests shared compute resources, but there is no evidence. An overlap of independent, coincidental events cannot be ruled out. As shown in OpenAI to Shut Down AI Browser Atlas, Move Features to Chrome Extension and App, OpenAI continues to restructure its application side. Stability on the infrastructure side and changes on the application side need to be evaluated separately.

Compute Lending Structure and Concentration Risk

The key point this time is the apology to “compute partners.” SpaceXAI apologized to users of its compute resources. No users were named. This highlighted the fact that lending resources externally exists as a business. Large-scale training and inference require huge facilities. Only a limited number of providers own the full capacity themselves.

Lending resources has advantages in cost and speed. Spare capacity can be shared to handle fluctuations in demand. Inference capacity can be increased without waiting for training of new models. However, dependence on a single site carries risks. A site failure can spread to multiple providers. This simultaneous failure recalled that structure.

Growing demand for compute is also supported by research. 3D reconstruction and video generation, such as in Streaming 3D Reconstruction Achieves SOTA with Geometric Context Transformer, require massive computation. The same applies to training and inference of language models. Expanding demand is driving consolidation of facilities. Consolidation increases both efficiency and vulnerability.

The opacity of contract structures is also an issue. Who uses which section of which facility is often undisclosed. Users cannot identify the cause during an outage. Decisions to switch to alternative routes are also delayed. Cross-referencing status notices becomes important. Comparing time records from each company is the starting point.

Cause Undisclosed, Focus on Prevention

SpaceXAI has not disclosed the cause. Musk’s explanation of corrective measures is also abstract. Specific recurrence-prevention measures are unknown. It has not indicated whether redundancy will be strengthened or operational procedures changed. There is also no means for third-party verification. Users have no choice but to accept the declaration of recovery.

The explanations from Anthropic and OpenAI are also limited. They only refer to elevated error rates. There is no distinction between an internal failure and an external resource issue. Neither company has responded to inquiries about the cause. Discrepancies between user-side records and official times also remain. Determining the cause will require waiting for future disclosures.

In large-scale outages, the granularity of information matters. Start and end times were disclosed. Affected model names were also shown on Claude’s side. However, the point of failure and the path of impact remain unknown. Whether there will be compensation or a report has also not been indicated. Contracts between providers may be constraining outage information.

Editorial Opinion

In the short term, we expect scrutiny of externally sourced compute resources to increase. The matching start times revealed where concentration lies. Securing alternative facilities and confirming switchover procedures will become key issues. We also see a need for standardized wording of times and affected targets in status notices.

In the long term, we see distributed design of compute facilities becoming a competitive advantage. Dependence on a single site is an operational weakness. We expect designs supporting multiple models and multiple infrastructures to spread among users as well. The scope of contract disclosure will also be subject to review.

The remaining question is where to set the level of disclosure for outage information. The point of failure and the scope of affected users remain unclear. No format for third-party verifiable reporting has been established. Without verifiable disclosure, maintaining trust will be difficult, in our view.

References

Frequently Asked Questions

When did the Grok outage occur and how long did it last?
It began at 6:30 a.m. Pacific Time and lasted about three and a half hours. Grok on X and the Android and iOS apps were affected. It was recorded as a models outage on the status page.
Are the Claude and ChatGPT issues related to SpaceXAI?
Claude saw elevated error rates at the same time and was marked resolved by 9:16 a.m. ChatGPT and Codex also saw increased issues from around 7:30 a.m. and were resolved by 9:55 a.m. None of the companies has indicated a causal relationship.
Did SpaceXAI explain the cause of the outage?
It said a failure at the Memphis compute facility was the trigger and declared recovery. It has not disclosed the specific point of failure or sequence. Musk said corrective measures will be taken to prevent recurrence.
Source: Engadget

Comments

← Back to Home