Elon Musk's Team Will Build a Startup in 72 Hours Using AI Agents on Livestream
Elon Musk's artificial intelligence company will attempt something rarely seen in public: building an entire startup in 72 hours using AI agents, with every decision streamed live to the world. Three SpaceXAI employees will start with a blank slate on September 15 and decide the company name, product idea, and business plan in real time as they work.
What Is Grok Bot Galaxy and How Will It Work?
The event, called Grok Bot Galaxy, runs September 15 through 17 at a San Francisco venue with a worldwide livestream. Sessions run from 8:30 a.m. to 6 p.m. Pacific time each day. The three builders are Matt Palmer, Lauren Tan, and Roshan Sadanani, and they will use Grok Bot for ideation, planning, product calls, and engineering work. Colleagues will drop in for separate sessions on sales, customer support, and marketing operations.
It's important to note that Grok Bot is not the same as the Grok chatbot that answers questions on X (formerly Twitter). Grok Bot is SpaceXAI's agent product, described as a team of AI agents that users can name, assign objectives, and set loose across apps and websites.
Why Is SpaceXAI Staging This Public Test?
SpaceXAI was formed after SpaceX acquired xAI in February in an all-stock deal that valued xAI at $250 billion. The company then bought coding firm Cursor for $60 billion in August. Musk has spent the year making large claims about Grok's capabilities, and rival AI labs have matched his assertions. A three-day unedited livestream build is a harder test to stage than a polished demo clip, making it a more credible demonstration of the technology's real-world performance.
However, the livestream remains a vendor showing off its own tool rather than an independent trial. The truly useful measure will be how much the three human employees steer, edit, and rescue the AI agents along the way, revealing whether the agents can truly operate autonomously or require constant human intervention.
How to Evaluate AI Agent Capabilities During the Livestream
- Human Intervention Level: Watch how often the three employees need to step in, correct, or redirect the AI agents versus how much work the agents complete independently without human guidance.
- Product Viability: Assess whether the final product is a working application with actual users or merely a company name and landing page, which would indicate very different levels of AI capability.
- Decision Quality: Evaluate whether the AI agents make coherent business decisions about product-market fit, pricing, and go-to-market strategy or if these require human judgment.
- Real-Time Problem Solving: Observe how the agents handle unexpected challenges, technical obstacles, or pivots that typically arise during startup development.
Questions about who answers for AI agents when they act alone remain unsettled in the broader tech industry. This livestream may provide some practical answers about accountability and autonomy.
What's the Timing Challenge for Grok?
The 72-hour build lands in an awkward week for agent products. Rival AI lab Anthropic published a report on September 11 documenting how people misused its Claude models, covering cyber operations, surveillance, fraud, and conventional weapons work. Anthropic said it removed the accounts involved.
Meanwhile, reports suggest that every side in the Gulf conflict involves a Claude user, a nuance that Elon Musk has acknowledged. Musk conceded that Grok is not yet the default choice in this arena, stating that he is "not sure how to feel about that". Absence from a rival's abuse report is not a safety record, since labs only publish what they catch on their own systems, and Grok Bot is barely a month old in public.
By Thursday, the livestream should show what the word "company" actually covers in this context. A working product with paying users would land differently from a company name and a landing page, and that distinction will matter significantly for how investors and competitors evaluate Grok Bot's real-world utility.