AI

Grok 4.6 and Hunyuan: The Co-design Loop Is Key to an AI Comeback

SpaceXAI has announced Grok 4.6. This report details how joint development with Cursor drove rapid AI model performance gains, and how Tencent adopted the same strategy with Hunyuan.

6 min read Reviewed & edited by the SINGULISM Editorial Team

Grok 4.6 and Hunyuan: The Co-design Loop Is Key to an AI Comeback
Photo by Igor Omilaev on Unsplash

SpaceXAI announced Grok 4.6 in the early hours of August 13, 2026. According to independent evaluations by Artificial Analysis, Grok 4.6’s intelligence index is 61 points, tied with ChatGPT-5.6 Sol and close behind Fable 5 Max’s 62 points. Its cost per task is $0.84, below Sol’s approximately $1.23 and Opus 5’s approximately $2.34.

A Sudden Surge After Five Months

Just a few months ago, Grok was on the verge of collapse. Grok 4.3 lagged far behind major models, and 10 of xAI’s 11 co-founders had resigned. Elon Musk publicly admitted, “The initial construction of xAI failed.” Grok was integrated into SpaceX, which was nearing its IPO, and restarted as SpaceXAI.

The Turning Point Brought by Joint

Development with Cursor

The rapid recovery of Grok was achieved through joint development with Cursor. Cursor is an AI code editor that Musk acquired for $60 billion. Around the same time in March when xAI’s founding team collapsed, two of Cursor’s senior product engineering leaders, Andrew Milich and Jason Ginsberg, joined xAI. In April, SpaceX and Cursor signed a contract for computing resources and joint development. In June, the option to acquire Cursor’s parent company was exercised.

According to documents filed with the SEC, software development is an important application scenario for AI. This is because it simultaneously provides high-quality structured data, rapid feedback cycles, and high-frequency usage. The work traces generated by AI programming are directly linked to model training and performance improvement.

How the Feedback Loop Works

Users work with Grok, and Cursor records how they work and where they fail. Engineers convert the behavioral traces and feedback results of real tasks into training data for large-scale models, strengthening the models through reinforcement learning and continuous training. Grok 4.5, released in July, caught up with GPT-5.5 in programming ability at half the price. One month later, Grok 4.6 achieved further performance gains.

This approach has also been validated by Anthropic. In February 2025, Claude released the AI programming assistant Claude Code alongside its new model, Sonnet 3.7. In less than a year, it overtook ChatGPT.

Tencent’s Similar Approach

At Tencent’s earnings call held just hours before the Grok 4.6 announcement, the co-design of WorkBuddy and Hunyuan was a key topic. WorkBuddy was developed by a small internal Tencent team, riding the wave of surging demand for AI office agents in March. Originally a modified version of the AI programming agent CodeBuddy, it reached 20.97 million monthly visits in June, ranking first among domestic desktop AI office agents.

Yao Shunyu (姚順雨), who is in charge of rebuilding Hunyuan, argued in his paper “The Second Half of AI” over a year ago that model evaluation would become more important than model training. CodeBuddy and WorkBuddy are rich in real-world tasks and evaluations. He proposed the co-design approach, deeply coordinating and optimizing the model and the product.

Tencent’s Martin Lau (劉熾平) said at the earnings call, “By continuously injecting real product usage and domain feedback into model training, Hunyuan has been able to verify model accuracy, identify and resolve edge cases, and achieve faster model iteration.” The Hy3 Preview was unveiled at the end of April, and the official version was released in July. It joined the top tier of domestic models in tasks such as agent, reasoning, and coding, achieving performance comparable to larger-parameter competitors with 295B parameters.

The Essence of the Co-design Loop

The mechanism of the Co-design Loop is simple. Users perform tasks, and the product records their traces. Engineers convert the data into model training. The model is strengthened and becomes capable of more advanced tasks. Advanced tasks generate more valuable data, which further strengthens the model. This positive feedback loop repeats.

In productivity scenarios such as programming and office work, records of failed tasks are more valuable than a single error on a benchmark, because they teach the model the actual causes of failure. Successful tasks can also be decomposed into reusable agent traces. This makes it possible to acquire rare data on real-world task distributions, operation traces, and failure cases.

Late Entrants and Industry Impact

After announcing Grok 4.6, Musk asserted, “Looking comprehensively at intelligence, speed, and cost, Grok 4.6 is objectively No. 1, and Grok 4.7 will surpass all current models.” He further stated, “SpaceX’s training corpus is extremely excellent and unique. If there is a model better than Grok 4.7 in real-world engineering, I would be surprised.”

Meta has also begun entering the Co-design Loop. Meta announced its first programming agent, “Muse Codem,” and the next-generation model “Muse Spark 1.2.” The contributor edition is 95% cheaper than the standard edition. The condition for activation is allowing Meta to use developers’ programming data for future model training.

The Co-design Loop has the potential to create, for the first time in AI large-scale models, network effects similar to those of internet platforms. More users bring data, data brings model capability, model capability produces more powerful products, and more powerful products further drive user growth.

Musk has long admired WeChat and wanted to turn Twitter into a WeChat-like super app. In an interview at the end of last year, he described X’s positioning as “WeChat++.” It can be said that in the AI field as well, he made the same strategic choice as Tencent.

Editorial Opinion

Grok 4.6 from SpaceXAI and Hunyuan from Tencent mark a strategic turning point in AI model development: a shift from conventional pre-training on massive data to continuous improvement through co-design with products. The joint development with Cursor not only solved xAI’s technical challenges but also secured access to software development, the most valuable data source. Achieving performance on par with major models in just five months demonstrates the effectiveness of the Co-design Loop. In the long term, this approach could change the competitive structure of the AI industry. As the limits of performance improvements through cost-no-object massive computation become clear, the Co-design Loop creates competition that emphasizes not only intelligence indices but also cost efficiency and real-world task execution capability. Meta’s free strategy with Muse Codem can be understood in this context. Once the data flywheel starts spinning, it will become difficult for latecomers to catch up. The fundamental challenge of the Co-design Loop is balancing user privacy with data utilization.

References

Source: 钛媒体

Comments

← Back to Home