ByteDance's New Data Army: Why AI's Next Battleground Isn't Models,It's Training Data
ByteDance has established a new top-level business unit dedicated to artificial intelligence data and security, consolidating scattered data teams across the company into a single organization led by Adam Wang. The move reflects a fundamental shift in how the world's leading AI companies compete: as model architectures converge and reasoning capabilities plateau, the quality and quantity of training data have become the primary differentiator between state-of-the-art systems and the rest.
The new unit sits at the same organizational level as Seed, Flow, Douyin, and other major ByteDance business divisions, underscoring how central data has become to the company's AI strategy. Before taking this role, Wang led platform responsibility and live streaming at TikTok, overseeing a business that once generated substantial revenue for the platform. His appointment signals ByteDance's commitment to treating data infrastructure with the same strategic importance as product development itself.
Why Is Data Suddenly More Important Than Model Architecture?
For years, AI competition focused on who could build the most sophisticated neural network architecture or secure the most computing power. That dynamic has shifted. As reasoning approaches converge across leading models and gains from architectural innovations become harder to achieve, the bottleneck has moved upstream: to the data that trains these systems.
The public internet, which powered the first generation of large language models (LLMs), is running dry. While the web contains vast amounts of published reports and documents, it captures far less of the actual process through which people produce results: working through ambiguous requests, gathering context, making mistakes, and correcting them. This gap is especially acute in specialized domains like medicine, law, and scientific research, where expert-level reasoning requires training data that simply doesn't exist in freely available form.
ByteDance's leadership has made this philosophy explicit. At a Seed all-hands meeting in late July, founder Zhang Yiming stated that ByteDance would "firmly reject distillation" in its efforts to advance foundation models. CEO Liang Rubo reinforced this message on August 5, saying, "ByteDance's large language models will firmly remain self-developed. We need to build the fundamentals well, accept falling behind in the short term, and continue optimizing for the long term".
Zhang Yiming
"ByteDance's large language models will firmly remain self-developed. We need to build the fundamentals well, accept falling behind in the short term, and continue optimizing for the long term. Most importantly, we must not lose our direction," stated Liang Rubo, CEO of ByteDance.
Liang Rubo, CEO at ByteDance
This commitment to building models from scratch, rather than copying or refining existing systems, makes data infrastructure non-negotiable. You cannot achieve state-of-the-art performance through shortcuts if you are building everything yourself.
How Does ByteDance's Data Operation Actually Work?
- Scale of Operation: More than 1,000 people within ByteDance were involved in evaluating model data for Seed alone, with more than ten data staff supporting each algorithm engineer on average. This contrasts sharply with video AI startups, where internal evaluation teams may number only several dozen people.
- Organizational Structure: The new unit consolidates several previously scattered teams, including a data team established in 2023 by Fu Yue, a TikTok founding team member, which initially numbered around 100 employees. The consolidated organization now spans product managers, data engineers, procurement specialists, quality control staff, operations personnel, and security and compliance experts.
- End-to-End Responsibilities: The unit handles the entire data production pipeline, including setting standards, sourcing and procurement, data synthesis and cleaning, and quality evaluation. Its core function is to provide cross-modal data services for all of ByteDance's foundation models.
- Competitive Methodology: ByteDance has adopted a system that assigns parallel teams to the same problem and compares their results. Teams are divided by area, including world models, code, and advanced disciplines. As model economics become clearer, individual data projects must now calculate their return on investment.
The integration of these teams began in early June, though ByteDance is still clarifying organizational structure and staffing as the consolidation spans multiple departments with overlapping responsibilities. The process reflects the complexity of coordinating data work across a company as large and diverse as ByteDance.
What Does This Mean for the Global AI Race?
ByteDance's move is not isolated. It reflects a broader shift across the global AI sector. Leading overseas foundation model companies are spending several times more on data than ByteDance and other major Chinese technology companies. External data budgets at leading Silicon Valley model companies are expanding by billions of dollars annually. Anthropic, for example, allocated more than USD 1 billion to reinforcement learning data in 2025 alone.
This spending has fueled rapid growth among data suppliers. Mercor, a startup founded only three years ago, saw its annualized revenue rise from USD 500 million last year to USD 2 billion by mid-2026. About 91 percent of its revenue came from leading model companies such as OpenAI and Anthropic, while its valuation reached USD 20 billion.
ByteDance has invested heavily in data compared with many other major Chinese technology companies. According to reports, ByteDance's data budget for training world models and coding models had already reached an eight-figure USD sum at the beginning of 2026, with instructions that the budget could be increased if necessary. The company continues to increase this investment.
Competition for data talent is intensifying across the industry. Tencent, which stepped up its work on foundation models last year, has recruited from ByteDance's data teams over the past six months, offering some candidates salaries as high as three times their previous compensation. Alibaba and Tencent have both increased their data procurement budgets. Major technology companies now employ various data exclusivity strategies, such as setting exclusive periods for datasets or temporarily securing exclusive access to key personnel from suppliers.
Yet the data market in 2026 faces a basic imbalance: AI companies have strong demand, while high-quality data remains in limited supply. That constraint extends from pretraining, where models learn general knowledge, through post-training, where they are refined for specific tasks. For ByteDance, a central challenge will be keeping a data organization of around 1,000 people responsive to rapid changes at the frontier of model development. As model capabilities continue to evolve, the types of data that are scarce today may be considerably less valuable six months from now.
The success of ByteDance's Seed 2.0 model has been described by multiple industry practitioners as "a victory for data," underscoring how thoroughly the competitive landscape has shifted. In an era where model architectures are converging and computing power is increasingly commoditized, the ability to source, synthesize, evaluate, and manage high-quality training data has become the true differentiator in AI competition.