Logo
FrontierNews.ai

ByteDance Is Building a 1,000-Person Data Army to Power Its AI Ambitions

ByteDance has established a new top-level business unit dedicated to artificial intelligence data and security, consolidating scattered data teams across the company into a single, coordinated force of roughly 1,000 people. The move reflects a fundamental shift in how the world's leading AI labs compete: as the supply of free, high-quality internet data dries up, companies are investing billions to generate, curate, and evaluate training data internally.

Why Is ByteDance Suddenly Obsessed With Data?

ByteDance's new AI data and security unit, led by Adam Wang (a former TikTok executive who previously oversaw the platform's lucrative live-streaming business), sits at the same organizational level as Seed, Flow, Douyin, and other major business divisions. This structural choice signals that data is no longer a supporting function; it is now a core competitive advantage.

The timing matters. ByteDance's founder Zhang Yiming recently declared that the company would "firmly reject distillation," meaning it refuses to train its models by copying knowledge from competitors' systems. Instead, ByteDance aims to build world-class models from scratch, using only its own data and training methods. That ambition requires an enormous, high-quality dataset.

To understand the scale: for every algorithm engineer working on Seed (ByteDance's foundation model), more than ten data specialists provided support during model evaluation. That ratio far exceeds what video AI startups typically deploy, where internal evaluation teams might number only dozens of people. For Seed alone, over 1,000 ByteDance employees were involved in evaluating model data.

What Does This New Unit Actually Do?

The consolidated organization will serve as a central data factory for all of ByteDance's foundation models. Its responsibilities span the entire data production pipeline:

  • Data Standards: Setting quality benchmarks and consistency rules across all model training projects.
  • Sourcing and Procurement: Acquiring datasets from external suppliers and managing exclusive access agreements with data providers.
  • Data Synthesis and Cleaning: Generating synthetic training examples and removing noise, errors, and low-quality samples from raw datasets.
  • Quality Evaluation: Testing finished datasets to ensure they meet performance standards before being used to train models.

The unit consolidates several previously scattered teams, including data teams that supported Dola (the international version of Doubao, ByteDance's consumer chatbot), the DMC platform, and the AI Data Platform (AIDP) under Flow. The integration began in early June 2026, and ByteDance is still clarifying the final organizational structure and staffing.

How Much Is ByteDance Spending on Data?

ByteDance's investment in data is substantial and growing. By the beginning of 2026, the company's data budget for training world models and coding models had already reached an eight-figure sum in US dollars, with instructions that the budget could be increased if necessary. For context, leading Silicon Valley AI companies like Anthropic allocated more than USD 1 billion to reinforcement learning data in 2025 alone, and their external data budgets are expanding by billions annually.

ByteDance has also adopted an internal system that assigns parallel teams to solve the same data problem and compares their results, ensuring quality and identifying the most effective approaches. Teams are organized by specialty, including world models, code, and advanced disciplines. As model economics become clearer, ByteDance now requires individual data projects to calculate their return on investment.

The company faces stiff competition for talent. Tencent, which stepped up its foundation model work last year, has recruited from ByteDance's data teams over the past six months, offering some candidates salaries as high as three times their previous compensation.

Why Is High-Quality Data Becoming Scarce?

The global AI sector faces a critical constraint: the pool of readily available, high-quality public internet data has become increasingly depleted. Over the past three years, the data required for model training has shifted from the public domain toward private sources. The public internet contains vast amounts of reports and documents, but it captures less of the messy, real-world process through which people actually produce results, such as working through ambiguous requests, gathering context, making mistakes, and correcting them.

Beyond coding, general-purpose AI models still struggle with highly specialized tasks in fields such as medicine, law, and scientific research. As developers seek to improve models on expert-level tasks, they need large amounts of specialized, proprietary data that simply does not exist on the open web. This constraint extends from pretraining (the initial phase where models learn from raw data) through post-training (the phase where models are fine-tuned for specific tasks).

The data market in 2026 faces a basic imbalance: AI companies have strong demand, while high-quality data remains in limited supply. Major technology companies, including Alibaba and Tencent, have increased their data procurement budgets and adopted various data exclusivity strategies, such as setting exclusive periods for datasets or temporarily securing exclusive access to key personnel from suppliers.

What Does This Mean for ByteDance's AI Models?

ByteDance is simultaneously pursuing an aggressive model development roadmap. The company is training a new AI model expected to reach up to 10 trillion parameters, significantly larger than Kimi K3, which is currently the largest AI model released in China. The training process is set to take between three to six months, after which the model will undergo fine-tuning before release.

For comparison, estimates suggest that Anthropic's Mythos 5 has around 8 trillion parameters, with Fable 5 at about 5 trillion. While the number of parameters is crucial for determining a model's capacity to store information, other factors like data quality and training methods also play significant roles in overall performance.

"ByteDance's large language models will firmly remain self-developed. We need to build the fundamentals well, accept falling behind in the short term, and continue optimizing for the long term. Most importantly, we must not lose our direction," stated Liang Rubo, CEO of ByteDance.

Liang Rubo, CEO at ByteDance

ByteDance's model development team, known as Seed and led by former Google DeepMind scientist Wu Yonghui, comprises around 2,000 members globally. The team has adopted a more independent approach to model development, avoiding the practice of distilling existing models from other labs. This strategy, in place for over a year, has contributed to a slower development pace compared to some competitors, but ByteDance leadership believes it is the only path to building a model that outperforms rivals.

How to Understand ByteDance's Data-First Strategy

  • The "No Distillation" Principle: ByteDance refuses to train models by copying knowledge from competitors. Instead, it builds from scratch using proprietary data, which requires enormous internal data infrastructure and expertise.
  • The Data-to-Engineer Ratio: For every algorithm engineer at ByteDance, more than ten data specialists provide support, a ratio far higher than most AI startups, reflecting the company's belief that data quality determines model quality.
  • The Organizational Signal: By elevating data to a top-level business unit, ByteDance is signaling that data is no longer a cost center but a strategic asset comparable to product development and business operations.
  • The Global Trend: Major AI labs worldwide, including Anthropic, OpenAI, and Chinese competitors like Alibaba and Tencent, are all increasing data spending as the supply of free, high-quality internet data becomes constrained.

For ByteDance, a central challenge will be keeping a data organization of around 1,000 people responsive to rapid changes at the frontier of model development. As model capabilities continue to evolve, the types of data that are scarce today may be considerably less valuable six months from now. The company's success will depend not just on the size of its data team, but on its ability to anticipate what kinds of data will matter most as AI capabilities advance.