AI Benchmarking Just Got Real: MLPerf Client v2.0 Adds Image Generation and Agent Testing
MLCommons has released MLPerf Client v2.0, a major update to the industry-standard benchmark for measuring how well personal computers run artificial intelligence workloads locally. The new version adds two significant capabilities: image generation testing and agentic AI evaluation, alongside refreshed language model benchmarks. This expansion reflects how rapidly AI capabilities are evolving beyond simple text processing into more complex, real-world tasks.
What's New in MLPerf Client v2.0?
MLPerf Client measures how effectively PCs, from laptops and desktops to workstations, execute AI tasks without relying on cloud servers. The v2.0 release introduces several major workload categories designed to keep pace with the rapid evolution of AI-enabled hardware and software. These additions represent a significant shift in how the industry evaluates AI performance on consumer and professional machines.
The benchmark now includes the following new and updated capabilities:
- Image Generation Category: Features Flux.2 klein 4B as an experimental test, enabling evaluation of generative visual capabilities on personal computers.
- Agentic AI Category: Benchmarks agentic AI performance through Software Engineering Agent and Data Analyst Agent scenarios, reporting end-to-end performance with breakdowns of language model inference and tool execution times.
- Updated Language Model Tests: Upgrades Phi 3.5 mini instruct to Phi 4 Mini Instruct in required workloads, while introducing Qwen 3 8B as an experimental test.
- Enhanced Summarization Tasks: Base tasks now include an Intermediate Summarization task featuring an input prompt of roughly 4,000 tokens, testing how well systems handle longer documents.
Why Should You Care About PC AI Benchmarking?
As AI capabilities move from data centers to personal devices, the ability to measure performance accurately becomes increasingly important. Benchmarks like MLPerf Client help hardware manufacturers, software developers, and consumers understand which systems can handle demanding AI workloads efficiently. The addition of image generation and agentic AI testing reflects a fundamental shift in what "AI on your PC" actually means in 2026.
Agentic AI, in particular, represents a new frontier. Unlike simple language model inference, agentic systems must reason through problems, call external tools, and execute multi-step workflows. Testing these capabilities on personal computers requires measuring not just inference speed but also the overhead of tool execution and decision-making processes. This is why MLPerf Client v2.0 reports both responsiveness and throughput metrics for these new workloads.
How to Understand MLPerf Client v2.0 Benchmarks
- Responsiveness Metrics: Measure how quickly a system responds to user input, critical for interactive applications like real-time image generation or code completion.
- Throughput Metrics: Measure how many tasks a system can complete in a given time period, important for batch processing and background AI operations.
- End-to-End Performance: For agentic AI, the benchmark breaks down time spent on language model inference versus tool execution, helping identify performance bottlenecks.
- Model-Specific Testing: Uses real, named models like Phi 4 Mini Instruct and Flux.2 klein 4B, allowing direct comparison across different hardware platforms.
Who's Behind This Benchmark?
MLPerf Client v2.0 is the product of collaboration among major technology leaders, including AMD, Intel, Microsoft, NVIDIA, Qualcomm Technologies, and top PC original equipment manufacturers (OEMs). These participants have contributed resources and expertise to ensure the benchmark reflects real-world use cases and hardware capabilities. The benchmark is freely available for download, with source code open for inspection and community contribution via the MLCommons GitHub repository.
MLCommons, the organization behind the benchmark, is an open engineering consortium dedicated to improving machine learning performance and transparency. The organization produces industry-leading benchmarks, datasets, and best practices spanning the full range of machine learning applications, from massive cloud training to resource-constrained edge devices. MLPerf has become the de facto standard for evaluating AI performance across the industry.
What Does This Mean for the Future of AI on PCs?
The expansion of MLPerf Client reflects a broader industry trend: AI workloads are becoming more diverse and more demanding. Early benchmarks focused primarily on language model inference because that was the dominant use case. But as image generation models, multimodal systems, and agentic AI become more prevalent, benchmarking needs to evolve accordingly. Version 2.0 represents a significant step toward a comprehensive, cross-platform benchmark for client AI computing.
For consumers and businesses evaluating AI-capable PCs, these benchmarks provide concrete, standardized metrics for comparison. Rather than relying on marketing claims, stakeholders can now run the same tests across different hardware configurations and directly compare results. This transparency is particularly important as AI capabilities become a key selling point for consumer laptops and professional workstations.