Logo
FrontierNews.ai

Which AI Models Actually Cost the Planet Less? A New Study Ranks 22 LLMs by Carbon Footprint

Not all artificial intelligence models are created equal when it comes to their environmental cost. A new comparative study of 22 large language models (LLMs), which are AI systems trained on vast amounts of text to generate human-like responses, found significant variation in how much carbon dioxide each one produces when answering user requests. The research, conducted by Devoteam, a global consulting firm, measured the operational carbon footprint of models from major players including OpenAI, Google, Anthropic, Meta, and others, revealing that the choice of which AI model to use matters far more for the planet than many organizations realize.

Why Does an AI Model's Carbon Footprint Vary So Much?

The carbon emissions from an LLM depend on multiple interconnected factors that go well beyond just the model's size or capability. When you send a request to an AI chatbot, that query travels through physical infrastructure: graphics processing units (GPUs), memory systems, cooling equipment, and power distribution networks. Each of these components draws electricity, and the amount varies dramatically depending on how the model is built and deployed.

The researchers identified several key levers that determine a model's environmental impact. The calculation chain works like this: the electrical power drawn by the hardware gets converted into total energy consumed based on how long it takes to process your request, and then that energy is converted into greenhouse gas emissions based on where the electricity comes from. A model that processes requests faster might draw more instantaneous power but consume less total energy overall, resulting in lower emissions per response.

How to Compare AI Models by Their Environmental Impact

  • Model Architecture and Efficiency: Different LLMs are built with different internal structures that affect how much computation is needed to generate each token, which is a unit of text roughly equivalent to a word or part of a word.
  • Infrastructure and Data Center Design: The same model running on different hardware or in data centers with varying cooling efficiency can produce vastly different carbon footprints, even when processing identical requests.
  • Electricity Carbon Intensity: Whether a data center is powered by renewable energy, natural gas, or coal dramatically changes the carbon emissions associated with running any given model, sometimes by a factor of several times.
  • Usage Volume and Speed: Models that generate responses faster consume less energy per request, and organizations that carefully control how much they use AI can reduce their total carbon impact significantly.

The Devoteam research team developed a standardized methodology to make fair comparisons across all 22 models. They normalized results to measure carbon emissions in grams of CO2 equivalent per 1,000 tokens processed, creating a common baseline that removes the distortion of studying different volumes of requests. This approach allows organizations to compare models on equal footing, regardless of how many users are querying them.

What Makes This Study Different from Previous AI Environmental Research?

Most discussions about AI's environmental impact focus on the enormous energy required to train these models in the first place, a one-time cost that gets spread across millions of future uses. This study takes a different angle by examining the recurring, operational carbon footprint during the inference phase, which is when the model is actually being used to answer questions and generate text. This distinction matters because it answers a practical question that organizations face daily: if I deploy this model today, how much carbon will I emit every time someone uses it ?

The research panel included both proprietary models from major tech companies and open-weight models that developers can download and run themselves. The diversity of the sample, covering eight major players in the LLM ecosystem, makes it possible to observe real differences in carbon footprint between models with varied sizes and capabilities. The study examined everything from relatively lightweight models designed for efficiency to large-scale models with advanced reasoning abilities.

"How can we design and deploy an AI that is performant enough to meet our needs, while minimising the resources required to run it?" noted Marwa Mokni, IT Project Lead at Devoteam.

Marwa Mokni, IT Project Lead at Devoteam

The research sits within Devoteam's broader E3S initiative, which focuses on efficient and sustainable smart solutions. The goal is not to crown a single "best" model, but rather to help organizations understand the trade-offs between performance and environmental cost, and to identify practical levers they can pull to reduce their AI-related emissions.

What Should Organizations Do With This Information?

The findings suggest that organizations have more control over their AI carbon footprint than they might assume. Rather than treating AI adoption as an all-or-nothing decision, companies can make deliberate choices about which models to deploy, where to run them, and how intensively to use them. A model that appears slightly less capable might produce substantially lower emissions, making it the more responsible choice for many applications. Similarly, running the same model in a data center powered by renewable energy versus fossil fuels can cut emissions by multiple times over.

The measurement methodology itself represents a step forward in AI accountability. By establishing a clear, comparable framework for calculating operational carbon footprints, the research creates a foundation for more transparent reporting. As AI adoption accelerates across industries, having standardized ways to measure and compare environmental impact becomes increasingly important for organizations trying to meet sustainability commitments.