Imagine that you want to hire a copywriter for your marketing agency. You have two options: Candidate A, who has an impeccable academic qualification and got a 10 on all of his theory exams, and Candidate B, who may not have as many diplomas, but has won the last ten creative writing tournaments in front of real audiences. Who would you choose?

In the world of Artificial Intelligence, something similar is happening to us. For years, we have been bombarded with acronyms like MMLU or GPQA, static exams where AI models get excellent grades.

But, for those of us who work in the “mud” of daily marketing, those notes tell us little about whether that AI is going to help us write an email that converts or design a brilliant campaign.

This is where “The AI ​​Hunger Games” comes in. Forget boring exams; What rules today is the ELO rating.

The origin: From the chess board to your screen

The ELO system was not born in a Silicon Valley laboratory. It was created in the 1960s by Arpad Elo, a physicist and chess master who was looking for a fairer way to measure the relative skill of players. Instead of giving a fixed score, the ELO is dynamic: points are gained or lost by “stealing” them from the opponent after each game.

If a player with a high rating loses against a rookie, the system takes away many points because the result was a surprise. If the favorite wins, it barely rises, because it is what was expected.

It’s a scale that never saturates and always tells you who is better in relation to others at that precise moment.

How does the AI ​​“Arena” work?

You’ve probably heard of GPT-4o, Claude or Gemini. But how do we know which one is really best for the average user? The answer lies in platforms like the Chatbot Arena (LMSYS).

Imagine a digital coliseum. You, as a user, ask a question or a task (a “prompt”). Two anonymous AI models (let’s call them Model A and Model B) generate a response at the same time. You read them and vote for the one you like the most, without knowing which is which.

It is a blind taste test, but with algorithms. After thousands of these “battles”, an ELO ranking is calculated that reflects the real preferences of humans. It is a “living” metric that captures what static reviews ignore: actual usefulness, tone, and the ability to connect with what we ask for.

Why this is pure gold for Marketing

If you are a CMO, copywriter or digital strategist, ELO should be your compass for several compelling reasons:

  • It is “cheat” proof: Traditional exams (static benchmarks) are easy to “train” specifically to pass, just like a student memorizing the answers to a test from previous years. ELO is unpredictable because it depends on what real users ask every day.
  • Reflects the “vibe” of the content: In marketing, we not only look for the information to be correct; We want it to sound good. The ELO reflects that human satisfaction that no mathematical formula can measure on its own.
  • Helps you optimize costs: For many marketing tasks, a cheaper model but with a competitive ELO can give you results almost identical to those of the market leader.

ELO in the daily life of companies

Don’t think this is just for tech geeks. Companies like Gong already use the ELO system internally.

Instead of an overall ranking, they use ELO to measure specific tasks, like summarizing a sales call or analyzing customer sentiment. They put different models to compete on their own data and only deploy the one that wins the efficiency “battle.”

This is the future of operational marketing: not choosing an AI because it is the most famous, but because it has proven to be the best in the ring of your specific need.

Not everything is perfect: The cracks of the coliseum

As in any competition, there are nuances. ELO in AI has its limitations that we must know:

  • The noise of the judges: Humans are subjective. Sometimes we vote for the answer that seems more polite or the one that is more beautifully formatted, even if the content is less precise.
  • It doesn’t work for everything: If you need an AI to program complex code or solve pure mathematical problems, ELO based on user votes can be volatile. There it is better to rely on direct verification metrics, such as code execution success rates.
  • The novelty trap: The system assumes that the competitors’ skills are stable, but the AI ​​evolves every week. A model that is the king of the arena today can become obsolete in a month if its training is not updated.

##Who will win the final battle?

The difference in quality between the most powerful AIs is becoming smaller, which forces us to be much more selective. The ELO rating takes the blinders off of Big Tech’s marketing and shows us who is really performing under pressure.

For your next campaign or to choose the tool that will automate your reports, don’t just look at the name of the company behind it. Go to the “Arena”, look at the ELO rankings and choose the gladiator who best knows how to connect with what your clients need.

In the end, in this competition, the real winners are us: the users who have at our disposal tools that are increasingly refined and aligned with what we consider valuable.

This post is also available in: Español Français Русский Italiano