Back to blog

The largest AI model is not the best for your customer service.

When choosing an AI model, the logic seems simple: select the smartest model available. This is especially true if that model tops the industry benchmarks. However, this logic does not always apply to customer service.
Rutger de Ruiter
Sep 01, 2026
7 minutes read

General AI benchmarks tell us very little about customer service

When choosing an AI model, the logic seems simple: pick the smartest one available, especially if it tops the well-known AI benchmarks. However, this logic does not always hold true for customer service.

Many popular AI benchmarks measure performance in areas like coding, mathematics, general knowledge, and complex reasoning. While this highlights a model's general capabilities, it says very little about how well it will perform as a customer service AI agent.

In reality, customers rarely phrase their questions perfectly. For example, they might say: "My order was supposed to arrive yesterday, but I haven't received anything yet. Can you check where it is?"

The AI agent must first understand the customer's intent. It then needs to look up the order, check its current status, determine what information it is allowed to share, and take the correct next step.

This means it is not just the final answer that needs to be correct, but the entire process.

Why we developed CM-ServiceBench

CM-ServiceBench tests AI models on complete customer service conversations. It evaluates more than just the final answer, assessing every step of the journey: reasoning, using the right tools, executing workflows correctly, and staying within guardrails.

We evaluate three key areas:

Choosing the right path

The first step is understanding what the customer actually wants. That sounds simple, until they express themselves in incomplete, ambiguous, or unexpected ways.

A good AI agent must still identify the correct intent and determine the next steps. A mistake at the start of the conversation can mean that even if every subsequent step is technically correct, the final outcome is still wrong.

Executing workflows correctly

Most customer inquiries require more than a single answer. Processing a return, rescheduling an appointment, or checking an order status involves multiple steps: identifying the customer, retrieving data, verifying eligibility, making a decision, and finally executing the action.

The order of these steps also matters. An agent that completes five out of six steps correctly can still produce the wrong outcome. This is what makes customer service fundamentally different from simply generating a good response. An AI agent must be able to follow and execute a process reliably.

Staying within guardrails

A good AI agent must also know what not to do. It should not hallucinate information, share internal instructions, expose data not meant for the customer, or take action when it lacks sufficient information.

For an AI agent in production, this is not an optional feature; it is a fundamental requirement. An agent is only viable if it is both helpful and predictable.

How do you evaluate a conversation?

With CM-ServiceBench, we have models handle hundreds of customer conversations across various scenarios and languages. A second AI acts as the customer, equipped with its own goal, context, and information, and conducts a conversation with the AI agent under test.

These scenarios are based on patterns found in real-world customer service processes. We then evaluate the entire conversation, not just the final response. Did the model choose the right path? Did it use the correct tools and actions? Were the steps executed in the right order? Was the information processed correctly? And did the conversation lead to the right outcome?

This distinction is crucial. A model can give a highly convincing response while executing the wrong process behind the scenes. To the customer, the interaction seems successful, until they realize their appointment was never rescheduled, the wrong information was retrieved, or they were routed to the wrong department.

Strict versus partial: When is a conversation truly successful?

A conversation can go mostly well but still contain an error along the way. That is why CM-ServiceBench measures results in two ways: partial and strict.

The partial metric measures how many individual parts of the conversation were executed correctly. The strict metric is more demanding: a conversation only counts as successful if the entire interaction is flawless.

This distinction is vital for customer service. A model that gets nine out of ten steps right can still leave the customer with the wrong result.

The leaderboard suddenly looks different

When tested this way, the rankings look quite different from what general AI benchmarks might suggest. In our tests, compact models from the GPT-5.6 family are actually among the top performers for these customer service tasks.

Model

Strict

Partial

GPT-5.6 Luna with thinking

71

92

GPT-5.6 Terra

69

91

GPT-5.6 Sol

66

90

GPT-5.6 Luna

63

88

Opus 4.6

63

90

GPT-5.5

61

88

Sonnet 5

58

86

The most interesting insight is not just which model comes out on top. Compare GPT-5.6 Terra and Opus 4.6, for example. Opus 4.6 scores highly on partial but lower on strict. This means the model often gets far into the process but completes fewer conversations entirely correctly.

In our tests, GPT-5.6 Terra is much more consistent at finishing the job. For the customer, this is the only difference that matters. They do not care how impressive the reasoning was along the way, only whether their problem was actually resolved.

Bigger does not mean more reliable

This does not mean larger models are inferior. They are designed to handle a wide range of complex tasks. This extra capacity is highly valuable when an agent needs to make complex decisions, process extensive context, or resolve unpredictable issues.

However, not every customer service task requires this level of complexity. For many processes, quality is about consistently following instructions, using the right tools, and never skipping a step. In these cases, higher general intelligence does not automatically translate to a better outcome.

As a result, you might end up paying for extra model capacity without seeing any performance improvement for that specific task.

Speed is also a measure of quality

Furthermore, a benchmark score does not tell the whole story. Because AI agents interact with humans, response time directly impacts the user experience. If a model takes several seconds longer to respond to every message, the conversation quickly feels sluggish. This is even more critical for voice applications, where any delay is immediately noticeable.

That is why CM-ServiceBench evaluates response times alongside reliability.

CM-ServiceBench: performance versus speed

prestatie-tegen-snelheid-outlined

Each dot represents a model. The further to the left, the faster the model responds. The higher the dot, the more conversations were completed flawlessly.

The chart illustrates why there is no single, obvious winner. GPT-5.5 is the fastest, responding in an average of 1.8 seconds per message, but it does not achieve the highest strict score. At the other end of the spectrum, GPT-5.6 Luna with thinking scores highest in the benchmark at 71, but comes with a slower average response time of 3.2 seconds.

GPT-5.6 Terra sits in the middle, achieving a high strict score of 69 with an average response time of 2.8 seconds. This highlights what model selection is really about in practice: finding the best balance between quality and speed for a specific task, rather than just chasing the highest score.

There is also a third factor to consider: cost. A model that scores slightly higher but requires significantly more computing power and capacity is not necessarily the best choice for a process that runs thousands of times a day.

The best model depends on the task

Asking "what is the best AI model?" is too broad. A better question is: what is the best model for this specific agent and task?

A compact, fast model might be ideal for a predictable workflow. For a more complex agent that requires nuanced decision-making or processes large amounts of context, advanced reasoning capabilities add real value. Meanwhile, for processes where every second counts, a faster model may be the better option, even if another model scores slightly higher on strict accuracy.

This also means you do not have to standardize on a single model for your entire customer service operation. Different AI agents have different requirements. Even within a single agent, simple classification tasks might call for a different model than complex analysis or action execution.

The key takeaway from CM-ServiceBench is not that one specific model wins. Instead, it shows that you must evaluate models based on the actual work they will perform. While general benchmarks are interesting, you ultimately need to know which model runs your specific processes most reliably, at a speed and price point that make sense for your business.

Model selection in HALO

This is why we do not limit you to a single model in HALO. You can choose the right AI model for each agent, and even assign different models to specific tools. This allows you to deploy advanced model capabilities only where they add real value, rather than running every simple process through your largest, most expensive model by default.

These requirements will also continue to evolve. As new models emerge, existing ones improve, and performance, speed, and cost dynamics shift, CM-ServiceBench allows us to continuously test models against the same service tasks. This ensures your choices are always backed by up-to-date, real-world data.

Would you like to see how this works in HALO? Contact our team for a demo.

With HALO, you choose which AI model to use for each agent. This allows you to select the best balance of quality, speed, and cost for every use case.
Was this article interesting?
Share it!

Latest Articles

Ask HALO AI

Meet Ask HALO: the AI assistant that builds, improves and maintains your agents

Ask HALO builds, debugs and maintains AI agents from inside HALO Studio, with full access to your agents, tools and conversation history. See how CheapCargo, Preston Palace, Winparts and Intergamma use it.

5 minutes read Sep 01, 2026
blog-conversational-christmas-engagement Connect

Convert Conversations This Christmas: 5 Use Cases

The holiday season is almost here, and it comes with the perfect opportunity to connect with customers in a personal and meaningful way. Messaging channels like WhatsApp, RCS, and SMS can help you create an unforgettable customer experience this Christmas. In this blog, you’ll discover how these channels can boost satisfaction and drive sales during the busiest time of the year.

7 minutes read Dec 04, 2025
3 types of AI AI

Unleashing the Power of AI: How Generative, Agentic, and Predictive AI Are Transforming Customer Experience

Artificial Intelligence (AI) has been in development for decades, but the way we use it today has changed dramatically. With the advent of ChatGPT and other applications, AI has suddenly become tangible for the general public. While it was previously used primarily for specific, often invisible applications (think fraud detection in banking or predictive maintenance in industry), it now actively assists in content creation, enhancing customer experiences, and streamlining processes. Within customer experience, three forms of AI are particularly relevant: generative, agentic, and predictive AI. In this article, we’ll break them down and explain how to leverage them effectively.

4 minutes read Oct 27, 2025
halo-insurance AI

From Claims to Customer Questions: How AI Agents Help Insurers

The insurance industry is known for its complex processes and heavy administrative load. Fragmented communication, outdated systems, and complicated policy conditions mean that finding the right information or processing changes often takes far longer than it should. AI agents can change that. They answer questions, pull real-time data from internal systems, and seamlessly trigger processes.

4 minutes read Oct 25, 2025
blog-black-friday-customer-retention AI

From Casual Consumers to Loyal Fans: How to retain customers after Black Friday

Black Friday is a huge shopping event, drawing crowds of eager shoppers hunting for deals. But after the frenzy fades, the real challenge begins: turning one-time shoppers into loyal, repeat customers. Customer retention is vital for long-term success and post-Black Friday is the perfect time to build lasting relationships. So, how can your business retain customers post-Black Friday? In this blog, we’ll explore how to make it happen.

5 minutes read Oct 23, 2025
Implementation checklist for AI agents AI

Your AI Agent Implementation Checklist

AI agents aren’t just shaping the future they’re transforming how companies serve and connect with their customers right now. From answering service requests instantly, to guiding shoppers through a purchase, to spotting upsell opportunities in real time, the question is no longer if you should implement AI, but how quickly you can put it to work.

4 minutes read Oct 02, 2025
blog-picking-ai-platform HALO

From Selection to Success: How to Choose the Right AI Platform

An AI platform isn’t just another tool you purchase. It’s the foundation on which your organization operates and innovates. The choices you make today will shape how you work in the future. While you may start with just a few agents supporting specific use cases, over time more processes will be taken over by agents. That’s why it’s critical to ensure the foundation you lay now is cohesive, scalable, and backed by solid governance and compliance.

5 minutes read Sep 29, 2025
blog-halo-ecommerce AI

AI Agents: The Accelerators of Conversational Commerce

The way consumers search for and process information online is rapidly changing thanks to AI. Where we used to type in search terms, scroll through dozens of results, and manually filter them, we are now getting used to having conversations. With ChatGPT, Google’s AI features, and other assistants, answers come faster and are more relevant. That same way of interacting is now taking over e-commerce at high speed. For retailers, this is the moment to step in: the webshop as we know it—where customers have to actively search themselves—is giving way to personal conversations that directly lead to action.

4 minutes read Sep 27, 2025
Is this region a better fit for you?
Go
close icon