Building an AI Support Assistant customers could trust

In December 2023, an automotive chatbot became famous for the wrong reason. A customer asked Chevrolet’s website chatbot whether they could buy a $70,000 SUV for $1. The chatbot appeared to agree.
The legal details mattered less than the signal it sent: when AI speaks on behalf of a company, customers experience it as the company’s judgement.
That was the leadership challenge behind Volvo’s AI Support Assistant.
Customer Care needed to reduce inbound contacts. The organization had fewer human agents, more EV-related questions, and more pressure to scale support. Chatbots can reduce cost in customer service, but Volvo could not treat this as a deflection problem alone.
For a premium brand, a poor AI experience can quickly damage trust. If the assistant gives generic answers, blocks escalation, invents information, or misunderstands the customer’s situation, the customer does not blame “the model.” They blame Volvo.
That expectation set the bar for the assistant. It could not send customers back to content they had already struggled to find or understand. It had to help them move forward, recognise when self-service had reached its limit, and keep a clear path to human support.
So the question became:
How do we move fast enough to reduce pressure on Customer Care, but safely enough to protect customer trust?
I helped frame this as a UX and research leadership challenge, not only a technology or cost-saving initiative. Together with Product and Engineering leadership, I connected the business pressure, customer research, UX risk, and operating constraints into a clearer direction for the team.
The analogy I used was the triage desk in a hospital emergency room, especially in Sweden, where triage determines whether you get immediate help, get routed onward, or wait.
Triage does not replace doctors or solve every case at the front door. It helps staff understand the situation, identify what they can handle safely, recognise when someone needs specialist attention, and route the person to the right help.
The leadership challenge was not to launch an AI assistant. It was to help a cross-functional team make responsible choices under pressure: where to use AI, where to constrain it, how to protect customer trust, and how to know when the experience was ready for customers.
The decisions that shaped the assistant
The team had discussed a chatbot for Customer Care for a long time. Even before the current AI wave, simple triage chatbots had become common in customer service. They could deflect repetitive contacts, collect context before handover, and reduce agents’ average handling time.
For years, the idea did not become a clear product direction.
The team focused on a known self-service model: improving the support website, making articles easier to find, and strengthening the content formats customers already used. That made sense. If customers could find the right article faster, self-service would improve without adding the risks of a conversational AI experience.
Then the context changed.
In 2025, Volvo went through several rounds of reorganization and resource reduction. Customer Care had fewer human agents. EV-related customer questions kept growing. The operational pressure became harder to ignore. We needed to help more customers solve simple problems before they reached an agent, without lowering the quality of care.
The chatbot became a serious option
The chatbot was no longer speculative. It became a serious option.
The team was not ready to move without concern.
The proposed shift asked the team to move from improving support pages and articles toward designing a new AI-enabled support experience. That changed the scope, the mindset, and the risk.
UX raised valid customer concerns:
- Would customers trust an AI assistant from Volvo?
- Would they feel helped, or blocked?
- Would AI become a barrier to human support?
- Could we keep the experience safe enough for a premium brand?
- Would the assistant understand real customer language, or only clean test questions?
This became one of the most important leadership moments in the project for me.
The business reason to explore AI was clear. If the team experienced the decision as purely cost-driven, we risked losing the customer perspective and the quality of the product direction.
Instead of trying to convince the team, we used research to make the tension visible and useful.
During summer 2025, the UX designer paired with user research to understand whether we were on the right path. The goal was not to prove that customers wanted a chatbot. The goal was to understand when AI support could help, and where it would create risk.
The research gave us a more nuanced answer.
Customers were open to AI when it helped them find information, understand a topic, compare options, or continue a journey without repeating themselves. They became more cautious when the issue felt emotional, high-stakes, safety-related, complaint-driven, or when they expected accountability.
That changed the conversation.
The question was no longer:
Will customers accept AI?
It became:
Where is AI safe and useful, and where must the experience hand over to a human?
That gave the team enough confidence to move forward. It also raised the bar. We were not building a chatbot to replace care. We were not using AI as a shortcut around customer support. We were designing a triage experience: something that could understand the customer’s situation, help with simple cases, recognise when an issue was too complex or sensitive, and make human handover clear when that was the safer path.
We could not expose generative answers
Engineering clarified an important constraint: we could not expose directly generative AI content to customers.
At first, that could have stopped the work. If the assistant could not freely generate answers, would it still be useful? Would it feel too rigid? Would we recreate a decision tree with a chat interface?
We compared three paths:
Fully generative AI
Flexible and natural, but harder to control and risky for trust.
A simple decision tree
Easier to control, but weak for real customer language and long-tail support questions.
Controlled AI decisioning
Use AI to understand the customer’s problem and route them to the right support path, while keeping answers grounded in approved Volvo content.
We chose the third path. That became the key product and UX decision: move fast, but inside trust boundaries.
The assistant could use AI to interpret intent and choose a resolution path. It would not invent answers. It would stay grounded in approved support content, keep fallback visible, and escalate when human care was needed.
That decision gave the team direction. It also changed the UX work.
We were not designing “a chatbot that sounds human.” We were designing the conditions that make customer-facing AI safe, useful, and understandable.
From clarification questions to option labels
Once we chose the safer middle path, the team had to make it work. The assistant could use AI to understand and route the customer’s question, but the answer still had to come from trusted Volvo content.
The system followed the triage logic. A customer described a problem in their own words. A supervisor bot interpreted the request, routed it to the right specialised task bot, collected simple context when needed, such as a licence plate for vehicle-related issues, and searched approved Volvo support content.

On paper, this was the right compromise.
The assistant would not invent the answer. It would search trusted content, choose the best match, and help the customer move forward.
Testing exposed the next problem.
Many customer questions did not point to one clear article. They pointed to several possible articles. Even with a filtering threshold, the system could find relevant content without knowing which path matched the customer’s need.
If the assistant was unsure, should it ask a follow-up question?
At first, that felt obvious. In a human conversation, you ask a clarifying question when someone is unclear. The team explored a guarded follow-up system. It would use the retrieved articles to create a short, predefined question that helped the customer narrow down the issue before seeing a solution. For example
”Are you trying to [topic 1], [topic 2] or [topic 3]?”
Then we hit a UX boundary.
Questions generated from article topics felt long and unnatural. Article content often duplicated or overlapped, so the assistant produced questions like:
“Are you trying to enable driving journal, troubleshoot trips not showing, or connect the Volvo Cars app to your car? “
That was an important UX moment. Clarification questions created more confusion
Technically, the system had done something reasonable. It found related content and tried to clarify. From the customer’s point of view, the assistant repeated itself and produced questions that were hard to process.
That distinction matters in customer care.
A wrong clarification question makes the customer work harder. It can also make the assistant feel less trustworthy before it gives an answer.
We stopped trying to make a controlled clarification question sound conversational. We removed the question. We replaced it with two short, scannable option labels, one per likely path, so the customer could recognise their situation and choose the direction that matched their need.
That small design change shifted the cognitive work.
The customer no longer had to interpret a vague or slightly wrong question. They could scan two possible meanings and choose one. We also kept the chat input visible next to the options, so the customer could continue in their own words if neither option felt right.

That became a reusable principle for the assistant:
- When AI confidence is not high enough, do not pretend to know. Give the customer a simple way to choose, correct, or continue.
For a moment, we had a workable pattern. The assistant could offer follow-up choices. The customer could scan and continue. The system stayed inside approved content boundaries.
Clean test questions hid real customer behaviour
The content designer on the team had built an evaluation set to measure accuracy and relevance. It helped us understand whether the system could find and route to the right answer. But the evaluation set used AI summaries of real questions. It did not preserve how customers had originally asked them.
Real customers do not speak in summaries.

They type keyword bursts. They make typos. They describe symptoms instead of naming the feature. They write emotional fragments when something is not working. They ask questions in a second language. On mobile, they often write the shortest version of what they mean.
So she tested the assistant against more realistic customer phrasing.
She tested 335 questions across 8 phrasing styles. The result showed a consistent 8–10 percentage point accuracy gap between clean phrasing and keyword or symptom phrasing.
ACCURACY OF RESPONSES PER TEXT VARIANT STYLEClean questions ███████████ 81%
Keyword bursts ████████ 73%
Symptom only. ███████ 71%
Verbose style ████████ 75%
From accuracy to experience quality
The team was not only testing the assistant against content. We were testing it against customer behaviour.
That pushed the team to change the evaluation approach. We needed to look beyond whether the syste could answer a well-written question. We needed to know whether it could support the messy, compressed, imperfect language customers use when they need help.
This changed the UX conversation again.
Accuracy still mattered, but it was not enough. The real question became:
- Can the assistant help customers make progress when their question is incomplete, messy, emotional, or unclear?
The work moved from chatbot design into AI service quality. The team was no longer only designing messages and flows.
They were shaping the conditions that make the system reliable enough to trust: better test data, clearer fallback paths, stronger content grounding, visible recovery, and evaluation based on real customer behaviour.
What changed because of this process
By the end of this phase, the team had improved the assistant and changed how we understood the work. We could no longer judge the product only by whether it retrieved the right article. It had to show that it could guide customers through uncertainty, recover when confidence was low, and behave consistently enough for Volvo to trust it in front of customers.
The people who made this possible
This was team work.
@Annelie Tinworth led the content design and AI prompting work, including the evaluation tests. She drove the UX solution for AI-generated options and created the phrasing guidelines that made the assistant clearer and safer for customers. She also went deep into Visual Studio and prompt engineering, showing the team how content design can pair with engineering in AI products.
@Johanna Bülling shaped the UX flows, interaction patterns, and fallback experience. She also co-drove the decision framework and metrics definition with her product counterpart, helping us measure contact deflection and experience quality. She used workshops, frameworks, and new processes to help the team move through change.
@Ines Anic drove the research work, creating a research plan that included secondary research, concept testing, and interviews. Her work helped us understand whether we were on the right path, where customers would accept AI, and where they would need more control or human support. At the beginning of the project, she worked closely with @Caisa Lundblad to close discovery and help the team choose the right direction.
Product and Engineering partnership made the work possible. Product leadership connected the assistant to business value, rollout, and prioritisation. Engineering translated trust boundaries into system architecture, routing logic, performance work, and evaluation infrastructure. Without that partnership, the UX principles would have stayed theoretical.
Leadership takeaway
The work started with a business need: reduce pressure on Customer Care by helping more customers solve simple problems themselves.
The leadership challenge was bigger than contact deflection. We had to define what responsible customer-facing AI should feel like for Volvo: useful, controlled, transparent, grounded in trusted content, and never experienced as a wall between the customer and human care.
The team moved through several turning points. We used research to understand where customers would accept AI. We chose controlled AI decisioning instead of unconstrained generation. We replaced confusing clarification questions with scannable option labels. We expanded testing from clean questions to real customer phrasing. We reframed quality from “did the system find an answer?” to “did the customer have a trustworthy next step?”
My leadership contribution was to help the team hold the whole problem together: business pressure, customer trust, technical constraints, content quality, research evidence, and service recovery. I created space for UX concerns instead of dismissing them, used research to turn debate into evidence, and helped Product, Engineering, Content, Research, and UX align around a direction that was useful and responsible.
This is the leadership lesson for me: leading AI innovation means creating the judgement, evidence, and operating conditions that let a team move fast without losing customer trust.
The final lesson was:
For customer-facing AI, the goal is not to make the system sound intelligent. The goal is to make the next step trustworthy.