Blog
Feb 11

Usability testing for Gen AI

AI is everywhere- helping us write emails, navigate complex systems and assist our workflows. As we build new products and services that integrate this technology it also brings us a new set of usability challenges. Unlike traditional software, AI behaves probabilistically, learns over time, and often operates as a ‘black box’. My work at Accenture, The Dock- Accenture’s global innovation and R&D hub involves coming up with new products or services while pushing the boundaries of technology. Our clients’ demand for Gen AI specifically has grown many-fold, requiring us to design AI solutions across practically all industries.

This unpack what makes usability testing for AI different, why it matters, and how we can evolve methods to keep up with increasingly intelligent systems. You will also get a resource pack at the end of this to equip yourself with tools and frameworks to build this skillset.

What makes usability testing for AI different?

Usability testing for AI is fundamentally about testing relationships; between users and systems. The variability of actions, outputs, expectations and interpretations add an interesting layer of complexity that as designers, we get to mediate. Traditional products and services are deterministic but AI tools by contrast, are probabilistic. Outputs vary based on input nuance, context, model state, training data, or even randomness (as in generative AI).

What factors can help test Gen AI tools?

As we test across variations and scenarios we need to assess the tools across various factors:

  • Trust & Explainability– Test how transparency, feedback, and failure-handling affect user confidence in the system. Additionally, can the user understand why the AI did what it did? Testing how well the system reveals its reasoning or source can help build trust for the user. This could include how the AI presents its sources & citations or how it informs users of any shortcomings or risk. Another aspect to think about is how do users calibrate trust over time. Usability testing should also surface cases of over-trust, where users defer too readily to the system despite uncertainty or errors.
  • Relevance & Tonality– Especially important in generative or conversational AI, responses must match the user’s intent and emotional tone. It should be contextual to themselves and the expected response and personality. Basically it should pass the vibe check. These behavioural configurations of a personality could meaningfully alter how a model thinks and responds. What is it emphasising on? How is it shaping its priorities? What is its degree of caution? At what level does it challenge its user? These nuances help a designer think deeply and work with data scientists and AI engineers to analyse a model’s personality.
  • Mental Effort– Most AI interfaces add a huge responsibility to users to construct their first prompt. We need to test the level of cognitive load added to a user and find ways of easing that by providing maybe example generations to educate or suggestion nudges to avoid the blank slate dilemma.
  • Control– Does the AI empower users with choice? Test if users feel in control or at the mercy of the system. The user should be able to guide, modify, or undo the AI’s decisions. Control helps users be in the pilot seat of the plane as opposed to a passenger. Things like allowing them to stop, pause or edit an action, control what data inputs are retained within the system and define the scope of the AI’s autonomy (such as specifying what the system can decide independently versus what requires explicit user approval.)

In addition to these one can look at factors like bias, adaptability, responsiveness, consistency, performance, social and ethical perception. Collaborating with data scientists, AI & ML engineers, psychologists and behavioural scientists will help bring more perspective of when we conduct testing for new AI tools and technologies.

Methodologies of testing that make sense:

Here are some methodologies that have worked for me and my teams at different stages of developing AI tools:

  • Wizard of Oz testing– This is where a user interacts with an AI tool interface where the responses are simulated by the researcher. This is a great way for early stage concepts to get an understanding of users expectation, personality and logic before developing the model.
  • Adversarial Testing– Borrowed from security and risks’ red teaming methodology this is one where you “break” the AI’s usability with various edge case scenarios. This is an efficient way of exposing biases or testing the AI consistency.
  • A/B testing – Testing across versions of a model or across various benchmarked tools help understand a baseline of how your product fares against being ‘helpful’ or ‘human like’ against others. Another way to do this would be to give various users a task and test how the AI responds to similar but different prompts.

This to me is an exciting and critical realm designers walk into- because now more than ever the user should be advocated for in the age of automation. Usability testing for Gen AI is not just about improving outputs; it’s about shaping relationships between people and increasingly intelligent systems and ensuring those relationships remain understandable, accountable, and human.


https://medium.com/design-bootcamp/usability-testing-for-gen-ai-1362e48ee998a>

Leave a reply

Your email address will not be published. Required fields are marked *