Project Overview

Validating AI Deployment Bots through Concept and Model Testing

The AI Deployment Application was an LLM-powered tool designed to support customers during software deployment. I led evaluative research to understand whether the experience was usable, useful, and trustworthy enough for broader adoption, resulting in a heuristic evaluation, unmoderated usability testing, and model-level feedback for the product team.

AI deployment application interface used for concept and model evaluation
Role
UX Researcher
Timeline
Multi-sprint evaluative study
Team
Product manager, ML stakeholders, UX, and technical subject matter experts
Users
Deployment practitioners with varied levels of product and domain experience
Initiated By
Product Manager
2 lenses

Evaluated both the user interface and the underlying AI model performance.

8 to 10 users

Tested realistic deployment scenarios with participants across experience levels.

+40%

Improved model effectiveness by the time the AI tool reached first release.

The challenge

Problem statement

The AI Deployment Application needed to be evaluated for both usability and effectiveness before broader adoption. Because it relied on an LLM to assist with deployment tasks, the team needed clarity on more than interface quality. They also needed to know whether responses were accurate, relevant, and helpful enough to support real user work.

Goals and objectives

  • Identify usability issues across the interface.
  • Assess the tool’s overall effectiveness during realistic deployment workflows.
  • Evaluate model output for accuracy, relevance, and response time.
  • Provide actionable recommendations to improve both the experience and the AI system.

Research Methods

To evaluate the product thoroughly, I looked at the experience through both a UX and AI-performance lens. That meant reviewing the interface structurally, then testing how the product performed in realistic usage scenarios.

Heuristic evaluation

  • Engaged 3 to 5 UX experts to independently review the interface using established usability heuristics.
  • Used this method to surface structural issues early and identify patterns that could undermine trust or comprehension before participant testing.

Unmoderated usability testing

  • Set up an unmoderated study in Maze with realistic deployment scenarios.
  • Recruited 8 to 10 participants with varied deployment experience to complete key tasks using the tool.

Model evaluation signals

  • Collected AI performance data including response accuracy, response time, and perceived relevance.
  • Paired behavioral observations with post-test questionnaire feedback to understand both task success and confidence in the assistant.

The Process

Balancing concept validation with product readiness.

The product was still evolving in real time, so the research needed to move fluidly between early evaluative signals and concrete release guidance.

01

Planning

  • Dissected the problem statement and aligned on the evaluation goals.
  • Selected methods and created a timeline that balanced product needs with research rigor.
02

Recruitment & setup

  • Identified key SMEs for expert review.
  • Set up an unmoderated Maze study and prepared deployment scenarios for participants.
03

Research execution

  • Ran heuristic reviews with UX experts.
  • Observed usability test performance across varied levels of deployment experience.
  • Captured post-test perceptions of usefulness, trust, and satisfaction.
04

Analysis

  • Compiled severity-ranked heuristic findings.
  • Analyzed task completion, time on task, satisfaction signals, and AI response quality.
  • Reviewed patterns in behavior and feedback to pinpoint the biggest breakdowns.
05

Reporting & recommendations

  • Prioritized issues based on impact to user experience and deployment success.
  • Created a comprehensive report for the product manager with recommended interface and model improvements.

Insights

The strongest insights came from looking at interaction quality and model behavior together. The interface could not be evaluated in isolation because trust in the tool depended on both what it looked like and what it said.

Trust was tied to clarity

Users did not just want answers quickly. They wanted clearer signals about what the bot could do, what it was uncertain about, and how much to rely on its outputs.

Usability issues weakened model confidence

Even strong responses lost value when the interface made next steps, context, or interaction states feel ambiguous.

Real workflows exposed where the concept held up

Scenario-based tasks showed where the assistant meaningfully supported deployment work and where it created friction or hesitation.

Cross-functional collaboration sharpened the research

Working closely with product and ML stakeholders helped surface risk areas like hallucination and confidence scoring early enough to influence the roadmap.

Signals reviewed

  • Task completion rates
  • Time on task
  • User satisfaction scores
  • Response accuracy
  • Response time
  • Perceived usefulness and relevance

How I framed the findings

I organized insights so product partners could clearly separate interface problems from model problems, while still seeing how they interacted in the full user experience.

Business Impact

What the work made possible

This evaluation gave the team actionable guidance for improving the product before wider adoption, while also helping the AI experience feel more intuitive, explainable, and reliable.

40%

Increase in model effectiveness by the time of deployment and first release.

2 layers

Recommendations targeted both UI design and AI model performance rather than treating them separately.

Clearer

Structured findings improved readability and made product tradeoffs easier for stakeholders to act on.

Earlier

Research helped the team optimize the experience before wider rollout instead of reacting after adoption issues grew.

Learning & Reflection

This project deepened my ability to design research that balances the complexity of AI systems with the everyday workflows of technical users. Working with a product still evolving in real time, I learned how to move between foundational inquiry and scrappy evaluative testing, often within the same week.

One of the biggest takeaways was that developers did not just want speed. They wanted clarity around what the bot could and could not do. Building user trust required transparency, not just better automation.

Collaborating with product and ML stakeholders early helped us identify risk areas around hallucination and confidence scoring before they became bigger product problems. It also reinforced how valuable mixed methods can be when the goal is to evaluate not just usability, but confidence in an AI system’s outputs.

Ultimately, this work reminded me that successful AI integration is not just about automation. It is about augmenting human work in ways that feel intuitive, explainable, and reliable.