/Evaluating the Trustworthiness of AI-Generated UX Feedback Across Multiple LLMs
Abstract

This study evaluates the trustworthiness of AI-generated User Experience (UX) feedback across multiple Large Language Models (LLMs), including ChatGPT 5.1, Claude Sonnet 4.5, and Gemini 3. Using real UI screenshots and interface videos, the study examines the accuracy, explainability, usefulness, heuristic alignment, hallucination frequency, and consistency of UX feedback generated under different prompting strategies. A mixed-method approach combines quantitative evaluation with qualitative analysis and cross-model verification based on Nielsen’s usability heuristics. The findings highlight variations in feedback quality, hallucination behaviour, prompt sensitivity, and modality support across the evaluated models. The study emphasizes the potential of LLMs as supportive tools for early-stage UX evaluation while highlighting the continued importance of human validation in UX decision-making.

RelatedView All
CitationsView All
Citing-
Cited By-