{"id":7029847,"date":"2026-08-06T08:09:07","date_gmt":"2026-08-06T08:09:07","guid":{"rendered":"https:\/\/peraltafinancing.com\/analytics\/does-ai-find-real-ui-problems-or-just-hallucinations-measuringu\/"},"modified":"2026-08-06T08:09:07","modified_gmt":"2026-08-06T08:09:07","slug":"does-ai-find-real-ui-problems-or-just-hallucinations-measuringu","status":"publish","type":"post","link":"https:\/\/fivemor.com\/?p=7029847","title":{"rendered":"Does AI Find Real UI Problems or Just Hallucinations? \u2013 MeasuringU"},"content":{"rendered":"<p> <br \/>\n<\/p>\n<div>\n<p><a href=\"https:\/\/measuringu.com\/wp-content\/uploads\/2026\/05\/052626-FeatureImage-scaled.png\"><img class=\"alignleft wp-image-47672 size-medium br-lazy\" src=\"https:\/\/measuringu.com\/wp-content\/uploads\/2026\/05\/052626-FeatureImage-300x169.png\" fetchpriority=\"high\" decoding=\"async\" alt=\"Feature image showing an AI robot, three documents each labeled &quot;Verified&quot;, &quot;Fake&quot;, or &quot;False alarm&quot;, and a researcher.\" width=\"300\" height=\"169\" data-brsrcset=\"https:\/\/measuringu.com\/wp-content\/uploads\/2026\/05\/052626-FeatureImage-300x169.png 300w, https:\/\/measuringu.com\/wp-content\/uploads\/2026\/05\/052626-FeatureImage-1024x576.png 1024w, https:\/\/measuringu.com\/wp-content\/uploads\/2026\/05\/052626-FeatureImage-768x432.png 768w, https:\/\/measuringu.com\/wp-content\/uploads\/2026\/05\/052626-FeatureImage-1536x864.png 1536w, https:\/\/measuringu.com\/wp-content\/uploads\/2026\/05\/052626-FeatureImage-2048x1152.png 2048w, https:\/\/measuringu.com\/wp-content\/uploads\/2026\/05\/052626-FeatureImage-600x338.png 600w\" data-brsizes=\"(max-width: 300px) 100vw, 300px\"\/><\/a>In a <a href=\"https:\/\/measuringu.com\/ai-vs-human-usability-problem-analysis-of-a-video\/\">previous\u00a0experiment<\/a>, AI identified roughly half the usability problems that trained researchers found in a video of a usability test session.<\/p>\n<p>That sounds promising. If AI can find usability issues, it can substantially increase the amount of usability testing that research teams can conduct.<\/p>\n<p>But in our analysis of that video, AI generated nearly as many <em>additional<\/em> problems that humans never flagged. Are these problems hidden gems missed by multiple researchers, or just AI hallucinations?<\/p>\n<p>For this article, we classified all the unique problems the AIs generated into one of three categories:<\/p>\n<ol>\n<li>a real problem humans missed<\/li>\n<li>a false alarm (a true observation misread as a usability problem)<\/li>\n<li>a hallucination (something the AI reported that simply never happened)<\/li>\n<\/ol>\n<p>What we found suggests that the new AI problems are mostly false alarms, but there are some notable exceptions.<\/p>\n<h2>Experimental Design: Four Researchers, Two LLMs, and One Video<\/h2>\n<p>For this study, we had four humans (professional UX researchers working at MeasuringU) review a video from a previous usability benchmark study of online dining reservation websites. Each researcher independently created a list of the usability issues they observed in the six-minute video.<\/p>\n<p>We ran the video through two LLMs (ChatGPT-5.4 Thinking and Gemini 3 Flash Thinking) four times <a href=\"https:\/\/measuringu.com\/ai-usability-problem-analysis-of-a-video\/\">using the same prompt<\/a> each time.<\/p>\n<p>So in this study, we held constant the video, the key elements of the prompt, and the LLM versions\/settings\u2014variables that we plan to vary in future studies. This time, we varied only the type of analyst: human, ChatGPT, and Gemini.<\/p>\n<h2>Gemini Finds a Jewel; ChatGPT Goes on a Tangent<\/h2>\n<p>Using the human-generated and -verified problem lists as the \u201cgold standard,\u201d Figure 1 shows a summary of what we found.<\/p>\n<p><a href=\"https:\/\/measuringu.com\/wp-content\/uploads\/2026\/05\/Figure2.png\"><img loading=\"lazy\" class=\"alignnone wp-image-47542 br-lazy\" src=\"https:\/\/measuringu.com\/wp-content\/uploads\/2026\/05\/Figure2-1024x901.png\" decoding=\"async\" alt=\"Venn diagram of usability problem discovery by humans, ChatGPT, and Gemini.\" width=\"700\" height=\"616\" data-brsrcset=\"https:\/\/measuringu.com\/wp-content\/uploads\/2026\/05\/Figure2-1024x901.png 1024w, https:\/\/measuringu.com\/wp-content\/uploads\/2026\/05\/Figure2-300x264.png 300w, https:\/\/measuringu.com\/wp-content\/uploads\/2026\/05\/Figure2-768x676.png 768w, https:\/\/measuringu.com\/wp-content\/uploads\/2026\/05\/Figure2-1536x1351.png 1536w, https:\/\/measuringu.com\/wp-content\/uploads\/2026\/05\/Figure2-2048x1802.png 2048w, https:\/\/measuringu.com\/wp-content\/uploads\/2026\/05\/Figure2-600x528.png 600w\" data-brsizes=\"(max-width: 700px) 100vw, 700px\"\/><\/a><\/p>\n<p class=\"wp-caption-text\" style=\"text-align: left;\"><strong>Figure 1: <\/strong>Venn diagram of usability problem discovery by humans, ChatGPT, and Gemini.<\/p>\n<p>We know Venn diagrams can generate some bad high school math memories, so here\u2019s a summary for all our sanity:<\/p>\n<ul>\n<li>Four human researchers found nine problems (3 + 2 + 3 + 1).<\/li>\n<li>Two AIs combined found 14 problems (6 + 1 + 3 + 4).<\/li>\n<li>Only three problems were found by researchers and both AIs (the 3 in the middle of the circles).<\/li>\n<li>ChatGPT matched five of the nine researcher-identified problems (3 + 2).<\/li>\n<li>Gemini matched four (3 + 1) of the researcher-identified problems.<\/li>\n<li>That leaves 11 problems the AIs flagged that no researcher identified (6 + 1 + 4).<\/li>\n<li>Of those 11 problems, six were unique to ChatGPT, four were unique to Gemini, and one was identified by both AIs.<\/li>\n<\/ul>\n<p>So, <strong>AIs generated 11 new problems<\/strong> not identified by any of the human researchers. Table 1 has details of those 11 problems, listed in chronological order using problem number codes from the previous article. Of the 11 problems no human flagged, one was a genuine find, seven were false alarms, and three were hallucinations. Here\u2019s more detail about each category.<\/p>\n<table id=\"tablepress-1047\" class=\"tablepress tablepress-id-1047\">\n<thead>\n<tr class=\"row-1\">\n<th class=\"column-1\">Prob #<\/th>\n<th class=\"column-2\">Problem Description<\/th>\n<th class=\"column-3\">Source<\/th>\n<th class=\"column-4\">Classification<\/th>\n<\/tr>\n<\/thead>\n<tbody class=\"row-striping\">\n<tr class=\"row-2\">\n<td class=\"column-1\"><strong>4b<\/strong><\/td>\n<td class=\"column-2\">Filters not helpful<\/td>\n<td class=\"column-3\">ChatGPT<\/td>\n<td class=\"column-4\">False alarm<\/td>\n<\/tr>\n<tr class=\"row-3\">\n<td class=\"column-1\"><strong>5b<\/strong><\/td>\n<td class=\"column-2\">Participant used Ctrl-F to search for &#8220;sushi&#8221; when it wasn&#8217;t in the 86-cuisine list<\/td>\n<td class=\"column-3\">Gemini<\/td>\n<td class=\"column-4\"><strong>Genuine find<\/strong><\/td>\n<\/tr>\n<tr class=\"row-4\">\n<td class=\"column-1\"><strong>6b<\/strong><\/td>\n<td class=\"column-2\">Search results for sushi included many non-sushi restaurants<\/td>\n<td class=\"column-3\">ChatGPT<\/td>\n<td class=\"column-4\">False alarm<\/td>\n<\/tr>\n<tr class=\"row-5\">\n<td class=\"column-1\"><strong>6c<\/strong><\/td>\n<td class=\"column-2\">Weak presentation of cuisine information in search results<\/td>\n<td class=\"column-3\">ChatGPT<\/td>\n<td class=\"column-4\">False alarm<\/td>\n<\/tr>\n<tr class=\"row-6\">\n<td class=\"column-1\"><strong>7b-Gem<\/strong><\/td>\n<td class=\"column-2\">Participant chose the highest price tier<\/td>\n<td class=\"column-3\">Gemini<\/td>\n<td class=\"column-4\"><strong><span>Hallucination<\/span><\/strong><\/td>\n<\/tr>\n<tr class=\"row-7\">\n<td class=\"column-1\"><strong>7b-GPT<\/strong><\/td>\n<td class=\"column-2\">Sorting by highest rated surfaced many non-sushi restaurants<\/td>\n<td class=\"column-3\">ChatGPT<\/td>\n<td class=\"column-4\">False alarm<\/td>\n<\/tr>\n<tr class=\"row-8\">\n<td class=\"column-1\"><strong>8b<\/strong><\/td>\n<td class=\"column-2\">UI pushes browsing without good decision support<\/td>\n<td class=\"column-3\">ChatGPT<\/td>\n<td class=\"column-4\">False alarm<\/td>\n<\/tr>\n<tr class=\"row-9\">\n<td class=\"column-1\"><strong>9b<\/strong><\/td>\n<td class=\"column-2\">Seating options only presented after selecting reservation time<\/td>\n<td class=\"column-3\">Gemini<\/td>\n<td class=\"column-4\">False alarm<\/td>\n<\/tr>\n<tr class=\"row-10\">\n<td class=\"column-1\"><strong>9c<\/strong><\/td>\n<td class=\"column-2\">Participant set time to 5:10 instead of 5:00<\/td>\n<td class=\"column-3\">Gemini<\/td>\n<td class=\"column-4\"><strong><span>Hallucination<\/span><\/strong><\/td>\n<\/tr>\n<tr class=\"row-11\">\n<td class=\"column-1\"><strong>10a<\/strong><\/td>\n<td class=\"column-2\">Selected restaurant labeled &#8220;seafood&#8221; rather than &#8220;sushi&#8221; by OpenTable<\/td>\n<td class=\"column-3\">Both<\/td>\n<td class=\"column-4\">False alarm<\/td>\n<\/tr>\n<tr class=\"row-12\">\n<td class=\"column-1\"><strong>10b<\/strong><\/td>\n<td class=\"column-2\">Task not completed\u2014participant never reached the reservation form<\/td>\n<td class=\"column-3\">ChatGPT<\/td>\n<td class=\"column-4\"><strong><span>Hallucination<\/span><\/strong><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p><!-- #tablepress-1047 from cache --><\/p>\n<p class=\"wp-caption-text\" style=\"text-align: left;\"><strong>Table 1: <\/strong>The 11 AI-generated problems not identified by any human researcher, classified by type.<\/p>\n<h3>Gemini\u2019s Genuine Find<\/h3>\n<p>Let\u2019s start with the good news: All four Gemini runs identified that after the participant expanded the cuisine filter to show all 86 cuisines, she used Ctrl-F to search the page for \u201csushi\u201d (5b-Gem)\u2014an event not reported by any of the human evaluators. It happened quickly, so it\u2019s possible that the search field was not in the visual focus of the humans who were likely examining the list of cuisines (Figure 2). We consider this a true usability problem because (1) this behavior was driven by poor filter design and (2) it was unsuccessful\u2014the word \u201csushi\u201d was not on the page even though the cuisine filter was fully expanded.<\/p>\n<p><a href=\"https:\/\/measuringu.com\/wp-content\/uploads\/2026\/05\/052626-F2V2.png\"><img loading=\"lazy\" class=\"alignnone size-full wp-image-47654 br-lazy\" src=\"https:\/\/measuringu.com\/wp-content\/uploads\/2026\/05\/052626-F2.png\" decoding=\"async\" alt=\"Frame from video showing Ctrl-F search field with first few letters of \u201csushi\u201d typed at the top right of the screen and 28 of the 86 cuisine types on the left. \" width=\"825\" height=\"648\" data-brsrcset=\"https:\/\/measuringu.com\/wp-content\/uploads\/2026\/05\/052626-F2.png 825w, https:\/\/measuringu.com\/wp-content\/uploads\/2026\/05\/052626-F2-300x236.png 300w, https:\/\/measuringu.com\/wp-content\/uploads\/2026\/05\/052626-F2-768x603.png 768w, https:\/\/measuringu.com\/wp-content\/uploads\/2026\/05\/052626-F2-600x471.png 600w\" data-brsizes=\"(max-width: 825px) 100vw, 825px\"\/><\/a><\/p>\n<p class=\"wp-caption-text\" style=\"text-align: left;\"><strong>Figure 2:<\/strong> Frame from video showing Ctrl-F search field with first few letters of \u201csushi\u201d typed at the top right of the screen and 28 of the 86 cuisine types on the left.<\/p>\n<h3>Seven False Alarms<\/h3>\n<p>Next, the not-so-good news. When a researcher identifies something that happened, but it\u2019s not really considered a problem, it\u2019s referred to as a false alarm. Sometimes things are literally a feature and not a bug! From our interpretation, AIs generated seven false alarms (not too different from what <a href=\"https:\/\/measuringu.com\/false-positives\/\">you sometimes see<\/a> with a group of human evaluators).<\/p>\n<h4><strong>The seafood\/sushi labeling issue (6b, 6c, 7b-GPT, 8b, 10a)<\/strong><\/h4>\n<p>Five of the seven false alarms (6b, 6c, 7b-GPT, 8b, 10a) were derived from ChatGPT taking the search for sushi restaurants too literally. After searching for sushi, many OpenTable results were labeled \u201cseafood.\u201d ChatGPT flagged this repeatedly across multiple runs in different ways (e.g., weak cuisine presentation, non-sushi results surfacing, poor decision support), but they all trace back to the same fundamental observation. ChatGPT only considered acceptable restaurants that OpenTable labeled as sushi restaurants, not restaurants that serve sushi on the menu, regardless of OpenTable\u2019s labeling.<\/p>\n<p>The restaurant the participant ultimately selected was labeled seafood, which led ChatGPT to declare task failure in three of four runs. The human reviewers took a more pragmatic view: the restaurant served sushi, so the participant successfully completed the task. Gemini flagged the same seafood\/sushi labeling issue once (10a) but didn\u2019t spiral into multiple variations of it.<\/p>\n<h4><strong>Seating options not shown until after time selection (9b-Gem)<\/strong><\/h4>\n<p>OpenTable withholds seating options until you pick a time. Given the range of possible seating configurations (inside, patio, bar, banquette, communal, high top, private, counter), showing them before a time is selected isn\u2019t really feasible. And if a seating option doesn\u2019t work out, the recovery path is low friction. Gemini flagged this as a problem. The human researchers recognized this as a design tradeoff rather than a usability problem.<\/p>\n<h4><strong>Filters not helpful (4b-GPT)<\/strong><\/h4>\n<p>We categorized this as a false alarm because it was overly vague. It\u2019s true that there were issues with some filters (e.g., cuisine), but that was not true of all filters.<\/p>\n<h3>Three Hallucinations<\/h3>\n<p>In contrast to false alarms, which we consider misinterpretations of events that happened, a hallucination is when a problem is associated with something that just didn\u2019t happen. We saw three of these.<\/p>\n<h4><strong>AI claimed the participant incorrectly selected the highest price tier (7b-Gem)<\/strong><\/h4>\n<p>From the narrative of the second Gemini run:<\/p>\n<blockquote>\n<p><em>The task required selecting a restaurant that was not the lowest or highest price point.<\/em><\/p>\n<p><em>Problem: The participant chose Ocean Prime, which is a restaurant (the highest tier on the platform). At 05:13, the participant verbally identified this as \u201cmid-range.\u201d<\/em><\/p>\n<p><em>User Impact: The participant technically failed this part of the task constraints.<\/em><\/p>\n<\/blockquote>\n<p><strong><em>This didn\u2019t happen<\/em><\/strong>. Ocean Prime had a mid-range price designation.<\/p>\n<h4><strong>AI claimed the participant set the reservation time for 5:10 pm (9c-Gem) <\/strong><\/h4>\n<p>From the narrative of the second Gemini run:<\/p>\n<blockquote>\n<p><em>The participant selected Ocean Prime at 05:10<\/em><\/p>\n<\/blockquote>\n<p><strong><em>This didn\u2019t happen<\/em><\/strong>. The participant, in accordance with the task instructions, selected 5:00 pm.<\/p>\n<h4><strong>AI claimed the participant did not reach the reservation form (10b-ChatGPT)<\/strong><\/h4>\n<p>From the narrative of the second ChatGPT run:<\/p>\n<blockquote>\n<p><em>By the end of the clip, they are still comparing list items and time slots; they do not appear to reach the restaurant detail\/reservation form step.<\/em><\/p>\n<\/blockquote>\n<p><strong><em>This isn\u2019t accurate<\/em><\/strong>. The clip ended with the participant selecting the reservation time, then standard dining room seating, then stopping before entering her personal information.<\/p>\n<p>The good news is that there were only three hallucinations out of 11 AI-generated problems. The bad news is you can\u2019t know which AI-generated problem descriptions were hallucinated without watching the video and reviewing all the problems yourself.<\/p>\n<h2>Summary and Discussion<\/h2>\n<p>In this article, we focused on qualitative similarities and differences in the usability problems listed by professional human UX researchers and two AIs (Gemini 3 Flash Thinking and ChatGPT-5.4 Thinking) after reviewing a video in which a participant made a restaurant reservation.<\/p>\n<p>Our key findings were:<\/p>\n<p><strong>False alarms and hallucinations dominate.<\/strong> Of the 11 problems the AIs generated that no human flagged, seven (64%) were false alarms, three (27%) were hallucinations, and one (9%) was a genuine find. That\u2019s a useful number to keep in mind: roughly nine out of ten AI-only problems in this study required either correction or dismissal.<\/p>\n<p><strong>AI adds value as a junior researcher, not a trusted expert.<\/strong> AI was able to find one problem (a participant had to use Ctrl-F) that was real and useful and not found by humans. But getting to it required reviewing ten other problems that ranged from technically true but irrelevant to simply fabricated. The ROI depends on how much that review costs you.<\/p>\n<p><strong>Most false alarms came from a single fixation.<\/strong> Five of the seven traced back to ChatGPT interpreting \u201csushi restaurant\u201d more literally than any human would. At least in this video and our criteria for what constitutes a problem, this is a systematic bias worth knowing about if you\u2019re using these models for task-based evaluations.<\/p>\n<p><strong>Hallucinations were infrequent but consequential.<\/strong> Three of the problems (27%) were hallucinations. Although nominally low, this is probably too high for most applications. You can\u2019t catch those without going back to the video, which means human review isn\u2019t optional.<\/p>\n<p><strong>Like humans, AI usability reviews of videos are prone to the \u201cevaluator effect.\u201d <\/strong>Just like with human evaluators, multiple runs of AI usability evaluations of videos are not perfectly consistent, so it\u2019s good practice to run these evaluations multiple times for consistency checks. Two of the three hallucinations came from the same Gemini run. Running multiple evaluations and looking for consistency across runs is a practical filter before any human review.<\/p>\n<p><strong>Bottom line: AI usability reviews of videos require human oversight. <\/strong>In their current form (what we tested), these AI products can add value to this type of UX research, but more as junior researchers whose actions and conclusions require expert human oversight rather than as trusted experts themselves.<strong><br \/><\/strong><\/p>\n<\/p><\/div>\n\n","protected":false},"excerpt":{"rendered":"<p>In a previous\u00a0experiment, AI identified roughly half the usability problems that trained researchers found in a video of a usability test session. That sounds promising. If AI can find usability issues, it can substantially increase the amount of usability testing that research teams can conduct. But in our analysis of that video, AI generated nearly [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":7029848,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[12033],"tags":[10959,60230,14041,13391,482],"dealstore":[],"offerexpiration":[],"class_list":["post-7029847","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-analytics","tag-find","tag-hallucinations","tag-measuringu","tag-problems","tag-real"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v26.4 - https:\/\/yoast.com\/wordpress\/plugins\/seo\/ -->\n<title>Does AI Find Real UI Problems or Just Hallucinations? \u2013 MeasuringU - Som2ny Network<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/fivemor.com\/?p=7029847\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Does AI Find Real UI Problems or Just Hallucinations? \u2013 MeasuringU - Som2ny Network\" \/>\n<meta property=\"og:description\" content=\"In a previous\u00a0experiment, AI identified roughly half the usability problems that trained researchers found in a video of a usability test session. That sounds promising. If AI can find usability issues, it can substantially increase the amount of usability testing that research teams can conduct. But in our analysis of that video, AI generated nearly [&hellip;]\" \/>\n<meta property=\"og:url\" content=\"https:\/\/fivemor.com\/?p=7029847\" \/>\n<meta property=\"og:site_name\" content=\"Som2ny Network\" \/>\n<meta property=\"article:published_time\" content=\"2026-08-06T08:09:07+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/08\/052626-FeatureImage-300x169.png\" \/>\n\t<meta property=\"og:image:width\" content=\"300\" \/>\n\t<meta property=\"og:image:height\" content=\"169\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/png\" \/>\n<meta name=\"author\" content=\"admin\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"admin\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"8 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\/\/fivemor.com\/?p=7029847#article\",\"isPartOf\":{\"@id\":\"https:\/\/fivemor.com\/?p=7029847\"},\"author\":{\"name\":\"admin\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371\"},\"headline\":\"Does AI Find Real UI Problems or Just Hallucinations? \u2013 MeasuringU\",\"datePublished\":\"2026-08-06T08:09:07+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\/\/fivemor.com\/?p=7029847\"},\"wordCount\":1725,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\/\/fivemor.com\/#organization\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/?p=7029847#primaryimage\"},\"thumbnailUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/08\/052626-FeatureImage-300x169.png\",\"keywords\":[\"Find\",\"Hallucinations\",\"MeasuringU\",\"problems\",\"Real\"],\"articleSection\":[\"Analytics\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\/\/fivemor.com\/?p=7029847#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\/\/fivemor.com\/?p=7029847\",\"url\":\"https:\/\/fivemor.com\/?p=7029847\",\"name\":\"Does AI Find Real UI Problems or Just Hallucinations? \u2013 MeasuringU - Som2ny Network\",\"isPartOf\":{\"@id\":\"https:\/\/fivemor.com\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\/\/fivemor.com\/?p=7029847#primaryimage\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/?p=7029847#primaryimage\"},\"thumbnailUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/08\/052626-FeatureImage-300x169.png\",\"datePublished\":\"2026-08-06T08:09:07+00:00\",\"breadcrumb\":{\"@id\":\"https:\/\/fivemor.com\/?p=7029847#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\/\/fivemor.com\/?p=7029847\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/?p=7029847#primaryimage\",\"url\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/08\/052626-FeatureImage-300x169.png\",\"contentUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/08\/052626-FeatureImage-300x169.png\",\"width\":300,\"height\":169},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\/\/fivemor.com\/?p=7029847#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\/\/fivemor.com\/?bp_activities=1\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Does AI Find Real UI Problems or Just Hallucinations? \u2013 MeasuringU\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\/\/fivemor.com\/#website\",\"url\":\"https:\/\/fivemor.com\/\",\"name\":\"Som2ny Network\",\"description\":\"Daily Deals\",\"publisher\":{\"@id\":\"https:\/\/fivemor.com\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\/\/fivemor.com\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\/\/fivemor.com\/#organization\",\"name\":\"Som2ny Network\",\"url\":\"https:\/\/fivemor.com\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/logo\/image\/\",\"url\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png\",\"contentUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png\",\"width\":300,\"height\":86,\"caption\":\"Som2ny Network\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/#\/schema\/logo\/image\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371\",\"name\":\"admin\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/image\/\",\"url\":\"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png\",\"contentUrl\":\"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png\",\"caption\":\"admin\"},\"sameAs\":[\"https:\/\/fivemor.com\"],\"url\":\"https:\/\/fivemor.com\/?author=1\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Does AI Find Real UI Problems or Just Hallucinations? \u2013 MeasuringU - Som2ny Network","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/fivemor.com\/?p=7029847","og_locale":"en_US","og_type":"article","og_title":"Does AI Find Real UI Problems or Just Hallucinations? \u2013 MeasuringU - Som2ny Network","og_description":"In a previous\u00a0experiment, AI identified roughly half the usability problems that trained researchers found in a video of a usability test session. That sounds promising. If AI can find usability issues, it can substantially increase the amount of usability testing that research teams can conduct. But in our analysis of that video, AI generated nearly [&hellip;]","og_url":"https:\/\/fivemor.com\/?p=7029847","og_site_name":"Som2ny Network","article_published_time":"2026-08-06T08:09:07+00:00","og_image":[{"width":300,"height":169,"url":"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/08\/052626-FeatureImage-300x169.png","type":"image\/png"}],"author":"admin","twitter_card":"summary_large_image","twitter_misc":{"Written by":"admin","Est. reading time":"8 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/fivemor.com\/?p=7029847#article","isPartOf":{"@id":"https:\/\/fivemor.com\/?p=7029847"},"author":{"name":"admin","@id":"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371"},"headline":"Does AI Find Real UI Problems or Just Hallucinations? \u2013 MeasuringU","datePublished":"2026-08-06T08:09:07+00:00","mainEntityOfPage":{"@id":"https:\/\/fivemor.com\/?p=7029847"},"wordCount":1725,"commentCount":0,"publisher":{"@id":"https:\/\/fivemor.com\/#organization"},"image":{"@id":"https:\/\/fivemor.com\/?p=7029847#primaryimage"},"thumbnailUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/08\/052626-FeatureImage-300x169.png","keywords":["Find","Hallucinations","MeasuringU","problems","Real"],"articleSection":["Analytics"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/fivemor.com\/?p=7029847#respond"]}]},{"@type":"WebPage","@id":"https:\/\/fivemor.com\/?p=7029847","url":"https:\/\/fivemor.com\/?p=7029847","name":"Does AI Find Real UI Problems or Just Hallucinations? \u2013 MeasuringU - Som2ny Network","isPartOf":{"@id":"https:\/\/fivemor.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/fivemor.com\/?p=7029847#primaryimage"},"image":{"@id":"https:\/\/fivemor.com\/?p=7029847#primaryimage"},"thumbnailUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/08\/052626-FeatureImage-300x169.png","datePublished":"2026-08-06T08:09:07+00:00","breadcrumb":{"@id":"https:\/\/fivemor.com\/?p=7029847#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/fivemor.com\/?p=7029847"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/?p=7029847#primaryimage","url":"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/08\/052626-FeatureImage-300x169.png","contentUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/08\/052626-FeatureImage-300x169.png","width":300,"height":169},{"@type":"BreadcrumbList","@id":"https:\/\/fivemor.com\/?p=7029847#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/fivemor.com\/?bp_activities=1"},{"@type":"ListItem","position":2,"name":"Does AI Find Real UI Problems or Just Hallucinations? \u2013 MeasuringU"}]},{"@type":"WebSite","@id":"https:\/\/fivemor.com\/#website","url":"https:\/\/fivemor.com\/","name":"Som2ny Network","description":"Daily Deals","publisher":{"@id":"https:\/\/fivemor.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/fivemor.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/fivemor.com\/#organization","name":"Som2ny Network","url":"https:\/\/fivemor.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/#\/schema\/logo\/image\/","url":"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png","contentUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png","width":300,"height":86,"caption":"Som2ny Network"},"image":{"@id":"https:\/\/fivemor.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371","name":"admin","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/#\/schema\/person\/image\/","url":"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png","caption":"admin"},"sameAs":["https:\/\/fivemor.com"],"url":"https:\/\/fivemor.com\/?author=1"}]}},"_links":{"self":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts\/7029847","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=7029847"}],"version-history":[{"count":0,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts\/7029847\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/media\/7029848"}],"wp:attachment":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=7029847"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=7029847"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=7029847"},{"taxonomy":"dealstore","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fdealstore&post=7029847"},{"taxonomy":"offerexpiration","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fofferexpiration&post=7029847"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}