Artificial intelligence isn’t ready to handle your holiday gift buying, fresh research indicates.
Shopping assistants powered by AI routinely deliver inconsistent information regarding pricing, product specs, and inventory status, per a report released this week by Product.ai, a startup that validates product claims against evidence.
The firm evaluated both complimentary and premium tiers of ChatGPT, Claude, Gemini, and Perplexity by posing 220 shopping-related queries covering everything from laptops and televisions to mattresses, sunscreen, and robotic vacuums. Each query was submitted five times per platform, yielding 8,794 total responses.
Product.ai discovered that 86% of the queries resulted in a reproducible factual discrepancy. The report classified a conflict as a verifiable inconsistency — like differing prices, models, or specifications — appearing across multiple answers.
“Simply put, these large language models haven’t reached the point of delivering a seamless end-to-end shopping experience,” stated Dakota Nunley, Product.ai’s head of search product.
AI has been steadily expanding its footprint in retail, with merchants and LLM developers alike rolling out features to help users locate and purchase items online. AI agents like Meta’s Muse or Instinct could eventually automate even more of the process.
Product.ai’s findings highlight a deeper problem: AI continues to have difficulty supplying the precise information that consumers or AI agents require.
“Perplexity is the only AI company relentlessly focused on achieving 100% accuracy, and we lead the industry in every measure of it,” a Perplexity spokesperson said. Representatives for the other three LLMs did not respond to requests for comment.
The research revealed that 97% of comparative questions — where models were asked to weigh two products against each other — led to a conflict. That figure exceeded the 75% conflict rate seen for basic factual or specification queries.
The models also had trouble supplying accurate pricing. Among the 913 responses that Product.ai could verify, 85% aligned with the current or listed price.
Performance varied across the LLMs. Gemini showed the highest proportion of questions with what Product.ai deemed a costly error: 56% on its free tier and 54% on its paid tier.
Claude’s results improved on its premium version, with costly errors dropping from 44% to 21%. Perplexity recorded the lowest costly-error rate at 14% for its paid tier, followed by ChatGPT’s paid version at 17%.
Product.ai additionally noted that the tools occasionally contradicted themselves when the same question was posed multiple times. Gemini’s free tier did this on 29% of questions, the highest rate observed in the study.
When answers were incorrect, the price discrepancy averaged a median of $300, according to Product.ai.
Such a significant gap should make consumers think twice about trusting AI agents, Nunley noted. “That’s a big miss, right there,” he said.
Rather than entrusting payment information to agents or accepting AI output as unquestionable truth, Nunley advised users to independently verify the details. That could involve consulting multiple AI platforms and cross-referencing results — or simply visiting the retailer’s website directly, he suggested.
“Use AI this upcoming holiday season as a discovery tool, as a thought partner,” Nunley said. “But be wary of going hands-off at the moment.”
Do you have a story idea about AI and shopping? Contact this reporter at abitter@Themoneytimes.com or via encrypted messaging app Signal at 808-854-4501. Use a personal email address, a nonwork WiFi network, and a nonwork device; here’s our guide to sharing information securely.

