General Tech
Business Insiderabout 3 hours ago
0

AI can't take over your holiday shopping quite yet, a new study suggests.

AI

A robotic hand holds a golden credit card between its index finger and thumb. A study from Product.ai found that major LLMs had high error rates when used to find products while shopping online.

AI can't take over your holiday shopping quite yet, a new study suggests.

Intelligence Insights

Context + impact, normalized for TechCulture.

The Big Picture
A robotic hand holds a golden credit card between its index finger and thumb. A study from Product.ai found that major LLMs had high error rates when used to find products while shopping online. Andriy Onufriyenko/Getty Images Don't rely on LLMs to do your holiday shopping this year, a new study says. A test of four AI engines found that they had issues providing accurate information about products. Problems ranged from giving the wrong price to not finding the latest model.
Why It Matters
A robotic hand holds a golden credit card between its index finger and thumb. A study from Product.ai found that major LLMs had high error rates when used to find products while shopping online.

Deepen your understanding

Use our AI to break down complex signals.

Select an AI action to generate more depth.

A robotic hand holds a golden credit card between its index finger and thumb.
A robotic hand holds a golden credit card between its index finger and thumb.
A study from Product.ai found that major LLMs had high error rates when used to find products while shopping online.

Andriy Onufriyenko/Getty Images

  • Don't rely on LLMs to do your holiday shopping this year, a new study says.
  • A test of four AI engines found that they had issues providing accurate information about products.
  • Problems ranged from giving the wrong price to not finding the latest model.

AI isn't able to take over your holiday shopping yet, a new study suggests.

AI-powered shopping assistants frequently give conflicting answers about prices, product specifications, and which items are current, according to a study published this week by Product.ai, a startup that checks product claims against evidence.

The company tested the free and paid versions of ChatGPT, Claude, Gemini, and Perplexity using 220 shopping questions about products ranging from laptops and TVs to mattresses, sunscreen, and robot vacuums. It ran each question five times on each service, capturing 8,794 responses.

Eighty-six percent of the questions produced a repeatable factual conflict, Product.ai found. The study defined a conflict as a checkable disagreement — such as a different price, product model, or specification — that appeared in multiple responses.

"In short, it says that these LLMs aren't there yet when it comes to this end-to-end experience," said Dakota Nunley, Product.ai's head of search product.

AI has been creeping further into shopping, with both retailers and LLMs offering tools to find and buy things online. AI agents such as Meta's Muse or Instinct could do even more of the work for you.

Product.ai's study points to a more fundamental issue: AI still struggles to provide accurate information that humans or AI agents need.

"Perplexity is the only AI company relentlessly focused on achieving 100% accuracy, and we lead the industry in every measure of it," a Perplexity spokesperson said. Representatives for the other three LLMs did not respond to requests for comment.

The study found that 97% of head-to-head questions, such as asking the AI models to compare two products, produced a conflict. That was higher than the 75% conflict rate for straightforward factual or specification questions.

The models also struggled to provide users with the correct product prices. Of the 913 answers that Product.ai could verify, 85% matched the current or listed price.

There were differences between the LLMs. Gemini had the highest share of questions with what Product.ai classified as a costly error: 56% on its free tier and 54% on its paid tier.

The results using Claude improved with its paid version, with costly errors falling to 21% from 44%. Perplexity had the lowest costly-error rate, at 14% for its paid tier, followed by ChatGPT's paid version at 17%.

Product.ai also found that the tools sometimes contradicted themselves when asked the same question repeatedly. Gemini's free tier did so on 29% of questions, the highest rate in the study.

When the answers were wrong, though, the price was off by a median of $300, Product.ai said.

Such a big difference should be enough to give shoppers pause about relying on AI agents, Nunley said. "That's a big miss, right there," he said.

Instead of handing over payment details to agents or taking whatever AI agents say as gospel, Nunley said, users need to verify the information themselves. That can mean asking multiple AI engines and comparing the results — or simply checking the seller's websites themselves, he said.

"Use AI this upcoming holiday season as a discovery tool, as a thought partner," Nunley said. "But be wary of going hands-off at the moment."

Do you have a story idea about AI and shopping? Contact this reporter at abitter@businessinsider.com or via encrypted messaging app Signal at 808-854-4501. Use a personal email address, a nonwork WiFi network, and a nonwork device; here's our guide to sharing information securely.

Read the original article on Business Insider
Startups Hardware AI

Intelligence Exchange

0

Log in to participate in the exchange.

Sign In

Syncing Discussions...

Finding Related Intelligence...
AI can't take over your holiday shopping quite yet, a new study suggests. | TechCulture