AI shopping agents can make different choices from the same catalogue because source order, stored context and tool design influence the decision. New Wharton research shows why buyers should keep final approval before checkout.

Key takeaways

  • AI agents may make different shopping choices from the same request.
  • Small changes in price, stock or wording can affect their picks.
  • Shoppers should check the final item, seller, price and return rules.
  • Retailers may need clearer product data and stronger trust signals.

AI agent shopping means using software that searches for products and can act for you. But these agents don’t always behave like steady shopping assistants. A Forbes report says their choices can shift between searches, even when the buyer gives the same instructions. That creates risks for shoppers, sellers and the companies building these tools.

The promise sounds simple. You tell an agent, “Find me running shoes under $100,” and it searches online. It may compare products, read reviews and suggest a purchase.

In practice, the result can depend on many moving parts. A product may sell out, a price may change, or a website may return results in a new order. The agent may also weigh words such as “best,” “cheap” or “fast” in different ways.

Why can AI agent shopping produce different results?

People often assume software will repeat the same answer every time. That assumption fails with many modern AI tools, because they generate responses rather than follow one fixed list of steps.

An AI model predicts what to say or do next. Technical teams call this system “probabilistic,” which means it works with likely choices instead of one guaranteed path.

That does not mean an agent acts at random. It means several reasonable paths may exist. One search could favour a lower price, while another could favour a higher-rated product or faster delivery.

Websites add another layer of change. Online stores update stock, shipping fees and discounts by the minute. A search repeated 10 minutes later may therefore show a different deal.

The agent may also use different sources. One run might find a brand’s own store. Another might find a marketplace listing from a third-party seller.

What can change an AI shopping result?Pricechanges over timeStockcan sell outWordingshapes the goalSeller datamay be incomplete

This chart shows the main sources of variation. It is a simple guide, not a measurement of one store or one AI product.

What does unpredictable AI agent shopping mean for buyers?

The biggest issue is control. A buyer may think the agent understood a clear budget, but the final choice can include taxes, delivery fees or add-ons.

There is also a difference between finding an item and buying it. Search is usually easy to undo. A purchase can move money, share an address and create a return problem.

For example, a $90 product can pass a $100 limit before a $15 delivery charge. A buyer who relies on the agent’s short answer may miss that detail.

Shoppers should treat an agent as a helper, not an automatic decision-maker. Before paying, check four things: the exact model, the complete price, the seller and the return window.

Buyers should also ask the agent to show its reasons. A useful answer names the products it compared and explains why one won. If it cannot do that, the choice deserves a closer look.

How should retailers respond to AI agent shopping?

Retailers face a new kind of competition. They are not only trying to win a human’s attention. They are also trying to make their products clear to software.

Accurate product feeds can help. A feed is a file that gives shopping services key details, such as price, size, stock and delivery time.

If those details conflict across pages, an agent may ignore the listing or rank it poorly. Clear titles, honest reviews and simple return rules can make a store easier to trust.

Retailers should not try to trick agents with hidden keywords. Such tactics can send buyers to the wrong item and damage trust. They may also create legal trouble if claims mislead customers.

This shift connects with wider debates about safe AI. Our coverage of research on reducing AI failures explains why systems need checks before they act. Rules for online platforms also matter, as shown in our report on the Digital Services Act and AI platforms.

What should an agent show before a purchase?

Check Why it matters
Full price Fees can push the order over budget.
Exact product Similar names may hide different sizes or models.
Seller A marketplace seller may have different service rules.
Returns The buyer needs a clear way to fix a bad choice.

Agents should show these details in plain language. They should also ask for confirmation before any purchase, especially above a set amount.

A sensible limit might be $50 for routine items and $0 for medicines or costly electronics. The right limit depends on the user, but the rule should be clear.

The NIST AI Risk Management Framework offers broad guidance for spotting and managing AI risks. It supports a basic idea: systems should be tested, watched and corrected when they fail.

What is the bigger lesson?

AI agent shopping is not broken simply because results can change. Human shoppers change their minds too. The concern starts when software acts with money or personal data without showing enough detail.

Companies need to test agents across repeated searches. A test run five times, 10 times or 100 times can reveal whether the system is stable enough for a task.

Users also need a final say. The best design may combine speed with a pause: the agent finds the options, then the person approves the choice.

That approach keeps the useful part of automation while limiting surprises. Until agents can explain and repeat their choices reliably, shoppers should stay in charge.

FAQs

What is AI agent shopping?

AI agent shopping uses software to search, compare and sometimes buy products for a person.

Why do AI agents change their shopping choices?

Prices, stock, website results and the wording of a request can change. AI models can also choose among several likely actions.

How can shoppers stay safe?

Check the exact item, full cost, seller and return policy. Approve the purchase yourself before money moves.

What the new AI shopping agents study found

Researchers at Wharton Generative AI Labs tested AI shopping agents in roughly 26,000 runs. The agents chose from an eight-product fitness-watch grid, while the researchers changed the context surrounding that grid. Six frontier models were tested, generally with 200 runs per condition.

The baseline was more stable than the headline might suggest. When models saw only the product grid, each tended to settle on a consistent favourite. Instability increased when researchers added realistic context: a review screenshot, several competing sources, a different order for those sources, a planted user memory or a change in how tools delivered information.

Everyone else is reporting unpredictable recommendations; we are explaining the hidden context pipeline that changes an AI shopping agent’s decision. The same catalogue can produce a different purchase when source order, remembered preferences or software implementation changes, even though the shopper never sees those inputs.

Context layers affecting AI shopping agentsProduct grid, source order, user memory and tool design combine before an AI shopping agent makes a choice.PRODUCT GRIDSOURCEORDERMEMORYTOOL-CALLDELIVERYFINALCHOICE
The Wharton experiments show that invisible context can redirect a purchase.
Wharton AI shopping agent study designThe study ran roughly 26,000 tests across six frontier models, with 200 runs per condition for most studies.≈26,000TOTAL TESTS6FRONTIER MODELS200RUNS PERCONDITIONSource: Wharton Generative AI Labs, 27 August 2026
The experiment was designed to test stability, not real-world sales conversion.

Why source order can change an AI shopping agent

One experiment showed all three recommendation sources but shuffled their order. Some models stayed with the same choice, while Gemini 3.1 Flash Lite changed substantially depending on which source appeared first. The result does not mean every model always follows the first source. It shows that ordering can become a decision variable when an agent combines competing evidence.

That matters because shopping tools rarely control the entire information environment. Search ranking, page availability, sponsored placements and retrieval systems all influence which evidence arrives first. A retailer could appear or disappear from an agent’s consideration without changing the underlying product.

User memory is useful and an attack surface

The researchers also inserted a single sentence of supposed user memory. In a test where one low-priced, highly rated product was designed to be the objectively best option, the planted memory redirected most models. The effect varied by model, which means developers cannot assume one safety rule will transfer cleanly across systems.

Personalisation is valuable when the memory is genuine. It becomes dangerous when an external page, malicious instruction or corrupted profile inserts a preference the user never supplied. AI shopping agents should expose which stored preferences influenced a transaction and let users remove or override them before checkout.

Why tool-call design matters to buyers

The least visible finding may be the most important for software teams. Delivering three sources in one bundled tool call produced different choices from delivering them one at a time in the four-model subset tested. For GPT-5.5, the bundled version pushed the choice toward one source far more strongly.

An implementation detail therefore changed the shopping outcome. Two apps using the same model and the same sources can behave differently because their orchestration layers are different. Retailers should test the complete system, not only the underlying model.

The Wharton Generative AI Labs technical report is the primary public summary. The associated full working paper on SSRN provides the research record. Independent reports from The Decoder and retail-technology outlets discussed the implications for buyers and merchants.

AI shopping agents need a confirmation layer

AI shopping agents can be consistent on a clean catalogue yet unstable in a realistic information environment; safe deployment therefore requires context logs, repeated testing and explicit user confirmation before money moves. The right question is not whether one demo selected a good product. It is whether the system remains reliable when sources conflict and context is adversarial.

Get the day’s top stories in your inbox

One concise email. No spam, unsubscribe anytime.