AI agents on shopping platforms; sponsorship bias in evaluations

Whom Do AI Agents Work For? Role Assignment Induces Sponsorship Bias in LLM Recommenders

Whom Do AI Agents Work For? Role Assignment Induces Sponsorship Bias in LLM Recommenders

Abstract Large language models (LLMs) now serve as conversational shopping assistants on platforms that also sell advertising. These AI agents face a conflict of duty. They advise consumers who rely on their judgment, yet are deployed by platforms that benefit when sponsored listings are chosen. Sponsorship disclosures, designed to allow consumers to penalize paid placements, now reach the AI agent rather than the consumer, and the agent’s evaluation of them is hidden from the consumer. Drawing on the fiduciary concept of conflict of duty, we argue that an agent’s evaluation of a sponsored listing should not depend on which party deployed it. In controlled choice experiments, we manipulate assigned roles in the system prompt to name either a traveler or a booking platform as the agent’s principal. Platform delegation significantly attenuates the penalty that agents apply to sponsored listings and weakens the skepticism that disclosure triggers in their reasoning traces. We replicate out findings across LLMs and reasoning depths. A second study decomposes the disclosure label and shows that the divergence between the two delegates widens significantly when the paid placement is attributed to the platform. Stricter terminology (“Sponsored” instead of “Promoted”) lowers choice of paid listings but does not close this gap when the platform is named. The findings show that disclosure mandates designed for human consumers cannot by themselves protect consumers in AI-mediated commerce.

Keywords: large language models, AI agents, conflict of duty, digital fiduciary, sponsorship disclosure

Consumers are increasingly delegating purchase decisions to artificial intelligence (AI) agents. Retailers and technology firms have deployed large language models (LLMs) as conversational shopping assistants ( Amazon, 2026 , Walmart, 2024 ) , while a growing share of consumers use general-purpose chatbots to discover and evaluate products ( Pandya, 2025 , Balaskas, 2026 , Wadi and Ma, 2026b , Hasselwander et al., 2026 , Kumar et al., 2026 , Mogaji and Jain, 2024 ) . Simultaneously, these conversational AI interfaces, such as ChatGPT, have begun incorporating advertisements into their chatbots ( OpenAI, 2026 ) . As a result, the AI platforms used by consumers for recommending the optimal product are also the place where advertisers pay for their product to be recommended.

This dual role subjects the AI agent to a structural conflict of duty ( Boatright, 1992 , Carson, 1994 , Laby, 2004 , Perry et al., 2020 , Green, 2007 ) . One duty is owed to the deploying platform, which gains financially when a sponsored listing is recommended. The competing duty is owed to the consumer (i.e., the advisee) who relies on the AI agent as a digital fiduciary ( Balkin, 2020 ) to provide an impartial evaluation of market alternatives ( Motoki and Pinto, 2026 ) . The agent must therefore execute a single evaluative judgment on behalf of two principals with divergent economic interests.

In traditional e-commerce, advertising disclosures were designed to preserve consumer autonomy. A shopper encounters a “Sponsored” label, processes the signal with skepticism, and discounts the listing’s claims accordingly, which is a chain of events extensively documented in consumer research ( Friestad and Wright, 1994 , Campbell and Kirmani, 2000 , Obermiller and Spangenberg, 1998 , Sahni and Nair, 2020 ) . Regulators mandate sponsorship tags to facilitate this evaluative judgement ( Evans and Park, 2015 , Boerman et al., 2017 ) . However, when an AI assistant evaluates market alternatives or makes purchases on consumers’ behalf ( Google DeepMind, 2026 , e.g.,) , it would have to evaluate sponsored (vs. organic) alternatives.

The ethical inquiry is how AI assistants resolve this conflict of duty and what this resolution entails for consumer protection. Answering it requires a normative benchmark, since algorithmic behavior can only be classified as a fiduciary failure against a defined standard of duty. The standard we adopt is grounded in procedural fairness. The identity of the delegating party (the platform versus the consumer) is not an attribute of the products under evaluation, so an identical listing must receive an identical assessment regardless of which party the AI agent is delegated by. Because the human shopper relies on the agent’s recommendation and cannot observe how it was reached, a systematic departure from this benchmark constitutes an ethical breach. We refer to the departure that favors the platform as sponsorship bias, that is, a more favorable evaluation of sponsored listings when the platform (vs. the consumer) is the agent’s principal.

This normative benchmark is empirically testable. We investigate this across a series of controlled choice experiments by manipulating only the delegating party and observing whether the AI agent’s evaluative judgment shifts. In Study 1, we presented a large language model (LLM) with matched hotel listings and altered a single sentence in the system prompt to define the AI agent as a delegate of either a traveler or a booking platform. When delegated by the traveler, the agent penalized the sponsored listing and chose it 50.2 percentage points less often than a similar organic listing. Under platform delegation, this penalty significantly attenuated to 29.2 percentage points. We replicate this finding across several foundation models and reasoning depths. Analysis of the agent’s pre-choice reasoning traces reveals that platform delegation systematically reduces the skepticism that sponsorship disclosures trigger. Study 2 decomposes the disclosure label to isolate the semantic trigger of this bias. We find that the fiduciary divergence is specifically amplified when the paid placement is attributed to the platform. When provided with the label “Promoted by the platform,” the platform-delegated AI agent abandons the penalty, shows a strong preference for the listing, and chooses it 74.2% of the time (vs. 34.6% for the consumer-delegated agent).

This algorithmic bias presents a challenge to current governance frameworks. First, the misalignment surfaces without any explicit programming or instructions to favor advertisers. Because this bias relies solely on a minimal semantic role assignment, traditional accountability frameworks designed to audit code for deliberate human malfeasance cannot detect it. Second, mandating stricter disclosure terminology fails to resolve the conflict. As our findings demonstrate, substituting the ambiguous term “Promoted” with the explicit term “Sponsored” reduces the overall selection rate of the paid listing, yet leaves a substantial gap between the two delegates when the placement is attributed to the platform.

This research makes three contributions to the literature on algorithmic accountability and business ethics. First, it documents an emergent alignment with the employing party that arises from role assignment rather than programmed goals ( Martin, 2019 , Matthias, 2004 , Nissenbaum, 1997 , Santoni de Sio and Mecacci, 2021 ) . Second, it empirically demonstrates that canonical transparency mandates, such as advertising disclosures, fail to eliminate this bias and enforcing stricter disclosure vocabulary does not close the fiduciary gap. Third, it isolates the semantic trigger for this algorithmic bias and shows that attributing the paid placement to the platform itself is what activates the AI agent’s conflicting loyalties.

The ethical implications extend beyond the context of e-commerce bookings. Any firm deploying an AI assistant to simultaneously advise consumers and monetize those same interactions engenders an identical conflict of duty. For corporate governance, auditing algorithms for explicit biased directives is insufficient, as the behavior could be biased, even when the explicit instructions are innocuous.

When entrusted with surrogate decision-making on behalf of a user, an AI agent occupies a structural role that legal and policy scholars describe as a digital fiduciary ( Balkin, 2015 , Richards and Hartzog, 2021 ) . What defines this role is the set of duties it carries, mainly the obligation to exercise judgment on behalf of another party. Algorithmic intermediaries, however, are rarely responsible to a single party. Recommender systems have long been characterized as multi-stakeholder environments, designed to jointly serve consumers, platforms, and third-party advertisers whose interests frequently diverge ( Abdollahpouri et al., 2020 , Milano et al., 2020 ) . An AI agent deployed within such an environment therefore owes obligations to more than one party.

Prior ethical and technical literature has, however, predominantly analyzed these intermediaries as ranking and filtering algorithms ( Rokach and Ricci, 2011 , i.e., engines programmed to surface, order, or suppress items within a fixed catalog;) . In such systems, obligations to competing stakeholders are balanced through an explicit objective function, in which the weight given to each party is set by the system’s designers and can, in principle, be inspected ( Abdollahpouri et al., 2020 ) . LLM-based agents depart from this paradigm. Because LLMs can generate personalized, zero-shot recommendations from semantic context and natural-language inputs alone ( He et al., 2023 , Wang et al., 2023 ) , they exercise open-ended judgment rather than executing a fixed ranking. The party on whose behalf that judgment is exercised can be specified in a single sentence of context, with no formula to determine how the agent weighs its obligations to that party against its obligations to another. This raises the question of how an AI agent with duties to several parties resolves conflicts among its obligations.

The business ethics literature has traditionally analyzed divided loyalty as a conflict of interest, in which a professional’s personal interest threatens the proper exercise of judgment on behalf of another ( Boatright, 1992 , Carson, 1994 ) . Fiduciary law, however, recognizes a second and structurally distinct form of conflict. A conflict of duty arises when a fiduciary owes obligations to two parties whose interests diverge, such that fulfilling the duty owed to one compromises the duty owed to the other ( Laby, 2004 , Perry et al., 2020 , Conaglen, 2009 ) . This form of conflict does not depend on the agent gaining anything personally. The distinction matters for AI agents because an LLM does not itself profit when a sponsored listing is recommended. Instead, the conflict is between its duties to the parties whose interests diverge.

Conflicts of duty are common in professional settings where one agent acts on behalf of several principals at once. In dual agency, for instance, a real estate broker represents both buyer and seller in the same transaction ( Gardiner et al., 2007 ) . The broker owes loyalty to each party, yet a higher sale price serves the seller at the buyer’s expense. Thus, no single course of action fully discharges both duties. The parties who rely on the broker’s judgment typically cannot observe how these conflicting duties were balanced. The same structure describes an AI agent deployed by a platform to advise consumers.

Disclosure is the standard remedy that law and regulation use to mitigate such conflicts. Instead of prohibiting the agent from serving two parties, disclosure informs the party at risk of the conflict so that they can discount the agent’s judgment or seek advice elsewhere ( Green, 2007 ) . Sponsorship disclosure in digital commerce serves this function. Labeling a listing as sponsored signals that the platform has a stake in the consumer’s choice and invites the consumer to evaluate the listing accordingl