← Day index · Block 6 of 10 · ← Previous: Noam Brown’s program · Next: Measurement →
Time budget
~60 minutes. Read Banchio & Skrzypacz and Kolumbus & Nisan in full (25 min); read the abstracts of NPGA, SODA, and Gaitonde et al. (15 min); skim the collusion section and keep the convergence table for reference. This block answers the question Block 5 left open: when every bidder in a sponsored-search market is a learning algorithm, what does the market do?
The question
Block 2 analyzed auctions among autobidders at equilibrium: value maximizers with ROAS constraints, uniform bid scaling, a price of anarchy of 2. Block 4 built the bidders: bandits, constrained RL, diffusion policies. Block 5 argued the right primitive for a bidder is regret minimization. Put those together and a hard question falls out. If hundreds of thousands of regret-minimizing, budget-paced, ROAS-targeting agents bid against each other on Amazon every hour, does the market converge to the equilibrium Block 2 analyzed? Does it converge at all? And if it converges, to what - the competitive outcome, or something that looks like a cartel nobody formed?
Three literatures answer, from three directions. Computational auction theory (the Munich group around Martin Bichler) asks whether learning can find Bayes-Nash equilibria of auctions where no closed form exists. Learning theory (Google Research, Cornell, Microsoft) asks what guarantees hold along the trajectory of no-regret dynamics, with or without convergence. And industrial organization asks whether learning agents collude. The answers are, respectively: yes but conditionally; yes and the guarantees do not need convergence; and it depends on the auction format in a way that matters for Amazon specifically.
Part I: Computing equilibria of auctions by learning
Bayes-Nash equilibria of auctions are known in closed form for a handful of symmetric cases. Position auctions with heterogeneous bidders, budgets, and relevance weights are not among them, so a bidder who wants to know “what would rational rivals do here” has no formula to consult. The Munich program treats this as a learning problem.
NPGA (Bichler, Fichtl, Heidekrüger, Kohring, Sutterer, Nature Machine Intelligence 2021) [1] parameterizes each bidder’s strategy as a neural network from valuation to bid and updates it in self-play by neural pseudogradient ascent - an evolution-strategies estimate of the gradient of ex-ante expected utility, needed because an auction’s allocation is a step function of the bid and ordinary gradients are zero almost everywhere. In every setting with a known analytical equilibrium (first-, second-, third-price, multi-unit, some combinatorial auctions), the learned strategies land on it, verified by estimating the ex-post utility loss against a computed best response. The open-source framework, bnelearn, ships those settings; it does not ship a position auction [1].
SODA (Bichler, Fichtl, Oberlechner, EC 2023 / Operations Research 2025) [2] takes the other route: discretize both values and bids, represent each bidder by a distributional strategy (a probability over bid levels for each value bucket), and run simultaneous online dual averaging. Because expected utility is linear in distributional strategies, convergence to a pure strategy is a certificate that you have an approximate equilibrium of the discretized game, and the paper proves the discretized equilibrium approximates the continuous one. Symmetric benchmarks are recovered in seconds [2]. Follow-ups extend to asymmetric bidders [3], first-order gradients through a smoothed surrogate allocation [4], and multi-stage auctions and contests solved by deep RL self-play with a verifier [5].
Why does gradient learning find these equilibria when the standard theory says it need not? Bichler and coauthors [6] reformulate the Bayes-Nash equilibrium of first- and second-price auctions as an infinite-dimensional variational inequality and show it is not monotone - the property that usually guarantees convergence of gradient dynamics. Second-price satisfies a weaker Minty condition; first-price does not. The equilibrium is nonetheless the unique solution among increasing bid functions, which is why the dynamics work when they work, and why the guarantee is conditional.
For a Sponsored Products bidder. A discretized bid grid per value bucket is exactly how production autobidders represent policies. SODA is a solver for that object and a certificate check: given a model of rival value distributions for a keyword cluster, compute the approximate equilibrium, then ask whether your current policy is near a best response to it. Amazon’s own auction-realism paper does the inverse step - inferring rival value distributions from aggregate bid data [7] - which is the input SODA needs. The gap the sweep found is real: no verified paper in this line handles a relevance-weighted position auction with budgets. That is a well-posed thesis problem.
Section takeaway. Equilibria of auctions with no closed form can be computed by self-play learning, with certificates, for the auction classes that have been tried. Position auctions with budgets and quality weights have not been tried in this line.
Part II: What no-regret dynamics guarantee without converging
The learning-theory literature asks a humbler question. Do not assume the market reaches equilibrium; ask what welfare it achieves while learning.
The single-bidder foundations are clean. Han, Zhou, and Weissman [8] give the minimax-optimal no-regret algorithm for repeated first-price auctions with censored feedback (you see the winning bid only when you lose), by treating the problem as a “partially ordered contextual bandit.” Han, Zhou, Flores, Ordentlich, and Weissman [9] drop the i.i.d. assumption entirely and achieve the same rate against adversarial values and competing bids, benchmarked against all Lipschitz bidding policies. Feng, Podimata, and Syrgkanis [10] model sponsored search directly: a bid determines a slot, value is revealed only on winning, and regret scales logarithmically in the number of bid levels. Aggarwal, Fikioris, and Zhao [11] add budget and ROI constraints in mixed first/second-price formats and prove an lower bound for bandit feedback in first price - the price of seeing only win/loss and price rather than the highest competing bid. Deng, Li, Tang, and Zhang [12] give the most recent single-bidder guarantee for exactly that feedback model: with win-only bandit feedback under a ROI constraint.
The market-level results are the ones that matter for policy. Balseiro and Gur [13] showed adaptive pacing - one multiplier on value, updated by dual descent on expenditure - is asymptotically optimal for an individual budget-constrained bidder against arbitrary rivals, and that when all bidders use it the dynamics converge to an approximate Nash equilibrium. Gaitonde, Li, Light, Lucier, and Slivkins [14] then proved that when all agents run a natural gradient-based multiplicative pacing rule, liquid welfare during learning is at least half of the optimum for any core auction - first-price, second-price, or GSP - without assuming convergence. Lucier, Pattathil, Slivkins, and Zhang [15] extended this to autobidders with both budget and ROI constraints under bandit feedback, keeping the guarantee “whether or not the bidding dynamics converges to an equilibrium.” Fikioris and Tardos [16] give the sequential-auction version: under a behavioral model where each budgeted buyer does at least a -fraction as well as value-shading, sequential first-price has liquid-welfare price of anarchy about 2.41 at , while sequential second-price can be arbitrarily inefficient.
Two 2026 preprints close the loop on complexity. Anagnostides, Gemp, Piliouras, and Spendlove [17] prove that finding an autobidding equilibrium within a factor of the best equilibrium’s welfare is NP-hard, so no learning dynamics should be expected to select the good equilibrium. Chen, Morgenstern, and Yang [18] reconcile PPAD-hardness of autobidding equilibria with the fast convergence seen in practice: the hardness needs atomic value distributions, and with continuous values the equilibrium becomes a separately monotone generalized Nash equilibrium with a solver that has last-iterate linear convergence, covering budget pacing and throttling as special cases.
For a Sponsored Products bidder. Design for guarantees along the trajectory, not at a fixed point. A pacing controller that is provably no-regret for you and provably half-efficient for the market is a better epistemic posture than a policy computed against an assumed equilibrium. This is also Block 5’s Pluribus lesson stated as theorems: strong play without an equilibrium guarantee is available, and the guarantees you can get are about the path.
Section takeaway. No-regret pacing gives individual regret with full information, worse with bandit feedback, and a market-wide liquid-welfare floor of that does not depend on convergence. Selecting a better-than-worst equilibrium is computationally hard.
Part III: Convergence, and to what
Whether learning bidders converge, and to what, depends on details a mechanism designer controls.
| Setting | Result | Source |
|---|---|---|
| Mean-based no-regret learners (Exp3, UCB, -greedy), second price / VCG | Iterates converge with high probability to truthful bidding | Feng, Guruganesh, Liaw, Mehta, Sethi, AAAI 2021 [19] |
| Same learners, first price | Converge to the Bayes-Nash shaded bid | [19] |
| Mean-based learners, repeated first price, fixed values | Nash convergence iff bidders share the highest value (time-average and last-iterate); two: time-average only; one: may fail | Deng, Hu, Lin, Zheng, WWW 2022 [20] |
| No-regret learners, symmetric IPV first price | Bayesian coarse correlated equilibrium unique only under a strictly concave prior or increasing strategies; otherwise low-price pooling is possible | Ahunbay & Bichler, SODA 2025 [21] |
| Common online learners, display first/second price, incl. ROI utilities | No systematic deviation from equilibrium; revenue equivalence between formats fails under ROI utilities | Bichler, Gupta, Oberlechner, ISR 2025/2026 [22] |
| Adaptive pacing, all bidders, second price with budgets | Converges to a tractable steady state, an approximate Nash equilibrium | Balseiro & Gur [13] |
| Multiplicative pacing equilibria, second price | Exist, can be multiple with very different welfare and revenue; welfare-optimal one NP-hard | Conitzer, Kroer, Sodomka, Stier-Moses, OR 2022 [23] |
| Pacing equilibria, first price | Unique, monotone, computable by Eisenberg-Gale convex program | Conitzer et al., Management Science 2022 [24] |
| Randomized truthful auctions, learning agents | Non-convergence to truthful bidding for general deterministic truthful auctions; learning-rate ratios matter; randomization can beat second price with reserves over long horizons | Aggarwal, Gupta, Perlroth, Velegkas, 2024 [25] |
Two rows deserve emphasis. Feng et al. [19] is reassuring for a second-price-like mechanism: generic learners drift toward truthful value bidding. Aggarwal et al. [25] qualifies it: for general deterministic truthful auctions, learning agents need not converge to truthfulness, and the relative learning rates of the agents change the outcome. Amazon’s relevance-weighted GSP with hidden reserves is neither a plain second-price auction nor a generic truthful mechanism, so neither row applies cleanly. The market is in a regime the theory has bracketed but not pinned.
Section takeaway. Second-price-like formats pull learners toward truthful bidding; first-price formats pull them toward shaded equilibria that may or may not be unique. Pacing under second price admits multiple equilibria; under first price it admits one.
Part IV: The delegation meta-game
There is a layer above the bidding agents that the literature only recently noticed. Advertisers do not bid; they set a target (a ROAS, a daily budget, a maximum CPC) and hand it to an algorithm - Amazon’s dynamic bidding, a third-party tool, their own. Kolumbus and Nisan [26] analyze exactly this: users delegate to regret-minimizing agents in repeated auctions. In second-price formats, users have an incentive to misreport their value to their own agent; in first-price formats, truthful reporting to the agent is dominant. The follow-up [27] generalizes it to a meta-game among users of learning agents across game classes, with equilibria unlike the standard predictions.
The implication for a platform is uncomfortable. “Value maximizer with truthfully reported ROAS target” - the modeling assumption of the entire Block 2 literature - is itself an equilibrium claim about the meta-game, and Kolumbus-Nisan say it fails in GSP-like formats. The implication for an advertiser is an opportunity: the target you feed your bidder is a strategic variable, and the platform cannot treat it as truthful. Block 5’s DiL-piKL analogy - the regularization strength is part of the strategy - is the same observation from the other side.
Guan, Zhang, Feng, and Lin [28] push one level further, to coordination within a bidding entity. For ROAS-constrained autobidders under common control - an agency, a multi-brand seller - having only the group’s highest-value bidder compete while the others stand aside provably dominates independent bidding in both constraint adherence and total value, for a broad class of algorithms, and the result holds on real data. The strategic unit is the account group, not the campaign.
Section takeaway. The inputs to autobidders are not truthful in second-price-like formats, and coordinated entities have a provable incentive to withhold all but their top bidder. Any equilibrium analysis that starts from “each campaign bids its true target” has assumed away a game.
Part V: Collusion without a cartel
Calvano, Calzolari, Denicolò, and Pastorello [29] showed in 2020 that tabular Q-learners in repeated Bertrand competition learn supracompetitive prices sustained by genuine reward-punishment schemes, with no communication and no instruction to collude. The question for auctions is whether bidders do the same thing in reverse: learn to bid low.
Banchio and Skrzypacz [30] gave the sharpest answer. Stateless Q-learners in repeated auctions with known identical values behave completely differently by format. In first price, bids ratchet down to tacitly collusive levels well below value: the incentive to win by a single increment, combined with exploration shocks, drives a downward drift that nobody corrects. In second price, the same learners converge to truthful bidding. Disclosing the minimum bid to win restores competition in first price. Banchio and Mantegazza [31] then showed collusion can arise even without memory, through “spontaneous coupling” - an endogenous correlation between the learners’ value estimates - and that asynchronous updates or richer feedback break it. Abada and Lambin [32] offer the deflationary reading: much apparent RL collusion is under-exploration, not sophistication, and forced exploration restores competition. Klein [33] shows asynchronous moves alone are no protection. Brown and MacKay [34] show that even without learning, asymmetric update frequencies plus commitment raise equilibrium prices. Fish, Gonczarowski, and Shorrer [35] find LLM pricing agents reach supracompetitive prices and that innocuous prompt wording changes the outcome; the mechanism, by their interpretability analysis, is fear of a price war.
Against this, Bichler, Gupta, and Oberlechner [22] find no systematic deviation from equilibrium for common online learners in display auctions, and Ahunbay and Bichler [21] pin down when no-regret dynamics can pool at low prices in first price: when the prior is not strictly concave and strategies are not forced to be increasing. The disagreement is not about facts; it is about which learners, which feedback, and which format.
The empirical side has shifted from pricing to advertising. Decarolis, Goldmanis, and Penta [36] show a common agency bidding for several clients in a GSP auction can coordinate bids - placing one client just below another to cut the other’s CPC - and that GSP is more vulnerable to this than VCG in both revenue and efficiency. Decarolis and Rovigatti [37], on roughly 40 million Google keyword auctions, find that merger-induced increases in agency concentration reduce platform revenue by about 11%, consistent with coordinated bid suppression. Zhao and Berman [38], with multi-agent RL simulations plus a large Amazon.com dataset, find that under high search costs learners coordinate on lower ad bids, which lowers consumer prices below competitive and expands demand; raising reserve prices does not restore platform profit but commission changes do. Hartline, Long, and Zhang [39] propose the audit: an approximately optimizing algorithm can retain enough of its own data to demonstrate non-collusion, a coordinated one cannot, and the platform that sees every bid can run the test per keyword cluster.
For Amazon Sponsored Products. The format matters in the right direction: a GSP-style CPC rule is structurally resistant to the Banchio-Skrzypacz ratchet, which is a first-price phenomenon. The exposure is elsewhere. GSP is more exploitable than VCG by common agencies [36], the third-party tools that bid for many competing sellers on the same keywords are exactly common agencies, stand-aside coordination is provably profitable for account groups [28], and Zhao-Berman’s Amazon-data result says coordination would show up as suppressed CPCs, so ad revenue rather than consumer price is the exposed margin [38]. The FTC complaint discussed in Block 8 alleges the platform used hidden reserves to raise prices; the collusion literature describes the mirror-image risk, bidders using shared tools to lower them. Both are consistent with the same auction.
Section takeaway. Learning bidders collude in first price and not in second price, when they collude at all; the mechanism can be memory, coupling, or under-exploration. For a GSP platform the live risk is common-agency coordination, which the platform can audit because it sees every bid.
Part VI: Amazon’s own view of learning bidders
Amazon’s auction group has published one paper that sits squarely in this block. Chen, Nabi, and Siniscalchi [7] argue real sponsored-search auctions violate GSP assumptions in four ways - values and CTRs vary by query while bids are set per targeting clause, competitors are unobserved and changing, feedback is partial and aggregated, and the pricing rule is only partly disclosed - and then model advertisers as adversarial-bandit learners (Hedge, EXP3-IX) on a bid grid facing allocation by bid times CTR and a soft floor. In the symmetric single-query benchmark the learners reproduce revenue equivalence, so soft floors are neutral, matching Zeithammer’s theory [40]; with multi-query targeting soft floors raise revenue even with symmetric bidders; under realistic asymmetric types, tuned hard reserves dominate soft floors. The paper also inverts the simulator to infer value distributions from aggregate bid data. A related industrial paper from Alibaba, MAAB [41], builds explicit anti-collusion machinery into a multi-agent RL bidder: “bar agents” set personalized bid floors to prevent collusive underbidding, an engineering acknowledgment of the Part V risk.
The two Amazon measurement papers that belong here are Jain, Hut, Islam, and Pan on cross-unit spillovers in ads A/B tests [42] - treated advertisers compete in the same auctions as controls, so SUTVA fails and spillovers must be measured through the “treatment intensity” of cluster peers - and the SERP interference network work that follows it. They are the bridge to Block 7: once you accept that bidders are learning agents, every experiment on them is an experiment on a dynamical system.
Section takeaway. Amazon has published a learning-bidder model of its own auction that shows mechanism conclusions flip depending on how rivals adapt. Treat it as the platform’s stated modeling stance, not as a disclosure of the production rule.
What to build
The block reduces to four engineering commitments for anyone building or auditing a bidder in this market.
- Compute the equilibrium you are supposedly in. Fit rival value distributions per keyword cluster (the inversion in [7]), solve with SODA-style dual averaging on a bid grid [2], and measure your policy’s utility loss against a best response. If the loss is large, you are either exploitable or exploiting; find out which.
- Design for trajectory guarantees. Use a pacing rule with an individual no-regret guarantee and a market-wide liquid-welfare floor [14][15]. Do not assume convergence, and do not assume the equilibrium you converge to is a good one [17].
- Treat targets as strategic. The ROAS target you hand a bidder is a move in a meta-game [26][27]; the account group, not the campaign, is the strategic unit [28].
- Audit for coordination. If you are the platform, run the Hartline-Long-Zhang test [39] per keyword cluster and watch agency concentration [37]. If you are an advertiser, know that the platform can.
Reading list for this block
- Banchio & Skrzypacz. Artificial Intelligence and Auction Design. EC 2022. 15 min. The format-dependence of learned collusion.
- Kolumbus & Nisan. Auctions between Regret-Minimizing Agents. WWW 2022. 15 min. Why autobidder inputs are strategic.
- Bichler et al. Learning equilibria in symmetric auction games using artificial neural networks. Nature Machine Intelligence 2021. 10 min (abstract and Figure 1). Neural equilibrium learning.
- Bichler, Fichtl & Oberlechner. Computing Bayes Nash Equilibrium Strategies in Auction Games via Simultaneous Online Dual Averaging. OR 2025. 10 min. Solver plus certificate.
- Gaitonde, Li, Light, Lucier & Slivkins. Budget Pacing in Repeated Auctions: Regret and Efficiency without Convergence. ITCS 2023. 10 min. The half-welfare floor along the path.
- Feng, Guruganesh, Liaw, Mehta & Sethi. Convergence Analysis of No-Regret Bidding Algorithms in Repeated Auctions. AAAI 2021. 5 min. Second price pulls learners to truth.
- Decarolis, Goldmanis & Penta. Marketing Agencies and Collusive Bidding in Online Ad Auctions. Management Science 2020. 10 min. GSP’s agency exposure.
- Chen, Nabi & Siniscalchi. Advancing Ad Auction Realism. AdKDD 2023. 10 min. Amazon’s learning-bidder model.
Questions to carry forward
- Which of the convergence rows in Part III actually applies to a relevance-weighted GSP with hidden hard and soft reserves - and can that be tested from Marketing Stream data alone?
- If the meta-game makes reported ROAS targets non-truthful, what does the platform’s “suggested bid” (Block 8) mean, and what should an advertiser do with it?
- Zhao-Berman predict coordination shows up as suppressed CPCs. Does the 50% fall in average winning bids that Amazon reports (Block 8) have a coordination component, a relevance component, or both, and what experiment separates them?
- Is there a SODA-style solver for position auctions with budgets and quality weights? Nobody has published one.
References
- Bichler, Fichtl, Heidekrüger, Kohring & Sutterer. Learning equilibria in symmetric auction games using artificial neural networks. Nature Machine Intelligence 2021. doi.org/10.1038/s42256-021-00365-4 · code: bnelearn
- Bichler, Fichtl & Oberlechner. Computing Bayes Nash Equilibrium Strategies in Auction Games via Simultaneous Online Dual Averaging. EC 2023; Operations Research 2025. arXiv:2208.02036 · doi.org/10.1287/opre.2022.0287
- Bichler, Kohring & Heidekrüger. Learning Equilibria in Asymmetric Auction Games. INFORMS Journal on Computing 2023. doi.org/10.1287/ijoc.2023.1281
- Kohring, Pieroth & Bichler. Enabling First-Order Gradient-Based Learning for Equilibrium Computation in Markets. ICML 2023 (venue per project README). arXiv:2303.09500
- Pieroth, Kohring & Bichler. Equilibrium Computation in Multi-Stage Auctions and Contests. 2023 preprint. arXiv:2312.11751
- Bichler, Lunowa, Oberlechner, Pieroth & Wohlmuth. On the Convergence of Learning Algorithms in Bayesian Auction Games. 2023 preprint. arXiv:2311.15398. Survey: Bichler, Durmann & Oberlechner. Agentic Markets: Game Dynamics and Equilibrium in Markets with Learning Agents. 2025. arXiv:2506.18571
- Chen, Nabi & Siniscalchi (Amazon). Advancing Ad Auction Realism: Practical Insights & Modeling Implications. AdKDD 2023. amazon.science · arXiv:2307.11732
- Han, Zhou & Weissman. Optimal No-regret Learning in Repeated First-price Auctions. Operations Research 2025. arXiv:2003.09795
- Han, Zhou, Flores, Ordentlich & Weissman. Learning to Bid Optimally and Efficiently in Adversarial First-price Auctions. 2020 preprint (rev. 2025; no archival venue verified). arXiv:2007.04568
- Feng, Podimata & Syrgkanis. Learning to Bid Without Knowing your Value. EC 2018. arXiv:1711.01333
- Aggarwal, Fikioris & Zhao. No-Regret Algorithms in non-Truthful Auctions with Budget and ROI Constraints. 2024 preprint. arXiv:2404.09832
- Deng, Li, Tang & Zhang. No-Regret Online Autobidding Algorithms in First-price Auctions. NeurIPS 2025 (per arXiv comments). arXiv:2510.16869
- Balseiro & Gur. Learning in Repeated Auctions with Budgets: Regret Minimization and Equilibrium. Management Science 2019. doi.org/10.1287/mnsc.2018.3174
- Gaitonde, Li, Light, Lucier & Slivkins. Budget Pacing in Repeated Auctions: Regret and Efficiency without Convergence. ITCS 2023 (extended 2026). arXiv:2205.08674
- Lucier, Pattathil, Slivkins & Zhang. Autobidders with Budget and ROI Constraints: Efficiency, Regret, and Pacing Dynamics. COLT 2024. arXiv:2301.13306
- Fikioris & Tardos. Liquid Welfare Guarantees for No-Regret Learning in Sequential Budgeted Auctions. EC 2023; Mathematics of Operations Research. arXiv:2210.07502
- Anagnostides, Gemp, Piliouras & Spendlove. Tight Inapproximability for Welfare-Maximizing Autobidding Equilibria. 2026 preprint. arXiv:2602.09110
- Chen, Morgenstern & Yang. Beyond the PPAD hardness of Auto-bidding Auctions. 2026 preprint. arXiv:2608.01889
- Feng, Guruganesh, Liaw, Mehta & Sethi. Convergence Analysis of No-Regret Bidding Algorithms in Repeated Auctions. AAAI 2021. arXiv:2009.06136
- Deng, Hu, Lin & Zheng. Nash Convergence of Mean-Based Learning Algorithms in First-Price Auctions. WWW 2022. arXiv:2110.03906
- Ahunbay & Bichler. On the Uniqueness of Bayesian Coarse Correlated Equilibria in Standard First-Price and All-Pay Auctions. SODA 2025. arXiv:2401.01185
- Bichler, Gupta & Oberlechner. Revenue in First- and Second-Price Display Advertising Auctions: Understanding Markets with Learning Agents. Information Systems Research (2025/2026). arXiv:2312.00243
- Conitzer, Kroer, Sodomka & Stier-Moses. Multiplicative Pacing Equilibria in Auction Markets. Operations Research 2022. arXiv:1706.07151
- Conitzer, Kroer, Panigrahi, Schrijvers, Sodomka, Stier-Moses & Wilkens. Pacing Equilibrium in First-Price Auction Markets. EC 2019; Management Science 2022. arXiv:1811.07166
- Aggarwal, Gupta, Perlroth & Velegkas. Randomized Truthful Auctions with Learning Agents. 2024 preprint. arXiv:2411.09517
- Kolumbus & Nisan. Auctions between Regret-Minimizing Agents. WWW 2022. arXiv:2110.11855
- Kolumbus & Nisan. How and Why to Manipulate Your Own Agent: On the Incentives of Users of Learning Agents. NeurIPS 2022. arXiv:2112.07640
- Guan, Zhang, Feng & Lin. On the Coordination of Value-Maximizing Bidders. ICML 2026 (per arXiv). arXiv:2511.04993
- Calvano, Calzolari, Denicolò & Pastorello. Artificial Intelligence, Algorithmic Pricing, and Collusion. American Economic Review 2020. doi.org/10.1257/aer.20190623
- Banchio & Skrzypacz. Artificial Intelligence and Auction Design. EC 2022. arXiv:2202.05947
- Banchio & Mantegazza. Artificial Intelligence and Spontaneous Collusion (a.k.a. Adaptive Algorithms and Collusion via Coupling). EC 2023. arXiv:2202.05946
- Abada & Lambin. Artificial Intelligence: Can Seemingly Collusive Outcomes Be Avoided? Management Science 2023. doi.org/10.1287/mnsc.2022.4623
- Klein. Autonomous algorithmic collusion: Q-learning under sequential pricing. RAND Journal of Economics 2021. doi.org/10.1111/1756-2171.12383
- Brown & MacKay. Competition in Pricing Algorithms. AEJ: Microeconomics 2023. doi.org/10.1257/mic.20210158
- Fish, Gonczarowski & Shorrer. Algorithmic Collusion by Large Language Models. EC 2026 (per arXiv). arXiv:2404.00806
- Decarolis, Goldmanis & Penta. Marketing Agencies and Collusive Bidding in Online Ad Auctions. Management Science 2020. doi.org/10.1287/mnsc.2019.3457
- Decarolis & Rovigatti. From Mad Men to Maths Men: Concentration and Buyer Power in Online Advertising. American Economic Review 2021. doi.org/10.1257/aer.20190811
- Zhao & Berman. Algorithmic Collusion of Pricing and Advertising on E-commerce Platforms. 2025 preprint. arXiv:2508.08325
- Hartline, Long & Zhang. Regulation of Algorithmic Collusion. CSLAW 2024. arXiv:2401.15794
- Zeithammer. Soft Floors in Auctions. Management Science 2019. doi.org/10.1287/mnsc.2018.3164
- Wen et al. (Alibaba). A Cooperative-Competitive Multi-Agent Framework for Auto-bidding in Online Advertising (MAAB). WSDM 2022. arXiv:2106.06224
- Jain, Hut, Islam & Pan (Amazon). Cross-Unit Spillovers in A/B Testing: Empirical Evidence from Ads. CODE@MIT 2023. amazon.science
Part of A Day on Amazon Ads. The parent note is Amazon Ads, Deeply; the interactive companion is The Ad Auction, From the Inside.