← Day index · Block 9 of 10 · ← Previous · Next →

This is the card catalogue for the day. Every paper mentioned across the ten blocks appears once, filed under the theme where it does the most work, with a one-line idea and a one-line reason it matters for a Sponsored Products bidder or for the platform running the auction. Where the underlying research record left an uncertainty - a venue that could not be confirmed, a headline number that two sources disagree on - the flag is kept in a parenthetical rather than smoothed over. [Brown] marks papers on which Noam Brown is an author. Part B reorders the load-bearing subset chronologically; Part C is the emergency version of the day. The map this catalogue serves is Amazon Ads, Deeply.

The reading convention: the block link after each item says where in the day the paper is actually discussed, so you can jump from the bibliography into the argument and back.

Part A - By theme

1. Mechanism and classical auction theory

  • Myerson. Optimal Auction Design. Mathematics of Operations Research 1981. doi - revenue-optimal single-item auctions are virtual-value maximizers with a reserve. ⟶ The reason every platform’s reserve price exists, and the benchmark learned mechanisms are measured against. read in Block 1
  • Edelman, Ostrovsky, Schwarz. Internet Advertising and the Generalized Second-Price Auction. AER 2007. aeaweb - GSP is not truthful but has locally envy-free equilibria whose payments coincide with VCG. ⟶ The spine of Amazon’s pricing; explains why the market does not collapse into constant re-shading. read in Block 1
  • Varian. Position Auctions. Int. J. Industrial Organization 2007. berkeley - slot value equals advertiser value times slot click factor; symmetric Nash equilibria and revenue bounds. ⟶ The cleanest model of why position, not just winning, is what a bid buys. read in Block 1
  • Lahaie, Pennock. Revenue Analysis of a Family of Ranking Rules for Keyword Auctions. EC 2007. acm - “squashing” the quality weight between pure bid-rank and pure quality-rank; neither extreme is revenue-optimal. ⟶ The relevance exponent in score = b·q^α is a tunable revenue lever, and Amazon’s “increasingly weighted relevance” is a move along this dial. read in Block 1
  • Ostrovsky, Schwarz. Reserve Prices in Internet Advertising Auctions: A Field Experiment. EC 2011 / JPE 2023. jpe · ec’11 · working paper - Yahoo randomized theory-based reserves; effects concentrated on high-volume, few-bidder keywords; the normalized +12.85% headline shrinks to roughly +8% in the journal version and the raw +3.8% revenue-per-keyword estimate is not significant. ⟶ Reserves work, unevenly, and the honest field evidence has wide error bars. (Working paper dated 2016, journal 2023.) read in Block 1
  • Athey, Ellison. Position Auctions with Consumer Search. QJE 2011. doi - consumers search top-down, so slot values are endogenous to who is shown. ⟶ Sponsored Products click curves are not fixed; the ad mix changes them. read in Block 1
  • Milgrom. Simplified Mechanisms with an Application to Sponsored-Search Auctions. Games and Economic Behavior 2010. doi - restricting the message space (one bid per keyword) removes bad equilibria. ⟶ Why a single per-keyword bid is a feature of the design, not a limitation. read in Block 1
  • Ghose, Yang. An Empirical Analysis of Search Engine Advertising. Management Science 2009. doi - first large-scale empirics on keyword auctions, CTR and CVR by position and keyword type. ⟶ The template for measuring what a keyword is worth from logs. read in Block 1
  • Yang, Ghose. Analyzing the Relationship Between Organic and Sponsored Search Advertising. ISR 2010. nyu - organic and paid listings positively interact. ⟶ Turning off ads on a term you rank for organically is not free. read in Block 7
  • Yang, Xiao, Wu. Learning and Pricing Models for Repeated Generalized Second-Price Auction in Search Advertising. EJOR 2020. doi - reserve-price learning with revenue regret in repeated GSP over heterogeneous slots. ⟶ The platform-side learning problem written in GSP terms. read in Block 1
  • Aggarwal, Goel, Motwani. Truthful Auctions for Pricing Search Keywords. EC 2006. doi - the laddered auction: a truthful position auction with the same allocation as GSP. ⟶ Truthfulness in position auctions is possible; GSP chose not to be. read in Block 1
  • Caragiannis, Kaklamanis, Kanellopoulos, Kyropoulou, Lucier, Paes Leme, Tardos. Bounding the Inefficiency of Outcomes in Generalized Second Price Auctions. JET 2015. doi - GSP price of anarchy 1.282 for pure Nash, 2.927 for Bayes-Nash. ⟶ The worst-case efficiency loss of the mechanism you are bidding into. read in Block 1
  • Wilkens, Cavallo, Niazadeh. GSP: The Cinderella of Mechanism Design. WWW 2017. doi - GSP is truthful for value maximizers with a ROAS constraint. ⟶ If you are a value maximizer, bidding your target-adjusted value in GSP is optimal; the autobidding-world theory starts here. read in Block 1
  • Bergemann, Dütting, Paes Leme, Zuo. Calibrated Click-Through Auctions. WWW 2022. arXiv - when the platform’s CTR signal is calibrated, welfare-optimal auction design changes. ⟶ The pCTR term in score = b·q is an information-design object. read in Block 1
  • Paes Leme, Pál, Vassilvitskii. A Field Guide to Personalized Reserve Prices. WWW 2016. doi - per-bidder reserves in position auctions, with practical algorithms. ⟶ Reserves are set per advertiser, not per keyword. read in Block 1
  • Mohri, Muñoz Medina. Learning Theory and Algorithms for Revenue Optimization in Second-Price Auctions with Reserve. ICML 2014. pmlr - learning-theoretic guarantees for learned reserves. ⟶ The platform’s reserve is a learned function of features. read in Block 1
  • Derakhshan, Golrezaei, Paes Leme. Linear Program-Based Approximation for Personalized Reserve Prices. Management Science 2022. doi - LP approximation for personalized reserves in eager second-price auctions. ⟶ How the reserve layer the FTC complaint describes could be computed. read in Block 1
  • Zeithammer. Soft Floors in Auctions. Management Science 2019. doi - soft floors (a reserve below which the auction becomes first-price) do not raise revenue in the standard model. ⟶ The theory of the exact mechanism at issue in the FTC complaint; read against Amazon’s realism paper, which finds soft floors help in richer environments. read in Block 1

2. Autobidding-world theory and pacing

  • Aggarwal, Badanidiyuru, Mehta. Autobidding with Constraints. WINE 2019. springer - optimal single-agent bidding under general affine constraints; equilibrium exists; PoA ≥ 1/2. ⟶ The founding paper of the “value maximizer with a ROAS target” model, which is what every Sponsored Products advertiser actually is. read in Block 2
  • Deng, Mao, Mirrokni, Zuo. Towards Efficient Auctions in an Auto-bidding World. WWW 2021. arXiv - boosts (additive score adjustments) improve welfare and revenue with ROAS and budget constraints. ⟶ A quality “boost” is not a hack; it is a welfare lever. read in Block 2
  • Balseiro, Deng, Mao, Mirrokni, Zuo. The Landscape of Auto-bidding Auctions: Value versus Utility Maximization. EC 2021. acm - first-best revenue is achievable for value maximizers when either values or targets are private, never when both are; never for utility maximizers. ⟶ Whether your bidder is a value or utility maximizer changes what the platform can extract from you. read in Block 2
  • Balseiro, Deng, Mao, Mirrokni, Zuo. Robust Auction Design in the Auto-bidding World. NeurIPS 2021. arXiv - reserves raise both revenue and welfare, robustly across bidder types and across VCG, GSP and first-price. ⟶ Reserve prices are the one intervention that helps the platform without the platform needing to know who you are. read in Block 2
  • Mehta. Auction Design in an Auto-bidding Setting: Randomization Improves Efficiency Beyond VCG. WWW 2022. arXiv - randomized allocation beats the deterministic PoA of 2. ⟶ A platform facing autobidders can be more efficient by being less predictable. read in Block 2
  • Liaw, Mehta, Perlroth. Efficiency of Non-Truthful Auctions in Auto-bidding: The Power of Randomization. arXiv 2022 (2207.03630) / WWW 2023. arXiv · acm - every deterministic mechanism has PoA ≥ 2; first-price is exactly 2; a randomized auction hits 1.8 for two bidders; the edge vanishes as bidders grow. ⟶ The efficiency numbers you should quote for an autobidding market. read in Block 2
  • Deng, Mao, Mirrokni, Zhang, Zuo. Efficiency of the First-Price Auction in the Autobidding World. NeurIPS 2024 (arXiv 2022). arXiv - PoA 1/2 with autobidders alone, ≈0.457 mixed; machine-learned seller advice moves it smoothly toward 1. ⟶ Platform-supplied “suggested bids” are, in this model, an efficiency instrument. read in Block 2
  • Lucier, Pattathil, Slivkins, Zhang. Autobidders with Budget and ROI Constraints: Efficiency, Regret, and Pacing Dynamics. COLT 2024. arXiv - gradient-based bandit autobidders satisfy constraints, have vanishing regret, and guarantee liquid welfare ≥ 1/2 of optimal without converging. ⟶ You do not need equilibrium for guarantees; you need everyone pacing sensibly. read in Block 2
  • Deng, Golrezaei, Jaillet, Liang, Mirrokni. Multi-channel Autobidding with Budget and ROI Constraints. arXiv 2023. arXiv - per-channel ROI targets can be arbitrarily bad; per-channel budgets can be globally optimal. ⟶ How to split a budget across Sponsored Products, Brands and DSP: by budget, not by ROAS target.
  • Deng, Golrezaei, Jaillet, Liang, Mirrokni. Individual Welfare Guarantees in the Autobidding World with Machine-learned Advice. WWW 2024. arXiv - ML advice to the auctioneer buys per-advertiser welfare guarantees, not just aggregate ones. ⟶ Suggested bids and predicted values as an instrument that protects individual bidders. read in Block 2
  • Golrezaei, Lobel, Paes Leme. Auction Design for ROI-Constrained Buyers. WWW 2021. doi - optimal mechanisms when buyers have return-on-investment constraints. ⟶ The platform-side design for the ACOS-constrained advertiser. read in Block 2
  • Balseiro, Gur. Learning in Repeated Auctions with Budgets: Regret Minimization and Equilibrium. Management Science 2019. doi - adaptive pacing via multiplicative dual updates is asymptotically optimal against arbitrary rivals and converges to an approximate equilibrium when everyone uses it. ⟶ The theoretical justification for the multiplicative pacing controller every production system runs. read in Block 2
  • Balseiro, Besbes, Weintraub. Repeated Auctions with Budgets in Ad Exchanges: Approximations and Design. Management Science 2015. doi - fluid mean-field approximation of budgeted repeated auctions. ⟶ The approximation that makes budget pacing tractable at all. read in Block 2
  • Conitzer, Kroer, Sodomka, Stier-Moses. Multiplicative Pacing Equilibria in Auction Markets. Operations Research 2022 (arXiv 2017). arXiv - pacing multipliers in [0,1] define an equilibrium that exists, may be multiple with very different outcomes, and is NP-hard to optimize. ⟶ Your pacing multiplier is a strategic variable in a game, not a delivery knob. read in Block 2
  • Conitzer, Kroer, Panigrahi, Schrijvers, Sodomka, Stier-Moses, Wilkens. Pacing Equilibrium in First-Price Auction Markets. EC 2019 / Management Science 2022. arXiv - first-price pacing equilibrium is unique, monotone, and computable via an Eisenberg-Gale convex program. ⟶ First-price markets are better behaved for pacing than second-price ones. read in Block 2
  • Balseiro et al. A Field Guide for Pacing Budget and ROS Constraints. ICML 2024. pmlr - catalogue of pacing algorithms; multiplicative feedback controllers dominate in practice. ⟶ The practitioner’s reference for the pacing layer. read in Block 4
  • Gaitonde, Li, Light, Lucier, Slivkins. Budget Pacing in Repeated Auctions: Regret and Efficiency without Convergence. ITCS 2023. arXiv - pacing dynamics give welfare guarantees even when they never settle. ⟶ Stop waiting for convergence; the guarantees do not need it. read in Block 6
  • Fikioris, Tardos. Liquid Welfare Guarantees for No-Regret Learning in Sequential Budgeted Auctions. EC 2023. arXiv - liquid welfare bounds for no-regret budgeted learners. ⟶ The right welfare notion when everyone is budget-capped. read in Block 6
  • Aggarwal and 25 coauthors. Auto-bidding and Auctions in Online Advertising: A Survey. SIGecom Exchanges 2024. arXiv · acm - the field map for autobidding theory, efficiency, boosts and learning. ⟶ Read first for vocabulary or last for synthesis. (Exact ACM venue label unverified in the research record.) read in Block 2
  • Google. An Update on First Price Auctions for Google Ad Manager. Blog, May 2019. google - the industry’s move to unified first-price. ⟶ Context for why bid shading became a literature; not evidence about Amazon’s format. read in Block 2
  • Despotakis, Ravi, Sayedi. First-Price Auctions in Online Display Advertising. Journal of Marketing Research 2021. doi - why exchanges switched and what it does to bidders. ⟶ The economics behind the shading problem. read in Block 2
  • Gligorijevic, Zhou, Shetty, Kitts, Pan, Pan, Flores. Bid Shading in the Brave New World of First-Price Auctions. CIKM 2020. arXiv - ML bid shading for non-censored first-price auctions. ⟶ The template if a placement ever bills you your own bid. read in Block 2
  • Zhou et al. (Yahoo). An Efficient Deep Distribution Network for Bid Shading in First-Price Auctions. KDD 2021. arXiv - predict the minimum-winning-price distribution and shade to maximize surplus. ⟶ Shading is a landscape-forecasting problem in disguise. read in Block 2
  • Qu, Kan. Double Distributionally Robust Bid Shading for First Price Auctions. arXiv 2024. arXiv - robustness to uncertainty in both value and competing-bid distribution. ⟶ The shading model that survives a market shift. read in Block 2
  • Calvano, Calzolari, Denicolò, Pastorello. Artificial Intelligence, Algorithmic Pricing, and Collusion. AER 2020. aeaweb - Q-learners in repeated Bertrand learn supra-competitive prices with punishment-and-forgiveness, without communicating. ⟶ The warning for a market where every bidder is an RL agent. read in Block 2
  • Banchio, Skrzypacz. Artificial Intelligence and Auction Design. EC 2022. arXiv - learning bidders collude more in first-price than second-price auctions. ⟶ Auction format changes how much algorithmic collusion you get. read in Block 2
  • Rawat. Algorithmic Collusion in Auctions: Evidence from Controlled Laboratory Experiments. arXiv 2023-2025. arXiv - randomized experiments over RL and bandit bidders, 500 trials each. ⟶ Lab evidence on bid suppression; generalization to live markets flagged. read in Block 2
  • Aggarwal, Gupta, Perlroth, Velegkas. Randomized Truthful Auctions with Learning Agents. arXiv 2024. arXiv - no-regret agents do not converge to truthful bidding in deterministic truthful auctions; randomized auctions can beat second-price-with-reserve revenue over long horizons. ⟶ “Truthful mechanism” does not imply truthful behavior once bidders learn. read in Block 6
  • Babaioff, Cole, Hartline, Immorlica, Lucier. Non-Quasi-Linear Agents in Quasi-Linear Mechanisms. ITCS 2021. arXiv - what happens when budget- and ROI-constrained agents face mechanisms designed for quasi-linear ones. ⟶ Every classical guarantee about GSP assumed a bidder type Sponsored Products advertisers are not. read in Block 2
  • Liaw, Mehta, Zhu. Efficiency of the Generalized Second-Price Auction for Value Maximizers. WWW 2024. arXiv - PoA of GSP with value-maximizing autobidders. ⟶ The efficiency number for Amazon’s actual mechanism with Amazon’s actual bidder type. read in Block 2
  • Chen, Kroer, Kumar. Throttling Equilibria in Auction Markets. WINE 2021. arXiv - pacing by probabilistic participation rather than bid multipliers; equilibria exist and differ from multiplicative pacing. ⟶ Amazon’s “budget not paced through the day” behavior is closer to throttling than to multiplicative pacing. read in Block 2
  • Balseiro, Kroer, Kumar. Contextual Standard Auctions with Budgets: Revenue Equivalence and Efficiency Guarantees. Management Science 2023. arXiv - revenue equivalence across standard auction formats survives budgets in a contextual setting. ⟶ First-price vs second-price matters less than you think once everyone paces. read in Block 2
  • Feng, Padmanabhan, Wang. Online Bidding Algorithms for Return-on-Spend Constrained Advertisers. WWW 2023. arXiv - regret bounds for ROS-constrained online bidding. ⟶ The ACOS-target bidder as an online learning problem with guarantees. read in Block 2
  • Kitts et al. Ad Serving with Multiple KPIs. KDD 2017. doi - a production controller trading off several KPI constraints. ⟶ The industrial ancestor of USCB. read in Block 2
  • Wang, Yang, Deng, Kong. Learning to Bid in Repeated First-Price Auctions with Budgets. ICML 2023. pmlr - dual-based budgeted bidding in first-price with regret guarantees. ⟶ Budget pacing and shading in one algorithm. read in Block 2
  • Paes Leme, Sivan, Teng. Why Do Competitive Markets Converge to First-Price Auctions? WWW 2020. doi - first-price is the stable outcome of competition among exchanges. ⟶ Why the industry drift toward first-price pricing is structural, not a fad. read in Block 2
  • Pan et al. Bid Shading by Win-Rate Estimation and Surplus Maximization. AdKDD 2020. arXiv - estimate win rate as a function of bid, then maximize expected surplus. ⟶ The simplest correct shading rule. read in Block 2
  • Banchio, Mantegazza. Adaptive Algorithms and Collusion via Coupling. EC 2023. arXiv - a mechanism for why learning algorithms couple into collusion. ⟶ The theory behind the Q-learning result. read in Block 2
  • Assad, Clark, Ershov, Xu. Algorithmic Pricing and Competition: Empirical Evidence from the German Retail Gasoline Market. JPE 2024. doi - margins rose after algorithmic pricing adoption. ⟶ Field evidence that algorithms soften competition. read in Block 2
  • Decarolis, Rovigatti. From Mad Men to Maths Men: Concentration and Buyer Power in Online Advertising. AER 2021. doi - agency concentration lowers prices in ad auctions. ⟶ Who bids matters, not just how. read in Block 2
  • Decarolis, Goldmanis, Penta. Marketing Agencies and Collusive Bidding in Online Ad Auctions. Management Science 2020. doi - agencies coordinating bids in GSP. ⟶ Collusion via a shared bidding tool. read in Block 2
  • Fish, Gonczarowski, Shorrer. Algorithmic Collusion by Large Language Models. EC 2026. arXiv - LLM pricing agents collude. ⟶ The LLM-bidder version of the Calvano result. read in Block 2
  • Abada, Lambin. Artificial Intelligence: Can Seemingly Collusive Outcomes Be Avoided? Management Science 2023. doi - collusion-like outcomes from imperfect exploration. ⟶ Not all supra-competitive pricing is collusion. read in Block 2
  • Klein. Autonomous Algorithmic Collusion: Q-Learning under Sequential Pricing. RAND 2021. doi - collusion under sequential moves. ⟶ Hourly repricing is sequential. read in Block 2
  • Brown, MacKay. Competition in Pricing Algorithms. AEJ Micro 2023. doi - faster algorithms soften competition by commitment. ⟶ Repricing frequency is a strategic variable. (Zach Brown, not Noam.) read in Block 2
  • Hartline, Long, Zhang. Regulation of Algorithmic Collusion. CSLAW 2024. arXiv - regulate via no-regret requirements. ⟶ What a regulator could demand of autobidders. read in Block 2
  • Zhao, Berman. Algorithmic collusion in auctions, 2025. arXiv 2025. arXiv - a 2025 study of algorithmic collusion. ⟶ Recent. read in Block 2
  • Guan, Zhang, Feng, Lin. Algorithmic collusion, 2026. ICML 2026. arXiv - a 2026 study of learning agents and collusion. ⟶ Recent. read in Block 2

3. Learned mechanisms

  • Dütting, Feng, Narasimhan, Parkes, Ravindranath. Optimal Auctions through Deep Learning. ICML 2019. pmlr - RegretNet: allocation and payment networks trained with a regret penalty for incentive violations. ⟶ Mechanism design as constrained learning; the incentive constraint is measured, not assumed. read in Block 1
  • Shen, Tang, Zuo. Automated Mechanism Design via Neural Networks. AAMAS 2019. arXiv - MenuNet, exactly IC by construction via menus. ⟶ The alternative to penalizing regret: build mechanisms that cannot be gamed. read in Block 1
  • Zhang et al. (Alibaba). Optimizing Multiple Performance Metrics with Deep GSP Auctions for E-commerce Advertising. WSDM 2021. arXiv - replace b·q with a learned rank score optimizing a constrained mixture of revenue, CTR, CVR, experience. ⟶ The platform’s relevance multiplier is now a neural network with business KPIs in its loss. read in Block 1
  • Liu et al. (Alibaba). Neural Auction: End-to-End Learning of Auction Mechanisms for E-Commerce Advertising. KDD 2021. arXiv - differentiable sorting makes the whole auction learnable; deployed at Taobao with online A/B wins. ⟶ Proof that a learned auction runs at e-commerce scale. read in Block 1
  • Bai, Xie, Wang (Alibaba). Practical Constrained Optimization of Auction Mechanisms in E-Commerce Sponsored Search Advertising. arXiv 2018 (1807.11790). arXiv - reserve prices and mechanism parameters as constrained optimization over business metrics. ⟶ The pre-neural version of the same idea. read in Block 1
  • Golrezaei, Lin, Mirrokni, Nazerzadeh. Boosted Second Price Auctions. KDD 2021. acm - per-bidder boosts on second-price; up to 6% over standard second price with monopoly reserves. ⟶ A “boost” is the additive cousin of the multiplicative quality weight. (KDD 2021 is the only venue; percentage conditions not fully audited.) read in Block 1
  • Jeunen, Stavrogiannis, Sayedi, Allison (Amazon). A Probabilistic Framework to Learn Auction Mechanisms via Gradient Descent. AAAI 2023 Workshop on AI for Web Advertising. amazon.science - Gumbel noise in a Plackett-Luce allocation gives a differentiable, incentive-compatible auction with a matching price rule. ⟶ Amazon’s own gradient-based mechanism learning; workshop status limits production inference. read in Block 1
  • Sankar et al. Deep Learning Meets Mechanism Design: A Survey. arXiv 2024. arXiv - survey of differentiable economics. ⟶ The index to everything after RegretNet. read in Block 1
  • Ni, Wang, Chen, Yin, Lu et al. (ByteDance). Ad Auction Design with Coupon-Dependent Conversion Rate in the Auto-bidding World. ACM 2023. acm - conversion rates depend on coupons, which changes the mechanism. ⟶ A response model that ignores promotions mis-prices the auction. (Venue/year flagged in the record.) read in Block 1
  • Wang et al. Autobidding Auctions with LLM-Powered Creatives. ICML 2026 (OpenReview). openreview - generated creatives interact with autobidding in a dynamic Stackelberg game. ⟶ Creative quality will eventually be a bidding variable. (Bibliographic status flagged.) read in Block 4
  • Dütting, Feng, Narasimhan, Parkes, Ravindranath. Optimal Auctions through Deep Learning: Advances in Differentiable Economics. JACM 2024. doi - the journal version of RegretNet with the field’s progress since. ⟶ Read this instead of the ICML paper if you have time for one. read in Block 1
  • Rahme, Jelassi, Bruna, Weinberg. A Permutation-Equivariant Neural Network Architecture for Auction Design. AAAI 2021. arXiv - equivariance to bidder and item permutations. ⟶ Learned mechanisms should not depend on who is listed first. read in Block 1
  • Rahme, Jelassi, Weinberg. Auction Learning as a Two-Player Game (ALGnet). ICLR 2021. arXiv - the regret penalty becomes an adversary. ⟶ Mechanism learning as a game, which is what it is. read in Block 1
  • Duan et al. A Context-Integrated Transformer-Based Neural Network for Auction Design (CITransNet). ICML 2022. arXiv - transformers over bidder and item contexts. ⟶ Contextual mechanisms, as a sponsored-search auction is. read in Block 1
  • Ivanov, Safiulin, Filippov, Balabaeva. Optimal-er Auctions through Attention (RegretFormer). NeurIPS 2022. arXiv - attention-based mechanism learning with a better regret budget. ⟶ The state of the art in the RegretNet line. read in Block 1
  • Curry, Sandholm, Dickerson. Differentiable Economics for Randomized Affine Maximizer Auctions. IJCAI 2023. arXiv - exactly strategyproof learned auctions via affine maximizers. ⟶ Truthfulness by construction, learned. read in Block 1
  • Curry, Chiang, Goldstein, Dickerson. Certifying Strategyproof Auction Networks. NeurIPS 2020. arXiv - certify a learned auction’s incentive properties. ⟶ You can audit a neural mechanism. read in Block 1
  • Dütting, Mirrokni, Paes Leme, Xu, Zuo. Mechanism Design for Large Language Models. WWW 2024. arXiv - token auctions: advertisers bid to influence LLM-generated text. ⟶ The auction for the ad in the answer, not beside it. read in Block 1
  • Soumalias, Curry, Seuken. Truthful Aggregation of LLMs with an Application to Online Advertising (MOSAIC). NeurIPS 2025. arXiv - truthful aggregation of LLM outputs for ads. ⟶ Same problem, incentive-compatible. read in Block 1
  • Hajiaghayi, Lahaie, Rezaei, Shin. Ad Auctions for LLMs via Retrieval Augmented Generation. NeurIPS 2024. arXiv - segment auctions inside RAG. ⟶ What sponsored placement looks like when the SERP is a paragraph. read in Block 1
  • Feizi et al. Online Advertisements with LLMs: Opportunities and Challenges. arXiv 2023. arXiv - the agenda paper for LLM advertising. ⟶ Framing for the previous three. read in Block 1
  • Shah et al. Language Models as Auction Participants. arXiv 2025. arXiv - LLMs in the bidder seat, evaluated against theory. ⟶ Do LLM bidders behave like the models in Block 2 assume? read in Block 4

4. RL and generative bidding

  • Zhang, Yuan, Wang. Optimal Real-Time Bidding for Display Advertising. KDD 2014. doi - the bid should be a concave function of predicted value under a budget, not linear. ⟶ The pre-RL optimal bidding function that every RL bidder is implicitly trying to rediscover. read in Block 4
  • Cai, Ren, Zhang, Malialis, Wang, Yu, Guo. Real-Time Bidding by Reinforcement Learning in Display Advertising. WSDM 2017. acm · arXiv - RTB as an MDP: state = (time left, budget left), model-based DP. ⟶ Where the sequential framing of bidding begins. read in Block 4
  • Jin, Song, Li, Gai, Wang, Zhang (Alibaba). Real-Time Bidding with Multi-Agent Reinforcement Learning in Display Advertising. CIKM 2018. acm · arXiv - thousands of advertisers clustered into jointly learning agents; cooperation and competition effects. ⟶ The competitors are part of the state. (MADDPG attribution from secondary source.) read in Block 4
  • Wu, Chen, Yang, Wang, Tan, Zhang, Xu, Gai (Alibaba). Budget Constrained Bidding by Model-free Reinforcement Learning in Display Advertising. CIKM 2018. acm · arXiv - DRLB: learn to adjust a multiplier λ, with RewardNet and adaptive ε-greedy; weaker on low-AUC data. ⟶ Act on the multiplier, not the bid; and calibration quality bounds what RL can do. read in Block 4
  • Zhao, Qiu, Guan, Zhao, He (Alibaba). Deep Reinforcement Learning for Sponsored Search Real-time Bidding. KDD 2018. doi · arXiv - robust MDP over hourly bidding models for e-commerce search; +35% purchase efficiency, +23.7% CVR, +21.4% ROI reported. ⟶ The one RL bidding paper set in sponsored search on a retail platform; the action is an hourly policy, not a per-auction bid. read in Block 4
  • Yang et al. Bid Optimization by Multivariable Control in Display Advertising. KDD 2019. doi - PID-style multivariable control for constrained bidding. ⟶ The control-theory competitor to RL. read in Block 4
  • He, Chen, Wu, Pan, Tan, Yu, Xu, Zhu (Alibaba). A Unified Solution to Constrained Bidding in Online Display Advertising. KDD 2021. acm - USCB: one Lagrangian bidding formula covers budget, CPC and ROI caps; RL sets the multipliers. ⟶ Every advertiser constraint becomes a dual variable. read in Block 4
  • Mou, Huo, Bai, Xie, Yu, Xu, Zheng (Alibaba). Sustainable Online Reinforcement Learning for Auto-bidding. NeurIPS 2022. arXiv - SORL: V-CQL (variance-suppressed conservative Q-learning) plus safe exploration bridges the virtual-system / real-system gap; real A/B reported. ⟶ Conservative offline RL is the deployment bridge when exploration costs real money. read in Block 4
  • Korenkevych, Cheng, Balakir, Nikulkov, Gao, Cen, Xu, Zhu. Offline Reinforcement Learning for Advertising. arXiv 2023 / ACM 2024. arXiv · acm - a hybrid: trusted base policy plus a neural policy tuning selected parameters; ~200k campaigns, 1.2B steps. ⟶ The migration path from a rules bidder to a learned one. (Exact venue label flagged.) read in Block 4
  • Mou, Xu, Chen, Bai, Yu, Xu. PE-MORL: Pessimistic Environment Model-based Offline RL for Auto-bidding. arXiv 2025. arXiv - a competitor-aware environment model with pessimism penalties; large MAE/MSE reductions vs a GSP simulator. ⟶ Model the other bidders, then distrust your model. (Venue unverified.) read in Block 4
  • Guo, Huo, Zhang, Wang, Yu, Xu, Zhang, Zheng (Alibaba). AIGB: Generative Auto-bidding via Diffusion Modeling. KDD 2024. arXiv - DiffBid: a conditional diffusion model generates whole bidding trajectories; GMV +2.81%, ROI +3.36% online. ⟶ The paradigm shift from “choose a bid” to “generate a plan”. read in Block 4
  • Gao, Li, Mao, Jiang et al. (Kuaishou). GAVE: Generative Auto-bidding with Value-guided Explorations. arXiv 2025. arXiv - decision-transformer generation steered by a value signal. ⟶ Coverage from the generator, direction from the value model. (Venue and metrics unverified.) read in Block 4
  • Li, Mao, Gao, Jiang et al. GAS: Generative Auto-bidding with Post-training Search. arXiv 2024 / WWW 2025 record. arXiv - generate candidates, then search over them for the objective. ⟶ Search as a safety layer on top of generation - the ReBeL pattern in an ad system. read in Block 4
  • Jiang, Tang, Zeng et al. Optimal Return-to-Go Guided Decision Transformer for Auto-Bidding. arXiv 2025. arXiv - return-to-go conditioning turns a target ACOS into a control input. ⟶ Tell the model the outcome you want. (Full author list and venue unverified.) read in Block 4
  • Mou et al. AIGB-Pearl: Enhancing Generative Auto-bidding with Offline Reward Evaluation and Policy Search. arXiv 2025. arXiv - a learned trajectory evaluator plus KL-Lipschitz-constrained score maximization for safe exploration beyond the dataset. ⟶ Generative bidding gets a critic. (GMV figure disputed: ~3% vs 5.1%.) read in Block 4
  • Peng et al. Expert-Guided Diffusion Planner for Auto-bidding. arXiv 2025 / CIKM 2025 record. arXiv - expert trajectories as a behavioral prior for diffusion planning. ⟶ The behavior-regularization idea, bidding edition. (Full authors unverified.) read in Block 4
  • Lei, Zhao, Zhao, Zhang, Cai, Xie, Wang (Meituan). GRAD: Generative Reward-driven Ad Bidding. KDD 2026. arXiv - Action-MoE generator plus causal-transformer value estimator; ROI +10.68% reported. ⟶ Mixture-of-experts for heterogeneous campaign regimes. (Verify publication status.) read in Block 4
  • Meng et al. (JD.com). JD-BP: A Joint-Decision Generative Framework for Auto-Bidding and Pricing. arXiv 2026. arXiv - generate bids and prices jointly. ⟶ Bid and price co-evolve; evaluate them together. (Venue unverified.) read in Block 4
  • Yang, Zuo, Kim. Constrained Auto-Bidding via Generative Response Modeling. KDD 2026. arXiv - learn a generative response model, then solve for feasible multipliers; 33.88 vs 31.43 on AuctionNet. ⟶ Auditable constraints instead of reward shaping. read in Block 4
  • Wu et al. (Tencent). GRB: A Generative Reinforcement Bidding Framework for Multi-Channel Online Advertising. KDD 2026. acm - generative RL across five channels with online A/B. ⟶ Multi-surface coordination. (Authors and metrics incomplete in record.) read in Block 4
  • Jiang, Zhou, Zhang, Chen, Hu, Choi. Risk-aware Reinforcement Learning for Real-time Bidding. arXiv 2022 / KDD Explorations record. arXiv - value = mean minus a tunable multiple of predicted-response standard deviation. ⟶ Decide how much uncertainty you are willing to buy. (Final venue flagged.) read in Block 4
  • Lin, Zheng, Wu. Robust Auto-bidding for Censored and Distribution-shifted Environments. KDD 2024. acm - worst-case surplus guarantees under censored feedback and shift. ⟶ Logs only show what you won; bid for the worst case. read in Block 4
  • Ren, Qin, Zheng, Yang, Zhang, Yu. Deep Landscape Forecasting for Real-time Bidding Advertising. KDD 2019. doi · arXiv - RNN plus survival analysis for the censored winning-price distribution. ⟶ The bid → win-probability → price map every bidder needs. read in Block 4
  • Ou, Chen, Yang et al. Deep Landscape Forecasting in Multi-Slot Real-Time Bidding. KDD 2023. acm - correlated slots on one page, position-specific price distributions. ⟶ A search page has several sponsored slots; model them jointly. read in Block 7
  • Hajiaghayi et al. Analysis of a Learning Based Algorithm for Budget Pacing. arXiv 2022. arXiv - regret analysis of an adaptive pacing rule. ⟶ Separates the pacing failure from the bid-quality failure. read in Block 4
  • Mystique (authors not recovered). A Budget Pacing System for Performance Optimization in Online Advertising. WWW Companion 2024. acm - soft throttling against a daily target curve. ⟶ A production pacing layer that constrains a learned bidder. (Author metadata missing.) read in Block 4
  • Xu et al. Smart Pacing for Effective Online Ad Campaign Optimization. KDD 2015. doi - pace to a spend curve while maximizing performance. ⟶ Where the daily curve idea comes from. read in Block 4
  • Chen, Yuan, Ye, Majumder, Richardson. AucArena: Put Your Money Where Your Mouth Is - Evaluating Strategic Planning and Execution of LLM Agents in an Auction Arena. NeurIPS 2024 Open-World Agents workshop (withdrawn from ICLR 2024). arXiv - LLM agents in ascending auctions as a planning testbed. ⟶ Strategic-reasoning appendix, not a high-frequency bidder. read in Block 4
  • Cai, He, Li et al. RTBAgent: A Large Language Model-based Agent for Real-Time Bidding. arXiv 2025. arXiv - memories, retrieval, two-step decisions and daily reflection; 9 campaigns over 10 days. ⟶ Campaign-level planning by an LLM; revenue disclosure deferred by the authors. read in Block 4
  • Jiang, Xiong, Liu. HARBOR: A Testbed for Large Language Model Agents in Auctions. arXiv 2025. arXiv - profit-seeking LLM agents under strategic competition. ⟶ Stress-test an LLM planner before it touches bids. read in Block 4
  • Yin et al. InfoBid: Information Disclosure in Auctions with LLM-based Agents. arXiv 2025. arXiv - how much the platform reveals changes agent behavior. ⟶ Marketing Stream is an information-design choice. (Details unverified.) read in Block 4

5. Benchmarks and simulators

  • Jeunen, Murphy, Allison (Amazon). Learning to Bid with AuctionGym. AdKDD 2022 (best paper). amazon.science · github - most learning-to-bid methods are value-based bandits; adds policy-based and doubly robust bidders and an open simulator. ⟶ Amazon’s own statement that offline logs cannot evaluate a bidder alone. read in Block 4
  • Jeunen, Murphy, Allison (Amazon). Off-Policy Learning-to-Bid with AuctionGym. KDD 2023. acm - the archival version with the unified value-based view. ⟶ Read for the estimators. (Announcement 2022, paper 2023.) read in Block 7
  • Amazon Science. Amazon Scientists Win Best-Paper Award for Ad Auction Simulator. Blog 2022. amazon.science - the plain-language account. ⟶ Skim for framing. read in Block 8
  • Su, Huo, Zhang, Dou, Yu et al. (Alibaba). AuctionNet: A Novel Benchmark for Decision-Making in Large-Scale Games. NeurIPS 2024 Datasets & Benchmarks. arXiv - 10M ad opportunities, 48 auto-bidding agents, 500M+ auction records, GSP module; >1,500 competition teams; online LP is the strongest baseline. ⟶ The shared stress test for budget-constrained bidders. (Full author list and exact competition title unverified.) read in Block 4
  • Khirianova, Solodneva, Pudovikov et al. BAT: Benchmark for Auto-bidding Task. arXiv 2025. arXiv - dataset plus baselines for auto-bidding. ⟶ A second benchmark to avoid overfitting AuctionNet. (Author list partial.) read in Block 7
  • Chen, Nabi, Siniscalchi (Amazon). Advancing Ad Auction Realism: Practical Insights & Modeling Implications. AdKDD 2023. amazon.science · arXiv - query-dependent values, unobserved changing competitors, partial feedback, partially specified payments; adversarial-bandit advertisers; soft floors help in rich environments; value distributions inferable from bids. ⟶ Do not evaluate a Sponsored Products policy in a static, fully observed toy. read in Block 1

6. Prediction stack: CTR/CVR, sequences, scaling, calibration, delayed feedback

  • McMahan et al. (Google). Ad Click Prediction: a View from the Trenches. KDD 2013. google - FTRL-Proximal logistic regression over billions of hashed features; calibration lessons. ⟶ Calibration matters as much as ranking because probabilities feed a price. read in Block 3
  • He et al. (Facebook). Practical Lessons from Predicting Clicks on Ads at Facebook. ADKDD 2014. doi - GBDT leaves as LR features; freshness beats cleverness. ⟶ Data recency is a first-order lever. read in Block 3
  • Cheng et al. (Google). Wide & Deep Learning for Recommender Systems. 2016. arXiv - memorization plus generalization. ⟶ The template for every later CTR net. read in Block 3
  • Guo et al. (Huawei). DeepFM. IJCAI 2017. arXiv - learned pairwise interactions, no feature engineering. ⟶ Compact baseline for sparse advertiser × keyword × placement crosses. read in Block 3
  • Wang et al. (Google). DCN V2. WWW 2021. arXiv - bounded-degree explicit crosses with MoE decomposition for serving. ⟶ Production lessons on serving cost. read in Block 3
  • Zhou et al. (Alibaba). Deep Interest Network. KDD 2018. arXiv - target attention over behavior history. ⟶ The user representation should change with the ad being scored. read in Block 3
  • Zhou et al. (Alibaba). Deep Interest Evolution Network. AAAI 2019. doi - GRU over interests with an auxiliary next-behavior loss. ⟶ Sessions disambiguate ambiguous queries. read in Block 3
  • Pi, Zhu, Zhou et al. (Alibaba). SIM: Search-based User Interest Modeling with Lifelong Sequential Behavior Data. CIKM 2020. arXiv - retrieve a candidate-relevant subsequence, then attend; sequences to 54,000; +7.1% CTR. ⟶ Long histories are affordable if you retrieve first. read in Block 3
  • Chen, Xu, Pei, Lv, Zhuang, Ge (Alibaba). ETA: Efficient Long Sequential User Data Modeling for CTR. DLP-KDD 2022 (arXiv preprint 2021). arXiv - SimHash fingerprints and Hamming retrieval; ~120k QPS. ⟶ The engineering that makes long-sequence attention fit a search auction’s latency budget. read in Block 3
  • Chang et al. (Kuaishou). TWIN: Two-Stage Interest Network for Lifelong User Behavior Modeling. KDD 2023. arXiv - consistent attention across the retrieval and ranking stages. ⟶ Stage inconsistency is where long-sequence models lose lift. read in Block 3
  • Si, Guan, Sun et al. (Kuaishou). TWIN V2: Scaling Ultra-Long User Behavior Sequence Modeling. CIKM 2024. arXiv - hierarchical clustering of lifecycle behaviors plus cluster-aware target attention. ⟶ Category and brand affinity from years of history, at serving cost. (Full author list unverified.) read in Block 3
  • Xia et al. (Pinterest). TransAct: Transformer-based Realtime User Action Model. KDD 2023. acm - realtime action sequence plus batch embeddings. ⟶ Fresh session intent with a stable fallback. read in Block 3
  • Zhai et al. (Meta). Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative Recommendations (HSTU). ICML 2024. arXiv - recommendation as sequential transduction; 1.5T parameters; +12.4% online. ⟶ The scaling-law era reaches ranking; auditability inside an auction is the open question. read in Block 3
  • Zhang et al. (Meta). Wukong: Towards a Scaling Law for Large-Scale Recommendation. ICML 2024. pmlr - stacked factorization machines scale across two orders of complexity. ⟶ The non-Transformer scaling baseline. read in Block 3
  • Deng et al. (Kuaishou). OneRec: Unifying Retrieve and Rank with Generative Recommender and Preference Alignment. arXiv 2025. arXiv - a single generative model replaces the cascade. ⟶ What happens to eligibility and budget constraints when the cascade disappears. (Venue not established.) read in Block 3
  • Lu, Zheng et al. (ByteDance). Large Memory Network for Recommendation. arXiv 2025. arXiv - compressed persistent user memory. ⟶ Long-horizon intent without long sequences. (Metadata incomplete.) read in Block 3
  • Zhu et al. (ByteDance). RankMixer: Scaling Up Ranking Models in Industrial Recommenders. CIKM 2025. arXiv - hardware-aware scaling for ranking on a trillion-scale production dataset. ⟶ Scaling that respects a millisecond auction budget. read in Block 3
  • Ma et al. (Alibaba). ESMM: Entire Space Multi-Task Model for Post-Click Conversion Rate. SIGIR 2018. arXiv - estimate pCTR and pCTCVR over all impressions; pCVR falls out. ⟶ The industry default for CVR selection bias. read in Block 3
  • Wen et al. ESM2: Entire Space Multi-Task Modeling via Post-Click Behavior Decomposition. SIGIR 2020. arXiv - insert cart/wishlist actions between click and purchase. ⟶ Less sparse, more timely conversion value. read in Block 3
  • Xi et al. AITM: Modeling the Sequential Dependence among Audience Multi-step Conversions. KDD 2021. arXiv - adaptive information transfer along the funnel. ⟶ Funnel steps inform each other. read in Block 3
  • Ma, Zhao, Yi, Chen, Hong, Chi (Google). MMoE: Modeling Task Relationships in Multi-task Learning with Multi-gate Mixture-of-Experts. KDD 2018. acm - per-task gates over shared experts. ⟶ CTR and CVR should not be forced into one representation. read in Block 3
  • Tang, Liu, Zhao, Gong (Tencent). PLE: Progressive Layered Extraction. RecSys 2020. acm - kills the seesaw effect in multi-task ranking. ⟶ Balance click, conversion and revenue heads. read in Block 3
  • Wang et al. (Alibaba). ESCM2: Entire Space Counterfactual Multi-Task Model. SIGIR 2022. arXiv - ESMM plus counterfactual correction. ⟶ Fixes ESMM’s residual bias. read in Block 3
  • Ma, Ispir, Li et al. (Google). An Online Multi-task Learning Framework for Google Feed Ads Auction Models. KDD 2022. acm - continuous multi-task training with multi-stage label-delay handling and learned loss weights. ⟶ The closest public production analogue to an auction-facing prediction stack. read in Block 3
  • Chapelle (Criteo). Modeling Delayed Feedback in Display Advertising. KDD 2014. doi - model the delay distribution jointly with CVR. ⟶ The original censoring correction. read in Block 3
  • Ktena et al. (Twitter). Addressing Delayed Feedback for Continuous Training with Neural Networks in CTR Prediction. RecSys 2019. arXiv - importance-weighted duplicated samples keep a continuously trained model unbiased. ⟶ Log the click as negative now, re-inject as positive later. read in Block 3
  • Yasui, Morishita, Fujita, Shibata. A Feedback Shift Correction in Predicting Conversion Rates under Delayed Feedback (FSIW). WWW 2020. arXiv - importance weighting for the shift between observed and true labels. ⟶ Another path to the same correction. read in Block 3
  • Yang, Li, Han, Zhuang, Zhan, Zeng, Tong. ES-DFM: Capturing Delayed Feedback via Elapsed-Time Sampling. AAAI 2021. arXiv - instance-level importance weights from elapsed time. ⟶ Do not treat a two-hour-old click as a confirmed non-conversion. read in Block 3
  • Gu, Sheng, Fan, Zhou, Zhu (Alibaba). DEFER: Real Negatives Matter. KDD 2021. acm - ingest duplicated real negatives with importance sampling; >6% CVR gains. ⟶ Continuous training without feature-distribution bias. read in Block 3
  • Yang, Zhan. GDFM: Generalized Delayed Feedback Model with Post-Click Information. NeurIPS 2022. neurips - post-click behaviors as stochastic early evidence. ⟶ Intermediate signals buy timeliness. read in Block 3
  • Chen et al. DEFUSE: Asymptotically Unbiased Estimation for Delayed Feedback Modeling via Label Correction. WWW 2022. arXiv - label correction with unbiasedness guarantees. ⟶ Theory for the duplicate-and-reweight trick. read in Block 3
  • Liu, Ao, He et al. MISS: Online CVR Prediction via Multi-Interval Screening and Synthesizing under Delayed Feedback. AAAI 2024. doi - multiple maturity windows screened and fused. ⟶ A compromise between waiting for labels and trusting fresh negatives. read in Block 3
  • Ding et al. IF-DFM: Addressing Delayed Feedback via Influence Functions. arXiv 2025. arXiv - influence functions approximate retraining when late conversions arrive; 14.8 s updates reported. ⟶ Hourly repricing wants sub-minute label repair. (Authors and venue incomplete; runtime claim to verify.) read in Block 3
  • Pan, Ao, Tang, Lu, Liu, Xiao, He. Field-aware Calibration. WWW 2020. arXiv - field-level calibration error and a neural post-hoc calibrator. ⟶ Calibrate by placement and query class, not globally. read in Block 3
  • Fan, Si, Zhang. Calibration Matters: Tackling Maximization Bias in Large-scale Advertising Recommendation Systems. ICLR 2023. arXiv - the auction selects the argmax, so selected predictions are biased upward; variance-adjusted debiasing. ⟶ The winner’s curse inside your own CTR model. read in Block 3
  • Yang, Yang, Zou, Xu, Yuan, Zeng. DESC: Deep Ensemble Shape Calibration. arXiv 2024. arXiv - value calibration and shape calibration across fields. ⟶ Post-hoc repair without retraining the ranker. (Results not exposed in record.) read in Block 3
  • Zhao, Wu, Jia et al. (Huawei). ConfCalib: Confidence-Aware Multi-Field Model Calibration. arXiv 2024. arXiv - Wilson intervals set calibration intensity per field value. ⟶ Do not over-correct sparse tail keywords. (Venue beyond arXiv not established.) read in Block 3
  • He et al. Rankability-enhanced Revenue Uplift Modeling Framework for Online Marketing. KDD 2024. arXiv - uplift with ranking-aware losses. ⟶ Bid for incremental, not attributed, purchases. read in Block 3
  • Wu, Jia, Dong, Tang. Customer Lifetime Value Prediction: Towards the Paradigm Shift of Recommender System Objectives. RecSys 2023. acm - LTV as the objective. ⟶ Repeat purchase changes the margin term in the bid. read in Block 3
  • Ardalani et al. (Meta). Understanding Scaling Laws for Recommendation Models. arXiv 2022. arXiv - power laws for CTR models in data, parameters and compute. ⟶ The first evidence that ranking scales like language. read in Block 3
  • Zhang et al. Scaling Law of Large Sequential Recommendation Models. RecSys 2024. arXiv - scaling laws for sequential recommenders. ⟶ Sequence length and model size as levers. read in Block 3
  • Wang et al. Scaling Laws for Online Advertisement Retrieval. arXiv 2024. arXiv - scaling laws for the ad retrieval stage specifically. ⟶ Where compute buys the most in an ads stack. read in Block 3
  • Kuaishou. OneRec Technical Report. arXiv 2025. arXiv - the full system account of the unified generative recommender. ⟶ Read with the OneRec paper. read in Block 3
  • Meituan. MTGR: Industrial-Scale Generative Recommendation Framework. CIKM 2025. arXiv - generative ranking with HSTU-style blocks in a food-delivery ads and recommendation stack. ⟶ Generative ranking in a transactional marketplace. read in Block 3
  • ByteDance. LONGER: Scaling Up Long Sequence Modeling in Industrial Recommenders. RecSys 2025. arXiv - efficient long-sequence transformers at production scale. ⟶ The 2025 state of long-history modeling. read in Block 3
  • Meta Engineering. Meta’s Generative Ads Model (GEM). Blog, November 2025. engineering.fb - a foundation model for ads ranking distilled into serving models. ⟶ The industrial shape of scaling for ads. (Blog, not a paper.) read in Block 3
  • Chitlangia, Kesari, Agarwal (Amazon). Scaling Generative Pre-training for User Ad Activity Sequences. AdKDD 2023. amazon.science - generative pre-training over ad activity sequences at Amazon, with scaling curves. ⟶ Amazon’s own scaling-law evidence for ads. read in Block 3
  • Xia et al. (Pinterest). TransAct V2. CIKM 2025. arXiv - lifelong sequences plus realtime actions in one ranker. ⟶ The follow-up that closes the long/short gap. read in Block 3
  • Zhu et al. Entire Space Cascade Delayed Feedback Modeling (ECAD). CIKM 2023. arXiv - delayed feedback across the whole funnel cascade. ⟶ ESMM meets delayed feedback. read in Block 3
  • Wang et al. Look Ahead: Improving the Accuracy of Time-Series Forecasting by Previewing Belief Space (PACC). SIGIR 2023. arXiv - position-aware calibration for entire-space CVR. ⟶ Position bias and selection bias handled together. read in Block 3
  • Xue et al. Multi-Task Learning for CVR with Delayed Feedback. arXiv 2023. arXiv - multi-task heads over delay windows. ⟶ Another route to maturity-aware CVR. read in Block 3
  • Huangfu et al. Unbiased Delayed Feedback Label Correction (ULC). KDD 2023. arXiv - unbiased correction of immature labels. ⟶ The 2023 state of the art before influence functions. read in Block 3
  • Luo et al. Delayed feedback modeling, 2026. WWW 2026. arXiv - the newest entry in the delayed-feedback line. ⟶ Check whether it evaluates at decision time. read in Block 3
  • Deng, Wang, Tan, Xu, Gai (Alibaba). Calibrating User Response Predictions in Online Advertising. ECML-PKDD 2020. doi - calibration for ads at Alibaba. ⟶ An industrial calibration recipe from a sponsored-search platform. read in Block 3
  • Huang et al. MBCT: Tree-Based Feature-Aware Binning for Individual Uncertainty Calibration. WWW 2022. arXiv - feature-aware binning for calibration. ⟶ Calibration that varies with the input. read in Block 3
  • Sheng et al. (Alibaba). Joint Optimization of Ranking and Calibration (JRC). KDD 2023. arXiv - train for ranking and calibration jointly. ⟶ Stop trading AUC against calibration. read in Block 3
  • Zhang et al. Self-Boosted Calibration for Ranking (SBCR). KDD 2024. arXiv - self-boosting calibration in ranking models. ⟶ Follow-up to JRC. read in Block 3
  • Yan et al. (Google). Scale Calibration of Deep Ranking Models. KDD 2022. doi - calibrated scales for LTR outputs. ⟶ Ranking scores that mean something in dollars. read in Block 3
  • Kweon, Kang, Yu. Obtaining Calibrated Probabilities with Personalized Ranking Models. AAAI 2022. arXiv - calibration for personalized rankers. ⟶ Same problem, recommendation side. read in Block 3
  • Chaudhuri, Bagherjeiran, Liu (Amazon). Ranking and Calibrating Click-Attributed Purchases in Performance Display Advertising. KDD 2017 (AdKDD). amazon.science - calibrating post-click purchase predictions at Amazon. ⟶ Amazon’s own account of CVR calibration for pricing. read in Block 3
  • Karra et al. (Amazon). Nudging Neural Click Prediction Models to Pay Attention to Position. CIKM 2023. amazon.science - make position an explicit, removable feature. ⟶ Amazon’s position-debiasing for the click model that prices Sponsored Products. read in Block 3
  • Muhamed et al. (Amazon). CTR-BERT: Cost-Effective Knowledge Distillation for Billion-Parameter Teacher Models. NeurIPS 2021 ENLSP workshop. pdf - distill a billion-parameter text model into a CTR model. ⟶ Language models reach the Amazon ads CTR stack via distillation. read in Block 3
  • Agrawal, Ahemad, Sembium (Amazon). Rationale-Guided Distillation for E-Commerce Relevance Classification. COLING 2025. amazon.science - LLM rationales distilled into lightweight cross-encoders. ⟶ Relevance at auction latency with LLM judgment. read in Block 3
  • Lin et al. ClickPrompt: CTR Models are Strong Prompt Generators for Adapting Language Models to CTR Prediction. WWW 2024. arXiv - CTR model outputs as prompts for an LM. ⟶ One way to marry ID features and text. read in Block 3
  • Li et al. CTRL: Connect Collaborative and Language Model for CTR Prediction. arXiv 2023. arXiv - contrastive alignment of tabular and language representations. ⟶ The other way. read in Block 3
  • Lin et al. How Can Recommender Systems Benefit from Large Language Models: A Survey. arXiv 2023. arXiv - taxonomy of LLM-for-recsys. ⟶ The map for LLM-in-ranking work. read in Block 3
  • INSPIRE. 2026. arXiv - a 2026 entry in the LLM-for-CTR line. ⟶ Check for auction-facing evaluation. read in Block 3
  • Moraes et al. Uplift Modeling: From Causal Inference to Personalization (tutorial). arXiv 2023. arXiv - the uplift tutorial. ⟶ Start here for incremental bidding. read in Block 3
  • Zhang, Li, Liu. A Unified Survey of Treatment Effect Heterogeneity Modelling and Uplift Modelling. ACM Computing Surveys. arXiv - the survey. ⟶ Reference. read in Block 3
  • Ke et al. Addressing Exposure Bias in Uplift Modeling for Large-scale Online Advertising. ICDM 2021. doi - exposure bias in uplift at ad scale. ⟶ Uplift models inherit the auction’s selection. read in Block 3
  • Lewis, Wong. Incrementality Bidding and Attribution. arXiv 2022. arXiv - bid on incremental value directly. ⟶ The reward-signal fix stated as a bidding rule. read in Block 7

7. Retrieval, relevance and LLMs

  • Nigam et al. (Amazon). Semantic Product Search. KDD 2019. arXiv · amazon.science - dense retrieval over a billion-product catalog with a three-way loss. ⟶ Keyword coverage is learned, not hand-built. read in Block 3
  • Chang et al. (Amazon). Extreme Multi-label Learning for Semantic Matching in Product Search. KDD 2021. arXiv - XMC over products as labels. ⟶ Matching at catalog scale. read in Block 3
  • Muhamed et al. (Amazon). Web-scale Semantic Product Search with Large Language Models. PAKDD 2023. amazon.science - four-stage training; a distilled 75M model gains up to 23% relevance at DSSM latency. ⟶ Distill the LLM; do not serve it in the auction path. read in Block 3
  • Amazon. Improving Ad Matching via Cluster-Adaptive Keyword Expansion and Relevance Tuning. arXiv 2025. arXiv - cluster-adaptive expansion for sponsored matching. ⟶ Tail keyword discovery as a growth lever. read in Block 3
  • Shi, Rao, Wu, Zhang, Wang (Amazon). Campaign Keyword Augmentation via Generative Methods. ECNLP 2021. amazon.science - seq2seq plus trie search for cold campaigns. ⟶ Defines which clauses enter the auction at all. read in Block 8
  • Amazon. Automated Query-Product Relevance Labeling using Large Language Models for E-commerce Search. arXiv 2025. arXiv - LLMs as relevance labelers at scale. ⟶ LLMs are mature for labels, not (publicly) for the live ranker. (Author list not exposed.) read in Block 3

8. Position bias and counterfactual learning to rank

  • Joachims, Swaminathan, Schnabel. Unbiased Learning-to-Rank with Biased Feedback. WSDM 2017. acm - propensity-weighted LTR from position-biased clicks. ⟶ Top-of-search clicks are not relevance labels. read in Block 3
  • Wang et al. (Google). Position Bias Estimation for Unbiased Learning to Rank in Personal Search. WSDM 2018. acm - RandTopN and RandPair intervention schemes. ⟶ How to estimate the examination curve cheaply. read in Block 3
  • Amazon. Learning to Rank in the Position-Based Model with Bandit Feedback. CIKM 2020. arXiv - bandit LTR under PBM. ⟶ Amazon’s own position-debiasing. read in Block 3
  • Amazon. Off-policy Evaluation for LTR via Interpolating the Item-Position Model and the Position-Based Model. 2022. amazon.science - interpolate between click models. ⟶ Estimator choice is a click-model bet. read in Block 7
  • Yu (Amazon). Unbiased Counterfactual Estimation of Ranking Metrics. 2021. amazon.science - evaluate a ranker from another ranker’s logs, unbiasedly. ⟶ Offline evaluation of ad ranking changes. read in Block 7
  • Block, Kidambi, Hill, Joachims, Dhillon (Amazon). Counterfactual Learning to Rank for Utility-Maximizing Query Autocompletion. SIGIR 2022. amazon.science - rank by downstream purchase utility. ⟶ Optimize shopping utility, not clicks. read in Block 7
  • Xiao, Kveton, Katariya, Gangwani, Rangi (Amazon). Towards Sequential Counterfactual Learning to Rank. SIGIR-AP 2023. amazon.science - session-level estimators under sequential PBM. ⟶ Value accrues across reformulations. read in Block 7
  • Jakimov, Buchholz, Stein, Joachims (Amazon). Unbiased Offline Evaluation for Learning to Rank with Business Rules. CONSEQUENCES @ RecSys 2023. amazon.science · arXiv - Birkhoff-von Neumann correction for post-processed rankings. ⟶ Evaluate the deployed policy, not the unconstrained model. read in Block 7
  • Buchholz, London, Di Benedetto, Lichtenberg, Stein, Joachims (Amazon). Counterfactual Ranking Evaluation with Flexible Click Models. SIGIR 2024; the INTERPOL estimator first appeared at the RecSys 2022 CONSEQUENCES workshop. amazon.science · arXiv - the INTERPOL estimator trades click-model assumptions against variance. ⟶ Sponsored and organic results share a page; brittle click models fail. read in Block 7

9. Measurement, interference and incrementality

  • Blake, Coey. Why Marketplace Experimentation Is Harder than it Seems: The Role of Test-Control Interference. EC 2014. doi - eBay emails: naive user-level A/B overstated lift by ~2×. ⟶ Your control group competes in the same auctions. read in Block 7
  • Li, Zhao, Johari, Weintraub. Interference, Bias, and Variance in Two-Sided Marketplace Experimentation: Guidance for Platforms. WWW 2022. acm · arXiv - which side to randomize and at what proportion; bias reduction can raise variance. ⟶ The unit of randomization is part of the estimator. read in Block 7
  • Johari, Li, Liskovich, Weintraub. Experimental Design in Two-Sided Platforms: An Analysis of Bias. Management Science 2022. doi · arXiv - the journal treatment. ⟶ Read with the WWW paper. read in Block 7
  • Liu, Mao, Kang (LinkedIn). Trustworthy and Powerful Online Marketplace Experimentation with Budget-split Design. KDD 2021. acm · arXiv - split budgets to create two counterfactual marketplaces. ⟶ The cleanest design for a bid-policy test. read in Block 7
  • Bojinov, Simchi-Levi, Zhao. Design and Analysis of Switchback Experiments. Management Science 2023. doi · arXiv - minimax optimal time-block designs with exact randomization inference. ⟶ Randomize over time when you cannot randomize over users. read in Block 7
  • Ni, Kalfountzou, Bojinov (P&G / HBS). Reliable Switchback Experiments with Rerandomization for Auction Environments at Procter & Gamble. HBS WP 26-012, 2025. hbs - rerandomize schedules until covariates balance; 70 experiments across 8 markets. ⟶ An industrial auction-experiment playbook. (Draft; lift figures not settled.) read in Block 7
  • Bright, Delarue, Lobel. Reducing Marketplace Interference Bias via Shadow Prices. EC 2023 / Management Science 2025. doi · arXiv - compare shadow prices, not raw group value. ⟶ The right first-order correction for capacity spillovers. read in Block 7
  • Holtz et al. Reducing Interference Bias in Online Marketplace Experiments Using Cluster Randomization: Evidence from a Pricing Meta-experiment on Airbnb. Management Science 2025. informs - cluster randomization measured against a meta-experiment. ⟶ Empirical size of the interference bias. read in Block 7
  • Johnson, Lewis, Nubbemeyer. Ghost Ads: Improving the Economics of Measuring Online Ad Effectiveness. JMR 2017. doi - log the would-have-won impression for control users; 432 experiments. ⟶ The analogue for a sponsored slot is an open design problem. read in Block 7
  • Lewis, Rao. The Unfavorable Economics of Measuring the Returns to Advertising. QJE 2015. doi - ad effects are tiny relative to sales variance; experiments need enormous samples. ⟶ Why incrementality is expensive. read in Block 7
  • Gordon, Zettelmeyer, Bhargava, Chapsky. A Comparison of Approaches to Advertising Measurement. Marketing Science 2019. doi - observational methods fail to recover RCT lift at Facebook. ⟶ Attribution is not incrementality. read in Block 7
  • Gordon, Moakler, Zettelmeyer. Predicted Incrementality by Experimentation (PIE) for Ad Measurement. arXiv 2023, rev. 2026. arXiv - learn a map from campaign features to RCT lift; R² 0.88 vs 0.19 for last-click. ⟶ Scale incrementality to campaigns that cannot afford holdouts. read in Block 7
  • Meta. GeoLift Methodology. Documentation. geolift - augmented synthetic controls over geographies. ⟶ When user-level holdouts are impossible. read in Block 7
  • Meta. Conversion Lift Testing for Incrementality Measurement. Documentation. meta - exposed vs unexposed lift. ⟶ The industry definition of lift. read in Block 7
  • Ohlinger, Nedyalkov (Google). Incrementality Testing. Think with Google 2023. google - practitioner framing of RCTs vs attribution vs MMM. ⟶ Cross-platform template. read in Block 7
  • Waisman, Nair, Carrion. Online Causal Inference for Advertising in Real-Time Bidding Auctions. Marketing Science 2025. doi · arXiv - auction structure identifies ad effects; Thompson sampling controls experimentation cost. ⟶ Bidding, measurement and exploration in one framework. read in Block 7
  • Dudík, Langford, Li. Doubly Robust Policy Evaluation and Learning. ICML 2011. arXiv - DR estimator for contextual bandits. ⟶ The estimator AuctionGym’s bidders use. read in Block 7
  • Jiang, Li. Doubly Robust Off-policy Value Evaluation for Reinforcement Learning. ICML 2016. pmlr - DR for sequential policies. ⟶ Evaluating a multi-step bidding policy from logs. read in Block 7
  • Swaminathan, Joachims. Counterfactual Risk Minimization. ICML 2015. pmlr - learn from logged bandit feedback with variance regularization. ⟶ Off-policy learning, not just evaluation. read in Block 7
  • Yeom, Shin, Min, Yoon, Yu, Kang. Breaking Determinism: Stochastic Modeling for Reliable Off-Policy Evaluation in Ad Auctions. arXiv 2025 / ACM 2026. arXiv - deterministic auctions give zero propensity to losing bids; repurpose landscape models as propensities for self-normalized IPS. ⟶ The support problem for bidding OPE, and a fix. (Emerging.) read in Block 7
  • Gordon, Moakler, Zettelmeyer. Close Enough? A Large-Scale Exploration of Non-Experimental Approaches to Advertising Measurement. Marketing Science 2023. arXiv - 663 experiments; observational methods still miss. ⟶ The follow-up that settles the 2019 finding at scale. read in Block 7
  • Zeng et al. Sequentially Rerandomized Switchback Experiments. arXiv 2026. arXiv - rerandomize the switchback sequentially as data arrives. ⟶ The P&G idea made adaptive. read in Block 7
  • Farias et al. Markovian Interference in Experiments. NeurIPS 2022. arXiv - interference through shared state, with a differences-in-Q estimator. ⟶ A budget is shared state. read in Block 7
  • Wager, Xu. Experimenting in Equilibrium. Management Science 2021. arXiv - local experimentation to estimate equilibrium effects of a policy change. ⟶ Estimate the marketplace-wide effect from small perturbations. read in Block 7
  • Jeunen. A Common Misassumption in Online Experiments with Machine Learning Models. SIGIR Forum 2023. arXiv - the trained model in treatment saw different data. ⟶ Your A/B test of a bidder is also a test of its training data. read in Block 7
  • Jeunen, Ustimenko. Learning Metrics that Maximise Power for Accelerated A/B-Tests. KDD 2024. arXiv - learn proxy metrics for power. ⟶ Faster bid-policy tests. read in Block 7
  • Jain, Hut, Islam, Pan (Amazon). Cross-Unit Spillovers in A/B Testing: Empirical Evidence from Ads. CODE@MIT 2023. amazon.science - measured spillovers in Amazon ads experiments. ⟶ Amazon’s own evidence that interference bites. read in Block 7
  • Hut et al. (Amazon). Value of Stratification in Cluster-Randomized Experiments. CODE@MIT 2023. amazon.science - stratify clusters to recover power. ⟶ How to afford cluster designs. read in Block 7
  • Jain, Appala (Amazon). SERP Interference Network and Its Applications in Search Advertising. AdKDD 2024. amazon.science - the interference graph over search results pages. ⟶ The answer to “what is the unit” on a retail SERP. read in Block 7
  • Meloni et al. (Amazon). Performance of Synthetic Diff-in-Diff Models for Geo-Randomized Experiments. CODE@MIT 2024. amazon.science - SDID for geo tests. ⟶ Amazon’s geo-incrementality toolkit. read in Block 7
  • Vaver, Koehler (Google). Measuring Ad Effectiveness Using Geo Experiments. 2011. google - the original geo-experiment methodology. ⟶ Where GeoLift comes from. read in Block 7
  • Kerman, Wang, Vaver (Google). Estimating Ad Effectiveness Using Geo Experiments in a Time-Based Regression Framework. 2017. google - time-based regression for geo tests. ⟶ The refinement. read in Block 7
  • Bottou et al. Counterfactual Reasoning and Learning Systems: The Example of Computational Advertising. JMLR 2013. arXiv - the foundational paper on counterfactual estimation for ad systems, at Bing. ⟶ Read before any OPE paper. read in Block 7
  • Swaminathan, Joachims. The Self-Normalized Estimator for Counterfactual Learning (SNIPS). NeurIPS 2015. nips - normalize IPS to kill propensity overfitting. ⟶ The estimator you actually use. read in Block 7
  • Saito, Joachims. Off-Policy Evaluation for Large Action Spaces via Embeddings (MIPS). ICML 2022. arXiv - marginalize over action embeddings. ⟶ Bids are a large action space. read in Block 7
  • Saito, Joachims. Counterfactual Learning and Evaluation for Recommender Systems (tutorial). RecSys 2021. doi - the OPE tutorial. ⟶ Start here. read in Block 7
  • Wu, Yeh, Chen. Predicting Winning Price in Real Time Bidding with Censored Data. KDD 2018. doi - censored regression for winning prices. ⟶ Landscape forecasting before deep models. read in Block 7
  • Ou et al. A Survey on Bid Optimization in Real-Time Bidding Display Advertising. ACM TKDD 2024. doi - the archival survey. ⟶ Replaces the un-URLed entry above. read in Block 4

10. Attribution

  • Lewis, Zettelmeyer, Gordon, Garib, Hermle, Perry, Romero, Schnaidt (Amazon Ads). Amazon Ads Multi-Touch Attribution. arXiv 2025. arXiv - RCT-calibrated ML over shopping signals; hundreds of thousands of RCTs. ⟶ The reward signal your bidder optimizes; fractional credit is still not marginal value. read in Block 7
  • Zhao, Mahboobi, Bagheri. Shapley Value Methods for Attribution Modeling in Online Advertising. arXiv 2018. arXiv - simplified and ordered Shapley credit. ⟶ Diagnostic allocation under a model. read in Block 7
  • Singal, Besbes, Desir, Goyal, Iyengar. Shapley Meets Uniform: An Axiomatic Framework for Attribution in Online Advertising. WWW 2019. acm - axioms over a Markovian funnel. ⟶ Make attribution assumptions inspectable. read in Block 7

11. Equilibrium learning and regret minimization in auctions

  • Bichler, Fichtl, Heidekrüger, Kohring, Sutterer. Learning Equilibria in Symmetric Auction Games Using Artificial Neural Networks. Nature Machine Intelligence 2021. doi - NPGA: neural strategies, pseudogradient self-play, local Bayes-Nash equilibria that match known analytic solutions. ⟶ The closest methodological bridge from Brown-style self-play to auctions. read in Block 6
  • Heidekrüger, Sutterer, Kohring, Fichtl, Bichler. Equilibrium Learning in Combinatorial Auctions via Pseudogradient Dynamics. arXiv 2021. arXiv - handles nondifferentiable ex post payoffs. ⟶ Multi-slot allocations are not smooth. read in Block 6
  • Bichler, Fichtl, Oberlechner. Computing Bayes Nash Equilibrium Strategies in Auction Games via Simultaneous Online Dual Averaging (SODA). EC 2023 / Operations Research 73(2) 2025. doi · arXiv - discretize types and actions, learn distributional strategies with online optimization. ⟶ An offline equilibrium unit test for any bidder. read in Block 6
  • Bichler, Kohring, Heidekrüger. Learning Equilibria in Asymmetric Auction Games. INFORMS J. Computing 2023. doi - asymmetric bidders whose equilibria are PDEs. ⟶ Advertisers differ in value, budget and CVR; symmetry is the wrong benchmark. read in Block 6
  • Pieroth, Kohring, Bichler. Equilibrium Computation in Multi-Stage Auctions and Contests. arXiv 2023. arXiv - deep RL self-play learns multi-stage equilibria with a verifier. ⟶ Budget pacing is a multi-stage game. read in Block 6
  • Han, Zhou, Weissman. Optimal No-regret Learning in Repeated First-price Auctions. Operations Research 73(1) 2025 (arXiv 2020). doi · arXiv - regret bounds when you only see the winning bid. ⟶ Learn with the feedback you actually get. read in Block 6
  • Han, Weissman, Zhou. Learning to Bid Optimally and Efficiently in Adversarial First-price Auctions. arXiv 2020. arXiv - adversarial-competition regret. ⟶ Rivals who adapt against you. read in Block 6
  • Zhang, Han, Zhou, Flores, Weissman. Leveraging the Hints: Adaptive Bidding in Repeated First-Price Auctions. NeurIPS 2022. neurips - point or interval hints about rivals’ max bids. ⟶ Suggested bids as hints; keep the interval. read in Block 6
  • Feng, Podimata, Syrgkanis. Learning to Bid Without Knowing Your Value. EC 2018. arXiv - no-regret bidding when value is only observed on allocation. ⟶ You learn your CVR by winning. read in Block 6
  • Nedelec, Calauzènes, El Karoui, Perchet. Learning in Repeated Auctions. Foundations and Trends in ML 2022. arXiv - monograph on both sides learning. ⟶ The textbook for Block 6. read in Block 6
  • Kolumbus, Nisan. Auctions between Regret-Minimizing Agents. WWW 2022. acm · arXiv - what regret minimizers converge to in second-price and first-price auctions. ⟶ Individually sensible learning changes market outcomes. read in Block 6
  • Kolumbus, Nisan. How and Why to Manipulate Your Own Agent. NeurIPS 2022. arXiv - misreporting to your own autobidder can be rational. ⟶ The advertiser games the tool that games the auction. read in Block 6
  • Deng, Hu, Panageas, Roberts, Zhou et al. Nash Convergence of Mean-Based Learning Algorithms in First Price Auctions. WWW 2022. arXiv - when mean-based learners converge in first-price. ⟶ Convergence is not guaranteed and depends on the algorithm class. read in Block 6
  • Waisman, Nair, Carrion, Xu. Online Inference for Advertising Auctions. Stanford GSB WP 2019. gsb - exploit bid-optimization structure for causal inference. ⟶ The counterfactual a search-at-inference bidder needs. read in Block 7
  • Kohring, Pieroth, Bichler. Enabling First-Order Gradient-Based Learning for Equilibrium Computation in Markets. ICML 2023. arXiv - smooth the discontinuous auction payoff so first-order methods work. ⟶ Differentiable auction simulators. read in Block 6
  • Bichler et al. On the Convergence of Learning Algorithms in Bayesian Auction Games. arXiv 2023. arXiv - variational-inequality view of when equilibrium learning converges. ⟶ Conditions under which self-play in an auction settles. read in Block 6
  • Ahunbay, Bichler. On the Uniqueness of Bayesian Coarse Correlated Equilibria in Standard First-Price and All-Pay Auctions. SODA 2025. arXiv - no-regret dynamics in first-price converge to the unique equilibrium. ⟶ In first-price, learning finds the answer. read in Block 6
  • Bichler, Gupta, Oberlechner. Learning to Bid in Multi-Unit Auctions. ISR. arXiv - equilibrium learning in multi-unit formats. ⟶ Several slots, one bidder. read in Block 6
  • Bichler, Durmann, Oberlechner. Agentic Markets: A Survey. arXiv 2025. arXiv - the survey of markets where every participant is an algorithm. ⟶ The Block 6 reading list in one place. read in Block 6
  • Daskalakis, Syrgkanis. Learning in Auctions: Regret is Hard, Envy is Easy. FOCS 2016. arXiv - no-regret is computationally hard in combinatorial auctions; no-envy is tractable. ⟶ Choose your learning target. read in Block 6
  • Aggarwal, Fikioris, Zhao. No-Regret Algorithms in Non-Truthful Auctions with Budget and ROI Constraints. arXiv 2024. arXiv - regret bounds under both constraints in non-truthful formats. ⟶ Your constraints, in GSP. read in Block 6
  • Deng, Li, Tang, Zhang. Learning in auctions, 2025. NeurIPS 2025. arXiv - a 2025 result on learning dynamics in auctions. ⟶ Recent. read in Block 6
  • Anagnostides et al. Learning in auctions, 2026. arXiv 2026. arXiv - a 2026 result on equilibrium learning. ⟶ Recent. read in Block 6
  • Chen, Morgenstern, Yang. 2026. arXiv 2026. arXiv - a 2026 result on learning agents in auctions. ⟶ Recent. read in Block 6
  • Feng, Guruganesh, Liaw, Mehta, Sethi. Convergence Analysis of No-Regret Bidding Algorithms in Repeated Auctions. AAAI 2021. arXiv - when no-regret bidders converge, and to what. ⟶ The convergence question for autobidders. read in Block 6
  • Deng, Schneider, Sivan. Prior-Free Dynamic Auctions with Low Regret Buyers. NeurIPS 2019. arXiv - the seller exploits low-regret buyers. ⟶ The platform can learn against your learner. read in Block 6
  • Wen et al. (Alibaba). A Cooperative-Competitive Multi-Agent Framework for Auto-bidding (MAAB). WSDM 2022. arXiv - multi-agent autobidding with cooperation and competition. ⟶ The RL side of the equilibrium question. read in Block 6

12. Noam Brown’s program: poker, Diplomacy, reasoning

  • Brown, Sandholm. Regret-Based Pruning in Extensive-Form Games. NeurIPS 2015. nips [Brown] - skip actions with sufficiently negative regret, revisit when they could matter. ⟶ Prune bid levels that are persistently dominated. read in Block 5
  • Brown, Sandholm. Reduced Space and Faster Convergence in Imperfect-Information Games via Pruning. ICML 2017. pmlr [Brown] - best-response pruning; 7× space reduction. ⟶ Same. read in Block 5
  • Brown, Kroer, Sandholm. Dynamic Thresholding and Pruning for Regret Minimization. AAAI 2017. doi [Brown] - thresholding Hedge with only a constant-factor cost. ⟶ Cheap real-time search without breaking guarantees. read in Block 5
  • Brown, Sandholm. Safe and Nested Subgame Solving for Imperfect-Information Games. NeurIPS 2017 (best paper). neurips · arXiv [Brown] - a subgame cannot be solved in isolation; improve it locally while never doing worse than the blueprint. ⟶ The template for a safe local bid re-solver. read in Block 5
  • Brown, Sandholm. Superhuman AI for Heads-Up No-Limit Poker: Libratus Beats Top Professionals. Science 2018. doi [Brown] - blueprint plus nested subgame solving plus self-improvement; 120,000 hands. ⟶ A general policy plus local response to a changing game. (CV lists 2017; journal 2018.) read in Block 5
  • Brown, Sandholm, Amos. Depth-Limited Solving for Imperfect-Information Games. NeurIPS 2018. arXiv [Brown] - let the opponent pick among several continuation strategies at the leaf. ⟶ Evaluate a bid against a portfolio of plausible competitor responses, not one. read in Block 5
  • Brown, Lerer, Gross, Sandholm. Deep Counterfactual Regret Minimization. ICML 2019. pmlr · arXiv [Brown] - networks replace tabular regrets and abstraction. ⟶ Neural regret over bid, budget and query features. read in Block 5
  • Brown, Sandholm. Solving Imperfect-Information Games via Discounted Regret Minimization. AAAI 2019. doi · arXiv [Brown] - discount early regrets; beats CFR+ everywhere tested. ⟶ Adapt to drifting competition without forgetting. read in Block 5
  • Farina, Kroer, Brown, Sandholm. Stable-Predictive Optimistic Counterfactual Regret Minimization. ICML 2019. pmlr [Brown] - use predictions of future regret with stability control. ⟶ A principled slot for your landscape forecast. read in Block 5
  • Brown, Sandholm. Superhuman AI for Multiplayer Poker (Pluribus). Science 2019. doi [Brown] - six-player poker; self-play blueprint plus depth-limited search. ⟶ Two-player intuitions do not transfer to many-bidder markets, and yet the method worked. read in Block 5
  • Brown, Bakhtin, Lerer, Gong. Combining Deep Reinforcement Learning and Search for Imperfect-Information Games (ReBeL). NeurIPS 2020. arXiv [Brown] - public belief states make search possible in imperfect information; converges in two-player zero-sum. ⟶ Posterior over rivals’ bids as the belief state. read in Block 5
  • Lerer, Hu, Foerster, Brown. Improving Policies via Search in Cooperative Partially Observable Games. AAAI 2020. doi [Brown] - search improves every Hanabi agent tested. ⟶ Search as a universal post-processor. read in Block 5
  • Gray, Lerer, Bakhtin, Brown. Human-Level Performance in No-Press Diplomacy via Equilibrium Search. ICLR 2021. arXiv [Brown] - imitation plus one-step equilibrium search. ⟶ Opponent-aware but behaviorally plausible. read in Block 5
  • Bakhtin, Wu, Lerer, Brown. No-Press Diplomacy from Scratch (DORA). NeurIPS 2021. neurips · arXiv [Brown] - 10^20 actions per turn; policy proposals plus double oracle. ⟶ Huge continuous bid spaces. read in Block 5
  • Hu, Lerer, Cui, Pineda, Brown, Foerster. Off-Belief Learning. ICML 2021. pmlr [Brown] - do not infer intent from fragile conventions. ⟶ Do not read competitor strategy from a narrow history. read in Block 5
  • Fickinger, Hu, Amos, Russell, Brown. Scalable Online Planning via Reinforcement Learning Fine-Tuning. NeurIPS 2021. arXiv [Brown] - replace tabular search with online RL fine-tuning of the policy. ⟶ Low-latency local planning for a bid policy. read in Block 5
  • Ling, Brown. Safe Search for Stackelberg Equilibria in Extensive-Form Games. AAAI 2021. doi [Brown] - leader-commitment search never worse than blueprint. ⟶ The platform commits to a mechanism; bidders respond. read in Block 5
  • Jacob, Wu, Farina, Lerer, Hu, Bakhtin, Andreas, Brown. Modeling Strong and Human-Like Gameplay with KL-Regularized Search (piKL). ICML 2022. pmlr · arXiv [Brown] - regret minimization regularized toward an imitation policy. ⟶ The principled cousin of conservative offline RL for bidding. read in Block 5
  • Meta FAIR Diplomacy Team (incl. Brown). Human-Level Play in the Game of Diplomacy by Combining Language Models with Strategic Reasoning (Cicero). Science 2022. doi [Brown] - a language model for dialogue, a planner for actions. ⟶ The modular pattern: LLM for context and communication, numeric planner for the bid. read in Block 5
  • Zhang, Lerer, Brown. Equilibrium Finding in Normal-Form Games via Greedy Regret Minimization. AAAI 2022. arXiv [Brown] - regret-weighted iterate averaging. ⟶ Better equilibrium estimates in auction simulations. (Title differs between CV and arXiv.) read in Block 5
  • Sokota, Hu, Wu, Kolter, Foerster, Brown. A Fine-Tuning Approach to Belief State Modeling. ICLR 2022. iclr [Brown] - specialize a belief model at inference time. ⟶ Specialize the market model to today’s query. (Archival venue to confirm.) read in Block 5
  • Bakhtin, Wu, Lerer, Gray, Jacob, Farina, Miller, Brown. Mastering the Game of No-Press Diplomacy via Human-Regularized Reinforcement Learning and Planning (Diplodocus). ICLR 2023. openreview · arXiv [Brown] - DiL-piKL: RL plus planning regularized toward human play; 200-game tournament with 62 humans. ⟶ Learned policy for regularity, planning for local improvement. read in Block 5
  • OpenAI. Learning to Reason with LLMs and OpenAI o1 System Card. 2024. openai · system card · arXiv - large-scale RL to reason with chain of thought; Brown is a named contributor to reasoning models. ⟶ Test-time compute as the current form of the search thesis. (Public technical documents, not auction papers.) read in Block 5
  • Brown. Curriculum Vitae and talks. noambrown.com - the primary record for what is and is not a Brown paper. ⟶ Use to resolve title and year discrepancies. read in Block 5
  • Brown. Equilibrium Finding for Large Adversarial Imperfect-Information Games. PhD thesis, CMU 2020. pdf [Brown] - the unified account of Libratus and Pluribus. ⟶ The single best long read for Block 5. read in Block 5
  • Sokota, D’Orazio, Ling, Wu, Kolter, Brown. Abstracting Imperfect Information Away from Two-Player Zero-Sum Games. ICML 2023. arXiv [Brown] - a reduction that removes imperfect information for search. ⟶ Search machinery that does not need the full belief state. read in Block 5
  • Sokota, Farina, Wu, Hu, Wang, Kolter, Brown. The Update-Equivalence Framework for Decision-Time Planning. ICLR 2024. arXiv [Brown] - decision-time planning that mirrors the training update. ⟶ A principled way to add search to any learned bidder. read in Block 5
  • OpenAI (incl. Brown). GPT-5 System Card. arXiv 2026. arXiv [Brown] - the current reasoning-model system card. ⟶ Where the test-time-compute thesis stands now. read in Block 5
  • Brown. ReBeL: Combining Deep RL and Search for Imperfect-Information Games. Simons Institute talk. simons [Brown] - the talk version. ⟶ Watch before reading the paper. read in Block 5

13. Amazon disclosures, Amazon Science and academic studies of Amazon

  • Amazon Ads. Sponsored Products across retailers - how the auction works. Help page. amazon - eligibility, expected relevance, bid, context match, predicted engagement; hard thresholds and reserves; price may exceed runner-up but never the adjusted max bid. ⟶ The primary description of the mechanism, richer than textbook GSP. read in Block 8
  • Amazon Ads. Adjust Sponsored Products bids by placement. Help page. amazon - up to 900% combined adjustment; the $1 → $1.50 → $3 → $4.50 worked example. ⟶ Base bid, adjusted bid and CPC must be logged separately. read in Block 8
  • Amazon Ads. Dynamic bidding - down only / up and down / fixed. Help and guide. guide · help · API bid controls - Amazon moves your bid up to ±100% by predicted conversion likelihood. ⟶ The submitted bid is a function Amazon computes from yours. read in Block 8
  • Amazon Ads. Rest of Search Bid Adjustment for Sponsored Products. What’s new, January 2024. amazon - a third placement lever. ⟶ Placement granularity keeps increasing. read in Block 8
  • Amazon Ads. Amazon Marketing Stream. API docs and product page. docs · product · blog - hourly traffic and conversion summaries pushed to Firehose/SQS; hourly bid automation named as a use case; processed, not raw auction data. ⟶ The feedback channel for the hourly loop, and its limits. read in Block 8
  • Amazon Ads. Theme-based bid suggestions - quick-start guide. API docs. amazon - three bids from recent winning bids of similar products with historical impact metrics; explicitly not per-keyword forecasts. ⟶ A market anchor, not a demand curve. read in Block 8
  • Amazon Ads. What is CPC? Library guide. amazon - the billing definition. ⟶ Skim. read in Block 8
  • Amazon. Amazon’s Response to the FTC’s Lawsuit Regarding Sponsored Ads. About Amazon, August 2026. aboutamazon - GSP form since 2006; hard and soft reserves; ~92% of 2024 placements not the highest bid; mean winner ≈ the 128th bid; winning bids down 50%; CVR up 24%; >$8B saved. ⟶ Amazon’s account of rank selection. (Internal date inconsistency: 2019-2025 vs 2019-2024.) read in Block 8
  • FTC. FTC, States Sue Amazon Over Secret Ad Surcharge Scheme and Complaint. August 31, 2026. press · complaint - alleges eOPS-based hidden soft reserves from 2022; First Price Rate from 4% to 52-64%; advertisers paying their own bid close to 80% of the time. ⟶ The other account: price formation. Allegations, not findings. read in Block 8
  • Amazon. Q4 2024 Earnings Release and 2025 10-K. q4 2024 · 10-K · Q2 2026 - advertising services $37.7B (2022) → $46.9B → $56.2B → $68.6B (2025); Q2 2026 $19.8B, +26%. ⟶ Scale, not mechanism. read in Block 8
  • Amazon. AI advertising benefits. Library news. amazon - billions of parameters, real-time shopping and streaming signals, bid and budget recommendations. ⟶ Establishes relevance of the ML literature, not which paper is used. read in Block 8
  • Mondal, Kandregula, Agrawal, Sembium (Amazon). OPTIMUS: Optimal Offline Bidding Strategy for Manual Targeting Advertising Campaigns. ECML-PKDD 2026. amazon.science · pdf - equalize marginal ROAS across clauses via a Lagrangian over forecast bid landscapes; +2-6% sales online. ⟶ The clearest public Amazon-linked bid recommender. (Venue per one report; another could not confirm.) read in Block 8
  • Amazon Science. Ad-related technologies tag. amazon.science - the index of Amazon ads papers. ⟶ Where to look for 2026 additions. read in Block 8
  • Farronato, Fradkin, MacKay. Self-Preferencing at Amazon: Evidence from Search Rankings. NBER WP 30894, 2023; AEA P&P. nber · doi - 228k results from 184 users; sponsored prominence ≈ 7 positions; Amazon-brand coefficient ≈ 60% of the sponsored one. ⟶ Position is part of what a bid buys; measure it separately from quality. (Working paper.) read in Block 8
  • Dash, Ghosh, Mukherjee, Chakraborty, Gummadi. Sponsored is the New Organic. arXiv 2024. arXiv - 4,800 searches across four marketplaces; ~30% of results sponsored; a sponsored result precedes the top organic one 85% of the time; top sponsored often 50% costlier. ⟶ An outside stress test of the relevance claim. (Preprint, observational.) read in Block 8
  • Rock, Strauss, O’Reilly, Mazzucato. Behind the Clicks: Can Amazon Allocate User Attention as it Pleases? SSRN 2023. ssrn - attention share predicts clicks despite worse price or rating. ⟶ Prominence buys behavior. (Final venue not identified.) read in Block 8
  • Yu. The Welfare Effects of Sponsored Product Advertising. Stanford / SSRN 2024. ssrn - structural model on Amazon searches, purchases and bids; ads help newer, differentiated products; commission policy flips the welfare sign. ⟶ Revenue-maximizing auctions and conversion-maximizing bidders can pull in different directions. (Working paper.) read in Block 8
  • Preuss, Jungbauer, Janssen, Williams. Search Platforms: Big Data and Sponsored Positions. Economic Journal 2026. doi · cornell - sponsored slots can improve experience while organic obfuscation raises revenue. ⟶ Ranking architecture as a coordinated monetization system. read in Block 8
  • Rel.ai. How Amazon’s Ad Auction Actually Works. Blog, Feb 2026. rel.ai - practitioner explainer leading with relevance weighting. ⟶ Skim for framing. read in Block 1
  • Karlsson (Amazon). Multivariable Feedback Control for Multi-Constraint Optimization in Online Advertising. CDC 2025. amazon.science - control-theoretic pacing under several constraints. ⟶ Amazon’s control-theory answer to USCB. read in Block 8
  • Amazon Science. MESOB: Balancing Equilibria and Social Optimality. KDD workshop 2023. amazon.science - trade off equilibrium and social optimum in multi-agent bidding. ⟶ Amazon thinking about the agents-against-agents problem. read in Block 8
  • Ge et al. (Amazon). Multi-Task Combinatorial Bandits for Budget Allocation. AdKDD 2024. amazon.science - bandits over budget splits. ⟶ Budget allocation as exploration. read in Block 8
  • Nabi et al. (Amazon). Bayesian Meta-Prior Learning Using Empirical Bayes. Management Science 2022. amazon.science - hierarchical priors learned across tasks. ⟶ The shrinkage machinery for tail keywords. (Ads application not stated in the source.) read in Block 8
  • Qin (Amazon). Lengthen Your Attribution Window: Which Digital Ads Have Most Long-Term Impact? 2023. amazon.science - long-horizon effects by ad type. ⟶ The 14-day window undervalues some formats. read in Block 8
  • Pauwels, Schnaidt, Caddeo (Amazon). Causal Impact of Digital Display Ads on Advertiser Performance. EMAC 2022. amazon.science - causal effects of display at Amazon. ⟶ Amazon’s own incrementality evidence. read in Block 8
  • Amazon Ads. Best Practices for Your Sponsored Products Ads. Guide. amazon - “Daily budgets are not paced throughout the day”; a $100/day budget “may receive up to $3,000 worth of clicks in that calendar month”; the final CPC “will never exceed your maximum adjusted bid”. ⟶ Three sentences that pin down the pacing model (throttling, not multiplicative) and the price ceiling. read in Block 8

14. Industry and vendor documentation

  • Perpetua. Product page. perpetua.io - “contextual, conversion-based bidding algorithms” for growth, profitability, brand defense. ⟶ Controls disclosed; model not. read in Block 8
  • Pacvue. Pacvue for Amazon. pacvue - rules, AI bidding, dayparting, Marketing Stream; claims 10%+ ROAS lift. ⟶ Vendor-reported. read in Block 8
  • Quartile. Quartile for Amazon PPC. quartile - hourly single-keyword bidding off Marketing Stream and AMC; claims +41% ROAS. ⟶ The strongest public hourly-control description. read in Block 8
  • Adbrew. Product page. adbrew - AI insights plus rule-based automation. ⟶ Rules are concrete; attribution is not. read in Block 8
  • Teikametrics. Intro to the Teikametrics Bidder. Help center. teikametrics - forecast AOV and CVR, compute an optimal bid under an ACOS limit. ⟶ OPTIMUS-shaped pipeline, without the optimality conditions. read in Block 8
  • wnzhang. rtb-papers and MobileTeleSystems. rlrtb. GitHub collections. rtb-papers · rlrtb - curated RTB and RL-for-RTB paper lists. ⟶ For anything this catalogue missed. read in Block 4

Part B - Chronological

The load-bearing subset, in order. The last column names the shift each paper represents: what someone reading the field could believe after it that they could not before.

YearPaperThemeOne-line shift it represents
1981Myerson, Optimal Auction DesignMechanismRevenue-optimal auctions are virtual-value maximizers with reserves
2006Aggarwal, Goel, Motwani, Truthful auctions for keywordsMechanismA truthful position auction exists
2007Edelman, Ostrovsky, Schwarz, GSPMechanismSponsored search has an equilibrium theory, and it is not VCG
2007Varian, Position AuctionsMechanismSlots, not items, are what is sold
2007Lahaie, Pennock, Ranking rulesMechanismThe quality exponent is a revenue dial
2009Ghose, Yang, Empirical analysisMechanismKeyword value can be measured from logs
2010Milgrom, Simplified mechanismsMechanismRestricting bids removes bad equilibria
2010Yang, Ghose, Organic vs sponsoredMeasurementPaid and organic complement each other
2011Ostrovsky, Schwarz, Reserve field experimentMechanismReserves raise revenue in the field, unevenly
2011Athey, Ellison, Consumer searchMechanismClick curves are endogenous to the ad mix
2011Dudík, Langford, Li, Doubly robustOPELog propensities and you can evaluate untried policies
2013McMahan et al., View from the trenchesPredictionCalibration is a pricing property
2014He et al., Practical lessons at FacebookPredictionFreshness beats architecture
2014Chapelle, Delayed feedbackPredictionConversions are censored; model the delay
2014Zhang, Yuan, Wang, Optimal RTBBiddingThe optimal bid is concave in value under a budget
2014Blake, Coey, Marketplace interferenceMeasurementNaive A/B in a marketplace can be off by 2×
2015Caragiannis et al., Inefficiency of GSPMechanismGSP PoA is 1.282 pure, 2.927 Bayesian
2015Balseiro, Besbes, Weintraub, Repeated auctions with budgetsPacingBudgeted auctions have a tractable fluid limit
2015Brown, Sandholm, Regret-based pruning [Brown]BrownRegret tells you what not to compute
2015Lewis, Rao, Unfavorable economicsMeasurementAd effects are too small to see without huge experiments
2016Cheng et al., Wide & DeepPredictionMemorize and generalize in one net
2017Cai et al., RTB by RLBiddingBidding is a sequential decision problem
2017Brown, Sandholm, Safe and nested subgame solving [Brown]BrownYou can improve locally without ever doing worse
2017Joachims et al., Unbiased LTRPosition biasClicks are examination × relevance
2017Johnson, Lewis, Nubbemeyer, Ghost adsMeasurementMeasure the impression you did not show
2017Wilkens, Cavallo, Niazadeh, GSP: CinderellaMechanismGSP is truthful for value maximizers
2017Conitzer et al., Multiplicative pacing equilibriaPacingPacing multipliers form a game with multiple equilibria
2018Brown, Sandholm, Libratus [Brown]BrownBlueprint plus search beats humans at imperfect information
2018Brown, Sandholm, Amos, Depth-limited solving [Brown]BrownLeaf values must be sets of continuations, not numbers
2018Ma et al., ESMMPredictionEstimate CVR over the whole impression space
2018Zhou et al., DINPredictionUser representation should depend on the candidate
2018Wu et al., Budget-constrained model-free RLBiddingAct on the multiplier, not the bid
2018Zhao et al., Sponsored search RTB by RLBiddingRL bidding in retail sponsored search, hourly policies
2018Jin et al., Multi-agent RTBBiddingCompetitors are part of the state
2018Feng, Podimata, Syrgkanis, Bid without knowing your valueEquilibrium learningYou learn value only by winning
2019Aggarwal, Badanidiyuru, Mehta, Autobidding with constraintsAutobiddingAdvertisers are value maximizers with ROAS targets
2019Balseiro, Gur, Repeated auctions with budgetsPacingMultiplicative pacing is asymptotically optimal and equilibrates
2019Conitzer et al., First-price pacing equilibriumPacingFirst-price pacing equilibrium is unique and convex
2019Dütting et al., RegretNetLearned mechanismsMechanisms can be learned with a regret penalty
2019Brown et al., Deep CFR [Brown]BrownRegret can live in a network
2019Brown, Sandholm, Discounted CFR [Brown]BrownForget early regrets to converge faster
2019Brown, Sandholm, Pluribus [Brown]BrownThe method survives six players
2019Ren et al., Deep landscape forecastingBiddingThe win curve is a censored survival problem
2019Nigam et al., Semantic product searchRelevanceMatching is dense retrieval
2019Ktena et al., Delayed feedback, continuous trainingPredictionDuplicate and reweight to stay unbiased online
2019Zeithammer, Soft floorsMechanismSoft floors do not raise revenue in the standard model
2019Google, Ad Manager to first priceIndustryThe exchange world goes first-price
2020Calvano et al., Algorithmic collusionCollusionQ-learners collude without talking
2020Brown et al., ReBeL [Brown]BrownPublic belief states make search work in imperfect information
2020Pan et al., Field-aware calibrationPredictionCalibrate per field, not globally
2020Pi et al., SIMPredictionRetrieve from lifelong history, then attend
2020Gligorijevic et al., Bid shadingFirst priceShading is a learned landscape problem
2021Babaioff et al., Non-quasi-linear agentsAutobiddingClassical guarantees assumed the wrong bidder
2021Deng, Mao, Mirrokni, Zuo, Efficient auctions in an autobidding worldAutobiddingBoosts raise welfare and revenue
2021Balseiro et al., Value vs utility maximizersAutobiddingBidder type determines extractable revenue
2021Balseiro et al., Robust auction designAutobiddingReserves help without knowing bidder type
2021Bichler et al., NPGAEquilibrium learningNeural self-play finds auction equilibria
2021Zhang et al., Deep GSPLearned mechanismsThe relevance multiplier is a trained network
2021Liu et al., Neural AuctionLearned mechanismsA learned auction runs at Taobao scale
2021He et al., USCBBiddingEvery constraint is a dual variable
2021Liu, Mao, Kang, Budget-split designMeasurementSplit budgets to build two marketplaces
2021Gu et al., DEFERPredictionReal negatives, reweighted
2021Bakhtin et al., Diplomacy from scratch [Brown]BrownDouble oracle over 10^20 actions
2022Banchio, Skrzypacz, AI and auction designCollusionLearners collude more in first-price
2022Bergemann et al., Calibrated click-through auctionsMechanismpCTR is an information-design object
2022Mehta, Randomization beyond VCGAutobiddingRandomize to beat PoA 2
2022Liaw, Mehta, Perlroth, Non-truthful auctionsAutobiddingDeterministic PoA ≥ 2; randomized 1.8
2022Mou et al., SORL / V-CQLBiddingConservative offline RL is the deployment bridge
2022Jeunen, Murphy, Allison, AuctionGymBenchmarksMost learning-to-bid is value-based bandits; simulate
2022Jacob et al., piKL [Brown]BrownRegularize search toward a behavior prior
2022FAIR, Cicero [Brown]BrownLLM for dialogue, planner for actions
2022Kolumbus, Nisan, Regret-minimizing agentsEquilibrium learningLearners change what the auction converges to
2022Li et al., Two-sided interference guidanceMeasurementThe randomization side is a design variable
2022Fan, Si, Zhang, Calibration MattersPredictionThe argmax inflates selected predictions
2023Gordon, Moakler, Zettelmeyer, Close Enough?Measurement663 experiments; observational methods still miss
2023Kohring, Pieroth, Bichler, First-order equilibrium learningEquilibrium learningSmooth the auction, then differentiate
2023Bakhtin et al., Diplodocus [Brown]BrownHuman-regularized RL plus planning
2023Chen, Nabi, Siniscalchi, Auction realismBenchmarksSimulate changing, unobserved competitors
2023Bojinov, Simchi-Levi, Zhao, SwitchbacksMeasurementOptimal time-block designs with exact inference
2023Ostrovsky, Schwarz, Reserve prices (JPE)MechanismThe field experiment, archival
2023Farronato, Fradkin, MacKay, Self-preferencingAmazonSponsored prominence ≈ 7 positions
2024Dütting et al., Differentiable economics (JACM)Learned mechanismsRegretNet, archival
2024Dütting, Mirrokni, Paes Leme, Xu, Zuo, Mechanism design for LLMsLearned mechanismsAuctions over generated text
2024Liaw, Mehta, Zhu, GSP for value maximizersAutobiddingThe PoA of Amazon’s mechanism with Amazon’s bidders
2024Jain, Appala, SERP interference networkMeasurementThe interference graph of a search page
2024Guo et al., AIGBBiddingGenerate a bidding trajectory with diffusion
2024Su et al., AuctionNetBenchmarksA shared 500M-record benchmark for bidders
2024Lucier et al., Autobidders with budget and ROIAutobiddingGuarantees without convergence
2024Zhai et al., HSTUPredictionRanking enters the scaling-law era
2024Aggarwal et al., Autobidding surveyAutobiddingThe field gets a map
2024Aggarwal, Gupta, Perlroth, Velegkas, Randomized truthful auctions with learning agentsEquilibrium learningTruthful mechanisms are not truthful under learning
2024OpenAI, o1BrownTest-time compute as the search thesis at scale
2025Ahunbay, Bichler, Uniqueness of BCCE in first-priceEquilibrium learningNo-regret learning in first-price finds the equilibrium
2025Meta, GEMPredictionA foundation model for ads ranking
2025Karlsson, Multivariable feedback control (Amazon)AmazonControl theory for multi-constraint pacing
2025Lewis et al., Amazon Ads MTAAttributionRCT-calibrated ML attribution at Amazon
2025Mou et al., AIGB-PearlBiddingGenerative bidding gets a critic
2025Yeom et al., Breaking DeterminismOPELandscape models as propensities in deterministic auctions
2025Bichler, Fichtl, Oberlechner, SODA (OR)Equilibrium learningDiscretize and dual-average to Bayes-Nash
2026Fish, Gonczarowski, Shorrer, Collusion by LLMsCollusionLLM pricing agents collude too
2026Mondal et al., OPTIMUSAmazonMarginal-ROAS equalization at Amazon scale
2026FTC complaint and Amazon responseAmazonTwo incompatible accounts of the same auction
2026Yang, Zuo, Kim, Generative response modelingBiddingAuditable constraints instead of reward shaping
2026Preuss et al., Search platformsAmazonSponsored and organic ranking as one monetization system

Part C - If you only have three hours

Twelve papers, 180 minutes. Read them in this order; the first six give you the auction, the last six give you the learning.

MinutesPaperWhy this one
20Edelman, Ostrovsky, Schwarz, GSP (2007)The mechanism and its equilibrium in one sitting
10Amazon Ads, How the auction works help page + Amazon’s FTC response (2026)What the platform says it does, and the numbers
10FTC complaint, sections on eOPS and soft reserves (2026)The other account; read the two back to back
15Aggarwal, Badanidiyuru, Mehta, Autobidding with constraints (2019)Why you are a value maximizer and what that costs
15Balseiro, Gur, Repeated auctions with budgets (2019)Why every pacer is a multiplicative dual update
10Liaw, Mehta, Perlroth, Non-truthful auctions in auto-bidding (2022)The PoA numbers
15Ma et al., ESMM (2018) + Fan, Si, Zhang, Calibration Matters (2023)The two ways a CVR model silently mis-prices
20Jeunen, Murphy, Allison, Learning to bid with AuctionGym (2022)Amazon’s own account of why logs cannot evaluate a bidder
20Guo et al., AIGB (2024)The generative turn in bidding
15Brown, Sandholm, Safe and nested subgame solving (2017)The safe-local-improvement idea you will reuse
15Jacob et al., piKL (2022)Behavior-regularized search, which is what conservative offline RL is reaching for
15Mondal et al., OPTIMUS (2026)The whole loop, assembled, with online results

Then go back to the day index and pick the block that hurt most.