← Day index · Block 10 of 10 · ← Previous · Back to the day index →

The day ends where the literature does. Everything before this page is settled enough to teach; everything on it is a place where two careful papers disagree, where the theory has a guarantee the practice cannot use, or where a production system is running ahead of anything published. Each problem below is written the same way: what it is, what is actually known, why it resists, a first experiment you could run this month with an Amazon Ads API account, Marketing Stream, and one of the open simulators, and which blocks to re-read before you start. Citations point into the bibliography; the map is Amazon Ads, Deeply.

A note on the shape of these problems. Almost every one of them is a version of the same tension: the auction is a game between learning systems, and every tool we have for studying it - equilibrium theory, offline RL, off-policy evaluation, A/B testing - was built for a world where at least one side holds still. Nobody holds still on a retail search page.

1. Agents against agents: does sponsored search collude?

The problem. When every advertiser on a keyword is an RL bidder and the platform’s ranker is a learned mechanism, the repeated auction is a multi-agent game among algorithms. Calvano et al. showed Q-learners in repeated price competition learn supra-competitive prices with a punishment-and-forgiveness structure, without communication [Calvano et al. 2020]; nobody has established whether sponsored-search bidders do the analogous thing - bid suppression - in a live market.

What is known. Banchio and Skrzypacz show learning bidders collude more readily in first-price than second-price formats [Banchio, Skrzypacz 2022; Banchio, Mantegazza 2023], and field evidence from German gasoline shows algorithmic pricing raised margins [Assad et al. 2024]. Rawat’s controlled experiments over RL and bandit bidders find suppression in some designs [Rawat 2023-2025]. Aggarwal, Gupta, Perlroth and Velegkas show no-regret agents fail to converge to truthful bidding even in deterministic truthful auctions, and that randomized auctions can outperform second-price-with-reserve against learners over long horizons [Aggarwal et al. 2024]. Kolumbus and Nisan show what regret minimizers converge to, and that manipulating your own agent can be rational [Kolumbus, Nisan 2022]. Amazon’s own auction-realism paper models advertisers as adversarial bandits and finds soft floors help the platform [Chen, Nabi, Siniscalchi 2023].

Why it is hard. Collusion is defined relative to a competitive benchmark that is itself unknown in a relevance-weighted auction with private ROAS targets; the pacing-equilibrium literature says such markets can have multiple equilibria with very different prices [Conitzer et al. 2022], so a low-price outcome may be an equilibrium rather than a cartel. And the platform’s response - reserves, randomization, relevance weighting - changes the game while you measure it.

First experiment. In AuctionNet, replace the 48 stock agents on a subset of opportunities with your own bidder family (a USCB-style Lagrangian bidder, a V-CQL bidder, a diffusion bidder) and run 100k-auction episodes under three mechanisms: GSP with hard reserve, GSP with a soft floor calibrated as in Chen et al., and a randomized allocation à la Mehta. Track clearing prices relative to the value-truthful benchmark and the Calvano signature: a price drop after one agent deviates, followed by gradual recovery. Then test whether a soft floor or randomization breaks the pattern.

Re-read: Block 2, Block 6.

2. Two accounts of one auction: an identification problem

The problem. Amazon says roughly 92% of 2024 Sponsored Products placements did not go to the highest bid, that the mean winner was around the 128th bid by amount, that winning bids fell 50% and conversion rates rose 24% [Amazon 2026]. The FTC alleges that from 2022 an eOPS-based system set hidden soft reserves for almost all search-page ads, that the “First Price Rate” rose from 4% to 52-64%, and that advertisers paid their own bid close to 80% of the time [FTC 2026]. Both can be true at once: one is about rank selection, the other about price formation.

What is known. Amazon’s own help page says the price may exceed the runner-up bid but never the adjusted maximum bid, and that reserves use sale likelihood, predicted ROAS, competing bids and placement [Amazon Ads help]. Amazon’s response describes a hard reserve and a soft reserve, with a winner above both paying the soft reserve [Amazon 2026]. The theory says reserves can raise both revenue and welfare robustly [Balseiro et al. 2021] and the one field experiment found effects concentrated on high-volume, thin keywords [Ostrovsky, Schwarz 2011/2023].

Why it is hard. From an advertiser’s logs you observe bid, adjusted bid, CPC, position and outcome - never the runner-up score or the reserve. Separating “I paid my bid because the soft reserve bound” from “I paid my bid because the runner-up’s score was close” requires either the platform’s internals or a design that varies your bid while holding the competition fixed, which a single advertiser cannot do.

First experiment. Take 200 keywords with stable hourly traffic on Marketing Stream. Randomize bids on a fine grid (±5%, ±10%, ±20% around current) by hour in a rerandomized switchback [Ni, Kalfountzou, Bojinov 2025]. Regress CPC on bid within (keyword, placement, hour-of-week). Under pure second pricing with a bound reserve, is ≈0 until you change rank; under first-price-like soft-reserve pricing it is ≈1 over a range. Report the fraction of clicks in the ≈1 regime by placement. Log base bid, dynamic-bidding strategy and placement multiplier separately or the estimate is meaningless.

Re-read: Block 1, Block 8.

3. Learned mechanism versus learned bidder: whose model of the other side is wrong?

The problem. Deep GSP and Neural Auction train the platform’s allocation and pricing against a population of modeled bidders [Zhang et al. 2021; Liu et al. 2021]. AIGB and its successors train the bidder’s trajectory generator against a simulator of the platform [Guo et al. 2024]. Each side optimizes against a model of the other; at most one of those models is right, and the surplus flows toward whoever is less wrong.

What is known. Deep GSP’s rank score is a network optimizing a constrained mixture of revenue, CTR, CVR and experience. The efficiency theory says the outcome depends on whether bidders are utility or value maximizers and on whether the mechanism is deterministic [Balseiro et al. 2021; Liaw, Mehta, Perlroth 2022]. PE-MORL is the first bidder-side work to explicitly model the competitor population as part of the environment and penalize model error pessimistically [Mou et al. 2025]. Nobody has run a learned mechanism against learned bidders in a closed loop and reported who adapts faster.

Why it is hard. The two learning problems have different timescales (the platform retrains its ranker on days, the bidder reprices hourly), different observability (the platform sees everything, the bidder sees censored outcomes), and different objectives. Standard equilibrium concepts assume both sides have converged; here the interesting behavior is the transient.

First experiment. In AuctionGym, replace the fixed GSP allocation with a small Deep GSP-style scorer retrained every N simulated days on logged outcomes, and let a population of V-CQL or DiffBid bidders retrain every simulated hour. Sweep the ratio of retraining frequencies. Measure platform revenue, advertiser liquid welfare and the drift in the scorer’s implicit quality exponent. The prediction from Block 2 is that the faster learner captures surplus until the slower one’s next update; check whether the system cycles or settles.

Re-read: Block 1, Block 4.

4. Search at inference time for a bid

The problem. Brown’s program says a policy improved by search at decision time beats a policy alone, provided leaf values are handled correctly under imperfect information [Brown, Sandholm 2017; Brown et al. 2020]. A bidder could, each hour, take its policy’s proposed bid vector and run a shallow lookahead against a portfolio of competitor continuations before committing. Whether that helps or hurts depends entirely on the simulator, and the simulator is the weakest component we have.

What is known. Depth-limited solving lets the opponent choose among several continuation strategies at the leaf rather than assigning a single value [Brown, Sandholm, Amos 2018]. GAS already does post-training search over generated bid plans [Li et al. 2024] and AIGB-Pearl adds a learned evaluator with a KL-Lipschitz constraint to keep search inside the data [Mou et al. 2025]. Scalable online planning replaces tree search with online RL fine-tuning of the policy [Fickinger et al. 2021]. ReBeL’s convergence guarantee is for two-player zero-sum; the ad auction is neither.

Why it is hard. The failure mode is precise: search optimizes against simulator artifacts. In poker the rules are known; in the auction the “rules” include a relevance model and reserve function you cannot see, and the competitor continuations you enumerate are guesses. Search amplifies whatever the simulator gets wrong.

First experiment. Build a portfolio of four competitor continuations from AuctionNet’s agent families (a pacer, a value-truthful bidder, an aggressive multiplier bidder, a copy of yourself). For each hourly decision, evaluate the policy’s top-k candidate bid vectors against the worst case over the portfolio (the depth-limited-solving rule) and commit to the maximin. Compare against the policy alone on held-out AuctionNet episodes with perturbed competitor populations. Report the gain as a function of simulator misspecification: swap the competitor families at test time and see where search starts hurting.

Re-read: Block 5, Block 4.

5. Behavior-regularized policies as the principled version of conservative offline RL

The problem. Conservative offline RL for bidding (V-CQL, the hybrid base-policy approach) keeps the learned bidder near the logged behavior because unsupported actions are dangerous [Mou et al. 2022; Korenkevych et al. 2024]. piKL and DiL-piKL keep a searched policy near a human imitation policy because unregularized optimization produces strong but alien play [Jacob et al. 2022; Bakhtin et al. 2023]. These are the same idea with different justifications, and the auction version has never been stated cleanly.

What is known. piKL regularizes regret minimization with a KL penalty toward an anchor policy and shows both stronger and more human-like play; the penalty weight is the dial. AIGB-Pearl’s KL-Lipschitz constraint is the closest bidding analogue. Expert-guided diffusion planning uses expert trajectories as a prior [Peng et al. 2025]. The offline RL literature’s justification is support; piKL’s is behavioral plausibility. In an auction the second justification has teeth: a bidder that behaves unlike the population invites a platform response (a reserve, a relevance penalty) that the offline data never saw.

Why it is hard. Choosing the anchor. Logged behavior is the incumbent policy, which may be bad; a “human” anchor does not exist for hourly bids; the population’s behavior is unobserved except through prices. And the right regularization strength is a function of how much the market has drifted since the logs were collected, which is itself unknown.

First experiment. On AuctionNet, train a DiffBid or decision-transformer bidder, then fine-tune with a KL penalty toward three anchors: the logged behavior policy, a USCB Lagrangian bidder, and a value-truthful bidder. Sweep the penalty weight over three orders of magnitude. Evaluate on episodes with (a) the training competitor population and (b) a shifted population. The hypothesis, from piKL: there is a that improves on both the unregularized bidder under shift and the anchor in-distribution, and the anchor that works is the Lagrangian bidder, not the logs.

Re-read: Block 5, Block 4.

6. Delayed feedback meets hourly repricing

The problem. Amazon attributes conversions up to 14 days after a click; the hourly repricer needs a conversion estimate now. The click from 90 minutes ago is a negative in your training set and an unresolved event in reality, and the bias is worst for exactly the fresh data the repricer weights most.

What is known. Chapelle modeled the delay distribution; Ktena et al. duplicate and reweight; ES-DFM weights by elapsed time; DEFER uses real negatives with importance sampling and reports >6% CVR gains; GDFM uses post-click actions as early evidence; MISS fuses multiple maturity windows; IF-DFM approximates retraining with influence functions in seconds [Chapelle 2014; Ktena et al. 2019; Yang et al. 2021; Gu et al. 2021; Yang, Zhan 2022; Liu et al. 2024; Ding et al. 2025]. Google’s feed-ads stack handles label delay with multi-stage training [Ma et al. 2022]. None of these is evaluated on the quantity that matters here: the pricing error of a bid set at hour using labels matured to hour .

Why it is hard. The decision cadence (hourly) is much shorter than the label horizon (weeks), so every bid is set on mostly-immature data by construction; the correction depends on a delay distribution that itself shifts with promotions, day of week and product category; and the tail keywords that need the correction most have too few conversions to estimate a delay curve.

First experiment. From Marketing Stream, build a dataset of (click hour, conversion hour) for your account over 90 days. Fit a per-category delay distribution. Simulate the hourly repricer three ways: naive (immature labels as negatives), ES-DFM-weighted, and posterior-with-delay (the Bayesian shrinkage of the existing note with the delay likelihood from Chapelle). Replay 30 days and score each on realized 14-day ROAS of the bids it would have set. Report the gap by category and by keyword volume decile.

Re-read: Block 3.

7. Calibration under maximization bias in the auction’s selected region

The problem. The auction picks the argmax of ; the selected predictions are the ones most likely to be over-estimated, so even a globally calibrated model is mis-calibrated exactly where it is priced [Fan, Si, Zhang 2023]. For an advertiser, the same winner’s curse applies to the internal CVR model: the keywords you bid up are the ones whose posterior mean is highest, and those means are biased upward.

What is known. Field-aware calibration corrects per placement or query class [Pan et al. 2020]; DESC and ConfCalib handle multi-field and sparse-field calibration [Yang et al. 2024; Zhao et al. 2024]; variance-adjusted debiasing targets maximization bias directly [Fan, Si, Zhang 2023]. On the advertiser side, the posterior-width exploration signal in the existing note is a partial answer, since Thompson sampling implicitly discounts uncertain winners.

Why it is hard. The selection is performed by the platform’s model, not yours, so the bias in your own estimates depends on the correlation between your errors and Amazon’s - which you cannot observe. And post-hoc calibration on logged outcomes is itself selected data.

First experiment. For 500 keywords, compute your pre-bid CVR estimate and the realized 14-day CVR. Bin by the rank of your estimate within its category (top decile, middle, bottom). Calibration error in the top decile relative to the middle is the maximization bias. Then re-run with variance-adjusted estimates (shrink each estimate toward the cluster mean by a factor proportional to its posterior variance) and check whether the top-decile gap closes without hurting the rest.

Re-read: Block 3.

8. Off-policy evaluation with zero support in a deterministic auction

The problem. IPS and doubly robust estimators need the logging policy to have put positive probability on the action you want to evaluate. A deterministic auction gives zero propensity to every bid you did not submit, and to every position you did not win [Yeom et al. 2025]. AuctionGym’s authors name this as the reason a simulator exists at all [Jeunen, Murphy, Allison 2022].

What is known. Breaking Determinism repurposes a bid-landscape model as a propensity model for self-normalized IPS [Yeom et al. 2025]. Landscape forecasting gives the win-probability curve from censored data [Ren et al. 2019; Ou et al. 2023]. Amazon’s LTR group has estimators for business-rule post-processing and flexible click models [Jakimov et al. 2023; Buchholz et al. 2024]. Waisman, Nair and Carrion use the auction structure itself for identification and Thompson sampling to bound experimentation cost [Waisman et al. 2025].

Why it is hard. A landscape-model propensity is a model, so the “IPS” estimate inherits its bias and you are back to model-based evaluation with a different name. Real support requires randomizing bids, which costs money and, in a marketplace, contaminates your own control.

First experiment. Log a small exploration budget: 5% of keyword-hours get a bid perturbed by a random factor in , and log the factor. After four weeks you have a logging policy with known non-degenerate propensities on a band around your bids. Evaluate three candidate policies (a +10% bidder, a USCB bidder, the current one) with IPS, SNIPS and DR on that band; compare against the landscape-propensity estimator from Yeom et al. on the same data. The disagreement between the two is the price of assuming determinism away.

Re-read: Block 7.

9. Marketplace interference for bid-policy experiments: what is the unit?

The problem. If half your keywords bid more aggressively, prices rise for the other half, and every within-account comparison is contaminated. Across accounts, your treatment raises prices for competitors who are somebody else’s control. Blake and Coey measured a factor-of-two overstatement from exactly this on eBay [Blake, Coey 2014].

What is known. Li, Zhao, Johari and Weintraub give guidance on which side to randomize and at what proportion, and show bias reduction costs variance [Li et al. 2022]. Budget-split designs build two counterfactual marketplaces [Liu, Mao, Kang 2021]. Switchbacks randomize over time and rerandomization balances covariates [Bojinov et al. 2023; Ni et al. 2025]. Shadow prices give a first-order interference correction [Bright, Delarue, Lobel 2025]. Cluster randomization has been sized against a meta-experiment at Airbnb [Holtz et al. 2025]. Amazon’s own SERP interference network is the first attempt to write down the graph for a retail search page [Jain, Appala 2024], and Amazon has measured cross-unit spillovers in its ads experiments [Jain et al. 2023].

Why it is hard. Keywords share queries (broad match), queries share pages, pages share shoppers, and shoppers share sessions; the interference graph is dense and unobserved. Time is the only unit you fully control, and switchbacks lose power to carryover in a market where competitors’ pacers respond within the hour.

First experiment. Run the same bid change three ways on matched keyword sets for two weeks each: keyword-level randomization, campaign-level (budget-split) randomization, and a daily rerandomized switchback. The three estimates of the same effect will disagree; the pattern of disagreement (keyword-level largest, switchback smallest) is the interference signature and its size tells you which design you can afford to trust for the next test.

Re-read: Block 7.

10. Incrementality versus attribution as the reward signal

The problem. The bidder optimizes attributed sales because that is what the platform reports. Attributed is not incremental: observational attribution fails to recover experimental lift [Gordon et al. 2019], and last-touch systematically overvalues bottom-funnel brand terms [Yang, Ghose 2010]. Amazon’s MTA is RCT-calibrated but still allocates credit under a model [Lewis et al. 2025]. A bidder whose reward is wrong is optimizing the wrong thing with great precision.

What is known. Ghost ads make control-side measurement cheap for display [Johnson, Lewis, Nubbemeyer 2017] but there is no published analogue for a sponsored slot on a retail search page. PIE predicts RCT lift from campaign features with R² 0.88 against 0.19 for last-click [Gordon, Moakler, Zettelmeyer 2023]. Lewis and Rao explain why the experiments are expensive [Lewis, Rao 2015]. Uplift models rank by treatment effect rather than outcome [He et al. 2024].

Why it is hard. For a product you also rank for organically, the counterfactual to “show the ad” is “show the organic listing a few positions lower”, and the shopper may buy anyway; the incremental effect is small relative to sales variance; and the only clean holdout is to stop advertising, which cedes the slot to a competitor and changes the counterfactual.

First experiment. Pick 40 keywords where you hold a top-3 organic position. Geo- or time-randomize ad presence (GeoLift-style synthetic control across regions if you have regional data; otherwise a weekly switchback) and measure total (organic plus paid) sales, not attributed sales. Fit the incremental-per-dollar estimate against the attributed ROAS; the ratio, by keyword type, is the correction factor to apply to the bidder’s reward. Then re-run the bidder with corrected rewards in replay and see how much spend moves from brand terms to discovery terms.

Re-read: Block 7.

11. LLM planners over numeric bidders

The problem. AucArena, RTBAgent and HARBOR show LLM agents can plan in auctions [Chen et al. 2024; Cai et al. 2025; Jiang, Xiong, Liu 2025]. RTBAgent reports beating RL baselines on 9 campaigns with memory, retrieval and daily reflection. Whether a language model should ever be in the loop of a system that sets thousands of hourly bids is unresolved, and Cicero’s architecture - language model for context, numeric planner for actions - is the obvious hypothesis nobody has tested on ads [FAIR 2022].

What is known. Numeric bidders are calibrated, constraint-safe and fast; LLM agents are neither calibrated nor fast but can read a product page, a competitor’s listing or a promotion calendar. InfoBid shows LLM agents’ behavior depends on what the platform discloses [Yin et al. 2025]. Amazon uses LLMs for relevance labels, not for the live ranker, as far as public papers show.

Why it is hard. Evaluation. An LLM planner’s value is in rare, contextual decisions (a competitor went out of stock; a holiday is coming) whose effect is swamped by the numeric bidder’s hourly variance, and there is no benchmark with that structure.

First experiment. Give an LLM planner read access to your catalog, Marketing Stream summaries and competitor listing snapshots, and write access to exactly one control: the target ACOS per campaign, revised daily. The numeric bidder does everything else. A/B against a fixed target ACOS by campaign in a budget-split design for six weeks. The LLM earns its place if and only if the daily target changes correlate with realized ROAS improvements; log its stated reasons so you can audit the ones that worked.

Re-read: Block 4, Block 5.

12. Generative rankers inside an auditable auction

The problem. HSTU treats recommendation as sequential transduction at 1.5T parameters; OneRec replaces the retrieve-then-rank cascade with a single generator [Zhai et al. 2024; Deng et al. 2025]. An auction needs a per-candidate score, a price that backs out of that score, eligibility rules, budget checks and an audit trail an advertiser can dispute. A generator that emits a ranked list has none of those by default.

What is known. Deep GSP and Neural Auction show a learned scorer can sit inside an auction with business constraints in its loss [Zhang et al. 2021; Liu et al. 2021]. Amazon’s business-rules estimator shows how post-processing changes what you are evaluating [Jakimov et al. 2023]. Scaling laws for ranking are real [Zhang et al. 2024; Zhu et al. 2025]. No public paper puts a generative ranker in a priced auction.

Why it is hard. Prices in GSP are defined by the runner-up’s score; a generative model has no natural runner-up. Determinism, latency and per-ad explainability are requirements the generative literature has not had to meet.

First experiment. In AuctionGym, replace the CTR model with a small sequence model that scores candidates autoregressively, and derive prices by the standard GSP inequality on its per-candidate log-probabilities. Measure calibration of the implied in the top position, revenue, and the fraction of auctions where price is undefined or exceeds the bid. Then add a Plackett-Luce randomization as in Amazon’s probabilistic mechanism paper and see whether it restores a well-defined price [Jeunen et al. 2023].

Re-read: Block 3, Block 1.

A one-week plan

The day was reading. The week is building, and each day spends the previous day’s output.

  • Monday - the loop. Stand up Marketing Stream to a queue and land hourly (keyword, placement) traffic and conversion rows. Write the Bayesian posterior repricer from the existing note with per-category delay correction (Problem 6). Do not connect it to bids yet.
  • Tuesday - the simulator. Install AuctionGym and AuctionNet. Reproduce one baseline each (a value-based bandit in AuctionGym; online LP in AuctionNet). Write the four-continuation competitor portfolio from Problem 4 as AuctionNet agents.
  • Wednesday - the estimators. Implement IPS, SNIPS and DR over logged bids with the 5% exploration band from Problem 8. Turn on the exploration band in the live account at the end of the day, on a small campaign.
  • Thursday - the bidder. Train a USCB-style Lagrangian bidder and a conservative offline bidder (V-CQL or the hybrid base-policy design) on AuctionNet. Fine-tune the offline bidder with a KL anchor to the Lagrangian bidder (Problem 5) and pick on shifted-population episodes.
  • Friday - the experiment design. Choose 200 keywords. Write the rerandomized switchback schedule for the CPC-versus-bid regression (Problem 2) and the three-design interference test (Problem 9). Pre-register the estimands.
  • Saturday - the measurement. Build the incrementality panel from Problem 10 on the 40 organic-strong keywords, and the maximization-bias diagnostic from Problem 7. Both are read-only.
  • Sunday - write it down. One page per problem you touched: what you expected from the block, what the data did, which paper turned out to be wrong for your account. That page is the next reading day.

References

Every citation on this page resolves to an entry in the annotated bibliography. The ones used most:

  1. Calvano, Calzolari, Denicolò, Pastorello. Artificial Intelligence, Algorithmic Pricing, and Collusion. AER 2020. aeaweb
  2. Banchio, Skrzypacz. Artificial Intelligence and Auction Design. EC 2022. arXiv
  3. Rawat. Algorithmic Collusion in Auctions. arXiv 2023-2025. arXiv
  4. Aggarwal, Gupta, Perlroth, Velegkas. Randomized Truthful Auctions with Learning Agents. arXiv 2024. arXiv
  5. Kolumbus, Nisan. Auctions between Regret-Minimizing Agents. WWW 2022. arXiv
  6. Chen, Nabi, Siniscalchi. Advancing Ad Auction Realism. AdKDD 2023. amazon.science
  7. Conitzer, Kroer, Sodomka, Stier-Moses. Multiplicative Pacing Equilibria in Auction Markets. OR 2022. arXiv
  8. Amazon. Response to the FTC’s Lawsuit Regarding Sponsored Ads. Aug 2026. aboutamazon
  9. FTC. Complaint, FTC v. Amazon (Sponsored Ads). Aug 2026. ftc.gov
  10. Amazon Ads. Sponsored Products across retailers. Help. amazon
  11. Balseiro, Deng, Mao, Mirrokni, Zuo. Robust Auction Design in the Auto-bidding World. NeurIPS 2021. arXiv
  12. Ostrovsky, Schwarz. Reserve Prices in Internet Advertising Auctions: A Field Experiment. JPE 2023. doi
  13. Ni, Kalfountzou, Bojinov. Reliable Switchback Experiments with Rerandomization for Auction Environments at P&G. HBS WP 2025. hbs
  14. Zhang et al. Deep GSP Auctions. WSDM 2021. arXiv
  15. Liu et al. Neural Auction. KDD 2021. arXiv
  16. Guo et al. AIGB: Generative Auto-bidding via Diffusion Modeling. KDD 2024. arXiv
  17. Balseiro, Deng, Mao, Mirrokni, Zuo. Value versus Utility Maximization. EC 2021. acm
  18. Liaw, Mehta, Perlroth. Efficiency of Non-Truthful Auctions under Auto-bidding. 2022. arXiv
  19. Mou et al. PE-MORL. arXiv 2025. arXiv
  20. Brown, Sandholm. Safe and Nested Subgame Solving. NeurIPS 2017. neurips
  21. Brown, Bakhtin, Lerer, Gong. ReBeL. NeurIPS 2020. arXiv
  22. Brown, Sandholm, Amos. Depth-Limited Solving. NeurIPS 2018. arXiv
  23. Li et al. GAS: Generative Auto-bidding with Post-training Search. 2024. arXiv
  24. Mou et al. AIGB-Pearl. arXiv 2025. arXiv
  25. Fickinger, Hu, Amos, Russell, Brown. Scalable Online Planning via RL Fine-Tuning. NeurIPS 2021. arXiv
  26. Mou et al. Sustainable Online RL for Auto-bidding. NeurIPS 2022. arXiv
  27. Korenkevych et al. Offline RL for Advertising. 2024. arXiv
  28. Jacob et al. piKL. ICML 2022. pmlr
  29. Bakhtin et al. Diplodocus. ICLR 2023. openreview
  30. Peng et al. Expert-Guided Diffusion Planner for Auto-bidding. 2025. arXiv
  31. Chapelle. Modeling Delayed Feedback. KDD 2014. doi
  32. Ktena et al. Delayed Feedback for Continuous Training. RecSys 2019. arXiv
  33. Yang et al. ES-DFM. AAAI 2021. arXiv
  34. Gu et al. DEFER. KDD 2021. acm
  35. Yang, Zhan. GDFM. NeurIPS 2022. neurips
  36. Liu et al. MISS. AAAI 2024. doi
  37. Ding et al. IF-DFM. arXiv 2025. arXiv
  38. Ma et al. Online Multi-task Learning for Google Feed Ads Auction Models. KDD 2022. acm
  39. Fan, Si, Zhang. Calibration Matters. ICLR 2023. arXiv
  40. Pan et al. Field-aware Calibration. WWW 2020. arXiv
  41. Yang et al. DESC. arXiv 2024. arXiv
  42. Zhao et al. ConfCalib. arXiv 2024. arXiv
  43. Yeom et al. Breaking Determinism. arXiv 2025. arXiv
  44. Jeunen, Murphy, Allison. Learning to Bid with AuctionGym. 2022. amazon.science
  45. Ren et al. Deep Landscape Forecasting. KDD 2019. arXiv
  46. Ou et al. Deep Landscape Forecasting in Multi-Slot RTB. KDD 2023. acm
  47. Jakimov et al. Unbiased Offline Evaluation for LTR with Business Rules. 2023. amazon.science
  48. Buchholz et al. Counterfactual Ranking Evaluation with Flexible Click Models. SIGIR 2024. amazon.science
  49. Waisman, Nair, Carrion. Online Causal Inference for Advertising in RTB Auctions. Marketing Science 2025. doi
  50. Blake, Coey. Why Marketplace Experimentation Is Harder than it Seems. EC 2014. doi
  51. Li, Zhao, Johari, Weintraub. Interference, Bias, and Variance in Two-Sided Marketplace Experimentation. WWW 2022. arXiv
  52. Liu, Mao, Kang. Budget-split Design. KDD 2021. acm
  53. Bojinov, Simchi-Levi, Zhao. Design and Analysis of Switchback Experiments. Management Science 2023. doi
  54. Bright, Delarue, Lobel. Reducing Marketplace Interference Bias via Shadow Prices. EC 2023 / Management Science 2025. doi
  55. Holtz et al. Reducing Interference Bias Using Cluster Randomization (Airbnb). Management Science 2025. informs
  56. Gordon, Zettelmeyer, Bhargava, Chapsky. A Comparison of Approaches to Advertising Measurement. Marketing Science 2019. doi
  57. Yang, Ghose. Organic and Sponsored Search: Interdependence. ISR 2010. nyu
  58. Lewis et al. Amazon Ads Multi-Touch Attribution. arXiv 2025. arXiv
  59. Johnson, Lewis, Nubbemeyer. Ghost Ads. JMR 2017. doi
  60. Gordon, Moakler, Zettelmeyer. PIE for Ad Measurement. arXiv 2023/2026. arXiv
  61. Lewis, Rao. The Unfavorable Economics of Measuring the Returns to Advertising. QJE 2015. doi
  62. He et al. Rankability-enhanced Revenue Uplift Modeling. KDD 2024. arXiv
  63. Chen, Yuan, Ye, Majumder, Richardson. AucArena. NeurIPS 2024 Open-World Agents workshop. arXiv
  64. Cai et al. RTBAgent. arXiv 2025. arXiv
  65. Jiang, Xiong, Liu. HARBOR. arXiv 2025. arXiv
  66. Meta FAIR. Cicero. Science 2022. doi
  67. Yin et al. InfoBid. arXiv 2025. arXiv
  68. Zhai et al. HSTU. ICML 2024. arXiv
  69. Deng et al. OneRec. arXiv 2025. arXiv
  70. Zhang et al. Wukong. ICML 2024. pmlr
  71. Zhu et al. RankMixer. arXiv 2025. arXiv
  72. Jeunen, Stavrogiannis, Sayedi, Allison. A Probabilistic Framework to Learn Auction Mechanisms via Gradient Descent. AAAI-W 2023. amazon.science
  73. Banchio, Mantegazza. Adaptive Algorithms and Collusion via Coupling. EC 2023. arXiv
  74. Assad, Clark, Ershov, Xu. Algorithmic Pricing and Competition: German Retail Gasoline. JPE 2024. doi
  75. Jain, Appala (Amazon). SERP Interference Network and Its Applications in Search Advertising. AdKDD 2024. amazon.science
  76. Jain, Hut, Islam, Pan (Amazon). Cross-Unit Spillovers in A/B Testing: Empirical Evidence from Ads. CODE@MIT 2023. amazon.science

← Back to the day index · Amazon Ads, Deeply