← Day index · Block 7 of 10 · ← Previous · Next →
Time budget
~90 minutes. Read Lewis & Rao and Blake & Coey first - together they say the signal is tiny and the naive experiment is biased. Then Li-Zhao-Johari-Weintraub for which way to be wrong instead, and the off-policy section with the IPS derivation in hand. Skim incrementality-versus-attribution and the Amazon experimentation papers once; keep the decision table and simulator notes as reference.
The problem in two sentences
An A/B test assumes that what you do to the treatment group does not change what happens to the control group. In an auction that assumption is false by construction: raise the treatment group’s bids and the control group’s prices rise; give the treatment group a better conversion model and the control group loses auctions it used to win; change the ranking rule for half of queries and budget-constrained advertisers shift spend to the other half. The unit you randomized is not the unit that experiences the treatment, because the clearing price couples every participant to every other. This block covers the three responses: redesign the experiment so the coupling is contained, evaluate offline from logs and accept that the counterfactual was never observed, or simulate and accept that the simulator is a model. Every claim in the other nine blocks was measured by one of these three methods and inherits its weaknesses.
Section takeaway. Interference is not a nuisance in auction experiments; it is the structure of the object being measured.
The noise floor, before the bias
Before interference there is variance, and the founding number is Lewis and Rao’s (QJE 2015) [1]. Across 25 large digital-advertising field experiments, with more than 10 million person-weeks of data, the median confidence interval on return on investment was wider than 100 percentage points. Ad effects are small relative to the variance of purchasing, and a campaign that is highly profitable is statistically indistinguishable from one that loses money at sample sizes most advertisers can afford. Everything downstream - the pull toward observational methods, the appeal of attribution, the need for variance-reducing surrogates - is a response to that floor.
Observational methods do not escape it; they trade variance for bias. Gordon, Zettelmeyer, Bhargava and Chapsky (Marketing Science 2019) compared 15 Facebook RCTs to the observational estimates a careful analyst would have produced from the same data and found lifts overstated by factors of two to ten [2]. Gordon, Moakler and Zettelmeyer’s “Close Enough?” (Marketing Science 2023) repeats the exercise with double machine learning across many more campaigns: where RCTs measure lifts of 29%, 18% and 5%, DML reports 83%, 58% and 24% [3]. Jeunen and Ustimenko’s “Learning Metrics that Maximise Power” (KDD 2024) is the constructive response: learn a surrogate metric whose treatment effect is maximally detectable, and cut required sample size to about 12% of the raw metric’s [4].
Section takeaway. The signal is small and the observational shortcuts are biased by 2x to 10x; power is the first constraint on any auction experiment, not the last.
Marketplace interference: the eBay factor of two
Blake and Coey’s EC 2014 paper is the founding document of interference, and it earned its place with a number [5]. Randomizing eBay marketing emails across users, the naive estimate said the campaign generated roughly 0.74% additional revenue. Reanalyzing at the level of the auction - separating auctions where treated and control users competed from those where they did not - the effect was not distinguishable from zero. The naive estimate was too large by about a factor of two: treated buyers bid more, control buyers competing in the same auctions lost more and paid more, and the control group was made to look worse by the treatment it was supposed to be isolated from. The supply-demand model shows the bias depends on elasticities on each side, which is the polite way of saying it is not a constant you can correct for afterward.
Two papers turned that warning into guidance. Johari, Li, Liskovich and Weintraub (Management Science 2022) characterize the bias of customer-side and listing-side randomization in a two-sided market [6]; Li, Zhao, Johari and Weintraub (WWW 2022) ask which side to randomize and what fraction to treat, and find that competition biases the estimate under either choice, that the less-biased side depends on relative supply and demand, and that reducing bias typically raises variance [7]. Wager and Xu’s “Experimenting in Equilibrium” (Management Science 2021) handles the case where the treatment is a continuous platform parameter - a reserve price, a quality-score weight - by applying small local perturbations and estimating the equilibrium gradient rather than a discrete effect [8]. Farias, Li, Peng and Zheng’s “Markovian Interference in Experiments” (NeurIPS 2022) model the shared state that makes budgets a source of interference - one arm’s spending depletes a pool the other arm draws from - and propose a differences-in-Q’s estimator that corrects for it [9]. Jeunen’s SIGIR Forum note (2023) identifies a version of the same problem that is almost universal and almost never acknowledged: when both arms of an experiment are machine-learned models retrained on pooled logs, the arms interfere through the training data [10].
Then the designs. Liu, Mao and Kang’s budget-split design (KDD 2021, LinkedIn) splits each advertiser’s budget between arms rather than splitting users, creating two counterfactual marketplaces with their own competition; the estimator is unbiased for finite or infinite buyer budgets and reports a power gain of more than 15x over other unbiased designs [11]. Holtz and coauthors’ Airbnb pricing meta-experiment (Management Science 2025) shows cluster randomization - whole clusters of competing listings - reduces interference bias at a measurable cost in power [12]. Bright, Delarue and Lobel (EC 2023; Management Science 2025) attack from the allocation side: in generalized matching markets, compare shadow prices - the duals of the platform’s allocation problem - between arms rather than raw value, which they argue is the correct first-order approximation and less biased under their assumptions [13]. The transfer to ad pricing is an inference, though a clearing price is itself a shadow price on a slot.
Amazon’s own experimentation group has published the version of this that is closest to Sponsored Products. Jain, Hut, Islam and Pan document cross-unit spillovers in ads A/B tests empirically (CODE@MIT 2023) [14]; Hut, Mason, Islam and Esquerra quantify what stratification buys in cluster-randomized experiments (CODE@MIT 2023) [15]; and Jain and Appala’s SERP interference network (AdKDD 2024) builds a bipartite graph from queries to the ads that compete on them, clusters it into randomization units that contain the interference, and uses those units to evaluate a paid-search bidding algorithm [16]. That last paper is the single most relevant item in the block for anyone building an Amazon bidder: it is the platform describing how it measures exactly the kind of system this reading day is about.
Section takeaway. The naive marketplace experiment can be off by a factor of two; the fixes - choosing the side, local perturbation, clustering on the interference graph, splitting budgets, comparing shadow prices - each trade bias for variance and none removes the problem.
Time as the experimental unit
When you cannot assign treatment independently to units, assign it to periods. Switchbacks alternate the whole market between arms over time blocks; Bojinov, Simchi-Levi and Zhao (Management Science 2023) model carryover explicitly, formulate the optimal design as a minimax problem, and provide exact randomization-based inference and a procedure for identifying the carryover order [17]. The cost is power: one observation per block, and carryover forces you to discard or model the boundaries.
The auction-specific case is a draft and should be read as one. Ni, Kalfountzou and Bojinov’s HBS working paper 26-012 (2025) describes rerandomized switchbacks at Procter & Gamble: generate many candidate schedules, accept those with covariate imbalance below a threshold, analyze over the accepted set; P&G reports 70 experiments across five auction hypotheses in eight markets since January 2025 [18]. Its business-lift figures are explicitly not peer-reviewed. Zeng and coauthors’ sequentially rerandomized switchbacks (Stanford/Airbnb, 2026) formalize the idea and are simulation-only [19]. What makes the P&G paper worth reading is that it is a bidder-side experiment - an advertiser testing bidding hypotheses against a live auction it does not control, which is anyone building an Amazon bidder.
Section takeaway. Switchbacks let you randomize a market you cannot split; rerandomization makes them credible; carryover makes them expensive.
Incrementality is not attribution
The second family asks not “is policy A better than B” but “what did this ad cause”. Johnson, Lewis and Nubbemeyer’s ghost ads (JMR 2017) solved the economics for display: log the moment the ad server would have served the ad to a control user, serve something else, and compare exposed treatment users to would-have-been-exposed controls; 432 experiments ran on the method [20]. On a search page the slot is scarce and the substitute is a competitor’s ad, so transfer is not trivial. Geo experiments are the platform-scale alternative: Vaver and Koehler (Google, 2011) and Kerman, Wang and Vaver (2017) give the time-based regression framework [21][22]; Meta’s GeoLift is a synthetic-control implementation that exists as a repository and documentation with no standalone paper [23]; Meloni, Hut and Islam (Amazon, CODE@MIT 2024) evaluate synthetic difference-in-differences for geo-randomized experiments [24]. Meta’s Conversion Lift and Google’s incrementality guidance are documentation, not papers [25][26].
PIE - Predicted Incrementality by Experimentation (Gordon, Moakler, Zettelmeyer; 2023, revised 2026) - quantifies how badly attribution does [27]. Across 2,226 Meta experiments, a model predicting RCT incrementality from campaign features reaches out-of-sample of 0.88 on incremental conversions per dollar; seven-day last-click reaches 0.19. Amazon’s Multi-Touch Attribution (arXiv 2508.08209, August 2025) sits between: hundreds of thousands of RCTs calibrate campaign-level causal effects, and a production ML system uses attribution models as features to score individual touchpoints [28]. It is more defensible than last-click and honest about the trade - observational ML precise but biased, RCTs unbiased but noisy - but a fractional touchpoint credit is still not the incremental value of winning one auction. Two more Amazon papers bear on the window: Pauwels, Schnaidt and Caddeo (EMAC 2022) estimate the causal impact of display ads on advertiser performance [29], and Qin’s “Lengthen Your Attribution Window” (2023) finds upper-funnel ads realize only 30-50% of their effect within two weeks [30] - which means a 14-day window systematically underprices discovery keywords. The older attribution literature - Shapley values with journey order (2018) [31], Singal and coauthors’ axiomatic framework (WWW 2019) [32] - distributes credit under a model and estimates no treatment effect.
Section takeaway. Attribution allocates observed credit; incrementality estimates causal effect; last-click explains a fifth of the variance in the latter, and a 14-day window misses half of upper-funnel impact.
Off-policy evaluation: the counterfactual you never logged
The third family answers “what would policy B have done” from logs generated by policy A. The founding text is Bottou and coauthors’ JMLR 2013 paper on counterfactual reasoning for Bing ad placement [33]. Inverse propensity scoring reweights each logged outcome by the ratio of target to logging probability; doubly robust estimators (Dudík, Langford and Li, ICML 2011; Jiang and Li for the sequential case, ICML 2016) add an outcome model and stay consistent if either is right [34][35]. Swaminathan and Joachims’ counterfactual risk minimization (ICML 2015) turns evaluation into learning, and their self-normalized estimator (NIPS 2015) fixes propensity overfitting, where the learner exploits large weights [36][37]. Saito and Joachims’ MIPS (ICML 2022) handles large action spaces via embeddings - the natural framing when the action is an (ASIN, bid) pair [38]; their RecSys 2021 tutorial is the entry point [39]. Joachims, Swaminathan and Schnabel (WSDM 2017) brought the framework to ranking, where the action is a position and the propensity is examination probability [40].
Here is the auction-specific problem, and it is fatal to naive IPS. With logged bid , logging policy , target and reward ,
unbiased only if wherever . A production bidder is deterministic: . Any target that would bid something else puts mass where is zero, the weight is undefined, and the logs contain no observation of what that bid would have won or paid - because the outcome depends on competitors’ bids you never saw. These are the two difficulties Jeunen, Murphy and Allison name in AuctionGym, whose KDD 2023 paper shows most “learning to bid” methods are value-based estimators of the outcome model rather than genuine off-policy learners [41][42].
The patch is a bid landscape: a model of the market-price distribution , so win probability at any bid is and expected cost is the truncated mean. Wu, Yeh and Chen (KDD 2018) and Ren and coauthors (KDD 2019) fit this with survival analysis, the right tool because losses are censored - you learn only that the price exceeded your bid [43][44]; Ou and coauthors extend it to correlated multi-slot pages (KDD 2023) [45], and Ou’s TKDD 2024 survey covers the bid-optimization field [46]. With a landscape you can replace the missing with a modeled propensity, which is what Yeom and coauthors’ “Breaking Determinism” (2025; ACM record 2026) does for self-normalized IPS in deterministic auctions [47] - principled and emerging. Waisman, Nair and Carrion (Marketing Science 2025) close the loop from the economics side: auction structure identifies ad effects from optimal bids, and Thompson sampling learns bids while controlling the cost of the exploration identification requires [48].
Amazon’s science org has a coherent line on the ranking half. Yu’s unbiased counterfactual estimation of ranking metrics (2021) [49]; Block, Kidambi, Hill, Joachims and Dhillon (SIGIR 2022) ranking autocompletions by downstream purchase utility [50]; Xiao and coauthors (SIGIR-AP 2023) extending to multi-query sessions [51]; Jakimov, Buchholz, Stein and Joachims (RecSys CONSEQUENCES 2023) correcting for business rules that post-process the displayed ranking, via a Birkhoff-von-Neumann decomposition [52]; and Buchholz and coauthors’ Interpol (RecSys CONSEQUENCES 2022), which interpolates between click-model assumptions so the estimator is not hostage to the position-based model being exactly right [53]. Together they are what counterfactual evaluation of a sponsored-results ranker looks like when the evaluators own the logs.
Section takeaway. A deterministic bidder has zero propensity on every bid it did not submit, so naive IPS cannot evaluate a new policy; landscape models restore support at the price of being models.
Simulators: the last line before customer money
AuctionGym (Amazon Science 2022; KDD 2023) is a simulation environment for bandit bidding that varies environmental conditions, introduces policy-based and doubly robust bidding, and is open source so offline validation does not need proprietary logs [41][42]. Chen, Nabi and Siniscalchi’s “Advancing Ad Auction Realism” (2023) says what a simulator must get right: query-dependent slot values, unobserved and changing competitors, partial aggregated feedback, incompletely specified payment rules, advertisers as adversarial bandits [54]. AuctionNet (NeurIPS 2024 Datasets and Benchmarks) is Alibaba’s large-scale alternative - 48 auto-bidding agents, more than 500 million auction records, a customizable GSP module, and an explicit acknowledgment of residual bias against real data [55]; BAT (2025) is a further benchmark with a thin public record [56]. Block 4 uses these as the arena; here the point is epistemic. A simulator ranks policies under its own assumptions and cannot tell you what the assumptions cost.
Section takeaway. Simulators screen; they do not certify.
Choosing a method
| Design | When to use it | What it identifies | Main failure mode |
|---|---|---|---|
| Randomize users | Treatment does not change prices faced by others (rare in auctions) | Individual-level effect | Interference through the clearing price; eBay bias ~2x [5]; pooled-log retraining [10] |
| Budget-split | Budgeted advertisers compete for shared inventory | Marketplace effect in two counterfactual markets; >15x power vs other unbiased designs [11] | Needs control of budgets; markets must be separable |
| Cluster / interference graph | Competition is local to a query niche or geography | Cluster-level effect with contained spillover [12][16] | Fewer units, less power; clusters leak |
| Local perturbation | Treatment is a continuous parameter (reserve, weight) | Equilibrium gradient [8] | Only local; needs equilibrium to re-form |
| Switchback | Whole-market treatment; bounded carryover | Time-averaged market effect [17][18] | Carryover; low power; needs rerandomization |
| Shadow-price comparison | Platform’s allocation problem is modelable | First-order value-function difference [13] | Transfer from matching is an inference |
| Ghost ads / lift / geo | Question is “what did this ad cause” | Incremental effect [20][21][27] | Scarce slots; competitor substitute; wide CIs [1] |
| Off-policy from logs | New policy overlaps the logging policy | Counterfactual policy value [33][36] | Zero support for unsubmitted bids; censored prices [42][47] |
| Simulator | Nothing else is safe yet | Relative ranking under assumptions [54][55] | The assumptions; competitor adaptation |
The stack that follows is layered. Define the estimand - individual effect, market effect, causal lift, policy value - because the designs identify different things. Screen in a simulator with adversarial bidders and partial feedback. Log propensities, bids, market prices and exposure with enough exploration that off-policy evaluation has support, and use landscape models where it does not. Test survivors with an interference-aware design - budget-split or interference-graph clusters where competition dominates, switchbacks where time dominates, local perturbations for continuous parameters. Use power-maximizing surrogates because the signal is small. Calibrate scalable attribution against lift, lengthen the window for upper-funnel terms, and price bids from lift. That is the sequence the deep note’s job posting describes - “simulations of auction dynamics to test a policy before it touches customer money,” then “experiments on live auctions, with real spend as the readout” - and it is the sequence every result in Block 4 and Block 2 was, or was not, measured by.
Reading list for this block
- Lewis, Rao - The Unfavorable Economics of Measuring the Returns to Advertising - QJE 2015 - 15 min - the noise floor; read the confidence-interval figure.
- Blake, Coey - Why Marketplace Experimentation Is Harder than it Seems - EC 2014 - 15 min - the factor-of-two result and the supply-demand model.
- Jain, Appala - SERP Interference Network - AdKDD 2024 - 15 min - Amazon measuring a paid-search bidding algorithm; the most on-topic paper of the day.
- Li, Zhao, Johari, Weintraub - Interference, Bias, and Variance in Two-Sided Marketplace Experimentation - WWW 2022 - 10 min - which side to randomize.
- Jeunen, Murphy, Allison - Off-Policy Learning-to-Bid with AuctionGym - KDD 2023 - 15 min - the two named difficulties.
- Gordon, Moakler, Zettelmeyer - PIE - 2023/2026 - 10 min - 0.88 versus 0.19.
- Bottou et al. - Counterfactual Reasoning and Learning Systems - JMLR 2013 - 10 min - read sections 1-3 for the framing; the rest is reference.
- Reference only: everything else in the references.
Questions to carry forward
- If an advertiser-side bidder can only run switchbacks against an auction it does not control, what carryover order should it assume when the platform’s dynamic bidding is adjusting to its changes?
- Landscape models restore support, but they are trained on censored logs from the current bidder. How much of the “counterfactual” is the landscape model’s prior?
- PIE shows last-click explains a fifth of incrementality variance at Meta, and Qin shows a 14-day window misses half of upper-funnel effect. What experiment establishes the equivalent numbers for Sponsored Products when the platform owns both the RCTs and the attribution?
- The SERP interference network clusters queries by shared competing ads. How stable are those clusters when the bidding algorithm under test changes which ads compete?
- Every simulator in this block was built by someone with logs. What is the smallest simulator an outsider could build that is still adversarial enough to be informative?
References
- Lewis, Rao. The Unfavorable Economics of Measuring the Returns to Advertising. QJE 2015. doi.org/10.1093/qje/qjv023
- Gordon, Zettelmeyer, Bhargava, Chapsky. A Comparison of Approaches to Advertising Measurement: Evidence from Big Field Experiments at Facebook. Marketing Science 2019. doi.org/10.1287/mksc.2018.1135
- Gordon, Moakler, Zettelmeyer. Close Enough? A Large-Scale Exploration of Non-Experimental Approaches to Advertising Measurement. Marketing Science 2023. arXiv:2201.07055
- Jeunen, Ustimenko. Learning Metrics that Maximise Power for Accelerated A/B-Tests. KDD 2024. arXiv:2402.03915
- Blake, Coey. Why Marketplace Experimentation Is Harder than it Seems: The Role of Test-Control Interference. EC 2014. doi.org/10.1145/2600057.2602837
- Johari, Li, Liskovich, Weintraub. Experimental Design in Two-Sided Platforms: An Analysis of Bias. Management Science 2022. doi.org/10.1287/mnsc.2021.4247 · arXiv:2002.05670
- Li, Zhao, Johari, Weintraub. Interference, Bias, and Variance in Two-Sided Marketplace Experimentation: Guidance for Platforms. WWW 2022. arXiv:2104.12222
- Wager, Xu. Experimenting in Equilibrium. Management Science 2021. arXiv:1903.02124
- Farias, Li, Peng, Zheng. Markovian Interference in Experiments. NeurIPS 2022. arXiv:2206.02371
- Jeunen. A Common Misassumption in Online Experiments with Machine Learning Models. SIGIR Forum 2023. arXiv:2304.10900
- Liu, Mao, Kang (LinkedIn). Trustworthy and Powerful Online Marketplace Experimentation with Budget-split Design. KDD 2021. doi.org/10.1145/3447548.3467193
- Holtz et al. Reducing Interference Bias in Online Marketplace Experiments Using Cluster Randomization: Evidence from a Pricing Meta-experiment on Airbnb. Management Science 2025. pubsonline.informs.org
- Bright, Delarue, Lobel. Reducing Marketplace Interference Bias via Shadow Prices. EC 2023; Management Science 2025. doi.org/10.1287/mnsc.2022.01881 · arXiv:2205.02274
- Jain, Hut, Islam, Pan (Amazon). Cross-Unit Spillovers in A/B Testing: Empirical Evidence from Ads. CODE@MIT 2023. amazon.science
- Hut, Mason, Islam, Esquerra (Amazon). Value of Stratification in Cluster-Randomized Experiments. CODE@MIT 2023. amazon.science
- Jain, Appala (Amazon). SERP Interference Network and Its Applications in Search Advertising. AdKDD 2024. amazon.science
- Bojinov, Simchi-Levi, Zhao. Design and Analysis of Switchback Experiments. Management Science 2023. arXiv:2009.00148
- Ni, Kalfountzou, Bojinov. Reliable Switchback Experiments with Rerandomization for Auction Environments at Procter & Gamble. HBS Working Paper 26-012, 2025 (draft). hbs.edu
- Zeng et al. (Stanford/Airbnb). Sequentially-Rerandomized Switchback Experiments. 2026 (simulation only). arXiv:2604.02489
- Johnson, Lewis, Nubbemeyer. Ghost Ads: Improving the Economics of Measuring Online Ad Effectiveness. JMR 2017. doi.org/10.1509/jmr.15.0297
- Vaver, Koehler (Google). Measuring Ad Effectiveness Using Geo Experiments. 2011. research.google
- Kerman, Wang, Vaver (Google). Estimating Ad Effectiveness Using Geo Experiments in a Time-Based Regression Framework. 2017. research.google
- Meta. GeoLift (repository and documentation; no standalone paper). github.com · methodology
- Meloni, Hut, Islam (Amazon). Performance of Synthetic Diff-in-Diff Models for Geo-Randomized Experiments. CODE@MIT 2024. amazon.science
- Meta for Business. Conversion Lift (documentation). facebook.com
- Ohlinger, Nedyalkov (Google). Incrementality Testing. 2023 (documentation). business.google.com
- Gordon, Moakler, Zettelmeyer. Predicted Incrementality by Experimentation (PIE) for Ad Measurement. 2023, revised 2026. arXiv:2304.06828
- Lewis, Zettelmeyer, Gordon, Garib, Hermle, Perry, Romero, Schnaidt (Amazon Ads). Amazon Ads Multi-Touch Attribution. 2025 (v1). arXiv:2508.08209
- Pauwels, Schnaidt, Caddeo (Amazon). Causal Impact of Digital Display Ads on Advertiser Performance. EMAC 2022. amazon.science
- Qin (Amazon). Lengthen Your Attribution Window: Which Digital Ads Have Most Long-Term Impact. Applied Marketing Analytics 2023. amazon.science
- Zhao, Mahboobi, Bagheri. Shapley Value Methods for Attribution Modeling in Online Advertising. 2018. arXiv:1804.05327
- Singal, Besbes, Desir, Goyal, Iyengar. An Axiomatic Framework for Attribution in Online Advertising. WWW 2019. doi.org/10.1145/3308558.3313731
- Bottou et al. Counterfactual Reasoning and Learning Systems: The Example of Computational Advertising. JMLR 2013. arXiv:1209.2355
- Dudík, Langford, Li. Doubly Robust Policy Evaluation and Learning. ICML 2011. arXiv:1103.4601
- Jiang, Li. Doubly Robust Off-policy Value Evaluation for Reinforcement Learning. ICML 2016. proceedings.mlr.press
- Swaminathan, Joachims. Counterfactual Risk Minimization. ICML 2015. proceedings.mlr.press
- Swaminathan, Joachims. The Self-Normalized Estimator for Counterfactual Learning. NIPS 2015. papers.nips.cc
- Saito, Joachims. Off-Policy Evaluation for Large Action Spaces via Embeddings (MIPS). ICML 2022. arXiv:2202.06317
- Saito, Joachims. Counterfactual Learning and Evaluation for Recommender Systems (tutorial). RecSys 2021. doi.org/10.1145/3460231.3473320
- Joachims, Swaminathan, Schnabel. Unbiased Learning-to-Rank with Biased Feedback. WSDM 2017. doi.org/10.1145/3018661.3018699
- Jeunen, Murphy, Allison (Amazon). Learning to Bid with AuctionGym. Amazon Science 2022 (AdKDD workshop). amazon.science
- Jeunen, Murphy, Allison (Amazon). Off-Policy Learning-to-Bid with AuctionGym. KDD 2023. doi.org/10.1145/3580305.3599877
- Wu, Yeh, Chen. Deep Censored Learning of the Winning Price in the Real Time Bidding. KDD 2018. doi.org/10.1145/3219819.3220066
- Ren, Qin, Zheng, Yang, Zhang, Yu. Deep Landscape Forecasting for Real-time Bidding Advertising. KDD 2019. doi.org/10.1145/3292500.3330870
- Ou et al. Deep Landscape Forecasting in Multi-Slot Real-Time Bidding. KDD 2023. doi.org/10.1145/3580305.3599799
- Ou et al. A Survey on Bid Optimization in Real-Time Bidding Display Advertising. TKDD 2024. doi.org/10.1145/3628603
- Yeom, Shin, Min, Yoon, Yu, Kang. Breaking Determinism: Stochastic Modeling for Reliable Off-Policy Evaluation in Ad Auctions. 2025 (ACM record 2026). arXiv:2512.03354
- Waisman, Nair, Carrion. Online Causal Inference for Advertising in Real-Time Bidding Auctions. Marketing Science 2025. doi.org/10.1287/mksc.2022.0406
- Yu (Amazon). Unbiased Counterfactual Estimation of Ranking Metrics. 2021. amazon.science
- Block, Kidambi, Hill, Joachims, Dhillon (Amazon). Counterfactual Learning to Rank for Utility-Maximizing Query Autocompletion. SIGIR 2022. amazon.science
- Xiao, Kveton, Katariya, Gangwani, Rangi (Amazon). Towards Sequential Counterfactual Learning to Rank. SIGIR-AP 2023. amazon.science
- Jakimov, Buchholz, Stein, Joachims (Amazon). Unbiased Offline Evaluation for Learning to Rank with Business Rules. CONSEQUENCES workshop, RecSys 2023. amazon.science
- Buchholz, London, Di Benedetto, Lichtenberg, Stein, Joachims (Amazon). Counterfactual Ranking Evaluation with Flexible Click Models (Interpol). CONSEQUENCES workshop, RecSys 2022. arXiv:2210.09512
- Chen, Nabi, Siniscalchi (Amazon). Advancing Ad Auction Realism: Practical Insights & Modeling Implications. 2023 (AdKDD workshop). amazon.science
- Su et al. (Alibaba). AuctionNet: A Novel Benchmark for Decision-Making in Large-Scale Games. NeurIPS 2024 Datasets and Benchmarks. arXiv:2412.10798
- Khirianova et al. BAT: Benchmark for Auto-bidding Task. 2025. arXiv:2505.08485
← Day index · ← Block 6 · Block 8 → · Parent note: Amazon Ads, Deeply