How does UberEats rank restaurant ads?
In a recent blog post, Uber describes a new architecture it has applied to ads ranking in UberEats. The new architecture comprises two innovations:
- sequence-based user modeling.
- a Hetero-MMoE layer.
Sequence Modeling:
Instead of relying on a fixed set of historical interactions like clicks, this approach creates a long, trailing set of chronological user-restaurant interactions on the platform (clicks, add-to-cart, checkout) as well as other features such as restaurant UUID, cuisine type, and daypart and weekpart data.
This sequence is processed by a transformer-based model to produce an ad-conditioned representation capturing the user's behavioral context. The transformer model uses a candidate ad as the query, where attention determines which representations from the sequence are most relevant to a potential interaction. This is an interesting part of the architecture because it conditions the behavioral representation on the candidate ad, making the encoder target-aware.
Because multi-head attention involves each token attending to every other token, scaling with O(N^2), and a user's behavioral sequence can be long, the architecture applies the multi-head latent attention (MLA) mechanism introduced by DeepSeek. MLA compresses event-level tokens into L latent representations, where L << N and attention complexity reduces to O(N*L).
Hetero-MMoE
Rather than using a traditional MMoE layer that uses identical model architectures for the mix of experts, this blends different types of expert models that capture different orders of feature interactions, with task-specific gating networks weighting their contributions when predicting outcomes. The ranking model evaluates against two tasks, click-through rate (CTR) and click-to-order (CTO), and the model architectures assigned to the experts are: traditional MLP, Deep Cross Network (DCN) for capturing low- to mid-order interactions (eg., order-one, order-two), and Compressed Interaction Network (CIN) for capturing a mix of low- and high-order interactions, adapted from the Lian et al. (2018) xDeepFM paper.
xDeepFM builds on the Guo et al. (2017) DeepFM paper, which pairs factorization machines (FM), which are only capable of processing order-two pairwise interactions, with a DNN to learn higher-order interactions. xDeepFM introduces the CIN, which can capture higher-order interactions, and retains the DNN component.
This sequence modeling + Hetero-MMoE architecture improved the performance of ad ranking for UberEats, resulting in a 0.93% gain in AUC and 0.7% gain in log loss for pCTR and a 0.66% gain in AUC and 2.0% gain in log loss for pCTO. Uber's blog post states that it expects to apply this architecture to other use cases, such as search ads or home feed recommendations.
Blog post linked below.