Search filters help guests navigate vast catalogs in two-sided marketplaces like Airbnb, and recommending the right filters can meaningfully lift booking conversion. Many such production filter-ranking systems, however, represent the guest through hand-engineered, pre-aggregated features generated by ETL pipelines. This makes it expensive to maintain and difficult to extend for new filter types or contextual dimensions (trip length, group size). We present SIFT (Search Intent-to-Filter Transformer), a ranking model built on transformers that learns guest preferences directly from raw behavioral sequences. SIFT replaces manual feature engineering with a unified guest representation that feeds multiple prediction tasks, including booking likelihood, filter engagement, and ordinal capacity thresholds (e.g., 2+ bedrooms) -- a general framework for filter ranking in two-sided marketplaces that accommodates both boolean and numeric-range filter types. Extending SIFT to new filters requires only adding a new head, not a new feature pipeline. To keep serving fast, this guest representation is computed offline on a daily cadence rather than at request time. Offline, SIFT improves booking and amenity-engagement PR-AUC by +51.9% and +62.8% respectively over the production baseline. In online A/B testing, SIFT increased engagement with recommended filters by +20.0%, overall filter usage among searchers by +0.72%, and usage of the newly-supported bedroom, bathroom, and bed filters by +3.9%, +10.7%, and +0.52% respectively. Demonstrating the system's extensibility, we rapidly integrated a novel hotel-intent filter using the same shared representation, driving a +3.8% lift in uncancelled hotel bookings and a +0.76% lift in overall marketplace bookings. SIFT is now fully deployed in production, serving scalable personalization to millions of guests.
Figures & tables
Figure 1. Recommending filters in “Recommended for you” and “Amenities” section in filter panel and beneath the search bar. Screenshots of the Airbnb search interface showing recommended filter chips beneath the search bar and in the Recommended for you and Amenities sections of the filter panel.
Metric
Value
Top-1 filter disagreement rate
60–65%
Avg. top-5 filter set overlap
53% (2.65 / 5)
Guests sharing ≤ 2 of top-5 filters
44%
Table 1. Guest-level filter preference drift between two disjoint, season-matched twelve-month periods.
Metric
Value
Volume-weighted average
24.2%
Unweighted average
35.5%
Maximum observed
88.3%
Table 2. Population-level FCR drift over the same two disjoint years, measured as ∣ relative FCR change ∣ per filter.
Figure 2. SIFT training data construction and multi-task architecture. Each guest’s historical actions — search, reservation-request, and booking-confirmation events, each tokenized from discrete, null-padded features — form a time-ordered sequence spanning the guest’s journey up to the current date, consumed by a transformer encoder that produces a user embedding U (i.e. eu ), computed offline in a daily scoring job and cached in an embedding store. At inference and training time, U is combined with the current search session’s query features Q , along with the candidate filter’s label context L2 (multi-hot engagement) and L3 (ordinal capacity) — to feed three types of task-specific MLP heads: T1 (booking probability P(B=1) , trained on forward-attributed booking labels L1 with position-discount example weights), T2 (per-filter engagement probability P(Fk=1) , trained on L2 ), and T3 (one ordinal head per numeric capacity filter, each predicting P(#beds≥i) -style thresholds, trained on L3 ). Diagram of the SIFT pipeline: a time-ordered sequence of guest search, reservation-request and booking events feeds a transformer encoder that outputs a user embedding, which is combined with query features and candidate filter context and passed to booking, engagement and ordinal capacity MLP heads.
Figure 3. Position-discount example weighting for the booking objective. Across a guest’s search journey, the same listing ( e ) is booked after appearing at different ranks depending on which filters were applied: 5th with no filter (Search #1, w=1.558 ), not shown at all with filter {f1} (Search #2, B=0 ), 4th with filter {f2} (Search #3, w=1.621 ), and 2nd with filters {f2, f3} (Search #4, w=1.910 ), using wi=1+1/ln(posi+1) . Bookings following a search where the listing ranked higher receive more weight, since they are more plausibly attributable to the filter’s effect on the result set rather than the guest scrolling past many alternatives. Four example searches from one guest journey, each with a different filter set, showing the rank at which the booked listing appeared and the resulting position-discount weight.
Figure 4. Frank & Hall label construction for the Beds filter ( K=9 levels, K−1=8 sub-classifiers P0–P7). A guest’s applied selection (e.g. beds 2+ ) expands into a monotone binary vector 1[Vk∗≥v] for each threshold v . Grid showing how a guest's Beds 2+ selection maps to a monotone binary label vector across eight threshold sub-classifiers P0 to P7.
Model
Booking
Eng. (Macro)
Eng.(Micro)
Bedrooms
Beds
Bathrooms
MLP Baseline
0.1388
0.1186
0.2875
—
—
—
SIFT (ours)
0.2108
0.1931
0.3896
0.2349
0.1667
0.1570
Rel. gain
+51.9%
+62.8%
+35.5%
new
new
new
Ablation
SIFT w/o journey encoder
0.1410
0.1168
0.2572
0.1529
0.0676
0.0504
Table 3. Offline PR-AUC comparison (held-out 7-day window) and journey-encoder ablation. Baseline is production MLP with hand-aggregated features. Bedrooms/Beds/Bathrooms report macro PR-AUC across threshold sub-classifiers.
Airbnb is a community based on connection and belonging -- many hosts on Airbnb are everyday people who share their worlds to provide guests with the feeling of connection and being at home; Airbnb strives to connect people and places. Among our efforts to connect guests and hosts, we provide tools to enable hosts to set competitive prices, which helps improve affordability for guests while helping hosts get more bookings. We also personalize the guest experience to show them the listings that match their needs. To help inform these efforts, we combine economic modeling and causal inference techniques to understand how guests book stays based on the prices hosts set, among other factors, and how that preference varies across different guests and listings. Such understanding helps us identify opportunities for Airbnb to support the marketplace and better connect guests and hosts. For example, understanding how much guests respond to different prices helps optimize the tools that we provide to hosts, in order to enable hosts to choose and set competitive prices that further balance demand and supply. As another example, understanding heterogeneity in guest preferences helps us personalize the guest experience and better match them with the listings that meet their needs, based on how much they respond to different prices and other factors.
Personalization in two-sided marketplaces relies heavily on user-level features, yet for platforms with infrequent, high-consideration purchases, a large fraction of users lack sufficient history for effective recommendation, spanning both paid and organic channels. At Airbnb, a substantial share of search requests comes from logged-out or first-time users, with this challenge especially pronounced on paid-channel landing pages, leaving traditional user-level features unavailable for a large fraction of traffic. Privacy regulations and increasing restrictions on third-party cookies further limit identifier-based tracking for non-essential use cases. This paper introduces Proximity Features, a privacy-compliant feature system that groups users by geographic proximity using geo-IP data and an adaptive clustering algorithm, producing aggregated user-level signals for groups of approximately 1,000 nearby users without requiring a persistent individual identifier at inference time. Privacy is preserved by design: the pipeline operates on consented, aggregated data only within consent-gated privacy controls. The system is deployed in production at Airbnb, serving multiple surfaces including marketing landing pages and destination recommendation, with engagement emails integration under way. Online A/B experiments demonstrate statistically significant lifts in bookings, with the largest gains observed among users with absent or stale history.
Sequence modeling has become increasingly popular in recommendation and ranking algorithms, owing to its capacity to model users' historical behaviors and infer user intentions. Despite its theoretical simplicity, the practical deployment of a sequence model in production is non-trivial due to complexity of the sequence and sparse labels. For example, in Airbnb, guest sequences are often long, exploratory and complex, and we focus on booking labels, which are sparse. As such, we are often required to make various design decisions regarding data and modeling to strike a balance between effectiveness and scalability. This work delved into these production challenges and deployed JourneyFormer, a sequence modeling solution for search ranking at Airbnb. We detail crucial design considerations, covering aspects such as guest event selection, ID embeddings, model architecture, and label attribution. Additionally, we describe several tailored strategies to accelerate model training and inference. JourneyFormer has been successfully deployed within Airbnb's production, where its effectiveness and impact have been evidenced not only by improved offline ranking metrics but also by significant gains in key business metrics through online A/B testing across 2 production surfaces.