Search filters help guests navigate vast catalogs in two-sided marketplaces like Airbnb, and recommending the right filters can meaningfully lift booking conversion. Many such production filter-ranking systems, however, represent the guest through hand-engineered, pre-aggregated features generated by ETL pipelines. This makes it expensive to maintain and difficult to extend for new filter types or contextual dimensions (trip length, group size). We present SIFT (Search Intent-to-Filter Transformer), a ranking model built on transformers that learns guest preferences directly from raw behavioral sequences. SIFT replaces manual feature engineering with a unified guest representation that feeds multiple prediction tasks, including booking likelihood, filter engagement, and ordinal capacity thresholds (e.g., 2+ bedrooms) -- a general framework for filter ranking in two-sided marketplaces that accommodates both boolean and numeric-range filter types. Extending SIFT to new filters requires only adding a new head, not a new feature pipeline. To keep serving fast, this guest representation is computed offline on a daily cadence rather than at request time. Offline, SIFT improves booking and amenity-engagement PR-AUC by +51.9% and +62.8% respectively over the production baseline. In online A/B testing, SIFT increased engagement with recommended filters by +20.0%, overall filter usage among searchers by +0.72%, and usage of the newly-supported bedroom, bathroom, and bed filters by +3.9%, +10.7%, and +0.52% respectively. Demonstrating the system's extensibility, we rapidly integrated a novel hotel-intent filter using the same shared representation, driving a +3.8% lift in uncancelled hotel bookings and a +0.76% lift in overall marketplace bookings. SIFT is now fully deployed in production, serving scalable personalization to millions of guests.
Figures & tables
Figure 1. Recommending filters in “Recommended for you” and “Amenities” section in filter panel and beneath the search bar. Screenshots of the Airbnb search interface showing recommended filter chips beneath the search bar and in the Recommended for you and Amenities sections of the filter panel.
Metric
Value
Top-1 filter disagreement rate
60–65%
Avg. top-5 filter set overlap
53% (2.65 / 5)
Guests sharing ≤ 2 of top-5 filters
44%
Table 1. Guest-level filter preference drift between two disjoint, season-matched twelve-month periods.
Metric
Value
Volume-weighted average
24.2%
Unweighted average
35.5%
Maximum observed
88.3%
Table 2. Population-level FCR drift over the same two disjoint years, measured as ∣ relative FCR change ∣ per filter.
Figure 2. SIFT training data construction and multi-task architecture. Each guest’s historical actions — search, reservation-request, and booking-confirmation events, each tokenized from discrete, null-padded features — form a time-ordered sequence spanning the guest’s journey up to the current date, consumed by a transformer encoder that produces a user embedding U (i.e. eu ), computed offline in a daily scoring job and cached in an embedding store. At inference and training time, U is combined with the current search session’s query features Q , along with the candidate filter’s label context L2 (multi-hot engagement) and L3 (ordinal capacity) — to feed three types of task-specific MLP heads: T1 (booking probability P(B=1) , trained on forward-attributed booking labels L1 with position-discount example weights), T2 (per-filter engagement probability P(Fk=1) , trained on L2 ), and T3 (one ordinal head per numeric capacity filter, each predicting P(#beds≥i) -style thresholds, trained on L3 ). Diagram of the SIFT pipeline: a time-ordered sequence of guest search, reservation-request and booking events feeds a transformer encoder that outputs a user embedding, which is combined with query features and candidate filter context and passed to booking, engagement and ordinal capacity MLP heads.
Figure 3. Position-discount example weighting for the booking objective. Across a guest’s search journey, the same listing ( e ) is booked after appearing at different ranks depending on which filters were applied: 5th with no filter (Search #1, w=1.558 ), not shown at all with filter {f1} (Search #2, B=0 ), 4th with filter {f2} (Search #3, w=1.621 ), and 2nd with filters {f2, f3} (Search #4, w=1.910 ), using wi=1+1/ln(posi+1) . Bookings following a search where the listing ranked higher receive more weight, since they are more plausibly attributable to the filter’s effect on the result set rather than the guest scrolling past many alternatives. Four example searches from one guest journey, each with a different filter set, showing the rank at which the booked listing appeared and the resulting position-discount weight.
Figure 4. Frank & Hall label construction for the Beds filter ( K=9 levels, K−1=8 sub-classifiers P0–P7). A guest’s applied selection (e.g. beds 2+ ) expands into a monotone binary vector 1[Vk∗≥v] for each threshold v . Grid showing how a guest's Beds 2+ selection maps to a monotone binary label vector across eight threshold sub-classifiers P0 to P7.
Model
Booking
Eng. (Macro)
Eng.(Micro)
Bedrooms
Beds
Bathrooms
MLP Baseline
0.1388
0.1186
0.2875
—
—
—
SIFT (ours)
0.2108
0.1931
0.3896
0.2349
0.1667
0.1570
Rel. gain
+51.9%
+62.8%
+35.5%
new
new
new
Ablation
SIFT w/o journey encoder
0.1410
0.1168
0.2572
0.1529
0.0676
0.0504
Table 3. Offline PR-AUC comparison (held-out 7-day window) and journey-encoder ablation. Baseline is production MLP with hand-aggregated features. Bedrooms/Beds/Bathrooms report macro PR-AUC across threshold sub-classifiers.