QCATS: Query Context-Aware Transformer Slicing for Efficient Predictive Query Processing
Authors: Yueying Li, Zhongle Xie, Ke Chen, Lidan Shou
Organizations: The State Key Laboratory of Blockchain and Data Security, Zhejiang University · Hangzhou High-Tech Zone (Binjiang) Institute of Blockchain and Data Security
In-database predictive query processing increasingly applies Transformer-based models within relational pipelines. However, existing in-database inference typically exposes only tuple-level model inputs to the inference runtime, leaving relational predicates and metadata statistics invisible to neural execution planning. In this paper, we propose QCATS, a query context-aware transformer slicing framework that enables efficient sparse inference inside database systems. QCATS executes at query granularity: instead of routing individual tokens or tuples during inference, it uses query predicates and metadata statistics to pre-select context-aligned FFN slices before model execution. The framework comprises offline expert construction and lightweight query-level routing that dynamically selects experts during execution. QCATS further introduces system optimizations, including asynchronous CPU-GPU pipelines and routing-aware batching. Experiments on four predictive-query workloads with BERT-base and Qwen-0.6B show that QCATS achieves up to 4.42x latency reduction while preserving prediction accuracy comparable to dense baselines.
Figures & tables
Fig. 1 : FFN Activation Varies Across Query Contexts.
Fig. 2 : Average FFN neuron activations for text tuples selected by eight queries ( Q1 – Q8 ).
Figure 3
Fig. 5 : Illustration of Expert Construction.
Figure 5
Fig. 8 : QCATS Processing with Pipeline.
Dataset
Total Size
Context Columns
Amazon
1,908,006
[main_category, sub_category, price]
Kickstarter
202,835
[main_category, sub_category, goal]
OSHA
42,645
[occupation, accident_type, age]
CVE
90,847
[cwe, privileges_required, product_count]
TABLE I : Summary of dataset statistics and task definitions.
Dataset
BERT-base Backbone
Qwen-base Backbone
BERT-base
MoEfication
MoEBERT
QCATS w/o pipeline
QCATS
Qwen-base
QCATS w/o pipeline
QCATS
Amazon
279.05
299.32 (0.93 × )
236.16 (1.18 × )
164.02 (1.70 × )
98.69 (2.83 × )
1499.22
517.49 (2.90 × )
463.30 (3.24 × )
Kickstarter
352.94
376.14 (0.94 × )
295.33 (1.20 × )
172.89 (2.04 × )
129.13 (2.73 × )
1739.19
874.38 (1.99 × )
780.19 (2.23 × )
OSHA
785.40
847.95 (0.93 × )
636.40 (1.23 × )
307.46 (2.55 × )
230.08 (3.41 × )
5241.14
2231.12 (2.35 × )
1981.82 (2.64 × )
CVE
798.29
862.21 (0.93 × )
653.82 (1.22 × )
231.51 (3.45 × )
180.51 (4.42 × )
4521.26
1564.17 (2.89 × )
1418.08 (3.19 × )
TABLE II : End-to-end execution time (s) comparison.
Fig. 9 : Peak Memory usage comparison.
Model
Amazon
Kickstarter
OSHA
CVE
(Acc / F1)
(Acc / F1)
(Acc / F1)
(Acc / F1)
BERT-base
89.7 / 88.6
81.4 / 81.0
93.0 / 92.9
77.1 / 77.0
MoEfication
89.5 / 88.3
80.8 / 80.4
92.7 / 92.4
76.7 / 76.4
MoEBERT
89.9 / 88.7
81.6 / 81.3
93.0 / 92.9
78.1 / 78.0
QCATS (BERT)
89.6 / 88.5
81.8 / 81.4
93.0 / 92.9
77.9 / 77.8
Qwen-base
89.2 / 87.5
80.1 / 80.0
92.2 / 91.9
74.1 / 73.7
TABLE III : Comparison of Accuracy and Weighted F1-Score. Boldface indicates the best performance.
Fig. 10 : Latency Breakdown.
Fig. 11 : Ablation study.
Fig. 12 : Hyperparameter experiments.
Fig. 13 : Predictive queries in the experiment.
Fig. 14 : Performance evaluation of QCATS across diverse SQL workloads.
Model
Source Domain (Amazon)
Target Domain (Flipkart)
(Acc / F1)
(Acc / F1)
BERT-base
89.7 / 88.6
88.7 / 86.5
QCATS (BERT)
89.6 / 88.5
88.7 / 86.6
Qwen-base
89.2 / 87.5
88.6 / 85.6
QCATS (Qwen)
89.5 / 87.5
88.8 / 85.6
TABLE IV : Transferability to Similar Domains. Boldface indicates the best performance.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
ID
Filter Predicates P
Q1
main_category = ‘technology’ AND sub_category = ‘Gadgets’
Q2
sub_category = ‘Video Games’
Q3
main_category = ‘music’ AND 1000 < goal < 2000
Q4
main_category = ‘publishing’ AND country = ‘GB’ AND 20 < duration < 30
Q5
main_category = ‘film & video’ AND sub_category = ‘Drama’
Q6
sub_category = ‘Drinks’
Appendix
TABLE V : Filter Predicates for the Queries.
Fig. 15 : Cross-Query Top-768 Neuron Overlap Matrix. This heatmap visualizes the pairwise overlap percentage of the top-768 most active neurons across 16 distinct query contexts ( Q1 - Q16 ). Darker blue indicates higher overlap (similarity), while lighter colors indicate distinct activation patterns. Diagonal elements are 100% by definition.
Department of Computing and Mathematics University of Sao Paulo Ribeirão Preto, São Paulo, Brazil · Department of Information Systems University of Münster Münster, Germany