Routing Without Embeddings: Fast And Interpretable Routing With Regular Expressions
Organizations: Rice University
Abstract
Large Language Model (LLM) routers commonly rely on neural query embeddings, with larger encoders expected to better capture query intent and difficulty. Yet scaling Qwen2.5 encoders from 0.5B to 72B parameters brings little improvement in routing accuracy (Figure 1b), suggesting that small encoders may already capture the query properties needed for routing. We therefore investigate which properties matter and whether they can be extracted directly from text without a neural encoder. We introduce REGEXROUTE, a pipeline that uses sparse autoencoders (SAEs) to discover interpretable regular-expression (regex) features. Using unlabeled text, an LLM turns descriptions of grouped SAE latents into regex extractors and refines them to match latent activation patterns. These extractors supply numerical features to a lightweight routing head, eliminating neural encoding at inference (Figure 1a). Across four benchmarks, one fixed set of 128 features achieves 76.43% average routing accuracy, comparable to 76.41% for the strongest neural text encoder baseline, with much smaller latency and strong robustness. These findings establish explicit, interpretable text features as a practical basis for designing and understanding LLM routers.
Figures & tables
| Representation | EmbedLLM | NineBy30k | R2Bench | CARROT | Avg. |
|---|---|---|---|---|---|
| Best single model | 60.73 | 67.34 | 81.10 | 84.82 | 73.50 |
| Sentence embeddings | |||||
| MiniLM-L6 ( Wang et al., 2020 ) | 64.87 | 67.47 | 81.56 | 86.18 | 75.02 |
| MPNet-base ( Song et al., 2020 ) | 64.73 | 67.43 | 81.39 | 86.28 | 74.96 |
| BGE-base ( Xiao et al., 2023 ) | 64.47 | 67.77 | 81.53 | 86.51 | 75.07 |
| GTE-base ( Li et al., 2023 ) | 64.37 | 67.63 | 81.77 | 86.60 | 75.09 |
| Procedure | SAE guidance | Refinement | Held alignment | Test (%) | Held alignment | Test (%) |
|---|---|---|---|---|---|---|
| LLM-only | — | — | 0.2343 | 74.68 | 0.2852 | 75.67 |
| Initial regexes | — | 0.3770 | 75.68 | 0.3865 | 76.16 | |
| Ours | 0.4470 | 76.06 | 0.4709 | 76.43 | ||
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
| benchmark | #models | train | validation | test |
|---|---|---|---|---|
| EmbedLLM | 112 | 29,673 | 3,000 | 3,000 |
| NineBy30k | 9 | 23,027 | 3,000 | 5,000 |
| R2Bench | 10 | 18,580 | 3,097 | 9,291 |
| CARROT | 13 | 30,968 | 6,636 | 6,637 |
| category | representation | checkpoint | dim | pooling | context length |
| Sentence-embedding models | MiniLM-L6 | all-MiniLM-L6-v2 | 384 | mean | 512 |
| MPNet-base | all-mpnet-base-v2 | 768 | mean | 512 | |
| BGE-base | BAAI/bge-base-en-v1.5 | 768 | [CLS] | 512 | |
| GTE-base | thenlper/gte-base | 768 | mean | 512 | |
| E5-base | intfloat/e5-base-v2 | 768 | mean | 512 | |
| Encoder-only LMs | ModernBERT-base | answerdotai/ModernBERT-base | 768 | [CLS] | 8,192 |
| representation | clean | typo | synonym | grammar | paraphrase | drop |
|---|---|---|---|---|---|---|
| 128 regexes (ours) | 76.43 | 76.43 | 75.74 | 75.41 | 74.92 | |
| TF-IDF bags | 75.70 | 75.27 | 74.43 | 75.06 | 74.27 | |
| Qwen2.5-0.5B | 76.33 | 75.80 | 75.28 | 75.71 | 75.18 | |
| ModernBERT-large | 75.91 | 74.84 | 74.76 | 75.12 | 74.81 | |
| ModernBERT-base | 75.58 | 75.07 | 74.70 | 75.06 | 74.44 | |
| GTE-base | 75.09 | 74.58 | 74.68 | 74.91 | 74.72 |
| representation | clean | 5% | 10% | 20% | 30% |
|---|---|---|---|---|---|
| 128 regexes (ours) | 76.43 | 76.13 | 75.97 | 75.41 | 75.00 |
| TF-IDF bags | 75.70 | 75.00 | 74.96 | 74.41 | 74.36 |
| Qwen2.5-0.5B | 76.33 | 75.30 | 74.88 | 74.33 | 74.05 |
| E5-base | 74.92 | 74.45 | 74.44 | 74.06 | 73.85 |
| MiniLM-L6 | 75.02 | 74.63 | 74.43 | 74.16 | 73.86 |
| Router | p50 (ms) | p99 (ms) | VRAM (MB) |
|---|---|---|---|
| GPU | |||
| MiniLM-L6 | 5.0 | 10.4 | 204 |
| E5-base | 8.2 | 13.6 | 832 |
| Qwen2.5-0.5B | 19.5 | 25.8 | 2,617 |
| ModernBERT-large | 21.6 | 27.7 | 1,609 |
| CPU | |||
| router | device | p50 | p90 | p99 |
|---|---|---|---|---|
| 32 regexes (ours) | CPU | 1.8 | 3.8 | 11.5 |
| 128 regexes (ours) | CPU | 3.0 | 8.6 | 33.6 |
| MiniLM-L6 | GPU | 4.9 | 5.4 | 7.8 |
| E5-base | GPU | 8.2 | 8.9 | 10.8 |
| Qwen2.5-0.5B | GPU | 19.4 | 20.1 | 22.6 |
| MiniLM-L6 | CPU | 13.2 | 30.7 | 36.3 |
| Source LM | EmbedLLM | NineBy30k | R2Bench | CARROT | Avg. | |
|---|---|---|---|---|---|---|
| Gemma-2-2B | 16,384 | 67.37 | 69.01 | 82.16 | 87.18 | 76.43 |
| Qwen3-1.7B | 32,768 | 67.53 | 68.91 | 82.15 | 87.18 | 76.44 |
| Llama-3.1-8B | 32,768 | 67.07 | 69.27 | 81.95 | 87.19 | 76.37 |
| Spectral | Spherical -means | Agglomerative | Facility location | |
|---|---|---|---|---|
| 1 | 73.48 | 73.64 | 73.66 | 73.65 |
| 2 | 73.80 | 73.82 | 73.80 | 73.96 |
| 4 | 75.08 | 74.41 | 73.73 | 73.89 |
| 8 | 75.24 | 74.54 | 74.95 | 74.95 |
| 16 | 75.62 | 74.80 | 75.35 | 75.66 |
| 32 | 76.06 | 75.23 | 76.12 | 75.69 |
| Latents per family | ||||
|---|---|---|---|---|
| Grouping method | Accuracy (%) | Minimum | Median | Maximum |
| Spectral clustering | 76.43 | 31 | 115.0 | 439 |
| Spherical -means | 76.10 | 22 | 95.5 | 531 |
| Agglomerative clustering | 76.03 | 1 | 15.0 | 4,904 |
| Facility-location selection | 76.03 | 1 | 56.5 | 1,831 |
| Representation | Tree | RF | ET | -NN | -means |
|---|---|---|---|---|---|
| 128 regexes (ours) | 75.17 | 76.37 | 76.43 | 75.01 | 75.09 |
| TF-IDF bags | 74.41 | 75.57 | 75.70 | 74.73 | 74.68 |
| Llama-3.1-8B | 74.37 | 76.18 | 76.41 | 75.77 | 75.61 |
| Gemma-2-2B | 74.31 | 76.11 | 76.39 | 75.85 | 75.58 |
| Qwen2.5-0.5B | 74.54 | 76.10 | 76.33 | 76.04 | 75.73 |
| ModernBERT-large | 74.08 | 75.68 | 75.91 | 75.40 | 75.41 |
| Index | Feature name Aggregation; flags |
|---|---|
| Script (12 features) | |
| f1 | cat_accent_elision_union count ; — |
| (?:[\u00e0\u00e8\u00e9\u00ed\u00f2\u00f3\u00fa\u00ef\u00e7\u00c0\u00c8\u00c9\u00cd\u00d2\u00d3\u00da\u00cf\u00c7\u00b7]|\b[LlDdSsNnMmTt]\u2019’) | |
| f2 | it_accent_ratio ratio ; — |
| [\u00e0\u00e8\u00e9\u00ec\u00f2\u00f9\u00c0\u00c8\u00c9\u00cc\u00d2\u00d9] | |
| f3 | cjk_code_split count ; — |
| Representation | EmbedLLM | NineBy30k | R2Bench | CARROT | Mean |
|---|---|---|---|---|---|
| Random | 3.91 | 6.72 | 30.08 | 29.48 | 17.55 |
| E5-base | 24.71 | 35.77 | 72.46 | 66.29 | 49.81 |
| ModernBERT-base | 38.78 | 58.57 | 81.67 | 79.08 | 64.53 |
| Qwen2.5-0.5B | 44.82 | 60.12 | 84.84 | 82.59 | 68.09 |
| RegexRoute | 40.02 | 64.11 | 85.12 | 83.72 | 68.24 |
| Representation | EmbedLLM | NineBy30k | R2Bench | CARROT | Mean |
|---|---|---|---|---|---|
| Random | 3.91 | 6.72 | 30.08 | 29.48 | 17.55 |
| E5-base | 44.75 | 58.33 | 83.94 | 82.56 | 67.40 |
| ModernBERT-base | 41.88 | 64.73 | 83.99 | 82.53 | 68.28 |
| Qwen2.5-0.5B | 49.46 | 69.51 | 85.60 | 84.53 | 72.27 |
| RegexRoute | 37.80 | 63.05 | 84.33 | 82.83 | 67.00 |
| Representation | EmbedLLM | NineBy30k | R2Bench | CARROT | Mean |
|---|---|---|---|---|---|
| Random | 53.34 | 53.04 | 62.72 | 61.83 | 57.74 |
| E5-base | 60.56 | 58.75 | 71.10 | 70.10 | 65.13 |
| ModernBERT-base | 59.56 | 58.86 | 70.91 | 70.19 | 64.88 |
| Qwen2.5-0.5B | 61.39 | 59.70 | 71.35 | 70.02 | 65.62 |
| RegexRoute | 59.01 | 59.16 | 70.37 | 69.95 | 64.62 |
| Ridge | MLP | |||||
|---|---|---|---|---|---|---|
| Encoder | Mean | Median | Wtd. | Mean | Median | Wtd. |
| Qwen2.5-0.5B | 0.507 | 0.603 | 0.612 | 0.734 | 0.823 | 0.820 |
| Qwen2.5-1.5B | 0.542 | 0.642 | 0.652 | 0.751 | 0.843 | 0.840 |
| Qwen2.5-3B | 0.569 | 0.673 | 0.683 | 0.754 | 0.848 | 0.842 |
| Qwen2.5-7B | 0.594 | 0.701 | 0.713 | 0.754 | 0.854 | 0.841 |
| Qwen2.5-14B | 0.615 | 0.722 | 0.734 | 0.768 | 0.857 | 0.852 |
| Encoder | EmbedLLM | NineBy30k | R2Bench | CARROT |
|---|---|---|---|---|
| Ridge (linear) | ||||
| Qwen2.5-0.5B | 0.641 | 0.674 | 0.543 | 0.589 |
| Qwen2.5-1.5B | 0.667 | 0.721 | 0.589 | 0.632 |
| Qwen2.5-3B | 0.694 | 0.764 | 0.610 | 0.664 |
| Qwen2.5-7B | 0.740 | 0.775 | 0.644 | 0.692 |
| Qwen2.5-14B | 0.759 | 0.787 | 0.670 | 0.721 |