Project Greenhouse: Progress Toward Fully Open and Sovereign Agentic Search
Organizations: David R. Cheriton School of Computer Science University of Waterloo
Abstract
Project Greenhouse represents our exploration of a simple thesis: We believe that it is possible to build fully open and sovereign models for agentic search with only modest computational resources. As a first milestone, we describe how to build a competitive pointwise decoder-only reranker using a simple two-step recipe comprising pre-training from scratch followed by supervised fine-tuning, starting only from commonly available datasets. Contrary to the dominant approach in the literature, we do not rely on existing open-weight backbones from third parties, and thus we are fully in control of model training, from end to end. We were able to accomplish the bulk of our experiments using no more than a handful of GPUs. This report articulates the importance and benefits of our approach, and we share artifacts that enable transparent, independent reproduction of all aspects of model training. Beyond data, code, and configurations that capture our efforts, we also release checkpoints for our family of Gaggle models, demonstrating the feasibility of our approach and providing a first step toward validating our broader thesis.
Figures & tables
| Model | Size | DL19 | DL20 | DL21 | DL22 | DL23 | Mean | |
|---|---|---|---|---|---|---|---|---|
| 1 | BM25 ( Lin et al., 2021 ) | — | 0.506 | 0.480 | 0.446 | 0.269 | 0.263 | 0.393 |
| 2 | Oracle reranker | — | 0.892 | 0.871 | 0.835 | 0.678 | 0.669 | 0.789 |
| Out-of-the-box listwise rerankers | ||||||||
| 3 | Gemma-4 ( Gemma Team, 2026 ) | 26B | 0.743 | 0.705 | 0.723 | 0.516 | 0.497 | 0.637 |
| 4 | Qwen3.8 ( Qwen Team, 2026b ) | 27B | 0.732 | 0.705 | 0.716 | 0.520 | 0.492 | 0.633 |
| 5 | Jev | ? | 0.736 | 0.715 | 0.721 | 0.505 | 0.476 | 0.631 |
| Model | Size | COVID | News | Robust | NFC | SciFact | SCID | FiQA | Mean | |
|---|---|---|---|---|---|---|---|---|---|---|
| 1 | BM25 ( Lin et al., 2021 ) | — | 0.595 | 0.395 | 0.407 | 0.322 | 0.679 | 0.149 | 0.236 | 0.398 |
| 2 | Oracle reranker | — | 0.975 | 0.827 | 0.837 | 0.544 | 0.927 | 0.443 | 0.590 | 0.735 |
| Out-of-the-box listwise rerankers | ||||||||||
| 3 | Gemma-4 ( Gemma Team, 2026 ) | 26B | 0.846 | 0.492 | 0.662 | 0.387 | 0.805 | 0.208 | 0.439 | 0.548 |
| 4 | Qwen3.8 ( Qwen Team, 2026b ) | 27B | 0.849 | 0.505 | 0.647 | 0.390 | 0.803 | 0.225 | 0.462 | 0.554 |
| 5 | GPT-6.1 Sol | ? | 0.837 | 0.526 | 0.668 | 0.398 | 0.821 | 0.237 | 0.520 | 0.572 |
| Model | Data | Params | Tokens | CORE | MMLU | MMLU-Pro | |
| 1 | gaggle-nanochat-pretrained | ClimbMix | 3.3B | 283B | 0.458 | 0.493 | 0.178 |
| 2 | gaggle-nanochat-pretrained-bidirectional ‡ | ClimbMix | 3.3B | 287B | 0.339 | 0.453 | 0.187 |
| Open pre-training data | |||||||
| 3 | nanochat-d34 | FineWeb-Edu | 2.2B | 89B | 0.338 † | 0.257 | 0.115 |
| 4 | OPEN-1B ( Donaghy et al., 2026 ) | OPEN-1B mix | 1.6B | 400B | 0.239 | 0.259 † | 0.117 † |
| 5 | Puro 2B ( Luo et al., 2026a ) | Puro mix | 2.0B | 1.4T | 0.464 | 0.552 | 0.235 |
| Backbone | Size | LR | TREC DL | BEIR | |
| 1 | gaggle-nanochat-pretrained | 3.2B | 0.641 | 0.544 | |
| 2 | nanochat-d34 | 2.1B | 0.621 | 0.521 | |
| 3 | Gemma-4-E2B | 4.6B | 0.627 | 0.537 | |
| 4 | Qwen3.5-4B-Base | 4.2B | 0.626 | 0.530 | |
| 5 | MiniCPM5-2B-Base | 2.2B | 0.627 | 0.530 | |
| 6 | Gemma-4-E2B | 4.6B | 0.640 | 0.546 |
| Training data | Instances | TREC DL | BEIR | |
|---|---|---|---|---|
| 1 | RLHN-250K | 247,534 | 0.641 | 0.544 |
| 2 | RLHN-680K | 648,762 | 0.646 | 0.545 |
| 3 | Tevatron MS MARCO | 399,600 | 0.618 | 0.470 |
| Training queries | |||||
|---|---|---|---|---|---|
| Benchmark | 50K | 100K | 150K | 200K | 250K |
| TREC DL | 0.628 | 0.616 | 0.634 | 0.637 | 0.638 |
| BEIR | 0.525 | 0.533 | 0.542 | 0.543 | 0.545 |
| Configuration | TREC DL | BEIR | |
|---|---|---|---|
| 1 | Causal, standard prompt | 0.641 | 0.544 |
| 2 | Causal, QPQ prompt | 0.647 | 0.551 |
| 3 | Causal, mean pooling | 0.642 | 0.543 |
| 4 | MNTP, standard prompt | 0.643 | 0.544 |
| 5 | MNTP, appended mask | 0.643 | 0.543 |
| 6 | MNTP, mean pooling | 0.638 | 0.538 |
| Training configuration | TREC DL | BEIR | |
|---|---|---|---|
| 1 | LCE | 0.641 | 0.544 |
| 2 | LCE + self-filtering | 0.641 | 0.548 |
| 3 | LCE + self-filtering + self-distillation | 0.640 | 0.547 |
| Objective | TREC DL | BEIR | |
|---|---|---|---|
| 1 | LCE | 0.641 | 0.544 |
| 2 | Pointwise CE | 0.628 | 0.528 |
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
| Setting | Value | |
|---|---|---|
| 1 | Fine-tuning | All retained parameters; no frozen layers or LoRA |
| 2 | Scoring parameters | Two pre-trained LM output rows: 4,352 parameters; no new head |
| 3 | Implementation | Native backbone; scaled dot-product attention (SDPA) |
| 4 | Training precision | FP32 weights, gradients, and optimizer states; TF32 matrix multiplication |
| 5 | Evaluation precision | BF16 backbone; FP32 scoring rows and scoring computation; final hidden states cast to FP32 |
| 6 | Gradient checkpointing | Enabled |