We introduce a new generation of small reasoning models for RAG, search, and source summarization. Pleias-RAG-350m and Pleias-RAG-1B are mid-trained on a large synthetic dataset emulating the retrieval of a wide variety of multilingual open sources from the Common Corpus. They provide native support for citation and grounding with literal quotes and reintegrate multiple features associated with RAG workflows, such as query routing, query reformulation, and source reranking. Pleias-RAG-350m and Pleias-RAG-1B outperform SLMs below 4 billion parameters on standardized RAG benchmarks (HotPotQA, 2wiki) and are competitive with popular larger models, including Qwen-2.5-7B, Llama-3.1-8B, and Gemma-3-4B. They are the only SLMs to date maintaining consistent RAG performance across leading European languages and ensuring systematic reference grounding for statements. Due to their size and ease of deployment on constrained infrastructure and higher factuality by design, the models unlock a range of new use cases for generative AI.
Figures & tables
Figure 1: Scores on HotPotQA evaluation versus model size. Both Pleias models are Pareto-optimal among SLMs for RAG.
Figure 2: Comparison of a reconstruction of Anthropic citation mode vs. our method of straight generated citation, allowing for more direct intervention, such as citation shortening
Figure 3: Main scenarios incorporated into the reasoning model: trivial question (with a shortened reasoning mode), standard question, and refusal due to lack of source backing
Figure 4: Standardized RAG workflow integrated within the model also featuring further options available for implementation (query re-submission, standard refusal)
Figure 5: Token reassignment strategy for the RAG specialized models.
Figure 6: Simplified workflow of our retrieval strategy.
Figure 7: Training run of Pleias-RAG-350M.
Figure 8: Results of standard evaluation on English benchmarks.
Figure 9: Estimate of language performance loss in four European languages (French, Spanish, Italian, and German). Pleias are the only models with a negligible impact.