cs.CLDate pending

Fine PT-PT Web: A High-Quality 41 Billion Tokens Data Collection of the European Portuguese Web

Authors: Gonçalo VinagreRui Pedro GuerraPedro GomesMiguel Moura RamosDuarte Miguel AlvesAfonso SimplícioDiogo TavaresDavid Semedo+2 more

Abstract

Curating Web corpora for regional language variants like European Portuguese (PT-PT) is heavily bottlenecked by dialectal overlap (mainly with PT-BR) and data processing scale. This paper presents an efficient pipeline to curate a production-ready PT-PT corpus from the Portuguese Web, spanning 411 TB of raw data from Arquivo.pt. We introduce a novel post-scraping block that removes boilerplate and line duplicates prior to filtering. This early-stage intervention increases final document yield by 19.04% by rescuing valid text that standard heuristic filters prematurely discard. Integrated with rigorous language identification, weighted fuzzy deduplication, and neural quality classification, our pipeline offers a scalable framework and a clean, representative corpus optimized for LLM pre-training.

Explore similar work

CardsList
  1. NorBERTo: A ModernBERT Model Trained for Portuguese with 331 Billion Tokens Corpus

    Apr 30, 2026Enzo S. N. Silva, Pablo B. Costa, Raphael C. Vlasman +8PortugueseModernbert

  2. Beyond Multilingual Averages: MTEB-PT, a Benchmark for Portuguese Sentence Encoders

    Jul 5, 2026Lucas Hideki Takeuchi Okamura, Alexandre Alcoforado, Anna Helena Reali CostaDomain-Invariant Text EmbeddingsPortuguese