SINAI/ALIA-es-legal-administrative
Dataset Introduction The ALIA Spanish Legal and Administrative Corpus constitutes a strategic data infrastructure to support research in social sciences, legal studies, and computational linguistics, ensuring systematic access to multiple official repositories in a single consolidated dataset. With over 7 million instances and more than 5 billion tokens, it represents the most comprehensive corpus of legal and administrative texts in Spanish, combining source heterogeneity and… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-legal-administrative.
Dataset Introduction
The ALIA Spanish Legal and Administrative Corpus constitutes a strategic data infrastructure to support research in social sciences, legal studies, and computational linguistics, ensuring systematic access to multiple official repositories in a single consolidated dataset. With over 7 million instances and more than 5 billion tokens, it represents the most comprehensive corpus of legal and administrative texts in Spanish, combining source heterogeneity and advanced technical curation with datatrove.
Dataset Details
Dataset Description
The ALIA Spanish Legal and Administrative Corpus is an open-access data resource that compiles and organizes an extensive collection of official documents from the Spanish legal and administrative domain. Its purpose is to provide a homogeneous, structured, and accessible documentary base for researchers, academics, legal professionals, and public administration practitioners interested in the analysis and exploitation of normative, legislative, and administrative texts in Spanish.
This corpus has been designed with an integrative approach that encompasses state, regional, and provincial official bulletins, specialized registries, ministerial documents in key areas such as energy, environment, climate change, defense, and national security, public tenders and contracts, as well as parliamentary proceedings from the Andalusian Parliament. This diversity allows for comprehensive coverage of the documentary ecosystem that regulates institutional, economic, and social activity in Spain.
The scope of the corpus, with over 7 million instances and more than 5 billion tokens, makes it an unprecedented source for academic study of Spanish regulations, comparative legislative analysis, development of natural language processing (NLP) tools applied to legal-administrative language, and research in institutional open data. Its open and processed nature facilitates both manual exploration by legal professionals and documentation specialists, as well as advanced utilization in text mining projects, semantic modeling, information retrieval, and construction of artificial intelligence systems specialized in law and public administration.
Uses
The primary purpose of this corpus is to serve as a foundation for training and evaluating language models specialized in the Spanish legal-administrative domain, with applications in:
- Training large language models (LLMs) specialized in Spanish legal-administrative text.
- Legal-administrative information retrieval systems.
- Question-answering systems about Spanish legal-administrative.
- Research in natural language processing applied to the legal-administrative domain.
Dataset Structure
Data Instances
Each instance in the corpus has the following structure:
{
"id": "BOE_2024_123456",
"text": "DISPONICIONES GENERALES. Artículo 1. Objeto y ámbito de aplicación. 1. La presente Ley tiene por objeto establecer las bases del régimen jurídico del sector público y de la actividad administrativa, así como regular los principios que deben inspirar la actuación de las...",
"source_id": "Boletin_Oficial_Estado"
}Data Fields
- id (string): Unique document identifier.
- text (string): Content of the document.
- source_id (string): Source of origin of the document.
Data Splits
The complete dataset contains the following main sources with their statistics:
Example Usage
To load the dataset:
from datasets import load_dataset
# Load the complete dataset
data = load_dataset("SINAI/ALIA-es-legal-administrative", trust_remote_code=True)
# Load with streaming (recommended for large corpora)
data = load_dataset("SINAI/ALIA-es-legal-administrative", trust_remote_code=True, streaming=True)Example of data access:
# Access an example
example = data["train"]
print(f"ID: {example["id"]}")
print(f"Source: {example["source_id"]}")
print(f"Text: {example["text"][:200]}...")Dataset Creation
Curation Rationale
This corpus was created to address the need for specialized linguistic resources in Spanish legal and administrative language, fundamental for the development of the ALIA foundational model within the Spanish Government's Artificial Intelligence Strategy 2024. Its design responds to the demand from researchers, legal professionals, and data scientists who require systematic access to Spanish official documentation for AI model training and specialized linguistic analysis.
Source Data
The corpus integrates documentation from multiple Spanish official repositories:
Official Bulletins
- Boletín Oficial del Estado (Spanish State Official Bulletin): Spanish state legislation.
- Regional Bulletins: Boletín Oficial de la Junta de Andalucía (Andalusia), Boletín Oficial de Castilla y León (Castile and León), Boletín Oficial de la Región de Murcia (Murcia), Boletín Oficial de Cantabria (Cantabria), Boletín Oficial de Canarias (Canary Islands), Boletín Oficial de Ceuta (Ceuta).
- Provincial Bulletins: Boletín Oficial Provincial de Granada, Boletín Oficial Provincial de Jaén, Boletín Oficial Provincial de Sevilla, Boletín Oficial Provincial de Cordoba.
Specialized Registries
- Registro Mercantil Borme: Commercial Registry.
- Biblioteca Jurídica: Legal codes and compilations.
- EuroPat: Parallel corpus of European patents.
Ministerial Documentation
- Ministerio para la Transición Ecológica y el Reto Demográfico: Includes subsections on Climate Change, Energy, Environmental Quality and Assessment, Demographic Challenge, and National Parks.
- Ministerio de Defensa: Defense-related regulations and documents.
- Ministerio de Vivienda y Agenda Urbana: Housing and Urban Planning.
- Departamento de Seguridad Nacional: National Security.
Other Documents
- Licitaciones: Public contracts and tenders.
- ParlaMint-ES-AN: Parliamentary proceedings from the Parliament of Andalusia (1982-2025).
- NORMA: Economic and financial regulations.
All data come from official and publicly accessible sources.
Data Collection and Processing
Preprocessing system
The corpus is based on a previous version of nearly 20 billion tokens that was processed with an advanced cleaning methodology based on datatrove. This system automates the cleaning and preparation of large volumes of text in Spanish, eliminating duplicate and low-quality content.
Step 1: Configuration and paths loading
- Loading YAML configuration files with parameters (language threshold, filters, etc.).
- Definition of work paths and processing environment preparation.
Step 2: Language filtering
- Automatic language analysis of each document.
- Selection of Spanish texts according to confidence threshold.
Step 3: MinHash deduplication
- Advanced detection of repeated or highly similar content.
- Use of scalable comparison algorithms for large volumes.
Step 4: Quality filters and final cleaning
- Application of multiple specialized filters.
- Correction of encoding errors.
- Removal of sensitive or identifiable information.
Token counting was performed using tiktoken.
The final result is a corpus of 5,225,643,900 tokens distributed across 7,039,921 instances, optimized for language model training.
Annotations
The dataset does not contain additional annotations beyond the structural metadata extracted during processing (dates, document types, source).
Personal and Sensitive Information
The corpus has been subjected to cleaning processes to remove sensitive or identifiable information according to data protection regulations. Documents come from public official sources, although some may contain references to names in official contexts (legislators, public officials in the exercise of their duties). Users are advised to apply additional controls depending on the specific use of the corpus.
Considerations for Using the Data
Social Impact of Dataset
This corpus represents a significant advance in democratizing access to legal and administrative information in Spanish, facilitating the development of AI tools that can improve access to justice and understanding of regulations by citizens and professionals. It contributes to the national strategic objective of developing foundational AI models in Spanish with ethical and transparency standards.
Discussion of Biases
The corpus reflects the legal-administrative language used in Spanish, which may present biases inherent to:
- Formal institutional language, not representative of colloquial Spanish
- Possible overrepresentation of certain autonomous communities depending on the availability of digitized data
- The EuroPat corpus may contain technical anglicisms due to patent translations
- Reflection of regulatory frameworks and institutional perspectives specific to the Spanish context
Other Known Limitations
- The quality of original texts depends on official digitization and publication, and may contain OCR errors in historical documents
- Temporal coverage varies by source, being more complete in recent years
- The vocabulary is limited to the legal-administrative domain and may not be representative of other Spanish domains
- Some technical documents may include tables, formulas, or structured elements that are lost in the textual version
- Token and instance statistics by dataset correspond to data prior to final cleaning with datatrove
Citation
@misc{ALIA-es-legal-administrative,
title={ALIA Spanish Legal and Administrative Corpus},
author={SINAI Research Group},
year={2026},
publisher={HuggingFace},
howpublished={\url{https://huggingface.co/datasets/SINAI/ALIA-es-legal-administrative}}
}Funding
This work is funded by the Ministerio para la Transformación Digital y de la Función Pública - Funded by EU – NextGenerationEU within the framework of the project ALIA.
Acknowledgments
This dataset has been generated thanks to CEATIC (Centro de Estudios Avanzados en Tecnologías de la Información y de la Comunicación) – UJA (Universidad de Jaén) which provided the needed computational resources on its clusters.
Contact: ALIA Project - SINAI Research Group - Universidad de Jaén
More Information: SINAI Research Group | ALIA-UJA Project
