gplsi/alia_boua
📘 ALIA_BOUA Dataset The ALIA_BOUA dataset is a multilingual resource designed for text generation. The dataset consists of textual documents formatted in Markdown (.md), each provided as structured JSONL entries. Each entry includes information about the text's language, format, text, and metadata. 🧾 Column Descriptions Field Type Description format string Indicates the text format. All entries use "md" (Markdown). language string Language of the… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_boua.
📘 ALIA_BOUA Dataset
The ALIA_BOUA dataset is a multilingual resource designed for text generation.
The dataset consists of textual documents formatted in Markdown (`.md`), each provided as structured JSONL entries. Each entry includes information about the text's language, format, text, and metadata.
🧾 Column Descriptions
🌍 Dataset Composition
- Languages: Spanish (es) and Valencian (va)
- Format: JSON Lines (
.jsonl)
⚠️ Notes
- The dataset is automatically curated and automatically anonymized using the **AnonymizationPipeline**.
- The
metadatafield includes a key named `spans`, which contains information about the anonymized text segments, including their `start`, `end`, `label`, and `rank`. - Content may include Markdown formatting for structure (e.g., headers, lists, emphasis).
💰 Funding
This work is funded by the Ministerio para la Transformación Digital y de la Función Pública, co-financed by the EU – NextGenerationEU, within the framework of the project Desarrollo de Modelos ALIA.
<!-- ## 🙏 Acknowledgments
We would like to express our gratitude to all individuals and institutions that have contributed to the development of this work.
Special thanks to:
- [Data providers]
- [Technological support providers]
We also acknowledge the financial, technical, and scientific support of the Ministerio para la Transformación Digital y de la Función Pública - Funded by EU – NextGenerationEU within the framework of the project Desarrollo de Modelos ALIA. -->
📚 Reference
Please cite this dataset using the following BibTeX format:
@misc{alia2025boua,
author = {Espinosa Zaragoza, Sergio and Sep{\'u}lveda Torres, Robiert and Mu{\~n}oz Guillena, Rafael and Consuegra-Ayala, Juan Pablo},
title = {ALIA_BOUA Dataset},
year = {2025},
institution = {Language and Information Systems Group (GPLSI) and Centro de Inteligencia Digital (CENID), University of Alicante (UA)},
howpublished = {\url{[https://huggingface.co/datasets/gplsi/alia_boua}}](https://huggingface.co/datasets/gplsi/alia_boua}})
}
⚠️ Disclaimer
Be aware that the data may contain biases or other unintended distortions. When third parties deploy systems or provide services based on this data, or use the data themselves, they bear the responsibility for mitigating any associated risks and ensuring compliance with applicable regulations, including those governing the use of Artificial Intelligence. The University of Alicante, as the owner and creator of the dataset, shall not be held liable for any outcomes resulting from third-party use.
📜 License
This work is licensed under a Creative Commons Attribution 4.0 International (CC BY 4.0) licence.
