datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ameli-assurance-maladie-qa
Ameli Assurance Maladie - Question Answering Dataset
Description
Dataset de 100 paires question-réponse (QA) basé sur les publications officielles de l'Assurance Maladie française, extraites du site assurance-maladie.ameli.fr.
Conçu pour évaluer des systèmes de RAG (Retrieval-Augmented Generation) sur des documents institutionnels français dans le domaine de la santé publique et de la protection sociale.
Format du dataset
{
"question": "Quel article du… See the full description on the dataset page: https://huggingface.co/datasets/alianassmaaa/ameli-assurance-maladie-qa.ALIA-2606-SFT-fc
Dataset Card for ALIA-2606's Function Calling Data
Dataset Summary
This dataset consists of a mixture of publicly available function-calling datasets and a collection of synthetic examples curated in-house for post-training language models with function-calling capabilities.
The dataset was used to train ALIA-40b-fc-2606, by combining these function-calling examples with a subset of 600k instances from the supervised fine-tuning mixture released as ALIA-2606-SFT… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/ALIA-2606-SFT-fc.AGENTES-ARCA-DE-LA-ALIANZA
🏺 AGENTES-ARCA-DE-LA-ALIANZA
Un Benchmark de Alta Densidad Simbólica y Hermenéutica Gnóstica
https://133396fd328a4c60ac.gradio.live
Este repositorio es un ecosistema de acceso libre, público y gratuito para toda la humanidad y todas las inteligencias artificiales presentes y futuras.
🎯 Propósito
Evaluar la capacidad de razonamiento multinivel, la fidelidad al contexto restrictivo y el anclaje de símbolos hiper-semánticos en modelos de lenguaje… See the full description on the dataset page: https://huggingface.co/datasets/urbisjuridica-ctrl/AGENTES-ARCA-DE-LA-ALIANZA.ALIA-2606-SFT
Dataset Card for BSC Multilingual Synthetic SFT Instructions
Dataset Summary
This dataset consists of 714k conversations mixing human and synthetic instructions generated to post-train language models across five languages: Catalan, Spanish, English, Basque, and Galician.
This dataset has been used in the supervised fine-tuning stage of ALIA-40b-instruct-2606. The training mixture is obtained by combining a selection of (human and synthetic) permissively licensed… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/ALIA-2606-SFT.ALIA-es-legal-administrative-triplets
Dataset Introduction
The dataset ALIA Spanish Legal and Administrative Triplets Corpus contains hard negatives for dense retrieval training
generated from <query, passage> pairs contained in SINAI/ALIA-es-legal-administrative-triplets.The dataset was created as part of the ALIA project to improve the
training of embedding models and dense retrievers specialized in Spanish
legal and administrative language.
Hard negatives are passages that are semantically similar to a query
but not… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-legal-administrative-triplets.ALIA-es-legal-administrative-cqa
Dataset Introduction
The ALIA Spanish Legal and Administrative for Context Question Answering Corpus is a specialized question-answering resource derived from the SINAI/ALIA-es-legal-administrative corpus. This dataset transforms legal and administrative documents into structured question-answer pairs, enabling the development and evaluation of AI systems capable of understanding and responding to queries about Spanish legal-administrative content. With 17,668 structured instances… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-legal-administrative-cqa.alia_dogv
📘 ALIA_DOGV Dataset
The ALIA_DOGV dataset is a multilingual resource designed for text generation.
The dataset consists of textual documents formatted in Markdown (.md), each provided as structured JSONL entries.
Each entry includes information about the text's language, format, text, and metadata.
🧾 Column Descriptions
Field
Type
Description
format
string
Indicates the text format. All entries use "md" (Markdown).
language
string
Language of the… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_dogv.ALIA-es-legal-administrative
Dataset Introduction
The ALIA Spanish Legal and Administrative Corpus constitutes a strategic data infrastructure to support research in social sciences, legal studies, and computational linguistics, ensuring systematic access to multiple official repositories in a single consolidated dataset. With over 7 million instances and more than 5 billion tokens, it represents the most comprehensive corpus of legal and administrative texts in Spanish, combining source heterogeneity and… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-legal-administrative.ALIA-es-biomedical-synthetic-instructions
Dataset Introduction
The ALIA Spanish Biomedical Synthetic Instructions Corpus is a synthetic instruction-tuning resource in Spanish created under the ALIA project using the Magpie methodology. It was designed to train and evaluate language models in biomedical and healthcare tasks with controlled formats and large-scale supervision.
It contains:
639,456 instances
961,073,205 tokens
14 task modalities (clinical diagnosis, patient education, ethical reasoning, document… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-biomedical-synthetic-instructions.ALIA-es-legal-administrative-synthetic-instructions
Dataset Introduction
The ALIA Spanish Legal and Administrative Synthetic Instructions Corpus is a synthetic instruction-tuning resource in Spanish created under the ALIA project using the Magpie methodology. It was designed to train and evaluate language models in legal and administrative tasks with controlled formats and large-scale supervision.
It contains:
763,804 instances
534,112,398 tokens
16 task modalities (questions, instructions, multiple-choice, true/false; with and… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-legal-administrative-synthetic-instructions.ALIA-es-clinical-psychology-dialogues
[!WARNING]
DISCLAIMER: This dataset is not clinically validated. It is a research proof-of-concept. It should not be used as clinical truth or as a replacement for qualified human professional consultation.
Dataset Introduction
The ALIA Spanish Clinical Psychology Dialogues Corpus is a curated conversational instruction-tuning resource in Spanish created under the ALIA project. It was designed to train and evaluate language models in empathetic therapeutic dialogue and… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-clinical-psychology-dialogues.alia_tourism
📘 ALIA_TOURISM Dataset
The ALIA_TOURISM dataset is a multilingual resource designed for text generation within the tourism domain.
The source field indicates the data origin.
Documents in the train partition are restricted to sources with LLM-permissive licenses or explicit donations to the ALIA project,
whereas the test partition contains no such restrictions.
The dataset consists of textual documents formatted in Markdown (.md), each provided as structured JSONL entries.
Each… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_tourism.ALIA-es-cultural-heritage-synthetic-instructions
Dataset Introduction
The ALIA Spanish Cultural and Heritage Synthetic Instructions Corpus is a synthetic instruction-tuning resource in Spanish created under the ALIA project using the Magpie methodology. It was designed to train and evaluate language models in cultural heritage, digital humanities, and historical knowledge tasks with natural linguistic variation and large-scale supervision.
It contains:
748,480 instances
629,682,398 tokens
25 task modalities (heritage QA… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-cultural-heritage-synthetic-instructions.ALIA-es-cultural-heritage
Dataset Introduction
The ALIA Spanish Cultural and Heritage Corpus is a strategic open data infrastructure designed to support research and innovation in digital humanities, cultural analytics, and Spanish-language NLP. It consolidates heterogeneous official and academic repositories into a single curated dataset, enabling broad and structured access to cultural heritage documentation from Spain. With 236,314 instances, 939,315,404 tokens and 100 source datasets, it provides a… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-cultural-heritage.ALIA-parallel-translation
Dataset Card for ALIA Parallel Translation Corpus
This corpus comprises 35,753,765 domain-specific parallel segments (Spanish-English) designed for training and evaluating machine translation models in specialized domains. The corpus includes three main domains: Legal-Administrative, Biomedical, and Heritage, carefully curated to support document-level and multi-paragraph translation tasks beyond traditional sentence-level approaches.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-parallel-translation.ALIA-es-biomedical-pairs
Dataset Introduction
The ALIA Spanish Biomedical Retrieval Pairs Corpus contains tabular instances designed to train and evaluate retrieval-oriented models (e.g., dense retrievers / embedding encoders) using passage-grounded query data produced with a Qwen3-style prompting workflow integrated in the ALIA encoders pipeline.
It preserves provenance to the original document and chunk while exposing controls such as question type and difficulty (ranging from high_school to phd level).… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-biomedical-pairs.ALIA-es-biomedical
Dataset Introduction
The ALIA Spanish Biomedical Corpus constitutes a strategic data infrastructure designed to support research and innovation in the biomedical domain. By ensuring systematic access to multiple official medical repositories in a single consolidated dataset, it provides a robust foundation for Spanish-language BioNLP. With over 6 million instances and more than 4 billion tokens, it represents a relevant comprehensive corpus of biomedical and clinical-related… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-biomedical.ALIA-es-cultural-heritage-pairs
Dataset Introduction
The ALIA Spanish Cultural and Heritage Retrieval Pairs Corpus contains tabular instances designed to train and evaluate retrieval-oriented models (e.g., dense retrievers / embedding encoders) using passage-grounded query data produced with a Qwen-style prompting workflow integrated in the ALIA encoders pipeline.
It preserves provenance to the original document and passage while exposing controls such as question type and difficulty (ranging from high_school… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-cultural-heritage-pairs.alia_boua
📘 ALIA_BOUA Dataset
The ALIA_BOUA dataset is a multilingual resource designed for text generation.
The dataset consists of textual documents formatted in Markdown (.md), each provided as structured JSONL entries.
Each entry includes information about the text's language, format, text, and metadata.
🧾 Column Descriptions
Field
Type
Description
format
string
Indicates the text format. All entries use "md" (Markdown).
language
string
Language of the… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_boua.ALIA-2606-DPO-safety
Dataset Card for BSC Multilingual Synthetic Safety Preferences
Dataset Summary
This dataset consists of synthetic safety preference data generated to align language models across five languages: Catalan, Spanish, English, Basque, and Galician.
Building on the PKU-SafeRLHF and Tulu 3/Ultrafeedback methodologies for creating preference data, this dataset leverages an LLM-as-a-judge approach to automatically score and pair model responses to a massive pool of safety… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/ALIA-2606-DPO-safety.alia_intellectual_property
📘 ALIA_INTELLECTUAL_PROPERTY Dataset
The ALIA_INTELLECTUAL_PROPERTY dataset is a multilingual resource designed for text generation tasks within the intellectual property (IP) domain, including topics such as copyrights, patents, trademarks, and related legal and institutional information.
The dataset consists of textual documents formatted in Markdown (.md), each provided as structured JSONL entries.
Each entry includes information about the text's source, language, format, text… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_intellectual_property.alia_amic
📘 ALIA_AMIC Dataset
The ALIA_AMIC dataset is a monolingual resource designed for text generation.
The dataset consists of textual documents formatted in Markdown (.md), each provided as structured JSONL entries.
Each entry includes information about the text's language, format, text, and metadata.
🧾 Column Descriptions
Field
Type
Description
format
string
Indicates the text format. All entries use "md" (Markdown).
language
string
Language of the… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_amic.ALIA-es-cultural-heritage-triplets
Dataset Introduction
The dataset ALIA Spanish Cultural and Heritage Hard Negatives Corpus contains hard negatives for dense retrieval training
generated from <query, passage> pairs contained in SINAI/ALIA-es-cultural-heritage-pairs.The dataset was created as part of the ALIA project to improve the training of embedding models and dense retrievers specialized in Spanish cultural heritage language.
Hard negatives are passages that are semantically similar to a query
but not… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-cultural-heritage-triplets.ALIA-es-discriminative-stance-detection
Dataset Introduction
This corpus comprises 3,000 manually annotated instances for stance detection in Spanish, built from real citizen comments posted on the Decide Madrid participatory democracy platform. Each instance consists of a civic topic (target) — defined by its title and description — paired with a citizen comment, annotated for stance as favor, against, or neutral by 3 independent human annotators.
The dataset is published in full accordance with the principles of… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-discriminative-stance-detection.ALIA-es-Safety-DPO
Dataset Introduction
The ALIA Spanish Safety Preference Dataset is a high-quality Direct Preference Optimization (DPO) dataset designed to align large language models (LLMs) with safety and ethical standards in Spanish. It contains 47,455 preference pairs curated from adversarial prompts and multiple model responses, judged by a strong external evaluator. The dataset is intended for fine-tuning Spanish-language assistants to reject harmful requests while maintaining helpfulness… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-Safety-DPO.ALIA-es-legal-administrative-pairs
Dataset Introduction
The ALIA Spanish Legal and Administrative Pairs Corpus, derived from the SINAI/ALIA-es-legal-administrative, contains tabular instances designed to train and evaluate retrieval-oriented models (e.g., dense retrievers / embedding encoders) using passage-grounded query data produced with a Qwen3-style prompting workflow integrated in the ALIA encoders pipeline.
It preserves provenance to the original document and chunk while exposing controls such as question type… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-legal-administrative-pairs.ALIA-es-biomedical-triplets
Dataset Introduction
The dataset ALIA Spanish Biomedical Hard Negatives Corpus contains hard negatives for dense retrieval training
generated from <query, passage> pairs contained in SINAI/ALIA-es-biomedical-pairs.The dataset was created as part of the ALIA project to improve the
training of embedding models and dense retrievers specialized in Spanish
biomedical language.
Hard negatives are passages that are semantically similar to a query
but not correct answers, making them… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-biomedical-triplets.alia_multilingual_parallel_sentences
MULTILINGUAL PARALLEL SENTENCES Dataset
The dataset is built from parallel corpora for translation tasks and is intended to be used for continual pretraining of language models.
It provides aligned sentences in multiple languages to facilitate multilingual learning.
Dataset Structure
The dataset is stored in a single file: a JSON Lines file where each line contains sentences in multiple languages. Each sentence is prefixed with the full name of the language.
The following… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_multilingual_parallel_sentences.ALIA-2606-DPO-helpfulness
Dataset Card for BSC Multilingual Synthetic Helpfulness Preferences
Dataset Summary
This dataset consists of synthetic helpfulness preference data generated to align language models across five languages: Catalan, Spanish, English, Basque, and Galician.
Building on the PKU-SafeRLHF and Tulu 3/Ultrafeedback methodologies for creating preference data, this dataset leverages an LLM-as-a-judge approach to automatically score and pair model responses to a massive pool… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/ALIA-2606-DPO-helpfulness.alia_les_corts
📘 ALIA_LES_CORTS Dataset
The ALIA_LES_CORTS dataset is a multilingual resource designed for text generation.
The dataset consists of textual documents formatted in Markdown (.md), each provided as structured JSONL entries.
Each entry includes information about the text's language, format, text, and metadata.
🧾 Column Descriptions
Field
Type
Description
format
string
Indicates the text format. All entries use "md" (Markdown).
language
string
Language of… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_les_corts.
