datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ameli-assurance-maladie-qa
Ameli Assurance Maladie - Question Answering Dataset
Description
Dataset de 100 paires question-réponse (QA) basé sur les publications officielles de l'Assurance Maladie française, extraites du site assurance-maladie.ameli.fr.
Conçu pour évaluer des systèmes de RAG (Retrieval-Augmented Generation) sur des documents institutionnels français dans le domaine de la santé publique et de la protection sociale.
Format du dataset
{
"question": "Quel article du… See the full description on the dataset page: https://huggingface.co/datasets/alianassmaaa/ameli-assurance-maladie-qa.ALIA-es-legal-administrative-triplets
Dataset Introduction
The dataset ALIA Spanish Legal and Administrative Triplets Corpus contains hard negatives for dense retrieval training
generated from <query, passage> pairs contained in SINAI/ALIA-es-legal-administrative-triplets.The dataset was created as part of the ALIA project to improve the
training of embedding models and dense retrievers specialized in Spanish
legal and administrative language.
Hard negatives are passages that are semantically similar to a query
but not… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-legal-administrative-triplets.ALIA-es-legal-administrative-cqa
Dataset Introduction
The ALIA Spanish Legal and Administrative for Context Question Answering Corpus is a specialized question-answering resource derived from the SINAI/ALIA-es-legal-administrative corpus. This dataset transforms legal and administrative documents into structured question-answer pairs, enabling the development and evaluation of AI systems capable of understanding and responding to queries about Spanish legal-administrative content. With 17,668 structured instances… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-legal-administrative-cqa.ALIA-es-biomedical-synthetic-instructions
Dataset Introduction
The ALIA Spanish Biomedical Synthetic Instructions Corpus is a synthetic instruction-tuning resource in Spanish created under the ALIA project using the Magpie methodology. It was designed to train and evaluate language models in biomedical and healthcare tasks with controlled formats and large-scale supervision.
It contains:
639,456 instances
961,073,205 tokens
14 task modalities (clinical diagnosis, patient education, ethical reasoning, document… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-biomedical-synthetic-instructions.ALIA-es-legal-administrative-synthetic-instructions
Dataset Introduction
The ALIA Spanish Legal and Administrative Synthetic Instructions Corpus is a synthetic instruction-tuning resource in Spanish created under the ALIA project using the Magpie methodology. It was designed to train and evaluate language models in legal and administrative tasks with controlled formats and large-scale supervision.
It contains:
763,804 instances
534,112,398 tokens
16 task modalities (questions, instructions, multiple-choice, true/false; with and… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-legal-administrative-synthetic-instructions.ALIA-es-cultural-heritage-synthetic-instructions
Dataset Introduction
The ALIA Spanish Cultural and Heritage Synthetic Instructions Corpus is a synthetic instruction-tuning resource in Spanish created under the ALIA project using the Magpie methodology. It was designed to train and evaluate language models in cultural heritage, digital humanities, and historical knowledge tasks with natural linguistic variation and large-scale supervision.
It contains:
748,480 instances
629,682,398 tokens
25 task modalities (heritage QA… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-cultural-heritage-synthetic-instructions.ALIA-es-cultural-heritage
Dataset Introduction
The ALIA Spanish Cultural and Heritage Corpus is a strategic open data infrastructure designed to support research and innovation in digital humanities, cultural analytics, and Spanish-language NLP. It consolidates heterogeneous official and academic repositories into a single curated dataset, enabling broad and structured access to cultural heritage documentation from Spain. With 236,314 instances, 939,315,404 tokens and 100 source datasets, it provides a… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-cultural-heritage.ALIA-es-biomedical-pairs
Dataset Introduction
The ALIA Spanish Biomedical Retrieval Pairs Corpus contains tabular instances designed to train and evaluate retrieval-oriented models (e.g., dense retrievers / embedding encoders) using passage-grounded query data produced with a Qwen3-style prompting workflow integrated in the ALIA encoders pipeline.
It preserves provenance to the original document and chunk while exposing controls such as question type and difficulty (ranging from high_school to phd level).… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-biomedical-pairs.ALIA-es-biomedical
Dataset Introduction
The ALIA Spanish Biomedical Corpus constitutes a strategic data infrastructure designed to support research and innovation in the biomedical domain. By ensuring systematic access to multiple official medical repositories in a single consolidated dataset, it provides a robust foundation for Spanish-language BioNLP. With over 6 million instances and more than 4 billion tokens, it represents a relevant comprehensive corpus of biomedical and clinical-related… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-biomedical.ALIA-es-cultural-heritage-pairs
Dataset Introduction
The ALIA Spanish Cultural and Heritage Retrieval Pairs Corpus contains tabular instances designed to train and evaluate retrieval-oriented models (e.g., dense retrievers / embedding encoders) using passage-grounded query data produced with a Qwen-style prompting workflow integrated in the ALIA encoders pipeline.
It preserves provenance to the original document and passage while exposing controls such as question type and difficulty (ranging from high_school… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-cultural-heritage-pairs.TR-DataAnalystBench
TR-DataAnalystBench
A Turkish-language benchmark for evaluating whether language models can perform
data-analyst style reasoning over tables and charts: reading a value,
finding the maximum/minimum, comparing two years, computing an average or a
(signed) percentage change, ranking, summarizing a trend, and — importantly —
abstaining when the data does not contain the answer.
Gold answers are computed and verified with Python (not produced by a language
model), so the benchmark… See the full description on the dataset page: https://huggingface.co/datasets/alialp207/TR-DataAnalystBench.ALIA-es-cultural-heritage-triplets
Dataset Introduction
The dataset ALIA Spanish Cultural and Heritage Hard Negatives Corpus contains hard negatives for dense retrieval training
generated from <query, passage> pairs contained in SINAI/ALIA-es-cultural-heritage-pairs.The dataset was created as part of the ALIA project to improve the training of embedding models and dense retrievers specialized in Spanish cultural heritage language.
Hard negatives are passages that are semantically similar to a query
but not… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-cultural-heritage-triplets.ALIA-es-legal-administrative-pairs
Dataset Introduction
The ALIA Spanish Legal and Administrative Pairs Corpus, derived from the SINAI/ALIA-es-legal-administrative, contains tabular instances designed to train and evaluate retrieval-oriented models (e.g., dense retrievers / embedding encoders) using passage-grounded query data produced with a Qwen3-style prompting workflow integrated in the ALIA encoders pipeline.
It preserves provenance to the original document and chunk while exposing controls such as question type… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-legal-administrative-pairs.ALIA-es-biomedical-triplets
Dataset Introduction
The dataset ALIA Spanish Biomedical Hard Negatives Corpus contains hard negatives for dense retrieval training
generated from <query, passage> pairs contained in SINAI/ALIA-es-biomedical-pairs.The dataset was created as part of the ALIA project to improve the
training of embedding models and dense retrievers specialized in Spanish
biomedical language.
Hard negatives are passages that are semantically similar to a query
but not correct answers, making them… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-biomedical-triplets.dsp-fft-sampling-aliasing
Synthetic DSP Dataset: FFT + Sampling / Aliasing
This repository contains synthetic instruction-style DSP samples
designed for numerical reasoning and conceptual understanding of
Digital Signal Processing (DSP) fundamentals.
The dataset focuses on:
FFT bin reasoning and frequency-domain interpretation
Sampling theory
Aliasing effects
Dataset Origin & Verification
This dataset was generated as part of the project:
Fine-Tuning Lightweight Large Language Models for a… See the full description on the dataset page: https://huggingface.co/datasets/Irfanuruchi/dsp-fft-sampling-aliasing.TurkishIdentityMini
TurkishIdentityMini
Dataset Description
TurkishIdentityMini is a small, template-based Turkish instruction dataset designed to help LLMs respond correctly to identity-related questions. It contains instruction–output pairs where a user asks a chatbot about its name, origin, or creator, and the model responds using customizable {{model_name}} and {{team_name}} placeholders.
This dataset is useful for fine-tuning or instruction-tuning Turkish language models to maintain a… See the full description on the dataset page: https://huggingface.co/datasets/aliarda/TurkishIdentityMini.enoch-ai-research-corpus
Enoch AI Research Corpus
This dataset contains 393 AI-generated research artifacts produced by the Enoch agentic research system.
System repository: https://github.com/alias8818/enoch-agentic-research-system
Source corpus repository: https://github.com/alias8818/enoch-ai-research-corpus
Launch site: https://alias8818.github.io/enoch-agentic-research-system/
Current release correction
Older launch posts may mention 120 artifacts. The current public corpus indexes… See the full description on the dataset page: https://huggingface.co/datasets/aliasocracy/enoch-ai-research-corpus.jev_turkish_mmlu_traces
JEV Turkish MMLU & MMLU-Pro Traces
Traces of Jev (jev-latest, TypeSafe System One) answering Turkish multiple-choice
questions from the turkish_mmlu and turkish-mmlu-pro-preview datasets.
Each source row becomes a choice question; rows are grouped by subject and sent as one
request per subject (the subject is the state). Every trace row records jev's chosen
option, confidence, probability distribution, and (when captured) the request id, token
usage, and latency.… See the full description on the dataset page: https://huggingface.co/datasets/aliarda/jev_turkish_mmlu_traces.
