datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ALIA-es-biomedical-synthetic-instructions
Dataset Introduction
The ALIA Spanish Biomedical Synthetic Instructions Corpus is a synthetic instruction-tuning resource in Spanish created under the ALIA project using the Magpie methodology. It was designed to train and evaluate language models in biomedical and healthcare tasks with controlled formats and large-scale supervision.
It contains:
639,456 instances
961,073,205 tokens
14 task modalities (clinical diagnosis, patient education, ethical reasoning, document… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-biomedical-synthetic-instructions.ALIA-es-legal-administrative-synthetic-instructions
Dataset Introduction
The ALIA Spanish Legal and Administrative Synthetic Instructions Corpus is a synthetic instruction-tuning resource in Spanish created under the ALIA project using the Magpie methodology. It was designed to train and evaluate language models in legal and administrative tasks with controlled formats and large-scale supervision.
It contains:
763,804 instances
534,112,398 tokens
16 task modalities (questions, instructions, multiple-choice, true/false; with and… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-legal-administrative-synthetic-instructions.ALIA-es-cultural-heritage-synthetic-instructions
Dataset Introduction
The ALIA Spanish Cultural and Heritage Synthetic Instructions Corpus is a synthetic instruction-tuning resource in Spanish created under the ALIA project using the Magpie methodology. It was designed to train and evaluate language models in cultural heritage, digital humanities, and historical knowledge tasks with natural linguistic variation and large-scale supervision.
It contains:
748,480 instances
629,682,398 tokens
25 task modalities (heritage QA… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-cultural-heritage-synthetic-instructions.bio-devops-synthetic-instructions
Bio-DevOps Synthetic Instructions
This dataset contains synthetic instruction-following examples for biomedical-style data-engineering and scientific-computing workflows.
It was created for educational and portfolio use as part of a LoRA/QLoRA fine-tuning project using Qwen/Qwen2.5-Coder-7B-Instruct.
Related model:
AiLLMBS/qwen25-coder-bio-devops-lora
Dataset Contents
The dataset includes synthetic examples for:
Python CSV validation
pandas duplicate checks
bash… See the full description on the dataset page: https://huggingface.co/datasets/AiLLMBS/bio-devops-synthetic-instructions.synthetic-instructions
Synthetic Instruction Dataset
Generated using Qwen/Qwen2.5-3B-Instruct.
Samples: 25
Topics: web development, python programming, databases, deep learning, data science, natural language processing, computer vision
synthetic-instructions-judged
Synthetic Instructions (LLM Judged)
Generated with Distilabel-style pipeline, curated with LLM Judge.
Dataset Info
Total Generated: 25
Passed Judge (>= 3/5): 25
Generator: Qwen/Qwen2.5-3B-Instruct
Judge: Qwen/Qwen2.5-3B-Instruct
Quality Distribution
Score 4: 10 samples
Score 3: 15 samples
Fields
instruction: The question/task
input: Additional context (empty)
output: The response
topic: Subject area
quality_score: LLM Judge rating (1-5)
synthetic-queries-and-ml-instructions
Synthetic Dataset: Queries and ML Instructions
Dataset Description
Dataset Summary
This is a synthetic dataset with queries in Slovak language and ML instructions. The dataset was designed to train a model for extracting structured machine learning task requirements from natural language user queries.
The dataset contains user queries in Slovak describing ML tasks paired with structured JSON outputs containing task attributes like dataset modality, task type… See the full description on the dataset page: https://huggingface.co/datasets/kinit/synthetic-queries-and-ml-instructions.synthetic-conversations-and-ml-instructions
Synthetic Dataset: Conversations and ML Instructions
Dataset Description
Dataset Summary
This is a synthetic dataset of 5,000 Slovak multi-turn ML advisory conversations paired with structured JSON outputs. The dataset was designed to train and evaluate models that extract machine learning task requirements from realistic, natural Slovak dialogue, including cases where requirements are revealed gradually or changed during the conversation.
Each… See the full description on the dataset page: https://huggingface.co/datasets/kinit/synthetic-conversations-and-ml-instructions.
