datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
airflow-dag-dataset
Airflow DAG Generation Dataset
This dataset combines Airflow-specific DAG generation examples with general Python coding instructions for fine-tuning code generation models.
Dataset Description
Total Samples: 8,238
Dataset Composition
The dataset includes two types of examples, identified by the source field:
Airflow Instructions (source: "airflow") - 6,738 samples (81.8%)
High-quality DAG generation examples with instruction variants
Domain-specific Airflow… See the full description on the dataset page: https://huggingface.co/datasets/andrea-t94/airflow-dag-dataset.chsa-triage-medical-bilingual
CHSA Triage — Corpus médical bilingue (SFT + DPO)
Corpus destiné au post-training d'un agent d'aide au triage médical (POC, Centre
Hospitalier Saint-Aurélien). Bilingue français / anglais, anonymisé (RGPD) et
versionné (empreintes SHA-256 dans manifest.json).
⚠️ Usage : aide à la décision destinée à du personnel soignant. Ne pose pas de
diagnostic et ne remplace pas un professionnel de santé.
Contenu
Fichier
Description
Format
sft_train.jsonl /… See the full description on the dataset page: https://huggingface.co/datasets/DagueGG/chsa-triage-medical-bilingual.DAG-Reasoning-DeepSeek-R1-0528Click here to support our open-source dataset and model releases!
DAG-Reasoning-DeepSeek-R1-0528 is a dataset focused on analysis and reasoning, creating directed acyclic graphs testing the limits of DeepSeek R1 0528's graph-reasoning skills!
This dataset contains:
4.08k synthetically generated prompts to create directed acyclic graphs in response to user input, with all responses generated using DeepSeek R1 0528.
All responses contain a multi-step thinking process to perform effective… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/DAG-Reasoning-DeepSeek-R1-0528.dagger
DAGGER Training Dataset
Dataset Description
Training data for DAGGER (Distractor-Aware Graph Generation for Executable Reasoning) models. This dataset contains Bangla mathematical word problems paired with computational graph solutions, formatted for both SFT and GRPO training pipelines.
Highlights
3,000 training examples with verified computational graphs
Two training configs: SFT (with validation) and GRPO formats
GPT-4.1… See the full description on the dataset page: https://huggingface.co/datasets/dipta007/dagger.tau2-infinity-dag
tau2-infinity
An adaptive benchmark for evaluating LLM tool-use agents on airline customer service tasks. Generated using EnvScaler by VibrantLabs.
Overview
Each task requires an agent to transform an initial database state S_0 into a golden final state S* by executing a sequence of tool calls (flight searches, bookings, cancellations, updates, etc.). Tasks were adaptively generated to target specific difficulty levels against a calibration model.
Property
Value… See the full description on the dataset page: https://huggingface.co/datasets/vibrantlabsai/tau2-infinity-dag.youth-conversations-dag
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Youth Conversations Dataset – Dagbani
Dataset Summary
This dataset contains 774 synthetic transcripts of conversations
between a counsellor and a young person in Ghana, rendered in Dagbani.
Each transcript was… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/youth-conversations-dag.
