datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cannabis-fda-extractive-pilot
FDA Cannabis Extractive Experimental Pilot
Experimental, automatically screened, unreviewed draft dataset. This dataset is not medical advice, is not production-ready, and must not be represented as clinician-reviewed, legally cleared, or suitable for patient-facing systems.
This small English conversational dataset was created to test an auditable Gemma 4 fine-tuning pipeline. It contains exact answer passages from captured FDA pages about CBD/cannabis safety, paired with… See the full description on the dataset page: https://huggingface.co/datasets/aznatkoiny/cannabis-fda-extractive-pilot.stage3-real-expansion-agent-teacher-separated-pilot
Teacher-Separated Expansion Agent Pilot
A 10-task inspection batch generated by Qwen3-235B-A22B-Instruct-2507 from real
CLAPNQ, PubMedQA, MAUD, ContractNLI, and FinQA source tasks.
The teacher-only trajectory-generation system prompt is recorded in
metadata/generation-manifest.json for auditability, but is absent from every
saved training trajectory. Each final messages list begins with the real
memory-wrapped task user message, followed by native assistant expand calls,
exact… See the full description on the dataset page: https://huggingface.co/datasets/leonli66/stage3-real-expansion-agent-teacher-separated-pilot.router-pilot-tasks
BenchGen Router Pilot: task pool
The prompt pool a BenchGen router head trains on: 1110 normalized questions drawn from
public benchmark sources, balanced across domain and difficulty with a fixed seed, so results
reproduce. This is the task half of a paired release — the reward half (how well each pool
agent actually scored on a subset of these) is published separately at
benchgen/router-pilot.
Only task_id is the join key between the two datasets — load both and match on it… See the full description on the dataset page: https://huggingface.co/datasets/benchgen/router-pilot-tasks.router-pilot
BenchGen Router Pilot: per-question agent rewards
A routing dataset: for each question, how every agent in a fixed pool actually performed.
It is the input a model-selection policy trains on — not a question-answering dataset.
Each row is one task. mean_reward[i] is the accuracy of agent agent_order[i] over
3 independent attempts. Ties are preserved explicitly rather than collapsed by argmax,
because on easy questions several agents are genuinely equal and pretending otherwise… See the full description on the dataset page: https://huggingface.co/datasets/benchgen/router-pilot.lemonseed-rl-tasks-cogen-pilot
lemonseed-rl-tasks-cogen-pilot
LemonSeed — RL co-gen pilot tasks.
Contents
rl_tasks_cogen_pilot.jsonl (61 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Synthetic, generated programmatically for the LemonSeed 1.5B project (by Geramy L. Loveless). Data authored by Michael Anthony Falabella.
omission-detection-realdata-pilot
Omission Detection — Real-Data Pilot
What is Omission Detection?
Large language models (LLMs) in agentic pipelines often omit information
present in their context window — they fail to surface a relevant fact even
when it is theoretically visible. This dataset captures 372 controlled
trials from the real-data pilot, extending the synthetic sweep to
real-world documents and agent frameworks.
Each trial fetches a document from a real source (PubMed, HAPI FHIR, SEC… See the full description on the dataset page: https://huggingface.co/datasets/Santhiyarajan/omission-detection-realdata-pilot.
