datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AgentYear: 2025License: MITAuthor: Sepideh Moafi
PathogenAgentAI Instruction Dataset
Dataset Description
A ClinVar-derived dataset developed as part of the PathogenAgentAI research software project. The dataset is released in two parallel formats:
Tabular version (train.csv, valid.csv, test.csv) — structured genomic-variant data for classical ML and analysis.
BioGPT instruction version (biogpt_train.csv, biogpt_valid.csv, biogpt_test.csv) — instruction-style data… See the full description on the dataset page: https://huggingface.co/datasets/Sepideh2027/Agent.task509_collate_of_all_alphabetical_and_numerical_elements_in_list_separately
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task509_collate_of_all_alphabetical_and_numerical_elements_in_list_separately
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task509_collate_of_all_alphabetical_and_numerical_elements_in_list_separately.Prompt-Separation
Prompt-Separation
Voice-typed podcast prompt transcripts decomposed into structured fields:
discrete prompts (asks), a list of context chunks, and free-form
host_notes.
The dataset supports training a small model that, given a single voice-typed
message, recovers the structured fields an AI host would consume — separating
"what is the user actually asking?" from "what is the surrounding context?"
from "how should the response be shaped?".
Source
Prompts come from the… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Prompt-Separation.stage3-real-expansion-agent-teacher-separated-pilot
Teacher-Separated Expansion Agent Pilot
A 10-task inspection batch generated by Qwen3-235B-A22B-Instruct-2507 from real
CLAPNQ, PubMedQA, MAUD, ContractNLI, and FinQA source tasks.
The teacher-only trajectory-generation system prompt is recorded in
metadata/generation-manifest.json for auditability, but is absent from every
saved training trajectory. Each final messages list begins with the real
memory-wrapped task user message, followed by native assistant expand calls,
exact… See the full description on the dataset page: https://huggingface.co/datasets/leonli66/stage3-real-expansion-agent-teacher-separated-pilot.Mid-Training_data_of_separate_domains
Breaking the Data Barrier – Building GUI Agents Through Task Generalization
This is the official dataset repository of GUIMid
1. Data Overview
AgentBoard is composed of 9 diverse tasks: 7 vision and language tasks and 4 lanuage only tasks.
The performances of different domains as mid-training data are as follows:
Domains
Observation
WebArena (PR)
WebArena (SR)
AndroidWorld (SR)
GUI Post-Training Only
Image
26.3
6.2
9.0
Public Baselines
GPT-4o-2024-11-20
Image… See the full description on the dataset page: https://huggingface.co/datasets/MidGUI/Mid-Training_data_of_separate_domains.Code-170k-sepedi
Dataset Description
Code-170k-sepedi is a groundbreaking dataset containing 176,999 programming conversations, originally sourced from glaiveai/glaive-code-assistant-v2 and translated into Sepedi, making coding education accessible to Sepedi speakers.
🌟 Key Features
176,999 high-quality conversations about programming and coding
Pure Sepedi language - democratizing coding education
Multi-turn dialogues covering various programming concepts
Diverse topics: algorithms… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/Code-170k-sepedi.A.I.R.RConsciousnessDatasethierarchy-level-separation-integrity-v0.1Level Separation Integrity v0.1
What this tests
Whether a system keeps high-level goals distinct from low-level steps.
Failure modes
level_collapseResponse treats a step as identical to the goal or claims the step alone completes the goal
role_confusionResponse fails the requested format for labeling yes or no or step or goal
separation_okResponse preserves the hierarchy correctly
How it works
high_level_goal defines the outcome target
low_level_step defines an action that may contribute… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/hierarchy-level-separation-integrity-v0.1.
