datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
headwater-b2b-data-purification-ai-ingestion-implementation-kit
HeadWater B2B Data Purification & AI Ingestion Implementation Kit
He will not suffer thy foot to be moved: he that keepeth thee will not slumber. Psalm 121:3
YOU CAN HAVE IT NOW.
About This Kit
The Headwater B2B Data Purification & AI Ingestion Implementation Kit is engineered for one primary purpose: to eliminate development delays and buy back your operational momentum.
Instead of wasting weeks building foundational data plumbing from scratch… See the full description on the dataset page: https://huggingface.co/datasets/headwaterai/headwater-b2b-data-purification-ai-ingestion-implementation-kit.chatgpt-python311-implementation-77
ChatGPT Python 3.11 Implementation 77
A 77-record synthetic Python 3.11 implementation dataset generated with ChatGPT.
The exact generator model variant was not preserved. Creator recollection favors ChatGPT LunaMax, but ChatGPT Terra Max remains possible, so the dataset does not attribute generation to a single exact model.
Dataset Size
Metric
Count
Final records
77
Unique records
77
Python prompts
77
Python 3.11 prompts
77
Fresh GPT-5.6 Sol… See the full description on the dataset page: https://huggingface.co/datasets/TaskPuppyAI/chatgpt-python311-implementation-77.System-Prompt-Instruction-Real-world-Implementation-Training-set
SPIRIT Dataset (System Prompt Instruction Real-world Implementation Training-set)
Dataset Summary
SPIRIT is a high-quality system prompt instruction dataset designed to enhance language models' ability to follow complex system prompts. The dataset comprises real-world system prompts collected from GitHub repositories and synthetically generated conversations, specifically curated to improve system prompt adherence in large language models.
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/EricLu/System-Prompt-Instruction-Real-world-Implementation-Training-set.ptdbench-verl-implementation-torch-functional-dataset
PTDBench dataset snapshot: torch_functional
This repository stores the immutable runtime dataset snapshot for one
materialized PTDBench task. It intentionally excludes model weights and
training checkpoints.
PTDBench family: verl_implementation
Source evaluation metric: val-core/taco/acc/mean@1
Provenance: Processed from local TACO EASY (drop picture_num != 0); 8368 train / 184 test rows; bytes identical to task_function_call.
License: Apache-2.0
The artifact manifest records… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-verl-implementation-torch-functional-dataset.ptdbench-llama-dapo-implementation-task-monkey-patch-011-dataset
PTDBench dataset snapshot: task_monkey_patch_011
This repository stores the immutable runtime dataset snapshot for one
materialized PTDBench task. It intentionally excludes model weights and
training checkpoints.
PTDBench family: llama_dapo_implementation
Source evaluation metric: val-core/math_dapo/acc/mean@1
Provenance: Processed from BytedTsinghua-SIA/DAPO-Math-17k; task-specific bytes are pinned.
License: Apache-2.0
The artifact manifest records every hydrated runtime path… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-llama-dapo-implementation-task-monkey-patch-011-dataset.ptdbench-verl-reward-implementation-task-engine-base-025-dataset
PTDBench dataset snapshot: task_engine_base_025
This repository stores the immutable runtime dataset snapshot for one
materialized PTDBench task. It intentionally excludes model weights and
training checkpoints.
PTDBench family: verl_reward_implementation
Source evaluation metric: val-core/openai/gsm8k/acc/mean@1
Provenance: Processed runtime snapshot of openai/gsm8k.
License: MIT
The artifact manifest records every hydrated runtime path, byte size, and
SHA-256. The task… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-verl-reward-implementation-task-engine-base-025-dataset.
