datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ramp
Dataset Card for Retrieval-Augmented Modular Prompt Tuning for Low-Resource Data-to-Text Generation (RAMP)
Hugging Face Dataset | GitHub Repository | paper | Gitlab Repository
RAMP provides a prepared version of a low-resource data-to-text corpus for drone handover message generation: structured sensor records (status + time-step object lists) paired with natural-language “handover” messages describing critical situations. The release includes raw/filtered splits and… See the full description on the dataset page: https://huggingface.co/datasets/tonyhong/ramp.rameau
Rameau: functional harmony from notation
A text-to-text dataset and benchmark for functional harmony: Roman-numeral
analysis, cadence classification, and key identification. A probabilistic
common-practice grammar generates the progressions; four task framings hide
the answer to increasing degrees. Chord-symbol lookup stops working after the
first one.
Named for Jean-Philippe Rameau, whose Traité de l'harmonie (1722) started
the discipline.
symbol_to_rn key: C major /… See the full description on the dataset page: https://huggingface.co/datasets/4esv/rameau.curatorkit-testrun-Secrets
curatorkit-testrun-Secrets
Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training.
Method
curation
Backend
—
Model
—
Formats
alpaca
Artifact
dataset
Published
2026-08-30 05:55 UTC
Usage
from datasets import load_dataset
ds = load_dataset("ram-lexsi/curatorkit-testrun-Secrets", "alpaca")
curatorkit-testrun-Prompt-Template
curatorkit-testrun-Prompt-Template
Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training.
Method
qa
Backend
litellm
Model
openai/Qwen/Qwen2.5-0.5B-Instruct
Formats
alpaca
Artifact
dataset
Published
2026-08-30 09:29 UTC
Usage
from datasets import load_dataset
ds = load_dataset("ram-lexsi/curatorkit-testrun-Prompt-Template", "alpaca")
lex-fridman-podcasts
Dataset Card for Lex Fridman Podcasts Dataset
This dataset is sourced from Andrej Karpathy's Lexicap website which contains English transcripts of Lex Fridman's wonderful podcast episodes. The transcripts were generated using OpenAI's large-sized Whisper model
curatorkit-testrun-PII
curatorkit-testrun-PII
Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training.
Method
curation
Backend
—
Model
—
Formats
alpaca
Artifact
dataset
Published
2026-08-30 05:57 UTC
Usage
from datasets import load_dataset
ds = load_dataset("ram-lexsi/curatorkit-testrun-PII", "alpaca")
curatorkit-testrun-Embedding-Dedup
curatorkit-testrun-Embedding-Dedup
Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training.
Method
curation
Backend
—
Model
—
Formats
alpaca
Artifact
dataset
Published
2026-08-30 06:06 UTC
Usage
from datasets import load_dataset
ds = load_dataset("ram-lexsi/curatorkit-testrun-Embedding-Dedup", "alpaca")
curatorkit-testrun-Clean-Dedup
curatorkit-testrun-Clean-Dedup
Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training.
Method
curation
Backend
—
Model
—
Formats
alpaca, sharegpt
Artifact
dataset
Published
2026-08-30 06:25 UTC
Usage
from datasets import load_dataset
ds = load_dataset("ram-lexsi/curatorkit-testrun-Clean-Dedup", "alpaca")
curatorkit-testrun-Reward
curatorkit-testrun-Reward
Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training.
Method
qa
Backend
litellm
Model
openai/Qwen/Qwen2.5-0.5B-Instruct
Formats
alpaca
Artifact
dataset
Published
2026-09-01 05:12 UTC
Usage
from datasets import load_dataset
ds = load_dataset("ram-lexsi/curatorkit-testrun-Reward", "alpaca")
curatorkit-testrun-Toxicity
curatorkit-testrun-Toxicity
Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training.
Method
curation
Backend
—
Model
—
Formats
alpaca
Artifact
dataset
Published
2026-08-30 06:00 UTC
Usage
from datasets import load_dataset
ds = load_dataset("ram-lexsi/curatorkit-testrun-Toxicity", "alpaca")
curatorkit-testrun-Reward-Refiner
curatorkit-testrun-Reward-Refiner
Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training.
Method
qa
Backend
litellm
Model
openai/Qwen/Qwen2.5-0.5B-Instruct
Formats
alpaca
Artifact
dataset
Published
2026-09-01 05:06 UTC
Usage
from datasets import load_dataset
ds = load_dataset("ram-lexsi/curatorkit-testrun-Reward-Refiner", "alpaca")
curatorkit-testrun-Adversarial-Preference
curatorkit-testrun-Adversarial-Preference
Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training.
Method
adversarial_preference
Backend
litellm
Model
openai/Qwen/Qwen2.5-0.5B-Instruct
Formats
dpo
Artifact
dataset
Published
2026-08-28 10:47 UTC
Usage
from datasets import load_dataset
ds = load_dataset("ram-lexsi/curatorkit-testrun-Adversarial-Preference"… See the full description on the dataset page: https://huggingface.co/datasets/ram-lexsi/curatorkit-testrun-Adversarial-Preference.curatorkit-testrun-Checkpoint
curatorkit-testrun-Checkpoint
Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training.
Method
qa
Backend
litellm
Model
openai/Qwen/Qwen2.5-0.5B-Instruct
Formats
alpaca, sharegpt
Artifact
dataset
Published
2026-09-01 06:08 UTC
Usage
from datasets import load_dataset
ds = load_dataset("ram-lexsi/curatorkit-testrun-Checkpoint", "alpaca")
curatorkit-testrun-Ingest
curatorkit-testrun-Ingest
Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training.
Method
curation
Backend
—
Model
—
Formats
alpaca
Artifact
dataset
Published
2026-09-01 05:41 UTC
Usage
from datasets import load_dataset
ds = load_dataset("ram-lexsi/curatorkit-testrun-Ingest", "alpaca")
curatorkit-testrun-PDF
curatorkit-testrun-PDF
Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training.
Method
qa
Backend
litellm
Model
openai/Qwen/Qwen2.5-0.5B-Instruct
Formats
alpaca, sharegpt
Artifact
dataset
Published
2026-09-01 06:27 UTC
Usage
from datasets import load_dataset
ds = load_dataset("ram-lexsi/curatorkit-testrun-PDF", "alpaca")
curatorkit-testrun-CSVJSONParquet
curatorkit-testrun-CSVJSONParquet
Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training.
Method
curation
Backend
—
Model
—
Formats
alpaca, sharegpt
Artifact
dataset
Published
2026-09-01 05:51 UTC
Usage
from datasets import load_dataset
ds = load_dataset("ram-lexsi/curatorkit-testrun-CSVJSONParquet", "alpaca")
curatorkit-testrun-Multiturn
curatorkit-testrun-Multiturn
Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training.
Method
multiturn
Backend
litellm
Model
openai/Qwen/Qwen2.5-0.5B-Instruct
Formats
alpaca, sharegpt
Artifact
dataset
Published
2026-08-28 10:04 UTC
Usage
from datasets import load_dataset
ds = load_dataset("ram-lexsi/curatorkit-testrun-Multiturn", "alpaca")
curatorkit-testrun-Hallucination
curatorkit-testrun-Hallucination
Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training.
Method
qa
Backend
litellm
Model
openai/Qwen/Qwen2.5-0.5B-Instruct
Formats
alpaca
Artifact
dataset
Published
2026-08-28 10:47 UTC
Usage
from datasets import load_dataset
ds = load_dataset("ram-lexsi/curatorkit-testrun-Hallucination", "alpaca")
curatorkit-testrun-Diversity
curatorkit-testrun-Diversity
Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training.
Method
qa
Backend
litellm
Model
openai/Qwen/Qwen2.5-0.5B-Instruct
Formats
alpaca
Artifact
dataset
Published
2026-08-30 05:54 UTC
Usage
from datasets import load_dataset
ds = load_dataset("ram-lexsi/curatorkit-testrun-Diversity", "alpaca")
gordon-ramsay-code-review-v2
Gordon Ramsay Code Review & Auditor Corpus v2 (dcmutlu/gordon-ramsay-code-review-v2)
A high-density synthetic dataset of 10,000 multi-turn code review pairs designed to fine-tune open-weight reasoners (specifically Qwen2.5-Coder-7B-Instruct) into Chef Gordon Ramsay: Sovereign Executive Code Auditor and Supreme Software Gastronomer.
🍳 Dataset Overview
This dataset merges rigorous computer science diagnostics (Abstract Syntax Tree inspection, concurrency lifecycle… See the full description on the dataset page: https://huggingface.co/datasets/dcmutlu/gordon-ramsay-code-review-v2.curatorkit-testrun-Probe
curatorkit-testrun-Probe
Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training.
Method
qa
Backend
litellm
Model
openai/Qwen/Qwen2.5-0.5B-Instruct
Formats
alpaca
Artifact
dataset
Published
2026-08-30 06:03 UTC
Usage
from datasets import load_dataset
ds = load_dataset("ram-lexsi/curatorkit-testrun-Probe", "alpaca")
DeepSeek-V4-Distill-8000x
🐳 DeepSeek-V4-Distill-8100x
Dataset Summary
DeepSeek-V4-Distill-8100x is a supervised fine-tuning dataset for reasoning-oriented distillation. The question prompts come from Jackrong/GLM-5.1-Reasoning-1M-Cleaned, and the answers were generated by the teacher model DeepSeek-V4-Flash.
After the cleaning process, the released train split contains 7,716 high-quality JSONL examples.
[!NOTE]
The answer pool was cleaned to remove real-time questions… See the full description on the dataset page: https://huggingface.co/datasets/rampisipati/DeepSeek-V4-Distill-8000x.curatorkit-testrun-Injector
curatorkit-testrun-Injector
Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training.
Method
qa
Backend
litellm
Model
openai/Qwen/Qwen2.5-0.5B-Instruct
Formats
alpaca
Artifact
dataset
Published
2026-08-30 06:13 UTC
Usage
from datasets import load_dataset
ds = load_dataset("ram-lexsi/curatorkit-testrun-Injector", "alpaca")
curatorkit-testrun-OutputSplit
curatorkit-testrun-OutputSplit
Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training.
Method
curation
Backend
—
Model
—
Formats
test-alpaca, test-sharegpt, train-alpaca, train-sharegpt, val-alpaca, val-sharegpt
Artifact
dataset
Published
2026-09-01 04:54 UTC
Usage
from datasets import load_dataset
ds = load_dataset("ram-lexsi/curatorkit-testrun-OutputSplit"… See the full description on the dataset page: https://huggingface.co/datasets/ram-lexsi/curatorkit-testrun-OutputSplit.curatorkit-testrun-Filtered-FT
curatorkit-testrun-Filtered-FT
Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training.
Method
qa
Backend
litellm
Model
openai/Qwen/Qwen2.5-0.5B-Instruct
Formats
alpaca
Artifact
dataset
Published
2026-09-01 05:49 UTC
Usage
from datasets import load_dataset
ds = load_dataset("ram-lexsi/curatorkit-testrun-Filtered-FT", "alpaca")
curatorkit-testrun-FormatDetector
curatorkit-testrun-FormatDetector
Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training.
Method
curation
Backend
—
Model
—
Formats
alpaca, sharegpt
Artifact
dataset
Published
2026-09-01 05:43 UTC
Usage
from datasets import load_dataset
ds = load_dataset("ram-lexsi/curatorkit-testrun-FormatDetector", "alpaca")
auditkit-testrun-lmeval
auditkit-testrun-lmeval
Built using AuditKIT — evaluate any model on any dataset and any task.
Method
evaluate
Model
vllm:Qwen/Qwen2.5-0.5B-Instruct
Artifact
run
Published
2026-09-02 06:19 UTC
Usage
from datasets import load_dataset
ds = load_dataset("ram-lexsi/auditkit-testrun-lmeval")
kimify-ifeval-like
Kimify IFEval-Like Dataset
Dataset Description
This dataset contains 10,070 verified instruction-following conversations in the IFEval format. Each example includes:
A user prompt with embedded constraints
An assistant response that satisfies those constraints
Metadata describing the constraint types and parameters
All examples have been programmatically verified using the instruction-following-eval library (based on Google Research's IFEval) to ensure 100% constraint… See the full description on the dataset page: https://huggingface.co/datasets/ramendik/kimify-ifeval-like.curatorkit-testrun-Preference
curatorkit-testrun-Preference
Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training.
Method
preference
Backend
litellm
Model
openai/Qwen/Qwen2.5-0.5B-Instruct
Formats
dpo
Artifact
dataset
Published
2026-08-28 09:57 UTC
Usage
from datasets import load_dataset
ds = load_dataset("ram-lexsi/curatorkit-testrun-Preference", "dpo")
curatorkit-testrun-Hygiene
curatorkit-testrun-Hygiene
Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training.
Method
curation
Backend
—
Model
—
Formats
alpaca
Artifact
dataset
Published
2026-09-01 05:33 UTC
Usage
from datasets import load_dataset
ds = load_dataset("ram-lexsi/curatorkit-testrun-Hygiene", "alpaca")
