CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tonyhong /ramp Dataset Card for Retrieval-Augmented Modular Prompt Tuning for Low-Resource Data-to-Text Generation (RAMP) Hugging Face Dataset | GitHub Repository | paper | Gitlab Repository RAMP provides a prepared version of a low-resource data-to-text corpus for drone handover message generation: structured sensor records (status + time-step object lists) paired with natural-language “handover” messages describing critical situations. The release includes raw/filtered splits and… See the full description on the dataset page: https://huggingface.co/datasets/tonyhong/ramp.texttext-generationn<1K0 likes300 downloads1y agoHugging Face024esv /rameau Rameau: functional harmony from notation A text-to-text dataset and benchmark for functional harmony: Roman-numeral analysis, cadence classification, and key identification. A probabilistic common-practice grammar generates the progressions; four task framings hide the answer to increasing degrees. Chord-symbol lookup stops working after the first one. Named for Jean-Philippe Rameau, whose Traité de l'harmonie (1722) started the discipline. symbol_to_rn key: C major /… See the full description on the dataset page: https://huggingface.co/datasets/4esv/rameau.texttext-generation10K<n<100K0 likes196 downloads2mo agoHugging Face03ram-lexsi /curatorkit-testrun-Secrets curatorkit-testrun-Secrets Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training. Method curation Backend — Model — Formats alpaca Artifact dataset Published 2026-08-30 05:55 UTC Usage from datasets import load_dataset ds = load_dataset("ram-lexsi/curatorkit-testrun-Secrets", "alpaca") texttext-generationn<1K0 likes172 downloads25d agoHugging Face04ram-lexsi /curatorkit-testrun-Prompt-Template curatorkit-testrun-Prompt-Template Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training. Method qa Backend litellm Model openai/Qwen/Qwen2.5-0.5B-Instruct Formats alpaca Artifact dataset Published 2026-08-30 09:29 UTC Usage from datasets import load_dataset ds = load_dataset("ram-lexsi/curatorkit-testrun-Prompt-Template", "alpaca") texttext-generationn<1K0 likes141 downloads25d agoHugging Face05RamAnanth1 /lex-fridman-podcasts Dataset Card for Lex Fridman Podcasts Dataset This dataset is sourced from Andrej Karpathy's Lexicap website which contains English transcripts of Lex Fridman's wonderful podcast episodes. The transcripts were generated using OpenAI's large-sized Whisper model texttext-classificationn<1K6 likes102 downloads4y agoHugging Face06ram-lexsi /curatorkit-testrun-PII curatorkit-testrun-PII Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training. Method curation Backend — Model — Formats alpaca Artifact dataset Published 2026-08-30 05:57 UTC Usage from datasets import load_dataset ds = load_dataset("ram-lexsi/curatorkit-testrun-PII", "alpaca") texttext-generationn<1K0 likes90 downloads25d agoHugging Face07ram-lexsi /curatorkit-testrun-Embedding-Dedup curatorkit-testrun-Embedding-Dedup Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training. Method curation Backend — Model — Formats alpaca Artifact dataset Published 2026-08-30 06:06 UTC Usage from datasets import load_dataset ds = load_dataset("ram-lexsi/curatorkit-testrun-Embedding-Dedup", "alpaca") texttext-generationn<1K0 likes86 downloads25d agoHugging Face08ram-lexsi /curatorkit-testrun-Clean-Dedup curatorkit-testrun-Clean-Dedup Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training. Method curation Backend — Model — Formats alpaca, sharegpt Artifact dataset Published 2026-08-30 06:25 UTC Usage from datasets import load_dataset ds = load_dataset("ram-lexsi/curatorkit-testrun-Clean-Dedup", "alpaca") texttext-generationn<1K0 likes83 downloads25d agoHugging Face09ram-lexsi /curatorkit-testrun-Reward curatorkit-testrun-Reward Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training. Method qa Backend litellm Model openai/Qwen/Qwen2.5-0.5B-Instruct Formats alpaca Artifact dataset Published 2026-09-01 05:12 UTC Usage from datasets import load_dataset ds = load_dataset("ram-lexsi/curatorkit-testrun-Reward", "alpaca") texttext-generationn<1K0 likes82 downloads23d agoHugging Face10ram-lexsi /curatorkit-testrun-Toxicity curatorkit-testrun-Toxicity Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training. Method curation Backend — Model — Formats alpaca Artifact dataset Published 2026-08-30 06:00 UTC Usage from datasets import load_dataset ds = load_dataset("ram-lexsi/curatorkit-testrun-Toxicity", "alpaca") texttext-generationn<1K0 likes80 downloads25d agoHugging Face11ram-lexsi /curatorkit-testrun-Reward-Refiner curatorkit-testrun-Reward-Refiner Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training. Method qa Backend litellm Model openai/Qwen/Qwen2.5-0.5B-Instruct Formats alpaca Artifact dataset Published 2026-09-01 05:06 UTC Usage from datasets import load_dataset ds = load_dataset("ram-lexsi/curatorkit-testrun-Reward-Refiner", "alpaca") texttext-generationn<1K0 likes80 downloads23d agoHugging Face12ram-lexsi /curatorkit-testrun-Adversarial-Preference curatorkit-testrun-Adversarial-Preference Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training. Method adversarial_preference Backend litellm Model openai/Qwen/Qwen2.5-0.5B-Instruct Formats dpo Artifact dataset Published 2026-08-28 10:47 UTC Usage from datasets import load_dataset ds = load_dataset("ram-lexsi/curatorkit-testrun-Adversarial-Preference"… See the full description on the dataset page: https://huggingface.co/datasets/ram-lexsi/curatorkit-testrun-Adversarial-Preference.texttext-generationn<1K0 likes77 downloads27d agoHugging Face13ram-lexsi /curatorkit-testrun-Checkpoint curatorkit-testrun-Checkpoint Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training. Method qa Backend litellm Model openai/Qwen/Qwen2.5-0.5B-Instruct Formats alpaca, sharegpt Artifact dataset Published 2026-09-01 06:08 UTC Usage from datasets import load_dataset ds = load_dataset("ram-lexsi/curatorkit-testrun-Checkpoint", "alpaca") texttext-generationn<1K0 likes77 downloads23d agoHugging Face14ram-lexsi /curatorkit-testrun-Ingest curatorkit-testrun-Ingest Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training. Method curation Backend — Model — Formats alpaca Artifact dataset Published 2026-09-01 05:41 UTC Usage from datasets import load_dataset ds = load_dataset("ram-lexsi/curatorkit-testrun-Ingest", "alpaca") texttext-generationn<1K0 likes73 downloads23d agoHugging Face15ram-lexsi /curatorkit-testrun-PDF curatorkit-testrun-PDF Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training. Method qa Backend litellm Model openai/Qwen/Qwen2.5-0.5B-Instruct Formats alpaca, sharegpt Artifact dataset Published 2026-09-01 06:27 UTC Usage from datasets import load_dataset ds = load_dataset("ram-lexsi/curatorkit-testrun-PDF", "alpaca") texttext-generationn<1K0 likes73 downloads23d agoHugging Face16ram-lexsi /curatorkit-testrun-CSVJSONParquet curatorkit-testrun-CSVJSONParquet Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training. Method curation Backend — Model — Formats alpaca, sharegpt Artifact dataset Published 2026-09-01 05:51 UTC Usage from datasets import load_dataset ds = load_dataset("ram-lexsi/curatorkit-testrun-CSVJSONParquet", "alpaca") texttext-generationn<1K0 likes73 downloads23d agoHugging Face17ram-lexsi /curatorkit-testrun-Multiturn curatorkit-testrun-Multiturn Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training. Method multiturn Backend litellm Model openai/Qwen/Qwen2.5-0.5B-Instruct Formats alpaca, sharegpt Artifact dataset Published 2026-08-28 10:04 UTC Usage from datasets import load_dataset ds = load_dataset("ram-lexsi/curatorkit-testrun-Multiturn", "alpaca") texttext-generationn<1K0 likes72 downloads27d agoHugging Face18ram-lexsi /curatorkit-testrun-Hallucination curatorkit-testrun-Hallucination Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training. Method qa Backend litellm Model openai/Qwen/Qwen2.5-0.5B-Instruct Formats alpaca Artifact dataset Published 2026-08-28 10:47 UTC Usage from datasets import load_dataset ds = load_dataset("ram-lexsi/curatorkit-testrun-Hallucination", "alpaca") texttext-generationn<1K0 likes72 downloads27d agoHugging Face19ram-lexsi /curatorkit-testrun-Diversity curatorkit-testrun-Diversity Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training. Method qa Backend litellm Model openai/Qwen/Qwen2.5-0.5B-Instruct Formats alpaca Artifact dataset Published 2026-08-30 05:54 UTC Usage from datasets import load_dataset ds = load_dataset("ram-lexsi/curatorkit-testrun-Diversity", "alpaca") texttext-generationn<1K0 likes72 downloads25d agoHugging Face20dcmutlu /gordon-ramsay-code-review-v2 Gordon Ramsay Code Review & Auditor Corpus v2 (dcmutlu/gordon-ramsay-code-review-v2) A high-density synthetic dataset of 10,000 multi-turn code review pairs designed to fine-tune open-weight reasoners (specifically Qwen2.5-Coder-7B-Instruct) into Chef Gordon Ramsay: Sovereign Executive Code Auditor and Supreme Software Gastronomer. 🍳 Dataset Overview This dataset merges rigorous computer science diagnostics (Abstract Syntax Tree inspection, concurrency lifecycle… See the full description on the dataset page: https://huggingface.co/datasets/dcmutlu/gordon-ramsay-code-review-v2.texttext-generation10K<n<100K0 likes71 downloads15d agoHugging Face21ram-lexsi /curatorkit-testrun-Probe curatorkit-testrun-Probe Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training. Method qa Backend litellm Model openai/Qwen/Qwen2.5-0.5B-Instruct Formats alpaca Artifact dataset Published 2026-08-30 06:03 UTC Usage from datasets import load_dataset ds = load_dataset("ram-lexsi/curatorkit-testrun-Probe", "alpaca") texttext-generationn<1K0 likes68 downloads25d agoHugging Face22rampisipati /DeepSeek-V4-Distill-8000x 🐳 DeepSeek-V4-Distill-8100x Dataset Summary DeepSeek-V4-Distill-8100x is a supervised fine-tuning dataset for reasoning-oriented distillation. The question prompts come from Jackrong/GLM-5.1-Reasoning-1M-Cleaned, and the answers were generated by the teacher model DeepSeek-V4-Flash. After the cleaning process, the released train split contains 7,716 high-quality JSONL examples. [!NOTE] The answer pool was cleaned to remove real-time questions… See the full description on the dataset page: https://huggingface.co/datasets/rampisipati/DeepSeek-V4-Distill-8000x.texttext-generation1K<n<10K0 likes67 downloads5mo agoHugging Face23ram-lexsi /curatorkit-testrun-Injector curatorkit-testrun-Injector Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training. Method qa Backend litellm Model openai/Qwen/Qwen2.5-0.5B-Instruct Formats alpaca Artifact dataset Published 2026-08-30 06:13 UTC Usage from datasets import load_dataset ds = load_dataset("ram-lexsi/curatorkit-testrun-Injector", "alpaca") texttext-generationn<1K0 likes67 downloads25d agoHugging Face24ram-lexsi /curatorkit-testrun-OutputSplit curatorkit-testrun-OutputSplit Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training. Method curation Backend — Model — Formats test-alpaca, test-sharegpt, train-alpaca, train-sharegpt, val-alpaca, val-sharegpt Artifact dataset Published 2026-09-01 04:54 UTC Usage from datasets import load_dataset ds = load_dataset("ram-lexsi/curatorkit-testrun-OutputSplit"… See the full description on the dataset page: https://huggingface.co/datasets/ram-lexsi/curatorkit-testrun-OutputSplit.texttext-generationn<1K0 likes67 downloads23d agoHugging Face25ram-lexsi /curatorkit-testrun-Filtered-FT curatorkit-testrun-Filtered-FT Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training. Method qa Backend litellm Model openai/Qwen/Qwen2.5-0.5B-Instruct Formats alpaca Artifact dataset Published 2026-09-01 05:49 UTC Usage from datasets import load_dataset ds = load_dataset("ram-lexsi/curatorkit-testrun-Filtered-FT", "alpaca") texttext-generationn<1K0 likes64 downloads23d agoHugging Face26ram-lexsi /curatorkit-testrun-FormatDetector curatorkit-testrun-FormatDetector Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training. Method curation Backend — Model — Formats alpaca, sharegpt Artifact dataset Published 2026-09-01 05:43 UTC Usage from datasets import load_dataset ds = load_dataset("ram-lexsi/curatorkit-testrun-FormatDetector", "alpaca") texttext-generationn<1K0 likes62 downloads23d agoHugging Face27ram-lexsi /auditkit-testrun-lmeval auditkit-testrun-lmeval Built using AuditKIT — evaluate any model on any dataset and any task. Method evaluate Model vllm:Qwen/Qwen2.5-0.5B-Instruct Artifact run Published 2026-09-02 06:19 UTC Usage from datasets import load_dataset ds = load_dataset("ram-lexsi/auditkit-testrun-lmeval") texttext-generationn<1K0 likes61 downloads22d agoHugging Face28ramendik /kimify-ifeval-like Kimify IFEval-Like Dataset Dataset Description This dataset contains 10,070 verified instruction-following conversations in the IFEval format. Each example includes: A user prompt with embedded constraints An assistant response that satisfies those constraints Metadata describing the constraint types and parameters All examples have been programmatically verified using the instruction-following-eval library (based on Google Research's IFEval) to ensure 100% constraint… See the full description on the dataset page: https://huggingface.co/datasets/ramendik/kimify-ifeval-like.texttext-generation10K<n<100K0 likes60 downloads8mo agoHugging Face29ram-lexsi /curatorkit-testrun-Preference curatorkit-testrun-Preference Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training. Method preference Backend litellm Model openai/Qwen/Qwen2.5-0.5B-Instruct Formats dpo Artifact dataset Published 2026-08-28 09:57 UTC Usage from datasets import load_dataset ds = load_dataset("ram-lexsi/curatorkit-testrun-Preference", "dpo") texttext-generationn<1K0 likes57 downloads27d agoHugging Face30ram-lexsi /curatorkit-testrun-Hygiene curatorkit-testrun-Hygiene Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training. Method curation Backend — Model — Formats alpaca Artifact dataset Published 2026-09-01 05:33 UTC Usage from datasets import load_dataset ds = load_dataset("ram-lexsi/curatorkit-testrun-Hygiene", "alpaca") texttext-generationn<1K0 likes57 downloads23d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.