datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rameau
Rameau: functional harmony from notation
A text-to-text dataset and benchmark for functional harmony: Roman-numeral
analysis, cadence classification, and key identification. A probabilistic
common-practice grammar generates the progressions; four task framings hide
the answer to increasing degrees. Chord-symbol lookup stops working after the
first one.
Named for Jean-Philippe Rameau, whose Traité de l'harmonie (1722) started
the discipline.
symbol_to_rn key: C major /… See the full description on the dataset page: https://huggingface.co/datasets/4esv/rameau.lex-fridman-podcasts
Dataset Card for Lex Fridman Podcasts Dataset
This dataset is sourced from Andrej Karpathy's Lexicap website which contains English transcripts of Lex Fridman's wonderful podcast episodes. The transcripts were generated using OpenAI's large-sized Whisper model
gordon-ramsay-code-review-v2
Gordon Ramsay Code Review & Auditor Corpus v2 (dcmutlu/gordon-ramsay-code-review-v2)
A high-density synthetic dataset of 10,000 multi-turn code review pairs designed to fine-tune open-weight reasoners (specifically Qwen2.5-Coder-7B-Instruct) into Chef Gordon Ramsay: Sovereign Executive Code Auditor and Supreme Software Gastronomer.
🍳 Dataset Overview
This dataset merges rigorous computer science diagnostics (Abstract Syntax Tree inspection, concurrency lifecycle… See the full description on the dataset page: https://huggingface.co/datasets/dcmutlu/gordon-ramsay-code-review-v2.DeepSeek-V4-Distill-8000x
🐳 DeepSeek-V4-Distill-8100x
Dataset Summary
DeepSeek-V4-Distill-8100x is a supervised fine-tuning dataset for reasoning-oriented distillation. The question prompts come from Jackrong/GLM-5.1-Reasoning-1M-Cleaned, and the answers were generated by the teacher model DeepSeek-V4-Flash.
After the cleaning process, the released train split contains 7,716 high-quality JSONL examples.
[!NOTE]
The answer pool was cleaned to remove real-time questions… See the full description on the dataset page: https://huggingface.co/datasets/rampisipati/DeepSeek-V4-Distill-8000x.kimify-ifeval-like
Kimify IFEval-Like Dataset
Dataset Description
This dataset contains 10,070 verified instruction-following conversations in the IFEval format. Each example includes:
A user prompt with embedded constraints
An assistant response that satisfies those constraints
Metadata describing the constraint types and parameters
All examples have been programmatically verified using the instruction-following-eval library (based on Google Research's IFEval) to ensure 100% constraint… See the full description on the dataset page: https://huggingface.co/datasets/ramendik/kimify-ifeval-like.gordon-ramsay-code-review
gordon-ramsay-code-review
Autonomous synthetic pretraining dataset synthesized by JESUS Sovereign Forge.
Synthesized via JESUS Sovereign Cloud Model Forge (hf-colab-forge) for native byte-level micro-transformers (Atom GPT) and LLM fine-tuning.
Dataset Summary
Metric
Value
Total Scenarios
500
Train Samples
450
Validation Samples
50
Total Byte Tokens
819,927
Train Tokens
737,852
Val Tokens
82,075
Vocab Size
258 (UTF-8 Bytes + BOS/PAD)… See the full description on the dataset page: https://huggingface.co/datasets/dcmutlu/gordon-ramsay-code-review.code.evol.instruct.wiz.oss_python.jsonkimify-short-20260131A conversational dataset generated by Kimi K2 0905 Instruct. The user prompts were taken from two datasets:
smoltalk multilingual - English prompts in the "advice-seeking" category
smoltalk - in the "smol-magpie-ultra-short" category. Note these involve three user/assistant turns.
System prompts were used to encourage brevity, for example: "You are Kimi K2, a versatile AI assistant. Be concise, clear, and punchy—aim for brief but helpful responses. Keep your distinctive voice but stay… See the full description on the dataset page: https://huggingface.co/datasets/ramendik/kimify-short-20260131.data-oss_instruct-decontaminated_python.jsonlkimify-20251115
Took a randomized selection of prompts from smoltalk-multilingual (several categories such as advice-seeking) and smoltalk-magpie-ultra. Got answers from Kimi K2. Pruned untrue/unverifiable statements.
This dataset is intended to teach K2's style to an LLM. Maximal length of a full sample: 6000 tokens; average ~1470 tokebs
Ask-The-Ramayana
📖 AskTheRamayana
A curated dataset of 90+ short, factual statements about the Ramayana — designed for QA fine-tuning, RAG systems, and cultural knowledge retrieval.
🕉️ Dataset Overview
Attribute
Description
Theme
Character-based focus (Ram, Sita, Laxman, Hanuman, Ravan, or general)
Chapter
Kand from Valmiki Ramayana (Bal Kand to Uttar Kand)
Content
One sentence (≤20 words) describing a key event or trait
Each entry is concise, factual, and rooted in… See the full description on the dataset page: https://huggingface.co/datasets/kartik3das/Ask-The-Ramayana.
