datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RAMDocs
RAMDocs
Data for the paper Retrieval-Augmented Generation with Conflicting Evidence.
RAMDocs is a dataset that simulates complex and realistic scenarios for conflicting evidence for a user query, including ambiguity, misinformation, and noise. We provide the raw data file RAMDocs_test.jsonl.
Data Fields
Each instance contains the following fields:
question: The question
documents: list of documents, where each document contains the following fields:
text: text of the… See the full description on the dataset page: https://huggingface.co/datasets/HanNight/RAMDocs.ramanv-image-editing
ramanv-image-editing
Image editing dataset for training FLUX.1-Kontext / InstructPix2Pix style models.
Size
592,141 total editing pairs
Sources: ultraedit
Schema
Each shard tar contains {uid}_src.jpg, {uid}_edit.jpg, {uid}_mask.png (where available).
Metadata per record: instruction, prompt, edit_type, caption_before/after, license, sha256.
Licenses
MagicBrush, InstructPix2Pix, Pico-Banana, HumanEdit: CC-BY-4.0
UltraEdit, AnyEdit… See the full description on the dataset page: https://huggingface.co/datasets/lingamvamshikrishnareddy/ramanv-image-editing.ramanv-image-textrenderrameau
Rameau: functional harmony from notation
A text-to-text dataset and benchmark for functional harmony: Roman-numeral
analysis, cadence classification, and key identification. A probabilistic
common-practice grammar generates the progressions; four task framings hide
the answer to increasing degrees. Chord-symbol lookup stops working after the
first one.
Named for Jean-Philippe Rameau, whose Traité de l'harmonie (1722) started
the discipline.
symbol_to_rn key: C major /… See the full description on the dataset page: https://huggingface.co/datasets/4esv/rameau.lm-eval-results-Kukedlc-Ramakrishna-7b-v3-private
Dataset Card for Evaluation run of Kukedlc/Ramakrishna-7b-v3
Dataset automatically created during the evaluation run of model Kukedlc/Ramakrishna-7b-v3
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-Kukedlc-Ramakrishna-7b-v3-private.lex-fridman-podcasts
Dataset Card for Lex Fridman Podcasts Dataset
This dataset is sourced from Andrej Karpathy's Lexicap website which contains English transcripts of Lex Fridman's wonderful podcast episodes. The transcripts were generated using OpenAI's large-sized Whisper model
gordon-ramsay-code-review-v2
Gordon Ramsay Code Review & Auditor Corpus v2 (dcmutlu/gordon-ramsay-code-review-v2)
A high-density synthetic dataset of 10,000 multi-turn code review pairs designed to fine-tune open-weight reasoners (specifically Qwen2.5-Coder-7B-Instruct) into Chef Gordon Ramsay: Sovereign Executive Code Auditor and Supreme Software Gastronomer.
🍳 Dataset Overview
This dataset merges rigorous computer science diagnostics (Abstract Syntax Tree inspection, concurrency lifecycle… See the full description on the dataset page: https://huggingface.co/datasets/dcmutlu/gordon-ramsay-code-review-v2.DeepSeek-V4-Distill-8000x
🐳 DeepSeek-V4-Distill-8100x
Dataset Summary
DeepSeek-V4-Distill-8100x is a supervised fine-tuning dataset for reasoning-oriented distillation. The question prompts come from Jackrong/GLM-5.1-Reasoning-1M-Cleaned, and the answers were generated by the teacher model DeepSeek-V4-Flash.
After the cleaning process, the released train split contains 7,716 high-quality JSONL examples.
[!NOTE]
The answer pool was cleaned to remove real-time questions… See the full description on the dataset page: https://huggingface.co/datasets/rampisipati/DeepSeek-V4-Distill-8000x.rampnet-stage1-inputs
RampNet Stage 1 Inputs
The inputs to the RampNet Stage 1 pipeline, from RampNet: A Two-Stage Pipeline for
Bootstrapping Curb Ramp Detection in Streetscape Images from Open Government Metadata
(O'Meara et al., ICCV'25 CV4A11y workshop, arXiv:2508.09415).
projectsidewalk/rampnet-dataset
is what Stage 1 produced — 214k annotated panoramas. This is what went in.
v1.0-iccv2025 shipped the Stage 1 code without these files. That is not a gap a re-download can
close: the city open-data… See the full description on the dataset page: https://huggingface.co/datasets/projectsidewalk/rampnet-stage1-inputs.ramanv-image-lifestylewinogrande.jsonlWikipedia-Pagesagentconflictbench
AgentConflictBench
AgentConflictBench is a research benchmark for evaluating silent semantic
conflicts among independently valid AI-generated code changes.
Most coding-agent benchmarks ask whether an agent can solve one task in
isolation. AgentConflictBench asks whether two independently valid patches
still work when composed.
Dataset Summary
Instances: 28
Positive silent semantic conflicts: 25
Clean-composition controls: 3
Upstream repositories: 7
Languages:… See the full description on the dataset page: https://huggingface.co/datasets/ramachandra1996/agentconflictbench.ramskimify-ifeval-like
Kimify IFEval-Like Dataset
Dataset Description
This dataset contains 10,070 verified instruction-following conversations in the IFEval format. Each example includes:
A user prompt with embedded constraints
An assistant response that satisfies those constraints
Metadata describing the constraint types and parameters
All examples have been programmatically verified using the instruction-following-eval library (based on Google Research's IFEval) to ensure 100% constraint… See the full description on the dataset page: https://huggingface.co/datasets/ramendik/kimify-ifeval-like.slm-reasoning-baseline-error-taxonomy
SLM Reasoning Research — Baseline failure taxonomy
Part of the SLM Reasoning Research project.
All 631 errors from Qwen3-0.6B-Base's zero-shot, zero-training GSM8K baseline (full 1,319-example test
set, 52.16% accuracy), classified into an 11-category failure taxonomy using GPT-OSS-20B as judge
(verified by hand against a 30-example sample first).
misread_semantics dominates at 55.5% (350/631) — far more than arithmetic_slip (10%) — meaning
this model's core weakness is… See the full description on the dataset page: https://huggingface.co/datasets/Ram20307/slm-reasoning-baseline-error-taxonomy.gordon-ramsay-code-review
gordon-ramsay-code-review
Autonomous synthetic pretraining dataset synthesized by JESUS Sovereign Forge.
Synthesized via JESUS Sovereign Cloud Model Forge (hf-colab-forge) for native byte-level micro-transformers (Atom GPT) and LLM fine-tuning.
Dataset Summary
Metric
Value
Total Scenarios
500
Train Samples
450
Validation Samples
50
Total Byte Tokens
819,927
Train Tokens
737,852
Val Tokens
82,075
Vocab Size
258 (UTF-8 Bytes + BOS/PAD)… See the full description on the dataset page: https://huggingface.co/datasets/dcmutlu/gordon-ramsay-code-review.Wiki_Faiss_Indexes
dataset_info:
features:
- name: text
dtype: string
- name: embeddings
dtype: float32
shape: [384]
configs:
- config_name: default
data_files: "*.parquet"
Wikipedia IVF-OPQ-PQ Vector Database (GPU-Optimized)
A high-performance, GPU-accelerated FAISS vector database built from Wikipedia articles with pre-computed embeddings. This dataset contains approximately 35 million Wikipedia articles with 384-dimensional embeddings using the all-MiniLM-L6-v2… See the full description on the dataset page: https://huggingface.co/datasets/Ram-G/Wiki_Faiss_Indexes.code.evol.instruct.wiz.oss_python.jsonRAMEN-phase1slm-reasoning-dpo-pairs
SLM Reasoning Research — DPO preference pairs
Part of the SLM Reasoning Research project.
Preference pairs for DPO training: chosen = Arm A's teacher trace (verified correct), rejected =
the untrained Qwen3-0.6B-Base model's own natural wrong attempt on that same GSM8K train question —
not a synthetic negative. 7,435 train questions attempted; pairs kept only for the questions the base
model got wrong (39.4%).
Each row: example_id, prompt, chosen, rejected, reference_answer.
See… See the full description on the dataset page: https://huggingface.co/datasets/Ram20307/slm-reasoning-dpo-pairs.RAMP_finetuned_data_Cora
Dataset Details
This dataset contains finetuning data constructed from the Cora citation network for downstream text-rich graph tasks. It is used for finetuning RAMP (Raw-text Anchored Message Passing), which recasts the LLM as a graph-native aggregation operator on text-rich graphs.
The dataset includes the following files:
finetuned_cora_v1.json — Training set
finetuned_cora_val_v1.json — Validation set
eval_cora_v1.json — Test set
This is a release from our paper LLM as Graph… See the full description on the dataset page: https://huggingface.co/datasets/JJYDXFS/RAMP_finetuned_data_Cora.scientific-posttrain-raman-eval
Scientific Post-Training Raman Evaluation Resolver v1
This repository is the canonical provenance manifest for the
raman-bioprocess Harbor task. It intentionally contains no Raman spectra or
labels because the upstream RamanBench mirror requires users to respect each
original dataset's terms and prohibits unapproved redistribution.
manifest.json pins the public Hugging Face source, exact commit, four Parquet
file hashes, deterministic split, and target definitions. During the… See the full description on the dataset page: https://huggingface.co/datasets/aashay96/scientific-posttrain-raman-eval.ramanv-image-captions-6medical-emails-producta-noncase-combo-dataset
Medical Emails Product A and Non-Case Combined Classification Dataset
This dataset contains 800 unique synthetic medical and operational emails in strict JSONL format for email classification training.
Dataset File
medical_emails_producta_noncase_combo_800.jsonl - 800 emails
Classification Categories
The dataset contains 200 unique emails for each combined classification:
Medical Information, Non-Case - Product A medical information request plus a separate… See the full description on the dataset page: https://huggingface.co/datasets/Ramesh10/medical-emails-producta-noncase-combo-dataset.kimify-short-20260131A conversational dataset generated by Kimi K2 0905 Instruct. The user prompts were taken from two datasets:
smoltalk multilingual - English prompts in the "advice-seeking" category
smoltalk - in the "smol-magpie-ultra-short" category. Note these involve three user/assistant turns.
System prompts were used to encourage brevity, for example: "You are Kimi K2, a versatile AI assistant. Be concise, clear, and punchy—aim for brief but helpful responses. Keep your distinctive voice but stay… See the full description on the dataset page: https://huggingface.co/datasets/ramendik/kimify-short-20260131.medical-emails-producta-nonrelevant-combo-dataset
Medical Emails Product A and Non-Relevant Combined Classification Dataset
This dataset contains 800 unique synthetic emails in strict JSONL format for email classification training.
Dataset File
medical_emails_producta_nonrelevant_combo_800.jsonl - 800 emails
Classification Categories
The dataset contains 200 unique emails for each combined classification:
Medical Information, Non-Relevant
Adverse Event, Non-Relevant
Product Complaint, Non-Relevant
Other… See the full description on the dataset page: https://huggingface.co/datasets/Ramesh10/medical-emails-producta-nonrelevant-combo-dataset.aligntune-testrun-alignment-auditmedical-email-dataset-800
Ramesh10/medical-email-dataset-800
Generated by ML Intern
This dataset repository was generated by ML Intern, an agent for machine learning research and development on the Hugging Face Hub.
Try ML Intern: https://smolagents-ml-intern.hf.space
Source code: https://github.com/huggingface/ml-intern
Usage
from datasets import load_dataset
dataset = load_dataset('Ramesh10/medical-email-dataset-800')
transformed_JSON_databricks-dolly-15k.jsonl
Transformed Databricks-Dolly-15k Dataset
Summary
The Transformed Databricks-Dolly-15k dataset is a modification of the original open-source dataset created by Databricks employees, designed to facilitate instruction-following abilities in large language models (LLMs). This version has been specifically adapted to include responses in a JSON format, enhancing its utility for tasks requiring structured output.
Modifications
The primary transformation applied to… See the full description on the dataset page: https://huggingface.co/datasets/ramachetan22/transformed_JSON_databricks-dolly-15k.jsonl.
