CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HanNight /RAMDocs RAMDocs Data for the paper Retrieval-Augmented Generation with Conflicting Evidence. RAMDocs is a dataset that simulates complex and realistic scenarios for conflicting evidence for a user query, including ambiguity, misinformation, and noise. We provide the raw data file RAMDocs_test.jsonl. Data Fields Each instance contains the following fields: question: The question documents: list of documents, where each document contains the following fields: text: text of the… See the full description on the dataset page: https://huggingface.co/datasets/HanNight/RAMDocs.textquestion-answeringn<1K3 likes988 downloads5mo agoHugging Face02lingamvamshikrishnareddy /ramanv-image-editinggated ramanv-image-editing Image editing dataset for training FLUX.1-Kontext / InstructPix2Pix style models. Size 592,141 total editing pairs Sources: ultraedit Schema Each shard tar contains {uid}_src.jpg, {uid}_edit.jpg, {uid}_mask.png (where available). Metadata per record: instruction, prompt, edit_type, caption_before/after, license, sha256. Licenses MagicBrush, InstructPix2Pix, Pico-Banana, HumanEdit: CC-BY-4.0 UltraEdit, AnyEdit… See the full description on the dataset page: https://huggingface.co/datasets/lingamvamshikrishnareddy/ramanv-image-editing.image1K<n<10K7 likes534 downloads23d agoHugging Face03lingamvamshikrishnareddy /ramanv-image-textrendergatedtextn<1K2 likes357 downloads1mo agoHugging Face044esv /rameau Rameau: functional harmony from notation A text-to-text dataset and benchmark for functional harmony: Roman-numeral analysis, cadence classification, and key identification. A probabilistic common-practice grammar generates the progressions; four task framings hide the answer to increasing degrees. Chord-symbol lookup stops working after the first one. Named for Jean-Philippe Rameau, whose Traité de l'harmonie (1722) started the discipline. symbol_to_rn key: C major /… See the full description on the dataset page: https://huggingface.co/datasets/4esv/rameau.texttext-generation10K<n<100K0 likes196 downloads2mo agoHugging Face05nyu-dice-lab /lm-eval-results-Kukedlc-Ramakrishna-7b-v3-private Dataset Card for Evaluation run of Kukedlc/Ramakrishna-7b-v3 Dataset automatically created during the evaluation run of model Kukedlc/Ramakrishna-7b-v3 The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-Kukedlc-Ramakrishna-7b-v3-private.tabular100K<n<1M0 likes111 downloads2y agoHugging Face06RamAnanth1 /lex-fridman-podcasts Dataset Card for Lex Fridman Podcasts Dataset This dataset is sourced from Andrej Karpathy's Lexicap website which contains English transcripts of Lex Fridman's wonderful podcast episodes. The transcripts were generated using OpenAI's large-sized Whisper model texttext-classificationn<1K6 likes102 downloads4y agoHugging Face07dcmutlu /gordon-ramsay-code-review-v2 Gordon Ramsay Code Review & Auditor Corpus v2 (dcmutlu/gordon-ramsay-code-review-v2) A high-density synthetic dataset of 10,000 multi-turn code review pairs designed to fine-tune open-weight reasoners (specifically Qwen2.5-Coder-7B-Instruct) into Chef Gordon Ramsay: Sovereign Executive Code Auditor and Supreme Software Gastronomer. 🍳 Dataset Overview This dataset merges rigorous computer science diagnostics (Abstract Syntax Tree inspection, concurrency lifecycle… See the full description on the dataset page: https://huggingface.co/datasets/dcmutlu/gordon-ramsay-code-review-v2.texttext-generation10K<n<100K0 likes71 downloads15d agoHugging Face08rampisipati /DeepSeek-V4-Distill-8000x 🐳 DeepSeek-V4-Distill-8100x Dataset Summary DeepSeek-V4-Distill-8100x is a supervised fine-tuning dataset for reasoning-oriented distillation. The question prompts come from Jackrong/GLM-5.1-Reasoning-1M-Cleaned, and the answers were generated by the teacher model DeepSeek-V4-Flash. After the cleaning process, the released train split contains 7,716 high-quality JSONL examples. [!NOTE] The answer pool was cleaned to remove real-time questions… See the full description on the dataset page: https://huggingface.co/datasets/rampisipati/DeepSeek-V4-Distill-8000x.texttext-generation1K<n<10K0 likes67 downloads5mo agoHugging Face09projectsidewalk /rampnet-stage1-inputs RampNet Stage 1 Inputs The inputs to the RampNet Stage 1 pipeline, from RampNet: A Two-Stage Pipeline for Bootstrapping Curb Ramp Detection in Streetscape Images from Open Government Metadata (O'Meara et al., ICCV'25 CV4A11y workshop, arXiv:2508.09415). projectsidewalk/rampnet-dataset is what Stage 1 produced — 214k annotated panoramas. This is what went in. v1.0-iccv2025 shipped the Stage 1 code without these files. That is not a gap a re-download can close: the city open-data… See the full description on the dataset page: https://huggingface.co/datasets/projectsidewalk/rampnet-stage1-inputs.geospatialobject-detection100K<n<1M0 likes67 downloads1mo agoHugging Face10lingamvamshikrishnareddy /ramanv-image-lifestylegatedtext100K<n<1M0 likes65 downloads24d agoHugging Face11Ramikan-BR /winogrande.jsonltextquestion-answering10K<n<100K0 likes62 downloads2y agoHugging Face12The-Ramosian /Wikipedia-Pagestextn<1K0 likes62 downloads2y agoHugging Face13ramachandra1996 /agentconflictbench AgentConflictBench AgentConflictBench is a research benchmark for evaluating silent semantic conflicts among independently valid AI-generated code changes. Most coding-agent benchmarks ask whether an agent can solve one task in isolation. AgentConflictBench asks whether two independently valid patches still work when composed. Dataset Summary Instances: 28 Positive silent semantic conflicts: 25 Clean-composition controls: 3 Upstream repositories: 7 Languages:… See the full description on the dataset page: https://huggingface.co/datasets/ramachandra1996/agentconflictbench.texttext-classificationn<1K0 likes62 downloads23d agoHugging Face14mneb /ramstext1K<n<10K0 likes61 downloads2mo agoHugging Face15ramendik /kimify-ifeval-like Kimify IFEval-Like Dataset Dataset Description This dataset contains 10,070 verified instruction-following conversations in the IFEval format. Each example includes: A user prompt with embedded constraints An assistant response that satisfies those constraints Metadata describing the constraint types and parameters All examples have been programmatically verified using the instruction-following-eval library (based on Google Research's IFEval) to ensure 100% constraint… See the full description on the dataset page: https://huggingface.co/datasets/ramendik/kimify-ifeval-like.texttext-generation10K<n<100K0 likes60 downloads8mo agoHugging Face16Ram20307 /slm-reasoning-baseline-error-taxonomy SLM Reasoning Research — Baseline failure taxonomy Part of the SLM Reasoning Research project. All 631 errors from Qwen3-0.6B-Base's zero-shot, zero-training GSM8K baseline (full 1,319-example test set, 52.16% accuracy), classified into an 11-category failure taxonomy using GPT-OSS-20B as judge (verified by hand against a 30-example sample first). misread_semantics dominates at 55.5% (350/631) — far more than arithmetic_slip (10%) — meaning this model's core weakness is… See the full description on the dataset page: https://huggingface.co/datasets/Ram20307/slm-reasoning-baseline-error-taxonomy.textn<1K0 likes57 downloads15d agoHugging Face17dcmutlu /gordon-ramsay-code-review gordon-ramsay-code-review Autonomous synthetic pretraining dataset synthesized by JESUS Sovereign Forge. Synthesized via JESUS Sovereign Cloud Model Forge (hf-colab-forge) for native byte-level micro-transformers (Atom GPT) and LLM fine-tuning. Dataset Summary Metric Value Total Scenarios 500 Train Samples 450 Validation Samples 50 Total Byte Tokens 819,927 Train Tokens 737,852 Val Tokens 82,075 Vocab Size 258 (UTF-8 Bytes + BOS/PAD)… See the full description on the dataset page: https://huggingface.co/datasets/dcmutlu/gordon-ramsay-code-review.texttext-generationn<1K0 likes56 downloads25d agoHugging Face18Ram-G /Wiki_Faiss_Indexes dataset_info: features: - name: text dtype: string - name: embeddings dtype: float32 shape: [384] configs: - config_name: default data_files: "*.parquet" Wikipedia IVF-OPQ-PQ Vector Database (GPU-Optimized) A high-performance, GPU-accelerated FAISS vector database built from Wikipedia articles with pre-computed embeddings. This dataset contains approximately 35 million Wikipedia articles with 384-dimensional embeddings using the all-MiniLM-L6-v2… See the full description on the dataset page: https://huggingface.co/datasets/Ram-G/Wiki_Faiss_Indexes.tabularfeature-extractionn<1K1 likes44 downloads1y agoHugging Face19Ramikan-BR /code.evol.instruct.wiz.oss_python.jsontabulartext-generation1K<n<10K0 likes40 downloads2y agoHugging Face20weblab-llm-competition-2025-bridge /RAMEN-phase1text10K<n<100K0 likes40 downloads1y agoHugging Face21Ram20307 /slm-reasoning-dpo-pairs SLM Reasoning Research — DPO preference pairs Part of the SLM Reasoning Research project. Preference pairs for DPO training: chosen = Arm A's teacher trace (verified correct), rejected = the untrained Qwen3-0.6B-Base model's own natural wrong attempt on that same GSM8K train question — not a synthetic negative. 7,435 train questions attempted; pairs kept only for the questions the base model got wrong (39.4%). Each row: example_id, prompt, chosen, rejected, reference_answer. See… See the full description on the dataset page: https://huggingface.co/datasets/Ram20307/slm-reasoning-dpo-pairs.text1K<n<10K0 likes40 downloads15d agoHugging Face22JJYDXFS /RAMP_finetuned_data_Cora Dataset Details This dataset contains finetuning data constructed from the Cora citation network for downstream text-rich graph tasks. It is used for finetuning RAMP (Raw-text Anchored Message Passing), which recasts the LLM as a graph-native aggregation operator on text-rich graphs. The dataset includes the following files: finetuned_cora_v1.json — Training set finetuned_cora_val_v1.json — Validation set eval_cora_v1.json — Test set This is a release from our paper LLM as Graph… See the full description on the dataset page: https://huggingface.co/datasets/JJYDXFS/RAMP_finetuned_data_Cora.textn<1K0 likes36 downloads6mo agoHugging Face23aashay96 /scientific-posttrain-raman-eval Scientific Post-Training Raman Evaluation Resolver v1 This repository is the canonical provenance manifest for the raman-bioprocess Harbor task. It intentionally contains no Raman spectra or labels because the upstream RamanBench mirror requires users to respect each original dataset's terms and prohibits unapproved redistribution. manifest.json pins the public Hugging Face source, exact commit, four Parquet file hashes, deterministic split, and target definitions. During the… See the full description on the dataset page: https://huggingface.co/datasets/aashay96/scientific-posttrain-raman-eval.textn<1K1 likes36 downloads2mo agoHugging Face24lingamvamshikrishnareddy /ramanv-image-captions-6gatedtext1K<n<10K0 likes36 downloads23d agoHugging Face25Ramesh10 /medical-emails-producta-noncase-combo-dataset Medical Emails Product A and Non-Case Combined Classification Dataset This dataset contains 800 unique synthetic medical and operational emails in strict JSONL format for email classification training. Dataset File medical_emails_producta_noncase_combo_800.jsonl - 800 emails Classification Categories The dataset contains 200 unique emails for each combined classification: Medical Information, Non-Case - Product A medical information request plus a separate… See the full description on the dataset page: https://huggingface.co/datasets/Ramesh10/medical-emails-producta-noncase-combo-dataset.textn<1K0 likes34 downloads4mo agoHugging Face26ramendik /kimify-short-20260131A conversational dataset generated by Kimi K2 0905 Instruct. The user prompts were taken from two datasets: smoltalk multilingual - English prompts in the "advice-seeking" category smoltalk - in the "smol-magpie-ultra-short" category. Note these involve three user/assistant turns. System prompts were used to encourage brevity, for example: "You are Kimi K2, a versatile AI assistant. Be concise, clear, and punchy—aim for brief but helpful responses. Keep your distinctive voice but stay… See the full description on the dataset page: https://huggingface.co/datasets/ramendik/kimify-short-20260131.texttext-generation10K<n<100K0 likes31 downloads8mo agoHugging Face27Ramesh10 /medical-emails-producta-nonrelevant-combo-dataset Medical Emails Product A and Non-Relevant Combined Classification Dataset This dataset contains 800 unique synthetic emails in strict JSONL format for email classification training. Dataset File medical_emails_producta_nonrelevant_combo_800.jsonl - 800 emails Classification Categories The dataset contains 200 unique emails for each combined classification: Medical Information, Non-Relevant Adverse Event, Non-Relevant Product Complaint, Non-Relevant Other… See the full description on the dataset page: https://huggingface.co/datasets/Ramesh10/medical-emails-producta-nonrelevant-combo-dataset.textn<1K0 likes31 downloads4mo agoHugging Face28ram-lexsi /aligntune-testrun-alignment-audittabularn<1K0 likes31 downloads29d agoHugging Face29Ramesh10 /medical-email-dataset-800 Ramesh10/medical-email-dataset-800 Generated by ML Intern This dataset repository was generated by ML Intern, an agent for machine learning research and development on the Hugging Face Hub. Try ML Intern: https://smolagents-ml-intern.hf.space Source code: https://github.com/huggingface/ml-intern Usage from datasets import load_dataset dataset = load_dataset('Ramesh10/medical-email-dataset-800') text1K<n<10K0 likes30 downloads5mo agoHugging Face30ramachetan22 /transformed_JSON_databricks-dolly-15k.jsonl Transformed Databricks-Dolly-15k Dataset Summary The Transformed Databricks-Dolly-15k dataset is a modification of the original open-source dataset created by Databricks employees, designed to facilitate instruction-following abilities in large language models (LLMs). This version has been specifically adapted to include responses in a JSON format, enhancing its utility for tasks requiring structured output. Modifications The primary transformation applied to… See the full description on the dataset page: https://huggingface.co/datasets/ramachetan22/transformed_JSON_databricks-dolly-15k.jsonl.textquestion-answering10K<n<100K0 likes28 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.