CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01aimosprite /prompt-swap-mixed12-5xlr-e1-mxfp4-mergedtabularn<1K0 likes541 downloads6mo agoHugging Face02aimosprite /prompt-swap-mixed12-5xlr-e2-mxfp4-mergedtabularn<1K0 likes537 downloads6mo agoHugging Face03NuTonic /sat-vl-sft-postprocessed-merged-v1 Dataset Summary NuTonic/sat-bbox-metadata-sft-v1 is a metadata-first, procedural VLM SFT dataset built from an existing “sat-bbox” style dataset tree (Sentinel‑2 chips + per-tile JSON metadata sidecars, optionally paired Mapbox stills). The goal is to create high-signal, production-shaped supervision for multimodal chat models: Captioning for satellite chips Grounding (bounding boxes in normalized coordinates) for land-cover regions Class-focused captions and absence checks for… See the full description on the dataset page: https://huggingface.co/datasets/NuTonic/sat-vl-sft-postprocessed-merged-v1.imagetext-generation100K<n<1M0 likes505 downloads5mo agoHugging Face04aimosprite /prompt-swap-medium12-e2-mxfp4-mergedtabularn<1K0 likes360 downloads6mo agoHugging Face05survivi /grad_clip0.28_mergedtext100K<n<1M0 likes342 downloads1y agoHugging Face06w4nn4b3M4ST3R /raw-mergedtabular100K<n<1M0 likes272 downloads2mo agoHugging Face07nyu-dice-lab /lm-eval-results-alnrg2arg-blockchainlabs_7B_merged_test2_4-private Dataset Card for Evaluation run of alnrg2arg/blockchainlabs_7B_merged_test2_4 Dataset automatically created during the evaluation run of model alnrg2arg/blockchainlabs_7B_merged_test2_4 The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-alnrg2arg-blockchainlabs_7B_merged_test2_4-private.tabular100K<n<1M0 likes185 downloads2y agoHugging Face08michaeldinzinger /merged-trecdltexttext-retrieval1M<n<10M0 likes158 downloads1y agoHugging Face09vkehfdl1 /banana-merged Banana-Merged A synthetic multi-page visual question answering dataset with hard negatives, designed for fine-tuning visual document retrievers like ColFlor and ColPali. Dataset Summary Banana-Merged contains 1,100 training samples and 10,054 images (positive pages + hard negative variants). Each sample pairs a multi-page analytical query with a set of document images that collectively contain the answer, plus one or more hard negative documents that look visually and… See the full description on the dataset page: https://huggingface.co/datasets/vkehfdl1/banana-merged.imagevisual-question-answering1K<n<10K0 likes109 downloads5mo agoHugging Face10khtsly /cache_stage1_mergedtabularn<1K0 likes103 downloads15d agoHugging Face11WithinUsAI /fable_5_distillation_merged_cleaned_25k Claude Fable 5 Distillation Dataset 25,719 high-quality distilled examples for training LLMs to mimic Claude Fable 5's reasoning style — featuring multi-step chain-of-thought with <think> tags across 23+ technical domains. This dataset captures the distinctive reasoning patterns of Claude Fable 5 (Anthropic's Mythos-class model released June 2026): systematic decomposition, first-principles analysis, self-verification, alternative consideration, and synthesis.… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/fable_5_distillation_merged_cleaned_25k.text10K<n<100K6 likes95 downloads3mo agoHugging Face12professorsynapse /claudesidian-behaviors-merged Claudesidian Merged Behavioral Dataset Dataset Description This dataset contains 1,852 synthetic training examples demonstrating 8 different behavioral patterns for training language models to use the Claudesidian-MCP toolset effectively with Obsidian vaults. The dataset is specifically formatted for KTO (Kahneman-Tversky Optimization) preference learning with properly interleaved positive and negative examples. Behavioral Categories This dataset includes… See the full description on the dataset page: https://huggingface.co/datasets/professorsynapse/claudesidian-behaviors-merged.texttext-generation1K<n<10K0 likes91 downloads10mo agoHugging Face13prashantss1404 /Matplotlib_Seaborn_merged_prompt_completion_10ktext1K<n<10K0 likes68 downloads1y agoHugging Face14flatlander1024 /math_merged_cot_sol_pair_mixedPair Type Breakdown: Correct-Incorrect (C-I) Pairs: 5500 C-I with correct first ('[1]'): 2750 C-I with correct second ('[2]'): 2750 Correct-Correct (C-C) Pairs (Target: 2750, Max Diff: 150): 2750 C-C pairs from 'all_correct' problems: 906 Incorrect-Incorrect (I-I) Pairs (Target: 2750, Max Diff: 150): 2750 I-I pairs from 'all_incorrect' problems: 1156 text10K<n<100K0 likes63 downloads1y agoHugging Face15aimosprite /prompt-swap-medium12-e1-mxfp4-mergedtabularn<1K0 likes57 downloads6mo agoHugging Face16flatlander1024 /math_mergedTraining dataset contains aime (excluding 2024), math/train, math/test, openai_math_splits/train, and KbsdJames/Omni-MATH/test. Total of 17521 lines of unique problems. Testing dataset contains aime_24 and math500 (i.e. openai_math_splits/test). Total of 530 lines of unique problems. textquestion-answering10K<n<100K0 likes54 downloads1y agoHugging Face17emanubiz /opus-agent-merged-v2text1K<n<10K0 likes53 downloads5mo agoHugging Face18AiAF /Cleaned-sharegpt_Merged-Opus-33159-ShareGPTtexttext-generation10K<n<100K2 likes52 downloads6mo agoHugging Face19open-llm-leaderboard /Dans-DiscountModels__mistral-7b-test-merged-detailsgated Dataset Card for Evaluation run of Dans-DiscountModels/mistral-7b-test-merged Dataset automatically created during the evaluation run of model Dans-DiscountModels/mistral-7b-test-merged The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Dans-DiscountModels__mistral-7b-test-merged-details.tabular10K<n<100K0 likes45 downloads2y agoHugging Face20pre-to-post-olmo /openthoughts_merged_think_39k openthoughts_merged_think_39k Merged think-format SFT dataset (ShareGPT-style: system + conversations with from/value), 39,874 examples, for OLMo SFT. Composition (concatenation of two decontaminated think-format sources): open-thoughts114k_math_20k_decontam_think — 20,000 examples sampled from OpenThoughts-114k math, decontaminated against the OpenThoughts3 set below. openthoughts3_math_decontam_resp_lt8192_think — 19,874 examples from… See the full description on the dataset page: https://huggingface.co/datasets/pre-to-post-olmo/openthoughts_merged_think_39k.texttext-generation10K<n<100K0 likes42 downloads3mo agoHugging Face21open-llm-leaderboard /FlofloB__40k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-detailsgated Dataset Card for Evaluation run of FlofloB/40k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit Dataset automatically created during the evaluation run of model FlofloB/40k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__40k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-details.tabular10K<n<100K0 likes41 downloads2y agoHugging Face22sabin1234 /Merged_Nepali_Health_FAQ_Dataset Merged Nepali Health FAQ Dataset (Multi-Source) Overview This dataset (merged_nepali_sharegpt_after_removing_11_groups_keep_ids.jsonl) is a merged collection of 65 instruction-following conversation pairs in Nepali, combining Q&A content from 8 distinct real-world Nepali health institutions and organizations into a single ShareGPT-style file. Each record is a single-turn human↔gpt exchange: a Nepali-language question followed by a factual Nepali-language answer.… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/Merged_Nepali_Health_FAQ_Dataset.textn<1K0 likes41 downloads15d agoHugging Face23open-llm-leaderboard /FlofloB__100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-detailsgated Dataset Card for Evaluation run of FlofloB/100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit Dataset automatically created during the evaluation run of model FlofloB/100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-details.tabular10K<n<100K0 likes40 downloads2y agoHugging Face24jpacifico /merged-admin-def-dataset-16ktext10K<n<100K0 likes40 downloads2y agoHugging Face25emanubiz /opus-claude-mergedtext10K<n<100K1 likes40 downloads6mo agoHugging Face26voidful /gemini-3.1-opus-4.6-reasoning-merged Merged from https://huggingface.co/datasets/Roman1111111/gemini-3.1-pro-hard-high-reasoning https://huggingface.co/datasets/reedmayhew/gemini-3.1-pro-2048-reasoning-1100x https://huggingface.co/datasets/crownelius/Opus-4.6-Reasoning-3300x text1K<n<10K6 likes39 downloads7mo agoHugging Face27QuixiAI /WizardLM_evol_instruct_V2_196k_unfiltered_merged_splittext100K<n<1M38 likes37 downloads3y agoHugging Face28hsanyyasyn97gmail /transcripts-mergedtextn<1K0 likes37 downloads14d agoHugging Face29prithivMLmods /PyThagoreans-Merged PyThagoreans Dataset Overview The PyThagoreans dataset is a comprehensive collection of math problems and their solutions, designed to assist in learning and practicing mathematical problem-solving. This dataset includes a variety of problems, expected answers, and predicted answers, making it a valuable resource for students, educators, and researchers. Dataset Details Modalities Text: The dataset primarily contains text data, including math… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/PyThagoreans-Merged.textquestion-answering1M<n<10M2 likes34 downloads2y agoHugging Face30flatlander1024 /math_merged_cot_solA dataset consists problems from flatlander1024/math_merged and cot solutions generated by Llama-3.1-8b-Instruct. The is_correct label indicates whether the solution is correct or not. Number of lines: 13864, Overall correct rate: 57.3% textquestion-answering10K<n<100K0 likes34 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.