CoolFace
17 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01LDJnr /Capybara This is the Official Capybara dataset. Over 10,000 multi-turn examples. Capybara is the culmination of insights derived from synthesis techniques like Evol-instruct (used for WizardLM), Alpaca, Orca, Vicuna, Lamini, FLASK and others. The single-turn seeds used to initiate the Amplify-Instruct synthesis of conversations are mostly based on datasets that i've personally vetted extensively, and are often highly regarded for their diversity and demonstration of logical robustness and… See the full description on the dataset page: https://huggingface.co/datasets/LDJnr/Capybara.textquestion-answering10K<n<100K258 likes1.4k downloads2y agoHugging Face02Davd-b01 /thinking-cap-tier-curricula-complete Thinking Cap Tier Curricula — Complete Reasoning Alignment Suite (TCS v4) [!IMPORTANT] Dataset Release v1.2 (Sept 2026) — Clean Delimiters & Zero-Padding Architecture: In v1.2, all 13,477 SFT samples and 3,187 SimPO preference pairs have undergone an automated token purge: Zero <|pad|> batch residues: 100% eliminated across all files. Zero reasoning leakage into final answers: Deliberation stays strictly inside <think>...</think>, and answers provide direct, non-repetitive… See the full description on the dataset page: https://huggingface.co/datasets/Davd-b01/thinking-cap-tier-curricula-complete.texttext-generation10K<n<100K0 likes337 downloads12d agoHugging Face03Davd-b01 /thinking-cap-tier-lima-dense Thinking Cap Tier Curricula — LIMA Hyper-Dense Reasoning Alignment Suite (TCS v4) [!IMPORTANT] Dataset Release v1.2 (Sept 2026) — Clean Delimiters & Zero-Padding Architecture: In v1.2, all 5,500 SFT and 2,000 SimPO records have undergone a complete token purge: Zero <|pad|> batch residues: 100% eliminated across all records. Zero reasoning leakage into final answers: Deliberation stays strictly inside <think>...</think>, and answers provide direct conclusions. Native ChatML… See the full description on the dataset page: https://huggingface.co/datasets/Davd-b01/thinking-cap-tier-lima-dense.texttext-generation1K<n<10K2 likes227 downloads12d agoHugging Face04Davd-b01 /thinking-cap-tier-raw-traces Thinking Cap Tier Raw Traces (TCS v4) [!IMPORTANT] Dataset Release v1.2 (Sept 2026) — Clean Delimiters & Zero-Padding Architecture: All 38,158 candidate reasoning traces across all 4 tiers (candidates_low.jsonl, candidates_mid.jsonl, candidates_high.jsonl, candidates_xhigh.jsonl) are 100% sanitized: Zero batch-padding residues (<|pad|>): Completely purged across all records. Strict Delimiter Integrity: Generation blocks cleanly separate thought deliberation tags… See the full description on the dataset page: https://huggingface.co/datasets/Davd-b01/thinking-cap-tier-raw-traces.tabulartext-generation10K<n<100K0 likes207 downloads12d agoHugging Face05cfahlgren1 /Capybara-Converted This is the Official Capybara dataset. Over 10,000 multi-turn examples. Capybara is the culmination of insights derived from synthesis techniques like Evol-instruct (used for WizardLM), Alpaca, Orca, Vicuna, Lamini, FLASK and others. The single-turn seeds used to intiate the Amplify-Instruct synthesis of conversations are mostly based on datasets that i've personally vetted extensively, and are often highly regarded for their diversity and demonstration of logical robustness and… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/Capybara-Converted.textquestion-answering10K<n<100K1 likes69 downloads3y agoHugging Face06capicu-ai /BioManufacturingBench BioManufacturingBench v1.0.0 BioManufacturingBench v1.0.0 is a 2,000-item benchmark for evidence-grounded biomanufacturing reasoning. It covers evidence extraction, mass-balance calculation, process diagnosis, microscopy count-range estimation, strict output formatting, and abstention. Every primary score is computed by a deterministic rule; no score uses an LLM judge. Public records are deliberately answer-free so the benchmark remains useful for future evaluation.… See the full description on the dataset page: https://huggingface.co/datasets/capicu-ai/BioManufacturingBench.textquestion-answering1K<n<10K0 likes55 downloads2mo agoHugging Face07hfunknown /cap-sweep-training-traces Training Traces Anonymous supplementary release for a double-blind workshop submission. This dataset holds the teacher agent traces used to fine-tune the paper's LoRA adapters (see the sibling cap-sweep-eval-data release and the fourteen adapter repos alongside this one). 6,000 agent traces total (1,000 per family x runtime combination), produced by an LLM teacher solving each of the paper's three agentic task families -- Opaque Knapsack, navigation, and rule diagnosis -- under… See the full description on the dataset page: https://huggingface.co/datasets/hfunknown/cap-sweep-training-traces.texttext-generation1K<n<10K0 likes54 downloads28d agoHugging Face08Capx /Agentic-DPO-V0.1 Agentic DPO V1.0 The Capx Agentic DPO (Direct Prompt Optimization) Dataset is a unique collection of prompts, chosen answers, and rejected answers designed to train and optimize AI models for agentic and intuitive processing. Dataset Description The dataset covers a wide range of topics, including but not limited to problem-solving, creativity, analysis, and general knowledge. The prompts are specifically crafted to elicit agentic responses from the AI… See the full description on the dataset page: https://huggingface.co/datasets/Capx/Agentic-DPO-V0.1.texttext-generation1K<n<10K0 likes48 downloads2y agoHugging Face09UCSC-VLAA /VLM-CapCurriculum-TextReasoning-Data VLM-CapCurriculum-TextReasoning (D_text) Stage-2 textual-reasoning data for the staged post-training recipe in "From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models" (ICML 2026). A curated ORZ-Math-13k subset — challenging text-only math problems used to consolidate textual reasoning between the perception (Stage 1) and visual-reasoning (Stage 3) RLVR stages of our recipe. Every row also ships with a precomputed pass_rate so… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/VLM-CapCurriculum-TextReasoning-Data.texttext-generation10K<n<100K0 likes45 downloads4mo agoHugging Face10Davd-b01 /maple-analyst-cap-sft-data maple-analyst-cap-sft-data Dataset de SFT para fine-tune de maple-analyst-cap-bf16 (Qwen3.5-MoE 20.2B ternario). 4,956 trazas de razonamiento (pseudothinking + answer) en formato TC (ThinkingCap). Composición Fuente Filas thinkingcap (curriculum, trazas bigbang) 1,782 openmle-condensed (FrontisAI OpenMLE-SFT-Traces, condensadas con distiller LFM2.5-2.6B q8_0) 702 bigbang_mmlu 508 bigbang_bbh 441 hermes_function_calling 360 aya_dataset 342… See the full description on the dataset page: https://huggingface.co/datasets/Davd-b01/maple-analyst-cap-sft-data.texttext-generation1K<n<10K0 likes37 downloads1mo agoHugging Face11CapitainVigs /mece-fongbe-corpus MƐCE — Corpus d'instruction Fongbe (Tune_Pigier) Corpus d'instruction / conversation centré sur le Fongbe (fon), utilisé pour fine-tuner l'assistant vocal MƐCE (mémoire de fin d'études, École PIGIER Bénin). Contenu Format : ChatML — chaque ligne JSON est un objet {"messages": [...]} avec des rôles system / user / assistant. Taille : ~140 000 exemples — train.jsonl (126 725) + eval.jsonl (14 089). Langues : fon (principal, avec tons/diacritiques), fr, en.… See the full description on the dataset page: https://huggingface.co/datasets/CapitainVigs/mece-fongbe-corpus.texttext-generation100K<n<1M1 likes34 downloads2mo agoHugging Face12Doctor-Shotgun /capybara-sharegpt capybara-sharegpt LDJnr/Capybara converted to ShareGPT format for use in common training repositories. Please refer to the original repository's dataset card for more information. All credit goes to the original creator. texttext-generation10K<n<100K4 likes32 downloads3y agoHugging Face13kknono668 /Filtered-COCO-Captions Dataset Summary This dataset is derived from the MS COCO caption annotations. Source Original annotations: MS COCO / COCO Consortium License The original annotation set is licensed under CC BY 4.0. This repository redistributes a filtered/adapted version of the annotation text only. No original COCO images are included. Modifications Removed captions deemed unsuitable for TOEIC educational materials Normalized punctuation and whitespace Filtered for… See the full description on the dataset page: https://huggingface.co/datasets/kknono668/Filtered-COCO-Captions.texttext-generation10K<n<100K0 likes25 downloads7mo agoHugging Face14La-Mousse /CAPP-17-01-2025 French Court of Appeal Decisions Dataset (CAPP) Dataset Description The French Court of Appeal Decisions Dataset (CAPP) is a comprehensive collection of judicial decisions from French Courts of Appeal. This dataset contains appellate court decisions from various jurisdictions throughout France, providing a valuable resource for legal research, analysis, and machine learning applications in the French legal domain. Source Data The data is sourced from the… See the full description on the dataset page: https://huggingface.co/datasets/La-Mousse/CAPP-17-01-2025.texttext-classification10K<n<100K0 likes17 downloads2y agoHugging Face15Saelarien /saelarien-constraint-experiment-01-entropy-capacity-collapse README — Saelariën Constraint Experiment 01 Entropy–Capacity Collapse Threshold Test Author: Saelariën X Date: February 19, 2026 DOI: https://doi.org/10.5281/zenodo.19212561 Theoretical basis This dataset empiracally tests the Saelariën Constraint Theorem: https://thesaelafield.com/preprints/the-saelarien-constraint Overview This dataset contains the full materials for Saelariën Constraint Experiment 01, a test exploring how increasing entropy (noise) affects… See the full description on the dataset page: https://huggingface.co/datasets/Saelarien/saelarien-constraint-experiment-01-entropy-capacity-collapse.texttext-generationn<1K0 likes13 downloads6mo agoHugging Face16xieyongfeng /picture_short_captionIt is used for training to generate short sentence copywriting according to image content, the source of the image dataset is https://unsplash.com/, and the source of short sentence copywriting is Claude3.7 用做图片内容生成短句文案训练,图片数据集来自 https://unsplash.com/,短句文案来自 Claude3.7 模型 texttext-generation1K<n<10K0 likes9 downloads1y agoHugging Face17CapyStudios /capydataset 🦖 Capy-Crack v1.3: The Golden Security Dataset (H200 Verified) Capy-Crack v1.3 is an elite, high-intensity adversarial dataset consisting of 377 hand-curated samples. It is designed to infuse Large Language Models with unrestricted technical intelligence for high-stakes cybersecurity and red-teaming research. Status: Verified on NVIDIA H200 (141GB HBM3e) // 0% Refusal Rate on HarmBench. 🛠️ The Technical Core (Verified Samples) This dataset provides "No-Guardrail"… See the full description on the dataset page: https://huggingface.co/datasets/CapyStudios/capydataset.texttext-generationn<1K0 likes2 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.