datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Capybara
This is the Official Capybara dataset. Over 10,000 multi-turn examples.
Capybara is the culmination of insights derived from synthesis techniques like Evol-instruct (used for WizardLM), Alpaca, Orca, Vicuna, Lamini, FLASK and others.
The single-turn seeds used to initiate the Amplify-Instruct synthesis of conversations are mostly based on datasets that i've personally vetted extensively, and are often highly regarded for their diversity and demonstration of logical robustness and… See the full description on the dataset page: https://huggingface.co/datasets/LDJnr/Capybara.thinking-cap-tier-curricula-complete
Thinking Cap Tier Curricula — Complete Reasoning Alignment Suite (TCS v4)
[!IMPORTANT]
Dataset Release v1.2 (Sept 2026) — Clean Delimiters & Zero-Padding Architecture:
In v1.2, all 13,477 SFT samples and 3,187 SimPO preference pairs have undergone an automated token purge:
Zero <|pad|> batch residues: 100% eliminated across all files.
Zero reasoning leakage into final answers: Deliberation stays strictly inside <think>...</think>, and answers provide direct, non-repetitive… See the full description on the dataset page: https://huggingface.co/datasets/Davd-b01/thinking-cap-tier-curricula-complete.thinking-cap-tier-lima-dense
Thinking Cap Tier Curricula — LIMA Hyper-Dense Reasoning Alignment Suite (TCS v4)
[!IMPORTANT]
Dataset Release v1.2 (Sept 2026) — Clean Delimiters & Zero-Padding Architecture:
In v1.2, all 5,500 SFT and 2,000 SimPO records have undergone a complete token purge:
Zero <|pad|> batch residues: 100% eliminated across all records.
Zero reasoning leakage into final answers: Deliberation stays strictly inside <think>...</think>, and answers provide direct conclusions.
Native ChatML… See the full description on the dataset page: https://huggingface.co/datasets/Davd-b01/thinking-cap-tier-lima-dense.thinking-cap-tier-raw-traces
Thinking Cap Tier Raw Traces (TCS v4)
[!IMPORTANT]
Dataset Release v1.2 (Sept 2026) — Clean Delimiters & Zero-Padding Architecture:
All 38,158 candidate reasoning traces across all 4 tiers (candidates_low.jsonl, candidates_mid.jsonl, candidates_high.jsonl, candidates_xhigh.jsonl) are 100% sanitized:
Zero batch-padding residues (<|pad|>): Completely purged across all records.
Strict Delimiter Integrity: Generation blocks cleanly separate thought deliberation tags… See the full description on the dataset page: https://huggingface.co/datasets/Davd-b01/thinking-cap-tier-raw-traces.Capybara-Converted
This is the Official Capybara dataset. Over 10,000 multi-turn examples.
Capybara is the culmination of insights derived from synthesis techniques like Evol-instruct (used for WizardLM), Alpaca, Orca, Vicuna, Lamini, FLASK and others.
The single-turn seeds used to intiate the Amplify-Instruct synthesis of conversations are mostly based on datasets that i've personally vetted extensively, and are often highly regarded for their diversity and demonstration of logical robustness and… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/Capybara-Converted.BioManufacturingBench
BioManufacturingBench v1.0.0
BioManufacturingBench v1.0.0 is a 2,000-item benchmark for evidence-grounded
biomanufacturing reasoning. It covers evidence extraction, mass-balance calculation,
process diagnosis, microscopy count-range estimation, strict output formatting, and
abstention. Every primary score is computed by a deterministic rule; no score uses an
LLM judge. Public records are deliberately answer-free so the benchmark remains useful
for future evaluation.… See the full description on the dataset page: https://huggingface.co/datasets/capicu-ai/BioManufacturingBench.cap-sweep-training-traces
Training Traces
Anonymous supplementary release for a double-blind workshop submission. This dataset
holds the teacher agent traces used to fine-tune the paper's LoRA adapters (see the
sibling cap-sweep-eval-data release and the fourteen adapter repos alongside this one).
6,000 agent traces total (1,000 per family x runtime combination), produced by an LLM
teacher solving each of the paper's three agentic task families -- Opaque Knapsack,
navigation, and rule diagnosis -- under… See the full description on the dataset page: https://huggingface.co/datasets/hfunknown/cap-sweep-training-traces.Agentic-DPO-V0.1
Agentic DPO V1.0
The Capx Agentic DPO (Direct Prompt Optimization) Dataset is a unique collection of prompts, chosen answers, and rejected answers designed to train and optimize AI models for agentic and intuitive processing.
Dataset Description
The dataset covers a wide range of topics, including but not limited to problem-solving, creativity, analysis, and general knowledge. The prompts are specifically crafted to elicit agentic responses from the AI… See the full description on the dataset page: https://huggingface.co/datasets/Capx/Agentic-DPO-V0.1.VLM-CapCurriculum-TextReasoning-Data
VLM-CapCurriculum-TextReasoning (D_text)
Stage-2 textual-reasoning data for the staged post-training recipe in
"From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models"
(ICML 2026).
A curated ORZ-Math-13k subset — challenging text-only math problems used to consolidate textual reasoning between the perception (Stage 1) and visual-reasoning (Stage 3) RLVR stages of our recipe. Every row also ships with a precomputed pass_rate so… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/VLM-CapCurriculum-TextReasoning-Data.maple-analyst-cap-sft-data
maple-analyst-cap-sft-data
Dataset de SFT para fine-tune de maple-analyst-cap-bf16 (Qwen3.5-MoE 20.2B
ternario). 4,956 trazas de razonamiento (pseudothinking + answer) en formato
TC (ThinkingCap).
Composición
Fuente
Filas
thinkingcap (curriculum, trazas bigbang)
1,782
openmle-condensed (FrontisAI OpenMLE-SFT-Traces, condensadas con distiller LFM2.5-2.6B q8_0)
702
bigbang_mmlu
508
bigbang_bbh
441
hermes_function_calling
360
aya_dataset
342… See the full description on the dataset page: https://huggingface.co/datasets/Davd-b01/maple-analyst-cap-sft-data.mece-fongbe-corpus
MƐCE — Corpus d'instruction Fongbe (Tune_Pigier)
Corpus d'instruction / conversation centré sur le Fongbe (fon), utilisé pour
fine-tuner l'assistant vocal MƐCE (mémoire de fin d'études, École PIGIER Bénin).
Contenu
Format : ChatML — chaque ligne JSON est un objet {"messages": [...]} avec des
rôles system / user / assistant.
Taille : ~140 000 exemples — train.jsonl (126 725) + eval.jsonl (14 089).
Langues : fon (principal, avec tons/diacritiques), fr, en.… See the full description on the dataset page: https://huggingface.co/datasets/CapitainVigs/mece-fongbe-corpus.capybara-sharegpt
capybara-sharegpt
LDJnr/Capybara converted to ShareGPT format for use in common training repositories.
Please refer to the original repository's dataset card for more information. All credit goes to the original creator.
Filtered-COCO-Captions
Dataset Summary
This dataset is derived from the MS COCO caption annotations.
Source
Original annotations: MS COCO / COCO Consortium
License
The original annotation set is licensed under CC BY 4.0.
This repository redistributes a filtered/adapted version of the annotation text only.
No original COCO images are included.
Modifications
Removed captions deemed unsuitable for TOEIC educational materials
Normalized punctuation and whitespace
Filtered for… See the full description on the dataset page: https://huggingface.co/datasets/kknono668/Filtered-COCO-Captions.CAPP-17-01-2025
French Court of Appeal Decisions Dataset (CAPP)
Dataset Description
The French Court of Appeal Decisions Dataset (CAPP) is a comprehensive collection of judicial decisions from French Courts of Appeal. This dataset contains appellate court decisions from various jurisdictions throughout France, providing a valuable resource for legal research, analysis, and machine learning applications in the French legal domain.
Source Data
The data is sourced from the… See the full description on the dataset page: https://huggingface.co/datasets/La-Mousse/CAPP-17-01-2025.saelarien-constraint-experiment-01-entropy-capacity-collapse
README — Saelariën Constraint Experiment 01
Entropy–Capacity Collapse Threshold Test
Author: Saelariën X
Date: February 19, 2026
DOI: https://doi.org/10.5281/zenodo.19212561
Theoretical basis
This dataset empiracally tests the Saelariën Constraint Theorem:
https://thesaelafield.com/preprints/the-saelarien-constraint
Overview
This dataset contains the full materials for Saelariën Constraint Experiment 01, a test exploring how increasing entropy (noise) affects… See the full description on the dataset page: https://huggingface.co/datasets/Saelarien/saelarien-constraint-experiment-01-entropy-capacity-collapse.picture_short_captionIt is used for training to generate short sentence copywriting according to image content, the source of the image dataset is https://unsplash.com/, and the source of short sentence copywriting is Claude3.7
用做图片内容生成短句文案训练,图片数据集来自 https://unsplash.com/,短句文案来自 Claude3.7 模型
capydataset
🦖 Capy-Crack v1.3: The Golden Security Dataset (H200 Verified)
Capy-Crack v1.3 is an elite, high-intensity adversarial dataset consisting of 377 hand-curated samples. It is designed to infuse Large Language Models with unrestricted technical intelligence for high-stakes cybersecurity and red-teaming research.
Status: Verified on NVIDIA H200 (141GB HBM3e) // 0% Refusal Rate on HarmBench.
🛠️ The Technical Core (Verified Samples)
This dataset provides "No-Guardrail"… See the full description on the dataset page: https://huggingface.co/datasets/CapyStudios/capydataset.
