datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CaT-Bench
Dataset Card for CaT-Bench
CaT-Bench is a benchmark dataset designed to evaluate large language models' (LLMs) understanding of causal and temporal dependencies in natural language plans, specifically in cooking recipes based on the English Recipe Flow Graph Corpus by Yamakata et al. (2020). It consists of questions that test whether one step must necessarily occur before or after another, requiring reasoning about preconditions, effects, and the overall structure of the plan.… See the full description on the dataset page: https://huggingface.co/datasets/vanyacohen/CaT-Bench.Gpt-5.4-Xhigh-Reasoning-2000x
Gpt-5.4-Xhigh-Reasoning-2750x
A premium-quality reasoning dataset containing 2,752 elite samples distilled from GPT-5.4 XHIGH (the highest reasoning effort tier of GPT-5.4). Each sample features deep, multi-step Chain-of-Thought traces that are significantly longer and more rigorous than standard GPT-5.4 outputs.
This dataset is specifically designed for Supervised Fine-Tuning (SFT) to transform general-purpose language models into powerful reasoning models with explicit thinking… See the full description on the dataset page: https://huggingface.co/datasets/vanty120/Gpt-5.4-Xhigh-Reasoning-2000x.Gpt-5.4-Xhigh-Reasoning-750x
Gpt-5.4-Xhigh-Reasoning-750x
A premium-quality reasoning dataset containing 721 elite samples distilled from GPT-5.4 XHIGH (the highest reasoning effort tier of GPT-5.4). Each sample features deep, multi-step Chain-of-Thought traces specifically targeting ultra-hard, expert-level problems across 60+ scientific and technical domains.
This dataset is specifically designed for Supervised Fine-Tuning (SFT) to transform general-purpose language models into powerful reasoning models with… See the full description on the dataset page: https://huggingface.co/datasets/vanty120/Gpt-5.4-Xhigh-Reasoning-750x.letters_in_word
Количество букв в слове
Автосгенерированный датасет чтобы научить модель считать количество букв в слове.
grok_answer_mail_ru
Датасет ответов на Маил.ру
В этом датасете собраны ответы от Grok-3-latest (и немного chatgpt-4o-latest) на вопросы с Ответы Маил.ру
jorj
