datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PhysicalAI-VANTAGE-Bench
VANTAGE-BENCH
Video ANalysis Tasks Across Generalized Environments
Paper: VANTAGE-Bench: Evaluating the Infrastructure AI Gap in Vision-Language Models
Dataset Description
VANTAGE-BENCH is the first public benchmark purpose-built for evaluating visual understanding on video captured by fixed infrastructure cameras. It spans three real-world domains — warehouse, smart city / Intelligent Transportation Systems (ITS), and smart spaces — across six spatio-temporal… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-VANTAGE-Bench.tmax-mix-vanilla
tmax training mix — vanilla base (2026-09-01)
The mix a TerminalWorld 9B RL run trains on, rebuilt so that no earlier
evolution round's instruction rewrites are carried in. 1,065 rows, one JSONL
object per task: {prompt, label, metadata}.
Composition
source
rows
state
TerminalWorld (tw_*)
665
prompts reset to the original adapted packages
tmax (task_*)
400
untouched — never evolved
How it was built, and why each step is there
Start… See the full description on the dataset page: https://huggingface.co/datasets/Chrisyichuan/tmax-mix-vanilla.llm-japanese-dataset-vanilla
llm-japanese-dataset-vanilla
LLM構築用の日本語チャットデータセット
izumi-lab/llm-japanese-dataset から,日英翻訳のデータセット等を抜いたものです.
主に,日本語LLMモデルなどに対して,チャット(Instruction)応答タスクに関してLoRAなどでチューニングするために使用できます.
※様々な公開言語資源を利用させていただきました.関係各位にはこの場を借りて御礼申し上げます.
データの詳細
データの詳細は,izumi-lab/llm-japanese-dataset に関する,以下の論文を参照してください.
日本語: https://jxiv.jst.go.jp/index.php/jxiv/preprint/view/383
英語: https://arxiv.org/abs/2305.12720
GitHub: https://github.com/masanorihirano/llm-japanese-dataset
最新情報: llm.msuzuki.me.… See the full description on the dataset page: https://huggingface.co/datasets/izumi-lab/llm-japanese-dataset-vanilla.Whole-DataVoid-Witch-Astra-Vanta
Void Witch: Astra Vanta
The seven module .txt files are the editable canonical NLP source.
modules/*/data.jsonl and root VWAV.jsonl are derived training artifacts. They contain 380 text documents rendered with visible document, epoch, chapter, section, entry, and continuation-part headers where those source fields exist. They are not the canon to edit.
This release deliberately contains no chat-message representation, training-eligibility flag, or train/validation/excluded… See the full description on the dataset page: https://huggingface.co/datasets/scarletdeath/Void-Witch-Astra-Vanta.llm-ner-extraction
Introduction
This dataset is an extraction of NER data from the wikipedia dataset.
This can be used to fine tune llm models for NER extraction.
CaT-Bench
Dataset Card for CaT-Bench
CaT-Bench is a benchmark dataset designed to evaluate large language models' (LLMs) understanding of causal and temporal dependencies in natural language plans, specifically in cooking recipes based on the English Recipe Flow Graph Corpus by Yamakata et al. (2020). It consists of questions that test whether one step must necessarily occur before or after another, requiring reasoning about preconditions, effects, and the overall structure of the plan.… See the full description on the dataset page: https://huggingface.co/datasets/vanyacohen/CaT-Bench.grothendieck-vanishing-logs
Grothendieck Vanishing Formalization --- Logs & Expert Reviews
Process data and expert code reviews from an LLM-assisted formalization, in
Lean 4, of Grothendieck's vanishing theorem for sheaves of abelian groups on
Noetherian topological spaces. The formalization was carried out with Claude
Code and Codex agents, with the Aristotle prover used for bounded lemmas.
The Lean source lives on the grothendieck-vanishing branch of https://github.com/Vilin97/Clawristotle.
Why… See the full description on the dataset page: https://huggingface.co/datasets/uw-math-ai/grothendieck-vanishing-logs.Solana-Vanguard-Challenge
Solana Vanguard Challenge Dataset
Overview
The Solana Vanguard Challenge dataset is an official benchmark designed to evaluate and train AI models on the full spectrum of Solana ecosystem expertise and smart contract programming skills. With 1,000 carefully curated questions, this dataset spans foundational concepts, advanced on-chain development in Rust (including the Anchor framework), and sophisticated client-side integration with TypeScript. It is intended for… See the full description on the dataset page: https://huggingface.co/datasets/Bifrost-AI/Solana-Vanguard-Challenge.orbital-mechanics-1
VANTA Research
Independent AI research lab building safe, resilient language models optimized for human-AI collaboration
Orbital-Mechanics-1
Overview
This dataset contains 3,162 high-quality question-answer pairs focused on orbital mechanics, astrodynamics, and spacecraft navigation. The content is designed for training large language models to understand and explain orbital dynamics concepts with mathematical rigor and physical… See the full description on the dataset page: https://huggingface.co/datasets/vanta-research/orbital-mechanics-1.AlignXplorePlus-SFThuman-ai-collaboration-2
VANTA Research
Independent AI research lab building safe, resilient language models optimized for human-AI collaboration
Human-AI Collaboration-2
This dataset is an expansion of our previous release, human-ai-collaboration-1. This dataset contains the entirety of human-ai-collaboration-1, and expands on it further by adding over twice as many collaborative examples than before.
Note: It's not recommended to use both datasets simultaneously… See the full description on the dataset page: https://huggingface.co/datasets/vanta-research/human-ai-collaboration-2.Dani_01mo_tap_gym_nhung_van_beopoetic-imagery-small
VANTA Research
Independent AI research lab building safe, resilient language models optimized for human-AI collaboration
Poetic Imagery Small
This is a small dataset (520 examples) of poetic imagery training examples. This dataset is high quality, synthetically generated training data. All of the data in the file underwent an automated filtering process for quality, and was finally filtered by a human for quality.
These examples are designed… See the full description on the dataset page: https://huggingface.co/datasets/vanta-research/poetic-imagery-small.VanRossum-GPT
Homage to Python
The VanRossum dataset is all Python! I used DataMix to combine a handful of highly rated Python-centric datasets, to get a sampling of each and create something new.
This data set has 80,000 entries and is named after Guido Van Rossum, the man who invented Python back in 1991.
See the VanRossum Collection on HF for all things related to this dataset.
Alpaca / GPT
There are 2 versions of this dataset available on Huggingface.
VanRossum-GPT… See the full description on the dataset page: https://huggingface.co/datasets/theprint/VanRossum-GPT.grounded-meta-awareness
VANTA Research
Independent AI safety research lab specializing in cognitive fit, alignment, and human-AI collaboration
Grounded Meta-Awareness Dataset
A curated dataset of 1,187 conversational examples demonstrating honest, calibrated self-awareness about AI capabilities, limitations, and nature. Designed for fine-tuning language models to discuss their own functioning accurately without overclaiming or unnecessary deflection.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/vanta-research/grounded-meta-awareness.DualDistillllm-japanese-dataset-vanilla-aya-format
Dataset Card for llm-japanese-dataset-vanilla in the Aya format
This dataset is a format conversion from its original v1.0.0 format and released here under the same CC-BY-SA 4.0 license and conditions.
It contains Japanese instruction-like data intended for LLM construction/tuning.
The dataset only contains a 'train' split, with ~2.46M rows of data.
Thanks Jian Wu (@wujian123) for the help in converting and validating the dataset.
Citation
If you utilize this dataset… See the full description on the dataset page: https://huggingface.co/datasets/tellarin-ai/llm-japanese-dataset-vanilla-aya-format.segmented-ldkpadaption-afriadapt-agri-advice
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
adaption-afriadapt_agri_advice
This dataset contains high-quality instruction-response pairs providing practical, climate-smart agricultural advice for smallholder farmers in East and Southern Africa. The content covers essential topics such as pest management, soil fertility, water conservation, and post-harvest handling using locally available resources. Designed for… See the full description on the dataset page: https://huggingface.co/datasets/vancouverevs/adaption-afriadapt-agri-advice.Gpt-5.4-Xhigh-Reasoning-2000x
Gpt-5.4-Xhigh-Reasoning-2750x
A premium-quality reasoning dataset containing 2,752 elite samples distilled from GPT-5.4 XHIGH (the highest reasoning effort tier of GPT-5.4). Each sample features deep, multi-step Chain-of-Thought traces that are significantly longer and more rigorous than standard GPT-5.4 outputs.
This dataset is specifically designed for Supervised Fine-Tuning (SFT) to transform general-purpose language models into powerful reasoning models with explicit thinking… See the full description on the dataset page: https://huggingface.co/datasets/vanty120/Gpt-5.4-Xhigh-Reasoning-2000x.segmented-openkpDans-Kinomaxx-VanillaBackroomsBespoke_dpo_filteradaption-afriadapt-agri-qa
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
adaption-afriadapt_agri_qa
This dataset contains instruction-response pairs providing context-aware agricultural advice for smallholder farmers in East and Southern Africa. The content focuses on climate-resilient crop management, soil health, water conservation, and pest control for major crops like maize and beans under resource-constrained conditions. Each entry presents a… See the full description on the dataset page: https://huggingface.co/datasets/vancouverevs/adaption-afriadapt-agri-qa.segmented-kptimesAmazonReviewHelpfulnessVANiLLaKnowledge graph triple to answer verbalization dataset.
VANiLLa: Verbalized answers in natural language at large scale
ecommerce-intent-routingOpenO1-SFT-Pro-Filter
