datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
atcoder_cot
Dataset Card for Atcoder-CoT
Dataset Description
Atcoder-CoT is a proof-of-concept dataset designed to demonstrate how a dataset like the one found here can be used to generate synthetic datasets for training reasoning models, particularly for Supervised Fine-Tuning (SFT) and Knowledge Distillation. It leverages human-created and debugged solutions, combined with LLM-generated text to create conversational turns. The approach can also be easily adapted to simulate human… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/atcoder_cot.atcoder_abc_contests
Notification
Atcoder is selling this data now. If you are interested in accessing it please contact them.
Dataset Summary
This dataset aims to facilitate the creation of sophisticated, multi-turn dialogue datasets focused on coding
that could be used for training reasoning Large Language Models (LLMs), particularly for Supervised Fine-Tuning (SFT) and Knowledge Distillation techniques.
It also serves as a robust foundation for problem-solving in Large Language… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/atcoder_abc_contests.atcoder_abc_contests_small
Dataset Summary
This dataset aims to facilitate the creation of sophisticated, multi-turn dialogue datasets focused on coding
that could be used for training reasoning Large Language Models (LLMs), particularly for Supervised Fine-Tuning (SFT) and Knowledge Distillation techniques.
It also serves as a robust foundation for problem-solving in Large Language Models (LLMs).
The dataset includes both accepted and failed solutions from Atcoders's (ABC) contests.
In total, it features… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/atcoder_abc_contests_small.agentsim-atc
AgentSim Agent-Trace Corpus (ATC)
Grounded reasoning traces of retrieval-augmented question-answering agents,
generated by the AgentSim platform. 103,567 reasoning steps spanning
three established IR benchmarks (Quasar-T 38,915 + CausalQA 36,192 +
MSMARCO 28,460), with 20,548 supervised query-document-answer triples
extracted for fine-tuning, and 199,968 unique retrieved documents.
Every reasoning step traces back to specific documents in the source
corpus, enabling step-level… See the full description on the dataset page: https://huggingface.co/datasets/searchsim/agentsim-atc.gulf-coast-ga-atc-communications
Gulf Coast General Aviation ATC Communications Dataset
A conversational dataset of pilot–ATC radio exchanges designed for student pilots training in the Gulf Coast region (Texas, Louisiana, Mississippi, Alabama, Florida panhandle). Built for fine-tuning language models as study aids for aviation radio communications.
Dataset Overview
Metric
Value
Total conversations
5,401
Train split
4,860
Test split
541
Gulf Coast synthetic scenarios
3,400
Real ATC… See the full description on the dataset page: https://huggingface.co/datasets/starlineventures/gulf-coast-ga-atc-communications.agentsim-atc-multihop
AgentSim Agent-Trace Corpus — Multi-hop
A multi-hop sibling of the AgentSim Agent-Trace Corpus
(agentsim-atc)
with an evolved schema designed for student model distillation.
1 490 accepted SFT trajectories plus 2 980 step-level DPO preference
pairs, generated over 5 multi-hop QA datasets through a 7-action agentic
executor with an Always-Search Policy filter.
This corpus accompanies a follow-up technical report to "AgentSim: A
Platform for Verifiable Agent-Trace Simulation"… See the full description on the dataset page: https://huggingface.co/datasets/searchsim/agentsim-atc-multihop.atcoder_arc_contests
Notification
Atcoder is selling this data now. If you are interested in accessing it please contact them.
Dataset Summary
This dataset aims to facilitate the creation of sophisticated, multi-turn dialogue datasets focused on coding
that could be used for training reasoning Large Language Models (LLMs), particularly for Supervised Fine-Tuning (SFT) and Knowledge Distillation techniques.
It also serves as a robust foundation for problem-solving in Large Language… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/atcoder_arc_contests.atcoder_agc_contests
Notification
Atcoder is selling this data now. If you are interested in accessing it please contact them.
Dataset Summary
This dataset aims to facilitate the creation of sophisticated, multi-turn dialogue datasets focused on coding
that could be used for training reasoning Large Language Models (LLMs), particularly for Supervised Fine-Tuning (SFT) and Knowledge Distillation techniques.
It also serves as a robust foundation for problem-solving in Large Language… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/atcoder_agc_contests.Atcgpt-Fixed2
Dataset Card for Atcgpt-Fixed2
This dataset contains instruction-input-output pairs converted to ShareGPT format, designed for instruction tuning and text generation tasks.
Dataset Description
The dataset consists of carefully curated instruction-input-output pairs, formatted for conversational AI training. Each entry contains:
An instruction that specifies the task
An optional input providing context
A detailed output that addresses the instruction
Usage… See the full description on the dataset page: https://huggingface.co/datasets/HappyAIUser/Atcgpt-Fixed2.ATC-ShareGPT
Dataset Card for ATC-ShareGPT
This dataset contains instruction-input-output pairs converted to ShareGPT format, designed for instruction tuning and text generation tasks.
Dataset Description
The dataset consists of carefully curated instruction-input-output pairs, formatted for conversational AI training. Each entry contains:
An instruction that specifies the task
An optional input providing context
A detailed output that addresses the instruction
Usage
This… See the full description on the dataset page: https://huggingface.co/datasets/HappyAIUser/ATC-ShareGPT.atcoder_awc_contests
Notification
Atcoder is selling this data now. If you are interested in accessing it please contact them.
Dataset Summary
This dataset aims to facilitate the creation of sophisticated, multi-turn dialogue datasets focused on coding
that could be used for training reasoning Large Language Models (LLMs), particularly for Supervised Fine-Tuning (SFT) and Knowledge Distillation techniques.
It also serves as a robust foundation for problem-solving in Large Language… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/atcoder_awc_contests.ATCgpt-Fixed
Dataset Card for ATCgpt-Fixed
This dataset contains instruction-input-output pairs converted to ShareGPT format, designed for instruction tuning and text generation tasks.
Dataset Description
The dataset consists of carefully curated instruction-input-output pairs, formatted for conversational AI training. Each entry contains:
An instruction that specifies the task
An optional input providing context
A detailed output that addresses the instruction
Usage
This… See the full description on the dataset page: https://huggingface.co/datasets/HappyAIUser/ATCgpt-Fixed.
