datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
the-pile-splitted
Dataset description
The pile is an 800GB dataset of english text
designed by EleutherAI to train large-scale language models. The original version of
the dataset can be found here.
The dataset is divided into 22 smaller high-quality datasets. For more information
each of them, please refer to the datasheet for the pile.
However, the current version of the dataset, available on the Hub, is not splitted accordingly.
We had to solve this problem in order to improve the user… See the full description on the dataset page: https://huggingface.co/datasets/ArmelR/the-pile-splitted.CircuitSense
CircuitSense
This dataset is a comprehensive multimodal circuit question-answering benchmark designed to evaluate visual reasoning and problem-solving capabilities across three main domains: Perception, Analysis, and Design. The dataset contains structured question-answer pairs with accompanying visual content, targeting different engineering cognitive levels and reasoning tasks.
Dataset Structure
The dataset is organized into three primary folders, each containing… See the full description on the dataset page: https://huggingface.co/datasets/armanakbari4/CircuitSense.claude-fable-5-claude-code
claude-fable-5 Agent Traces
It's worth noting that our team was working with Glint-Research to collect as much fable data as possible.
These are just the anonymized raw traces of both of our teams combined. This means that Glint-Research/Fable-5-traces was created from formatting and splitting up this same dataset. If you use one for your tune, don't use the other (it's the same exact data).
For training on this dataset I recommend using the teich package to convert to openai… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/claude-fable-5-claude-code.arm_O3
REBench — arm / O3
Binary analysis dataset extracted with Ghidra 11.x from the REBench benchmark suite.
Features per row (one row = one function)
Column
Description
arch / opt_level
Architecture & optimization flag
package / binary_name
Source package and executable
original_function_name
Real symbol name (from unstripped binary)
stripped_function_name
Generic name used in stripped binary
original_code
Decompiled C with original names… See the full description on the dataset page: https://huggingface.co/datasets/Xtest/arm_O3.arm_O0
REBench — arm / O0
Binary analysis dataset extracted with Ghidra 11.x from the REBench benchmark suite.
Features per row (one row = one function)
Column
Description
arch / opt_level
Architecture & optimization flag
package / binary_name
Source package and executable
original_function_name
Real symbol name (from unstripped binary)
stripped_function_name
Generic name used in stripped binary
original_code
Decompiled C with original names… See the full description on the dataset page: https://huggingface.co/datasets/Xtest/arm_O0.full_checkbox_dropdown_radiobuttonUnique_Armygpt-5.5-agentThis dataset was generated using teich by TeichAI
Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below.
gpt 5.5 Agent Traces
This directory contains raw agent trace files generated by teich. (I also dropped in some of my own personal traces)
All assistant responses were generated by openai/gpt-5.5.
JSONL files: 88
Training-ready tools
A complete configured tools schema snapshot is embedded in the… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/gpt-5.5-agent.paired_arm_risc_augmented
Dataset Card for "paired_arm_risc_augmented"
More Information needed
gpn-star-p-uniform-v1-enhancer-arm-a
marin-dna/gpn-star-p-uniform-v1-enhancer-arm-a
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment.
This draft covers the enhancer region cohort with all species scope and preserves source FASTA/2bit letter case.
Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus.
Anchor eligibility uses calibrated entropy from the primate… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/gpn-star-p-uniform-v1-enhancer-arm-a.kimi-k2.6-claude-code-tracesThis dataset was generated using teich by TeichAI
Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below.
Kimi K2.6 Claude Code Traces
This directory contains raw agent trace files generated by teich.
All assistant responses were generated by moonshotai/kimi-k2.6.
JSONL files: 36
Format
Each file is newline-delimited JSON representing a single captured agent session.
The trace schema is designed for… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/kimi-k2.6-claude-code-traces.qwen3.7-max-pi-tracesThis dataset was generated using teich by TeichAI
Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below.
Qwen3.7 Max Pi Traces
This directory contains raw agent trace files generated by teich.
All assistant responses were generated by qwen/qwen3.7-max.
JSONL files: 47
Training-ready tools
A complete configured tools schema snapshot is embedded in the collapsed section at the bottom of this README.
Use it… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/qwen3.7-max-pi-traces.armbench_defect_datasetThis is data from the Amazon Armbench dataset (https://armbench.s3.amazonaws.com/index.html).
Each image is labeled with the failure mode {'nominal', 'package_defect', 'multi_pick'} in the 'label' field.
Failures are further specified in the 'sublabel' field {book_jacket', 'open_book_jacket', 'open_book', 'partial_box', 'empty_bag', 'torn_bag', 'open_box', 'crush_box'}.
Each image also contains a 'polygon' highlighting the area of interest.
To cite this dataset, please use… See the full description on the dataset page: https://huggingface.co/datasets/correll/armbench_defect_dataset.armnet-demo-leaderboardarm_O1
REBench — arm / O1
Binary analysis dataset extracted with Ghidra 11.x from the REBench benchmark suite.
Features per row (one row = one function)
Column
Description
arch / opt_level
Architecture & optimization flag
package / binary_name
Source package and executable
original_function_name
Real symbol name (from unstripped binary)
stripped_function_name
Generic name used in stripped binary
original_code
Decompiled C with original names… See the full description on the dataset page: https://huggingface.co/datasets/Xtest/arm_O1.arm_O2
REBench — arm / O2
Binary analysis dataset extracted with Ghidra 11.x from the REBench benchmark suite.
Features per row (one row = one function)
Column
Description
arch / opt_level
Architecture & optimization flag
package / binary_name
Source package and executable
original_function_name
Real symbol name (from unstripped binary)
stripped_function_name
Generic name used in stripped binary
original_code
Decompiled C with original names… See the full description on the dataset page: https://huggingface.co/datasets/Xtest/arm_O2.pubmed-rct20kThe small 20K version of the Pubmed-RCT dataset by Dernoncourt et al (2017).
@article{dernoncourt2017pubmed,
title={Pubmed 200k rct: a dataset for sequential sentence classification in medical abstracts},
author={Dernoncourt, Franck and Lee, Ji Young},
journal={arXiv preprint arXiv:1710.06071},
year={2017}
}
Note: This is the cleaned up version by Jin and Szolovits (2018).
@article{jin2018hierarchical,
title={Hierarchical neural networks for sequential sentence classification in… See the full description on the dataset page: https://huggingface.co/datasets/armanc/pubmed-rct20k.ScienceQAThis is the ScientificQA dataset by Saikh et al (2022).
@article{10.1007/s00799-022-00329-y,
author = {Saikh, Tanik and Ghosal, Tirthankar and Mittal, Amish and Ekbal, Asif and Bhattacharyya, Pushpak},
title = {ScienceQA: A Novel Resource for Question Answering on Scholarly Articles},
year = {2022},
journal = {Int. J. Digit. Libr.},
month = {sep}
}
textlatent_zebra_thinkmorph_armAB
Text-Latent (Arm A) vs All-Latent (Arm B) — Zebra-CoT + ThinkMorph
35638 samples/arm, 18 categories. Schema = ULVR/williamium style (sample_id, category, source_dataset,
question, answer, input_image, intermediate_image_N, num_intermediate_steps, messages_json).
armA_text_latent: real decoded text CoT + latent visual blocks (intermediate_image_1..3).
armB_render_latent: reasoning text RENDERED to images, all-latent baseline (intermediate_image_1..17).
messages_json = full Monet… See the full description on the dataset page: https://huggingface.co/datasets/RuoliuYang/textlatent_zebra_thinkmorph_armAB.minimax-m3-claude-code-tracesThis dataset was generated using teich by TeichAI
Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below.
Minimax M3 Claude Code Traces
This directory contains raw agent trace files generated by teich.
All assistant responses were generated by minimax/minimax-m3.
JSONL files: 31
Format
Each file is newline-delimited JSON representing a single captured agent session.
The trace schema is designed for… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/minimax-m3-claude-code-traces.phylop-uniform-v1-enhancer-arm-a
marin-dna/phylop-uniform-v1-enhancer-arm-a
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment.
This draft covers the enhancer region cohort with all species scope and preserves source FASTA/2bit letter case.
Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus.
Anchor eligibility uses the pipeline's pinned phyloP conservation… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/phylop-uniform-v1-enhancer-arm-a.qrpo-paper-llama-nosft-magpieair-armorm-temp1-ref50-offpolicy2best-armorm
qrpo-paper-llama-nosft-magpieair-armorm-temp1-ref50-offpolicy2best-armorm
Dataset with reference completions and rewards for a specific model and reward model, ready for training with the QRPO reference codebase (https://github.com/CLAIRE-Labo/quantile-reward-policy-optimization).
Part of the dataset collection for the paper Quantile Reward Policy Optimization: Alignment with Pointwise Regression and Exact Partition Functions (https://arxiv.org/pdf/2507.08068).
ArmenianParaphrasePC
ArmenianParaphrasePC
An MTEB dataset
Massive Text Embedding Benchmark
asparius/Armenian-Paraphrase-PC
Task category
t2t
Domains
News, Written
Reference
https://github.com/ivannikov-lab/arpa-paraphrase-corpus
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["ArmenianParaphrasePC"])
evaluator = mteb.MTEB(task)
model = mteb.get_model(YOUR_MODEL)… See the full description on the dataset page: https://huggingface.co/datasets/mteb/ArmenianParaphrasePC.icon_arm-rolloutsarmnetbench_v01_robometer
ArmnetBench v0.1
The ArmnetBench v0.1 core benchmark contains 50 human-teleoperated reference
trajectories per task plus 2,518 evaluation trajectories from 7 policies trained or
fine-tuned on those reference datasets, across 8 single-arm and 4 bimanual tasks performed
on the low-cost SO-101 robot arm. It was collected using the
Armnet arm farm.
The full release contains 3,718 labelled episodes: the 3,118 core benchmark
episodes plus 600 episodes from extra runs. The extra runs… See the full description on the dataset page: https://huggingface.co/datasets/armnet/armnetbench_v01_robometer.dora-armsqrpo-paper-llama-nosft-magpieair-armorm-temp1-ref50-offline-armorm
qrpo-paper-llama-nosft-magpieair-armorm-temp1-ref50-offline-armorm
Dataset with reference completions and rewards for a specific model and reward model, ready for training with the QRPO reference codebase (https://github.com/CLAIRE-Labo/quantile-reward-policy-optimization).
Part of the dataset collection for the paper Quantile Reward Policy Optimization: Alignment with Pointwise Regression and Exact Partition Functions (https://arxiv.org/pdf/2507.08068).
llama3-ultrafeedback-armorm
Dataset Card for llama3-ultrafeedback-armorm
This dataset was used to train princeton-nlp/Llama-3-Instruct-8B-SimPO-v0.2.
If you are interested in training other model types (e.g., Mistral, Gemma-2), please refer to their corresponding datasets: princeton-nlp/mistral-instruct-ultrafeedback, and princeton-nlp/gemma2-ultrafeedback-armorm.
Dataset Structure
This dataset contains around 60k training samples and 2k testing samples, following the original splits in… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/llama3-ultrafeedback-armorm.dataset8teich-test-v1
hy3-preview coding agent traces
This directory contains raw agent trace files generated by teich.
All assistant responses were generated by tencent/hy3-preview:free.
Training-ready tools
Use this tools payload when rendering converted examples through your training chat template.
The same structure is emitted on each converted example as the tools field.
[
{
"type": "function",
"function": {
"name": "bash",
"description": "Execute bash… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/teich-test-v1.
