datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CircuitSense
CircuitSense
This dataset is a comprehensive multimodal circuit question-answering benchmark designed to evaluate visual reasoning and problem-solving capabilities across three main domains: Perception, Analysis, and Design. The dataset contains structured question-answer pairs with accompanying visual content, targeting different engineering cognitive levels and reasoning tasks.
Dataset Structure
The dataset is organized into three primary folders, each containing… See the full description on the dataset page: https://huggingface.co/datasets/armanakbari4/CircuitSense.claude-fable-5-claude-code
claude-fable-5 Agent Traces
It's worth noting that our team was working with Glint-Research to collect as much fable data as possible.
These are just the anonymized raw traces of both of our teams combined. This means that Glint-Research/Fable-5-traces was created from formatting and splitting up this same dataset. If you use one for your tune, don't use the other (it's the same exact data).
For training on this dataset I recommend using the teich package to convert to openai… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/claude-fable-5-claude-code.scientific_papersScientific papers datasets contains two sets of long and structured documents.
The datasets are obtained from ArXiv and PubMed OpenAccess repositories.
Both "arxiv" and "pubmed" have two features:
- article: the body of the document, pagragraphs seperated by "/n".
- abstract: the abstract of the document, pagragraphs seperated by "/n".
- section_names: titles of sections, seperated by "/n".full_checkbox_dropdown_radiobuttongpt-5.5-agentThis dataset was generated using teich by TeichAI
Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below.
gpt 5.5 Agent Traces
This directory contains raw agent trace files generated by teich. (I also dropped in some of my own personal traces)
All assistant responses were generated by openai/gpt-5.5.
JSONL files: 88
Training-ready tools
A complete configured tools schema snapshot is embedded in the… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/gpt-5.5-agent.gpn-star-p-uniform-v1-enhancer-arm-a
marin-dna/gpn-star-p-uniform-v1-enhancer-arm-a
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment.
This draft covers the enhancer region cohort with all species scope and preserves source FASTA/2bit letter case.
Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus.
Anchor eligibility uses calibrated entropy from the primate… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/gpn-star-p-uniform-v1-enhancer-arm-a.kimi-k2.6-claude-code-tracesThis dataset was generated using teich by TeichAI
Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below.
Kimi K2.6 Claude Code Traces
This directory contains raw agent trace files generated by teich.
All assistant responses were generated by moonshotai/kimi-k2.6.
JSONL files: 36
Format
Each file is newline-delimited JSON representing a single captured agent session.
The trace schema is designed for… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/kimi-k2.6-claude-code-traces.textlatent_zebra_thinkmorph_armAB
Text-Latent (Arm A) vs All-Latent (Arm B) — Zebra-CoT + ThinkMorph
35638 samples/arm, 18 categories. Schema = ULVR/williamium style (sample_id, category, source_dataset,
question, answer, input_image, intermediate_image_N, num_intermediate_steps, messages_json).
armA_text_latent: real decoded text CoT + latent visual blocks (intermediate_image_1..3).
armB_render_latent: reasoning text RENDERED to images, all-latent baseline (intermediate_image_1..17).
messages_json = full Monet… See the full description on the dataset page: https://huggingface.co/datasets/RuoliuYang/textlatent_zebra_thinkmorph_armAB.qwen3.7-max-pi-tracesThis dataset was generated using teich by TeichAI
Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below.
Qwen3.7 Max Pi Traces
This directory contains raw agent trace files generated by teich.
All assistant responses were generated by qwen/qwen3.7-max.
JSONL files: 47
Training-ready tools
A complete configured tools schema snapshot is embedded in the collapsed section at the bottom of this README.
Use it… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/qwen3.7-max-pi-traces.phylop-uniform-v1-enhancer-arm-a
marin-dna/phylop-uniform-v1-enhancer-arm-a
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment.
This draft covers the enhancer region cohort with all species scope and preserves source FASTA/2bit letter case.
Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus.
Anchor eligibility uses the pipeline's pinned phyloP conservation… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/phylop-uniform-v1-enhancer-arm-a.ScienceQAThis is the ScientificQA dataset by Saikh et al (2022).
@article{10.1007/s00799-022-00329-y,
author = {Saikh, Tanik and Ghosal, Tirthankar and Mittal, Amish and Ekbal, Asif and Bhattacharyya, Pushpak},
title = {ScienceQA: A Novel Resource for Question Answering on Scholarly Articles},
year = {2022},
journal = {Int. J. Digit. Libr.},
month = {sep}
}
pubmed-rct20kThe small 20K version of the Pubmed-RCT dataset by Dernoncourt et al (2017).
@article{dernoncourt2017pubmed,
title={Pubmed 200k rct: a dataset for sequential sentence classification in medical abstracts},
author={Dernoncourt, Franck and Lee, Ji Young},
journal={arXiv preprint arXiv:1710.06071},
year={2017}
}
Note: This is the cleaned up version by Jin and Szolovits (2018).
@article{jin2018hierarchical,
title={Hierarchical neural networks for sequential sentence classification in… See the full description on the dataset page: https://huggingface.co/datasets/armanc/pubmed-rct20k.dataset8minimax-m3-claude-code-tracesThis dataset was generated using teich by TeichAI
Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below.
Minimax M3 Claude Code Traces
This directory contains raw agent trace files generated by teich.
All assistant responses were generated by minimax/minimax-m3.
JSONL files: 31
Format
Each file is newline-delimited JSON representing a single captured agent session.
The trace schema is designed for… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/minimax-m3-claude-code-traces.dataset5real_cartpole_200kThis dataset contains sequences of actions, motor angles and pendulum angles as well as velocities for a rotary inverted pendulum robot. The dataset was collected while training the robot to swing up and balance (wandb run).Angles are in radian. Velocities were computed from the angles and fed to the policy. Control frequency is 75Hz.The action maps to the motor voltage with:
deadzone = 0.1
center = 0.05
max_act = 0.9
if abs(action) > center:
V = np.sign(action) * (… See the full description on the dataset page: https://huggingface.co/datasets/armandpl/real_cartpole_200k.guppylm-60k-generic
GuppyLM Chat Dataset
Training data for GuppyLM — a ~9M parameter LLM that talks like a small fish.
Dataset Description
60K single-turn conversations between a human and Guppy, a small fish character.
Guppy speaks in short, lowercase sentences about water, food, light, and tank life.
It doesn't understand human abstractions.
Example
Input: are you hungry
Output: yes. always yes. i will swim to the top right now.
Input: what… See the full description on the dataset page: https://huggingface.co/datasets/arman-bd/guppylm-60k-generic.test1_arman_20260829_141011This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
7
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_yaw.pos"… See the full description on the dataset page: https://huggingface.co/datasets/makermods/test1_arman_20260829_141011.teich-test-v1
hy3-preview coding agent traces
This directory contains raw agent trace files generated by teich.
All assistant responses were generated by tencent/hy3-preview:free.
Training-ready tools
Use this tools payload when rendering converted examples through your training chat template.
The same structure is emitted on each converted example as the tools field.
[
{
"type": "function",
"function": {
"name": "bash",
"description": "Execute bash… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/teich-test-v1.Fable-5-Chat
TheFusionCube Fable-5 Chat Conversion
Source dataset: TheFusionCube/Fable-5-CoT-Traces
Output file: train.jsonl
Source rows: 468
Kept rows: 353
Dropped category == "decoy" rows: 115
Dropped blank prompt/response rows: 0
Each row has:
{
"prompt": "...",
"messages": [
{"role": "user", "content": "..."},
{"role": "assistant", "content": "..."}
],
"tools": [],
"metadata": {"trace_type": "chat", "category": "..."}
}
qwen37-pi-qwen36-27b-topk40-logprobs
Qwen3.7 Pi Trace Top-40 Teacher Logprobs
Offline top-40 teacher logprobs for cumulative assistant-turn rows from
armand0e/qwen3.7-max-split-formatted.
These files are intended to be loaded with snapshot_download, not
datasets.load_dataset.
Contents
manifest.json: shard metadata and filtering counts
shard-*.pt: tokenized examples with labels, target positions, top-k token ids,
and top-k teacher logprobs
chat_template.jinja: the exact chat template used for… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/qwen37-pi-qwen36-27b-topk40-logprobs.ur3-3task-lerobot
EmbodyX UR3 — three bimanual manipulation tasks (LeRobot v2.1)
Real-world teleoperated demonstrations on a bimanual dual-arm UR3, released as the three
tasks used to fine-tune armanakbari4/imagewam-ur3-3task.
task
episodes
frames
fps
instruction
blue_basket_lerobot
100
29,653
15
"put the medicine then the measuring tape inside the blue basket"
drawer_lerobot
100
32,982
15
"open the drawer, put the white box inside the drawer then close the drawer"… See the full description on the dataset page: https://huggingface.co/datasets/armanakbari4/ur3-3task-lerobot.minimax-m2.7-agent
Agentic Training Traces
This directory contains raw agent trace files generated by agentic-datagen.
All assistant responses were generated by minimax/minimax-m2.7.
Trace files: 20
Training-ready tools
Use this tools payload when rendering converted examples through your training chat template.
The same structure is emitted on each converted example as the tools field.
[
{
"type": "function",
"function": {
"name": "bash",
"parameters": {… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/minimax-m2.7-agent.gpt-5.5-chatThis dataset was generated using teich by TeichAI
Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below.
gpt 5.5 Chat Traces
This directory contains newline-delimited JSON training examples generated by teich.
All assistant responses were generated by openai/gpt-5.5.
Rows: 133
Format
Each file is newline-delimited JSON where every line is already a training example.
Chat-only datasets include messages… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/gpt-5.5-chat.Vidcodeipcc-testbadlogicgames-pi-mono-opus-filteredFiltered version of badlogicgames/pi-mono - Only opus traces, dropped invalid sessions as well.
All traces present are training safe and teich compatible
kimi-k2.6-agentThis dataset was generated using teich by TeichAI
Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below.
Kimi K2.6 Agent Traces
This directory contains raw agent trace files generated by teich.
All assistant responses were generated by moonshotai/kimi-k2.6.
JSONL files: 15
Training-ready tools
A complete configured tools schema snapshot is embedded in the collapsed section at the bottom of this README.… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/kimi-k2.6-agent.kazakh-tts-testkazakh-tts-val
