datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
stack-v2-starcoder2-3bRULER-8192-Qwen2.5-3B-tokenizerfull-math-private-n256-Qwen2.5-3B-Instruct-boneasyr1-grounding-dataset-30k-not_grounded-SE-GUI-3B-2MPQwen3.6-35B-A3B-mcr-stage-b
Qwen3.6-35B-A3B — MCR Stage B Corpus (Distributed Reasoning Localization)
First systematic mechanistic-intervention corpus on a hybrid MoE + GDN + Gated-Attention architecture.
📄 Paper: Loop-Intolerance Profiling: Localizing Distributed Reasoning in a Hybrid MoE Architecture via Nine Convergent Intervention Experiments — submitted to arXiv (2026-04-20, in moderation). Final arXiv ID will be added here once approved.
This dataset contains per-token residual-stream activations at… See the full description on the dataset page: https://huggingface.co/datasets/caiovicentino1/Qwen3.6-35B-A3B-mcr-stage-b.full-math-private-n256-Llama-3.2-3B-Instruct-bonQwen3.6-35B-A3B-Tool-Calling
Qwen3.6-35B-A3B Tool-Calling Dataset
This repository presents a function and tool-calling preference and supervised fine-tuning dataset constructed from Nemotron-RL agentic prompt corpora.
For each source prompt, the model was sampled four times with thinking mode enabled. Each resulting candidate trajectory was then evaluated against the dataset’s ground-truth action using exact matching on both the function name and the parsed function arguments.
Overview… See the full description on the dataset page: https://huggingface.co/datasets/zake7749/Qwen3.6-35B-A3B-Tool-Calling.preprocessed-full-math-private-n256-Llama-3.2-3B-Instruct-bonfull-math-private-Qwen2.5-3B-Instruct-bonstratified-solvable-1k-math-private-Qwen2.5-3B-Instruct-bonterminal_bench_2_tasktrove_dq_stack_pytest_step25_30b_a3b_20260730_053956
TaskTrove stack-pytest — training rollout traces (Qwen3-Coder-30B-A3B, step 25)
Terminus-2/Harbor rollouts recorded while training
laion/tasktrove-dq-stack-pytest-step25-30b-a3b
with SkyRL on the TaskTrove stack-pytest source.
One row per trial, holding that trial's last episode as an OpenAI-style conversations list, the
task instruction, the reward the verifier assigned (result), and the verifier's own stdout
(verifier_output).
Source run… See the full description on the dataset page: https://huggingface.co/datasets/laion/terminal_bench_2_tasktrove_dq_stack_pytest_step25_30b_a3b_20260730_053956.openthoughts4-code-9168-prompts-qwen3-30b-a3b-thinking-2507-n16-flattened-logprobs-k16
OpenThoughts-4 Code SDG: Qwen3-30B-A3B-Thinking-2507 (n=16, top-16 logprobs)
Synthetic generations from
Qwen/Qwen3-30B-A3B-Thinking-2507
on the Marin OpenThoughts-4 code SDG prompt
set.
Each prompt is sampled n=16 times, and for every generated token the dataset
stores the chosen-token log probability plus the top-16 log probabilities
over the vocabulary, enabling distillation, KL-style fine-tuning,
reranking, and uncertainty analysis.
Generation setup
Field… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-code-9168-prompts-qwen3-30b-a3b-thinking-2507-n16-flattened-logprobs-k16.AI21-Jamba2-3B
juiceb0xc0de/AI21-Jamba2-3B
A brain atlas for ai21labs/AI21-Jamba2-3B, a 28-layer hybrid Mamba/transformer from AI21 Labs. This is not a chat dataset or a benchmark. It is an internal-mechanics map built by running activations through a corpus of prompts and scoring what each layer, component, head, and direction is doing.
Jamba is an interesting subject because it is mostly not attention. Of the 28 layers, only 2 carry attention, and both of those run a single KV head.… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/AI21-Jamba2-3B.browsecomp-qwen35-35b-a3b-think
browsecomp-qwen35-35b-a3b-think
Deep research agent evaluation on data/browsecomp.jsonl (normal split).
Results
Metric
Value
pass@4
43.0%
avg@4
24.8%
Trajectory accuracy
24.8% (1258/5064)
Questions
1266
Trajectories
5064 (4 per question)
Avg tool calls
41.1
Full conversations
❌
Model & Setup
Model
Qwen3.5-35B-A3B
Judge
gpt-4o
Max tool calls
50
Temperature
0.7
Blocked domains
huggingface.co
Tool… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/browsecomp-qwen35-35b-a3b-think.SlimPajama-3Bterminal_bench_2_tasktrove_dq_unix_step10_30b_a3b_20260730_014756
terminal_bench_2_tasktrove_dq_unix_step10_30b_a3b
OpenCode agent trajectories from the TaskTrove DQ unix arm of a Qwen3-Coder-30B-A3B
agentic RL sweep, exported from the complete Harbor rollout artifact set.
Coverage
Built from the full trace_jobs prefix of run rl-tasktrove-dq-sweep-30b-qwen3-coder-30-20260727-082204-e42f1d
(12034 trial directories, 11937 of them scored).
quantity
value
scored trials (result.json)
11937
rows published
11937
coverage… See the full description on the dataset page: https://huggingface.co/datasets/laion/terminal_bench_2_tasktrove_dq_unix_step10_30b_a3b_20260730_014756.llama-3b-gold-15M-student-generations_SNIS_2048_tune422v1llama-3b-gold-15M-student-generationsRULER-32768-Qwen2.5-3B-tokenizerqwen3.6-35B-A3B_resultsMATH_train_llama3.2-3b-instructllm-jp-4-32b-a3b-thinking-dpo-data
llm-jp-4-32b-a3b-thinking-dpo-data
Overview
This dataset is a Direct Preference Optimization (DPO) dataset used to train llm-jp-4-32b-a3b-thinking.
It is constructed by pairing multiple candidate responses for a given prompt and selecting preferred (chosen) and non-preferred (rejected) responses. The splits reasoning_low, reasoning_medium, and reasoning_high correspond to different reasoning effort settings used during response generation.
The fields chosen_analysis… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/llm-jp-4-32b-a3b-thinking-dpo-data.details_Lansechen__Qwen2.5-3B-Open-R1-GRPO-math-selected-default
Dataset Card for Evaluation run of Lansechen/Qwen2.5-3B-Open-R1-GRPO-math-selected-default
Dataset automatically created during the evaluation run of model Lansechen/Qwen2.5-3B-Open-R1-GRPO-math-selected-default.
The dataset is composed of 3 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 11 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/Lansechen/details_Lansechen__Qwen2.5-3B-Open-R1-GRPO-math-selected-default.uniagent-qwen3-30b-a3b-r2e-rolloutslonghealth-llama-3.2-3bGuideRNA-3B
Dataset Card for GuideRNA-3B
Dataset Summary
GuideRNA-3B is a large transcriptome sequence corpus consisting of over 3.7 billion paired sequences extracted from the specific transcriptome of 23 cell lines and over 200 segmented genomes of RNA virus.
Supported Tasks
Based on this nucleotide sequence corpus, we are able to establish a foundation model to characterize the manifold of CRISPR guide RNA targeting regions in order to undertake further downstreaming… See the full description on the dataset page: https://huggingface.co/datasets/michaelm16/GuideRNA-3B.in1k_clip_qwen25vl_3b_224res_64tokens_new_ptin1k_clip_qwen25vl_3b_448res_256tokens_new_merged_ptswesmith-qwen3.6-35b-a3b
SWE-smith trajectories from Qwen3.6-35B-A3B
Multi-turn coding-agent trajectories (issue → tool-using rollout → patch) produced by
Qwen3.6-35B-A3B on SWE-smith tasks, stored untokenized.
This is the exact SFT corpus used for the harbor arm of the
nanoswe teacher-distillation experiments.
101,901 trajectories over 45,242 unique SWE-smith task instances (3 sampled rollouts
per task, ~2.25 surviving filtering), 53 parquet shards, ~1.4 GB.
≈1.96B training tokens = exactly one epoch… See the full description on the dataset page: https://huggingface.co/datasets/nanoswe/swesmith-qwen3.6-35b-a3b.llama-3b-gold-15M-student-generations_PRESAMPLING_2048_tune422v1
