datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
copyright_gpt_neo_1_3Bforge-3b-pretrain-data
FORGE-3B Pretraining Data
Tokenized and packed pretraining data for the FORGE-3B language model.
Stats
Total tokens: 51.4070B
Domains: 10/10
Sequence length: 2048 tokens
Format: .npy shards of shape (N, 2048) with dtype uint32
Tokenizer: CRAYON (xerv-crayon, standard profile)
Domain Breakdown
Domain
Weight
Tokens (B)
Status
fineweb_edu
30%
15.0008
✓
thestack
16%
8.0011
✓
wikipedia
8%
4.2791
✓
openwebmath
8%
3.9654
✓
books
7%… See the full description on the dataset page: https://huggingface.co/datasets/Phase-Technologies/forge-3b-pretrain-data.all_pile_gpt_neo_1_3Bstack-v2-starcoder2-3bllama-3b-residualsRULER-8192-Qwen2.5-3B-tokenizerfull-math-private-n256-Qwen2.5-3B-Instruct-bonllama-3b-embedseasyr1-grounding-dataset-30k-not_grounded-SE-GUI-3B-2MPfineweb-llama3b-residualsqwen3-30b-a3b-base-reasoning-sft-nemotron-math-v4-cot4k12k-500m-supervised
Qwen3-30B-A3B Reasoning SFT Prepacked Nemotron Math v4 CoT 4k-12k
This dataset is a train-ready, offline-prepacked SFT corpus for full supervised
fine-tuning of Qwen/Qwen3-30B-A3B-Base into a math reasoning model.
Source And Filtering
Source dataset: nvidia/Nemotron-SFT-Math-v4
Source revision: a94e56aeddcf6e75d28c8bd210f40fa62309288d
Source split: train
Intended subset: cot
Preferred source during selection: AoPS
Length filter: 4,000 to 12,000 supervised… See the full description on the dataset page: https://huggingface.co/datasets/ar0cket1/qwen3-30b-a3b-base-reasoning-sft-nemotron-math-v4-cot4k12k-500m-supervised.Qwen3.6-35B-A3B-mcr-stage-b
Qwen3.6-35B-A3B — MCR Stage B Corpus (Distributed Reasoning Localization)
First systematic mechanistic-intervention corpus on a hybrid MoE + GDN + Gated-Attention architecture.
📄 Paper: Loop-Intolerance Profiling: Localizing Distributed Reasoning in a Hybrid MoE Architecture via Nine Convergent Intervention Experiments — submitted to arXiv (2026-04-20, in moderation). Final arXiv ID will be added here once approved.
This dataset contains per-token residual-stream activations at… See the full description on the dataset page: https://huggingface.co/datasets/caiovicentino1/Qwen3.6-35B-A3B-mcr-stage-b.full-math-private-n256-Llama-3.2-3B-Instruct-bonQwen3.6-35B-A3B-Tool-Calling
Qwen3.6-35B-A3B Tool-Calling Dataset
This repository presents a function and tool-calling preference and supervised fine-tuning dataset constructed from Nemotron-RL agentic prompt corpora.
For each source prompt, the model was sampled four times with thinking mode enabled. Each resulting candidate trajectory was then evaluated against the dataset’s ground-truth action using exact matching on both the function name and the parsed function arguments.
Overview… See the full description on the dataset page: https://huggingface.co/datasets/zake7749/Qwen3.6-35B-A3B-Tool-Calling.shapleymcg-qwen3-30b-a3b-reproducibility
ShapleyMCG Qwen3-30B-A3B reproducibility artifacts
This dataset preserves the calibration statistics, exact corrected-R10
EXL3/MCG candidates, source and corpus identities, BF16 teacher/student logits,
tokenwise KLD, allocations, attribution ledgers, hashes, and publication
receipts for the Qwen3-30B-A3B experiments in
brandonmmusic-max/shapleymcg.
The complete cross-checkpoint
results ledger
and
method specification
distinguish the predecessor routed-p2 allocator from the full… See the full description on the dataset page: https://huggingface.co/datasets/brandonmusic/shapleymcg-qwen3-30b-a3b-reproducibility.preprocessed-full-math-private-n256-Llama-3.2-3B-Instruct-bonfull-math-private-Qwen2.5-3B-Instruct-bondetails_openlm-research__open_llama_3b
Dataset Card for Evaluation run of openlm-research/open_llama_3b
Dataset Summary
Dataset automatically created during the evaluation run of model openlm-research/open_llama_3b on the Open LLM Leaderboard.
The dataset is composed of 122 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 6 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_openlm-research__open_llama_3b.stratified-solvable-1k-math-private-Qwen2.5-3B-Instruct-bongwosc-o3b-strain
GWOSC O3b 16 kHz strain
Upload in progress. Verified span shards are being added while source spans finish downloading and conversion.
This dataset contains the public Gravitational Wave Open Science Center O3b
strain release at 16,384 Hz for H1, L1, V1. Each detector is represented
independently. A contiguous source span produces three Parquet files:
Strain, DQmask, and Injmask.
Licence and acknowledgement
Creative Commons Attribution 4.0 International
Data… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/gwosc-o3b-strain.llm-jp-4-32b-a3b-thinking-dpo-data
llm-jp-4-32b-a3b-thinking-dpo-data
Overview
This dataset is a Direct Preference Optimization (DPO) dataset used to train llm-jp-4-32b-a3b-thinking.
It is constructed by pairing multiple candidate responses for a given prompt and selecting preferred (chosen) and non-preferred (rejected) responses. The splits reasoning_low, reasoning_medium, and reasoning_high correspond to different reasoning effort settings used during response generation.
The fields chosen_analysis… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/llm-jp-4-32b-a3b-thinking-dpo-data.Ornith-1.5-35B-A3B-GGUF-metricsterminal_bench_2_tasktrove_dq_stack_pytest_step25_30b_a3b_20260730_053956
TaskTrove stack-pytest — training rollout traces (Qwen3-Coder-30B-A3B, step 25)
Terminus-2/Harbor rollouts recorded while training
laion/tasktrove-dq-stack-pytest-step25-30b-a3b
with SkyRL on the TaskTrove stack-pytest source.
One row per trial, holding that trial's last episode as an OpenAI-style conversations list, the
task instruction, the reward the verifier assigned (result), and the verifier's own stdout
(verifier_output).
Source run… See the full description on the dataset page: https://huggingface.co/datasets/laion/terminal_bench_2_tasktrove_dq_stack_pytest_step25_30b_a3b_20260730_053956.details_togethercomputer__RedPajama-INCITE-Base-3B-v1
Dataset Card for Evaluation run of togethercomputer/RedPajama-INCITE-Base-3B-v1
Dataset Summary
Dataset automatically created during the evaluation run of model togethercomputer/RedPajama-INCITE-Base-3B-v1 on the Open LLM Leaderboard.
The dataset is composed of 122 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_togethercomputer__RedPajama-INCITE-Base-3B-v1.openthoughts4-code-9168-prompts-qwen3-30b-a3b-thinking-2507-n16-flattened-logprobs-k16
OpenThoughts-4 Code SDG: Qwen3-30B-A3B-Thinking-2507 (n=16, top-16 logprobs)
Synthetic generations from
Qwen/Qwen3-30B-A3B-Thinking-2507
on the Marin OpenThoughts-4 code SDG prompt
set.
Each prompt is sampled n=16 times, and for every generated token the dataset
stores the chosen-token log probability plus the top-16 log probabilities
over the vocabulary, enabling distillation, KL-style fine-tuning,
reranking, and uncertainty analysis.
Generation setup
Field… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-code-9168-prompts-qwen3-30b-a3b-thinking-2507-n16-flattened-logprobs-k16.llama3.2_3b_tokenizingdataAI21-Jamba2-3B
juiceb0xc0de/AI21-Jamba2-3B
A brain atlas for ai21labs/AI21-Jamba2-3B, a 28-layer hybrid Mamba/transformer from AI21 Labs. This is not a chat dataset or a benchmark. It is an internal-mechanics map built by running activations through a corpus of prompts and scoring what each layer, component, head, and direction is doing.
Jamba is an interesting subject because it is mostly not attention. Of the 28 layers, only 2 carry attention, and both of those run a single KV head.… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/AI21-Jamba2-3B.SlimPajama-3Bdetails_Azure99__blossom-v2-3b
Dataset Card for Evaluation run of Azure99/blossom-v2-3b
Dataset Summary
Dataset automatically created during the evaluation run of model Azure99/blossom-v2-3b on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Azure99__blossom-v2-3b.terminal_bench_2_tasktrove_dq_unix_step10_30b_a3b_20260730_014756
terminal_bench_2_tasktrove_dq_unix_step10_30b_a3b
OpenCode agent trajectories from the TaskTrove DQ unix arm of a Qwen3-Coder-30B-A3B
agentic RL sweep, exported from the complete Harbor rollout artifact set.
Coverage
Built from the full trace_jobs prefix of run rl-tasktrove-dq-sweep-30b-qwen3-coder-30-20260727-082204-e42f1d
(12034 trial directories, 11937 of them scored).
quantity
value
scored trials (result.json)
11937
rows published
11937
coverage… See the full description on the dataset page: https://huggingface.co/datasets/laion/terminal_bench_2_tasktrove_dq_unix_step10_30b_a3b_20260730_014756.
