datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RULER-8192-Qwen2.5-3B-tokenizerfull-math-private-n256-Qwen2.5-3B-Instruct-bonfull-math-private-n256-Llama-3.2-3B-Instruct-bonpreprocessed-full-math-private-n256-Llama-3.2-3B-Instruct-bonfull-math-private-Qwen2.5-3B-Instruct-bonstratified-solvable-1k-math-private-Qwen2.5-3B-Instruct-bonopenthoughts4-code-9168-prompts-qwen3-30b-a3b-thinking-2507-n16-flattened-logprobs-k16
OpenThoughts-4 Code SDG: Qwen3-30B-A3B-Thinking-2507 (n=16, top-16 logprobs)
Synthetic generations from
Qwen/Qwen3-30B-A3B-Thinking-2507
on the Marin OpenThoughts-4 code SDG prompt
set.
Each prompt is sampled n=16 times, and for every generated token the dataset
stores the chosen-token log probability plus the top-16 log probabilities
over the vocabulary, enabling distillation, KL-style fine-tuning,
reranking, and uncertainty analysis.
Generation setup
Field… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-code-9168-prompts-qwen3-30b-a3b-thinking-2507-n16-flattened-logprobs-k16.AI21-Jamba2-3B
juiceb0xc0de/AI21-Jamba2-3B
A brain atlas for ai21labs/AI21-Jamba2-3B, a 28-layer hybrid Mamba/transformer from AI21 Labs. This is not a chat dataset or a benchmark. It is an internal-mechanics map built by running activations through a corpus of prompts and scoring what each layer, component, head, and direction is doing.
Jamba is an interesting subject because it is mostly not attention. Of the 28 layers, only 2 carry attention, and both of those run a single KV head.… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/AI21-Jamba2-3B.browsecomp-qwen35-35b-a3b-think
browsecomp-qwen35-35b-a3b-think
Deep research agent evaluation on data/browsecomp.jsonl (normal split).
Results
Metric
Value
pass@4
43.0%
avg@4
24.8%
Trajectory accuracy
24.8% (1258/5064)
Questions
1266
Trajectories
5064 (4 per question)
Avg tool calls
41.1
Full conversations
❌
Model & Setup
Model
Qwen3.5-35B-A3B
Judge
gpt-4o
Max tool calls
50
Temperature
0.7
Blocked domains
huggingface.co
Tool… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/browsecomp-qwen35-35b-a3b-think.llama-3b-gold-15M-student-generations_SNIS_2048_tune422v1RULER-32768-Qwen2.5-3B-tokenizerdetails_Lansechen__Qwen2.5-3B-Open-R1-GRPO-math-selected-default
Dataset Card for Evaluation run of Lansechen/Qwen2.5-3B-Open-R1-GRPO-math-selected-default
Dataset automatically created during the evaluation run of model Lansechen/Qwen2.5-3B-Open-R1-GRPO-math-selected-default.
The dataset is composed of 3 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 11 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/Lansechen/details_Lansechen__Qwen2.5-3B-Open-R1-GRPO-math-selected-default.uniagent-qwen3-30b-a3b-r2e-rolloutsin1k_clip_qwen25vl_3b_224res_64tokens_new_ptin1k_clip_qwen25vl_3b_448res_256tokens_new_merged_ptllama-3b-gold-15M-student-generations_PRESAMPLING_2048_tune422v1openthoughts4-science-26041-prompts-qwen3-30b-a3B-thinking-2507-n8-flattened-logprobs-k16
OpenThoughts-4 Science SDG: Qwen3-30B-A3B-Thinking-2507 (n=8, top-16 logprobs)
Synthetic generations from
Qwen/Qwen3-30B-A3B-Thinking-2507
on the Marin OpenThoughts-4 science SDG prompt
set.
Each prompt is sampled n=8 times, and for every generated token the dataset
stores the chosen-token log probability plus the top-16 log probabilities
over the vocabulary, enabling distillation, KL-style fine-tuning,
reranking, and uncertainty analysis.
Generation setup
Field… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-science-26041-prompts-qwen3-30b-a3B-thinking-2507-n8-flattened-logprobs-k16.qwen36-35b-a3b-fp8-two-blackhole-tt-cache
Qwen3.6-35B-A3B-FP8 two-Blackhole TT cache
This dataset contains the generated same-source compressed owner-bank cache used by a public Qwen/Qwen3.6-35B-A3B-FP8 two-Blackhole runtime project.
Project repo:
https://github.com/PMZFX/TT-qwen36-35b-a3b-fp8-two-blackhole
The GitHub repo contains the runtime code, TT-Lang spike, reliability harnesses, release notes, and helper scripts. This dataset supplies the generated TT cache that is too large for the GitHub repo.
Contents… See the full description on the dataset page: https://huggingface.co/datasets/katostrofik/qwen36-35b-a3b-fp8-two-blackhole-tt-cache.llama-3.2-3b-atlas
llama-3.2-3b-atlas
details_Qwen__Qwen3-30B-A3B-Thinking-2507_v2
Dataset Card for Evaluation run of Qwen/Qwen3-30B-A3B-Thinking-2507
Dataset automatically created during the evaluation run of model Qwen/Qwen3-30B-A3B-Thinking-2507.
The dataset is composed of 116 configuration, each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Qwen__Qwen3-30B-A3B-Thinking-2507_v2.preprocessed-full-math-private-Llama-3.2-3B-Instruct-bonDeepHermes-3-Llama-3-3B-Preview_eval_2e29
mlfoundations-dev/DeepHermes-3-Llama-3-3B-Preview_eval_2e29
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
AIME25
HLE
LiveCodeBenchv5
Accuracy
0.0
2.5
5.2
18.2
3.5
2.7
2.3
1.4
3.8
0.0
8.0
1.5
AIME24
Average Accuracy: 0.00% ± 0.00%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
0.00%
0
30… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/DeepHermes-3-Llama-3-3B-Preview_eval_2e29.full-gsm8k-private-n256-Qwen2.5-3B-Instruct-bonfull-math-private-Llama-3.2-3B-Instruct-bonllama3.2-3b-instruct-atlas
juiceb0xc0de/llama3.2-3b-instruct-atlas
A brain atlas for meta-llama/Llama-3.2-3B-Instruct, the 3B instruction-tuned member of the Llama 3.2 family. This is not a chat dataset or a benchmark - it is an internal-mechanics map built by running activations through a corpus of prompts and scoring what each layer, component, head, and feature direction is doing.
If you want to know where an instruction-tuned model keeps its register machinery, which directions survive a causal test… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/llama3.2-3b-instruct-atlas.fineweb-edu-3Bpreprocessed-full-aime_2023-n256-Qwen2.5-3B-Instruct-bonopen-thoughts-4-30k-code-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8
Open Thoughts 4 - Code (Qwen3-30B-A3B-Thinking-2507, 32K tokens, n=8)
This dataset contains code reasoning problems with 8 independent responses generated by Qwen3-30B-A3B-Thinking-2507.
Overview
Source: marin-community/open-thoughts-4-30k-code-qwen3-32b-annotated (prompts only)
Model: Qwen/Qwen3-30B-A3B-Thinking-2507
Temperature: 0.8
Max tokens: 32,768
Columns
Column
Description
instruction_seed
The code problem prompt
_source
Source dataset… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-code-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8.preprocessed-full-gsm8k-private-n256-Qwen2.5-3B-Instruct-bondetails_Lansechen__Qwen2.5-3B-Open-R1-GRPO-math-selected-cosine-v2
Dataset Card for Evaluation run of Lansechen/Qwen2.5-3B-Open-R1-GRPO-math-selected-cosine-v2
Dataset automatically created during the evaluation run of model Lansechen/Qwen2.5-3B-Open-R1-GRPO-math-selected-cosine-v2.
The dataset is composed of 3 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 10 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/Lansechen/details_Lansechen__Qwen2.5-3B-Open-R1-GRPO-math-selected-cosine-v2.
