datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Feedback-Collection
Dataset Card
Dataset Summary
The Feedback Collection is a dataset designed to induce fine-grained evaluation capabilities into language models.\
Recently, proprietary LLMs (e.g., GPT-4) have been used to evaluate long-form responses. In our experiments, we found that open-source LMs are not capable of evaluating long-form responses, showing low correlation with both human evaluators and GPT-4.\
In our paper, we found that by (1) fine-tuning feedback generated by GPT-4… See the full description on the dataset page: https://huggingface.co/datasets/prometheus-eval/Feedback-Collection.peerreview-bench
PeerReview Bench
CMU Paper Reviewer:https://prometheus-eval.github.io/cmu-paper-reviewer/
Repository:https://github.com/prometheus-eval/cmu-paper-reviewer
Paper:https://arxiv.org/abs/2605.20668
Point of Contact:seungone@kaist.ac.kr
Expert-annotated review items from scientific papers, organized for three
complementary evaluation tasks. All data in this dataset is intended
for evaluation, not training. All configs reference a shared, deduplicated
file store (submitted_papers)… See the full description on the dataset page: https://huggingface.co/datasets/prometheus-eval/peerreview-bench.BiGGen-Bench
BIGGEN-Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models
Dataset Description
BIGGEN-Bench (BiG Generation Benchmark) is a comprehensive evaluation benchmark designed to assess the capabilities of large language models (LLMs) across a wide range of tasks. This benchmark focuses on free-form text generation and employs fine-grained, instance-specific evaluation criteria.
Key Features:
Purpose: To evaluate LLMs on diverse capabilities… See the full description on the dataset page: https://huggingface.co/datasets/prometheus-eval/BiGGen-Bench.Preference-Collection
Dataset Card
Dataset Summary
The Preference Collection is a dataset designed to induce fine-grained evaluation capabilities into language models.
Recently, proprietary LLMs (e.g., GPT-4) have been used to evaluate long-form responses. In our experiments, we found that open-source LMs are not capable of evaluating long-form responses, showing low correlation with both human evaluators and GPT-4.\
In our paper, we found that by (1) fine-tuning feedback generated by GPT-4… See the full description on the dataset page: https://huggingface.co/datasets/prometheus-eval/Preference-Collection.microagent-train-v3
microagent-train-v3
SFT corpus for training a 4B-class terminal-agent model (Qwen3-4B-Thinking) in the microagent XML protocol. v3 = the 26,627-trajectory v2 corpus plus 3,951 synthetic failure-recovery trajectories that target the specific execution weaknesses found in v1 evaluation.
Why v3 exists
The v1 model (prometheus04/qwen3-4b-thinking-microagent-v1-merged) scored 1/89 (1.12%) on Terminal-Bench 2.0. Trajectory analysis showed the model reasoned correctly but failed… See the full description on the dataset page: https://huggingface.co/datasets/prometheus04/microagent-train-v3.microagent-train-v2
microagent-train-v2
Curated SFT corpus for training a terminal/bash agent. Derived from
nvidia/Nemotron-Terminal-Corpus
with a custom code-specific filter that recovers parse-error trajectories.
Quick numbers
26,627 trajectories
~244M tokens (avg 36.7k chars/trajectory)
94.9% <finish> endings (successful completion)
5.1% <give_up> endings (Nvidia-style informative failures)
81.7% multi-turn (≥6 turns), avg ~8.5 turns
Math-free (math.parquet dropped — 4B base already… See the full description on the dataset page: https://huggingface.co/datasets/prometheus04/microagent-train-v2.prometheus-bfsi-tier1-triage
Pioneer BFSI / Fintech Tier-1 Autonomous Triage Dataset
This dataset contains high-grade multi-turn conversational traces (ChatML format) curated according to the Pioneer paper data curation methodology for training and evaluating specialized 8B Small Language Models (SLMs) in the Banking, Financial Services, and Insurance (BFSI) vertical.
Dataset Structure
Train Split (train): 350 traces
75% Gold Standard Tasks: Standard operational workflows across 30… See the full description on the dataset page: https://huggingface.co/datasets/StarsMakeGalaxy/prometheus-bfsi-tier1-triage.SWE-Prometheus
SWE-Prometheus Public Tasks
Public question package for SWE-Prometheus from CosmosMind AI Lab.
This release contains the 22 public tasks used in the benchmark. Each task
includes the task statement, the fixed repository revision, the environment
definition, the evaluation entry point, and the characterization tests used as
the behavior gate. Reference scores, model results, traces, and treated evidence
are intentionally excluded. A further 38 tasks are held out and not… See the full description on the dataset page: https://huggingface.co/datasets/CosmosMind/SWE-Prometheus.matilda-smollm-mix-15b-gpt2
matilda-smollm-mix-15B-gpt2
15 B GPT-2-BPE tokens drawn from a 5:1 token-balanced mix of
HuggingFaceTB/smollm-corpus:
Source
Share
Tokens
fineweb-edu-dedup
83.33 %
12.50 B
cosmopedia-v2
16.67 %
2.50 B
Total: 15,000,349,569 tokens across 151 shards (shard_*.bin, uint16,
100 M tokens per shard).
The full SmolLM recipe is 75 / 15 / 10 fineweb-edu / cosmopedia-v2 / python-edu.
python-edu was dropped because the HuggingFaceTB/smollm-corpus subset
ships only blob_id… See the full description on the dataset page: https://huggingface.co/datasets/prometheus04/matilda-smollm-mix-15b-gpt2.traj
TerminalTraj-Terminus2
Packaged from m-a-p/TerminalTraj
into clean Terminus 2 format for fine-tuning nvidia/Nemotron-Terminal-8B.
Conversion Summary
Metric
Value
Source trajectories
20,000
Kept trajectories
19,848
Dropped
152 (0.76%)
Token p50
5884
Token p90
17046
Token p99
31885
Token max
70476
Over 8192 tokens
6919
Format
Each row:
conversations: ChatML list. Assistant turns with commands carry:{"analysis": "..."… See the full description on the dataset page: https://huggingface.co/datasets/prometheus04/traj.
