datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vicgalle__Merge-Mixtral-Prometheus-8x7B-details
Dataset Card for Evaluation run of vicgalle/Merge-Mixtral-Prometheus-8x7B
Dataset automatically created during the evaluation run of model vicgalle/Merge-Mixtral-Prometheus-8x7B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/vicgalle__Merge-Mixtral-Prometheus-8x7B-details.chargoddard__prometheus-2-llama-3-8b-details
Dataset Card for Evaluation run of chargoddard/prometheus-2-llama-3-8b
Dataset automatically created during the evaluation run of model chargoddard/prometheus-2-llama-3-8b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/chargoddard__prometheus-2-llama-3-8b-details.taming-modern-prometheus-assurance
Taming the Modern Prometheus — Agentic Financial Assurance Benchmark
A small, transparent benchmark for evidence-grounded agentic workflows in financial assurance.
Contents
13 labelled cases.
3 public-derived Microsoft aggregate financial checks.
10 synthetic audit, ICFR, CAM, governance, ESG, adversarial, and reproducibility cases.
Required evidence, red-team challenge, target assertion, and non-compensatory gate for each case.
The public-derived rows are based… See the full description on the dataset page: https://huggingface.co/datasets/SADHON/taming-modern-prometheus-assurance.SWE-Prometheus
SWE-Prometheus Public Tasks
Public question package for SWE-Prometheus from CosmosMind AI Lab.
This release contains the 22 public tasks used in the benchmark. Each task
includes the task statement, the fixed repository revision, the environment
definition, the evaluation entry point, and the characterization tests used as
the behavior gate. Reference scores, model results, traces, and treated evidence
are intentionally excluded. A further 38 tasks are held out and not… See the full description on the dataset page: https://huggingface.co/datasets/CosmosMind/SWE-Prometheus.lm-eval-results-AiMavenAi-AiMaven-Prometheus-private
Dataset Card for Evaluation run of AiMavenAi/AiMaven-Prometheus
Dataset automatically created during the evaluation run of model AiMavenAi/AiMaven-Prometheus
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-AiMavenAi-AiMaven-Prometheus-private.matilda-smollm-mix-15b-gpt2
matilda-smollm-mix-15B-gpt2
15 B GPT-2-BPE tokens drawn from a 5:1 token-balanced mix of
HuggingFaceTB/smollm-corpus:
Source
Share
Tokens
fineweb-edu-dedup
83.33 %
12.50 B
cosmopedia-v2
16.67 %
2.50 B
Total: 15,000,349,569 tokens across 151 shards (shard_*.bin, uint16,
100 M tokens per shard).
The full SmolLM recipe is 75 / 15 / 10 fineweb-edu / cosmopedia-v2 / python-edu.
python-edu was dropped because the HuggingFaceTB/smollm-corpus subset
ships only blob_id… See the full description on the dataset page: https://huggingface.co/datasets/prometheus04/matilda-smollm-mix-15b-gpt2.QE-wmt21-1k-prometheus
data: WMT21 Metrics Shared Task EN-DE validation set (1k)
QE model: prometheus-eval/prometheus-7b-v2.0
vicgalle__Merge-Mistral-Prometheus-7B-details
Dataset Card for Evaluation run of vicgalle/Merge-Mistral-Prometheus-7B
Dataset automatically created during the evaluation run of model vicgalle/Merge-Mistral-Prometheus-7B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/vicgalle__Merge-Mistral-Prometheus-7B-details.
