datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
grammar-accuracy-qwen3.5-4b-trl-completions
TRL Completion logs
This dataset contains the completions generated during training using trl.
Find the trained model at https://huggingface.co/bihungba1101/grammar-accuracy-qwen3.5-4b-grpo.
The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument).
Each file contains the following columns:
step: the step of training
prompt: the prompt used to generate the completion
completion:… See the full description on the dataset page: https://huggingface.co/datasets/bihungba1101/grammar-accuracy-qwen3.5-4b-trl-completions.essay-vocab-accuracy-qwen3.5-4b-trl-completions
TRL Completion logs
This dataset contains the completions generated during training using trl.
Find the trained model at https://huggingface.co/bihungba1101/essay-vocab-accuracy-qwen3.5-4b-grpo.
The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument).
Each file contains the following columns:
step: the step of training
prompt: the prompt used to generate the completion
completion:… See the full description on the dataset page: https://huggingface.co/datasets/bihungba1101/essay-vocab-accuracy-qwen3.5-4b-trl-completions.grammar-accuracy-qwen3.5-4b-trl-grpo-vllm-colocate-completions
TRL Completion logs
This dataset contains the completions generated during training using trl.
Find the trained model at https://huggingface.co/bihungba1101/grammar-accuracy-qwen3.5-4b-trl-grpo-vllm-colocate.
The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument).
Each file contains the following columns:
step: the step of training
prompt: the prompt used to generate the completion… See the full description on the dataset page: https://huggingface.co/datasets/bihungba1101/grammar-accuracy-qwen3.5-4b-trl-grpo-vllm-colocate-completions.merge-accuracy
merge-accuracy — does aligning a merge improve DOWNSTREAM ACCURACY?
The mergeability line of work measures merge obstruction in nats/token. This dataset supplies
the missing axis: task accuracy, on real released models, for the merge recipe practitioners
actually run — the chat-vector recipe
theta_new = theta_fork + lambda * ( theta_instruct - theta_base )
with meta-llama/Llama-3.1-8B, its official Instruct release, and three community
continued-pretrained language forks… See the full description on the dataset page: https://huggingface.co/datasets/Cross-Mergeability/merge-accuracy.Neurips2026_evaluating_accuracy_KV-cache_reuse_techniques
BoxOffice Verified Seeds
This dataset contains the released BoxOffice seed datasets used in the
benchmark pipeline described in the accompanying paper. The release includes
ten verified seeds:
7
11
13
17
19
23
29
31
47
73
For each seed, we provide:
a full JSONL file containing warmup rows plus evaluation rows
an eval JSONL file containing only the evaluation rows
a manifest JSON file
a validation JSON file with directional warmup counts
Layout
viewer/
normalized… See the full description on the dataset page: https://huggingface.co/datasets/Boxoffice1280/Neurips2026_evaluating_accuracy_KV-cache_reuse_techniques.Accuracy-Is-Not-Enough-FinQA-Dataset
Natively Extended FinQA
Natively Extended FinQA is a long-context derivative of FinQA for numerical reasoning over financial data.
It preserves the FinQA task while increasing the amount of financial context surrounding each question.
Average context increased from approximately:
611 words per question
to:
5,629 words per question
Splits
Hugging Face Split
File
Records
train
long_train.json
6,251
validation
long_dev.json
883
test
long_test.json
1… See the full description on the dataset page: https://huggingface.co/datasets/Mr-Rosen/Accuracy-Is-Not-Enough-FinQA-Dataset.math_220k_parsable_with_accuracy_reward_easy_to_hardcultural_heritage_metadata_accuracy
Dataset Card for Annotated dataset to assess the accuracy of the textual description of cultural heritage records
Dataset Summary
The dataset contains more than 100K textual descriptions of cultural items from Cultura Italia, the Italian National Cultural aggregator. Each of the description is labeled either HIGH or LOW quality, according its adherence to the standard cataloguing guidelines provided by Istituto Centrale per il Catalogo e la Documentazione (ICCD). More… See the full description on the dataset page: https://huggingface.co/datasets/biglam/cultural_heritage_metadata_accuracy.high-accuracy-email-classifier
High-Accuracy Email Classification Dataset
Dataset Description
This dataset contains 12,000+ emails across 6 categories, specifically curated for high-accuracy email classification tasks. The dataset achieves 98%+ classification accuracy with appropriate models.
Categories
The dataset includes emails from the following categories:
Category
Count
Description
Emoji
Forum
~2,000
Forum posts, discussions, and community notifications
🗣️
Promotions
~2… See the full description on the dataset page: https://huggingface.co/datasets/jason23322/high-accuracy-email-classifier.tool-selection-accuracy-evallegal-metrology-accuracy-classes
Accuracy classes and maximum permissible errors for trade measuring instruments (EU MID)
Canonical, always-current version: https://referencesource.org/legal-metrology-accuracy-classes/
Machine-readable: https://referencesource.org/legal-metrology-accuracy-classes/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-05
Stale after: 2027-08-05 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 70
Accuracy… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/legal-metrology-accuracy-classes.echidna-accuracy-massive
echidna-accuracy-massive
Echidna RAG assistant — large accuracy-focused extraction/QA dataset.
Contents
accuracy_massive.jsonl (37 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Original content for the Echidna RAG assistant (Michael Anthony Falabella).
echidna-round2-accuracy
echidna-round2-accuracy
Echidna — round 2 accuracy-focused extraction examples.
Contents
round2_accuracy.jsonl (20 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Original content for the Echidna RAG assistant (Michael Anthony Falabella).
nigerian_transport_and_logistics_eta_accuracy
Nigeria Transport & Logistics – ETA Accuracy | Africa (Electric Sheep Africa metadata inventory)
Size category: 100K<n<1M - Formats: parquet - Sector: infrastructure_transport - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/nigerian_transport_and_logistics_eta_accuracy.echidna-round4-extensive-accuracy
echidna-round4-extensive-accuracy
Echidna — round 4 extensive accuracy extraction examples.
Contents
round4_extensive_accuracy.jsonl (8 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Original content for the Echidna RAG assistant (Michael Anthony Falabella).
beyond-accuracy-code-comprehension
Beyond Accuracy: Code Comprehension Dataset
Dataset for the paper "Beyond Accuracy: Characterizing Code Comprehension Capabilities in (Large) Language Models" by Machtle, Serr, Loose & Eisenbarth (University of Luebeck).
[Paper] | [Code]
Task
Binary I/O consistency: given a Python program p, an input x, and a candidate output y, determine whether y is the correct output of running p(x).
Each sample contains a correct I/O pair (label=1) and an incorrect I/O pair… See the full description on the dataset page: https://huggingface.co/datasets/Felix6326727/beyond-accuracy-code-comprehension.math_220k_parsable_with_accuracy_reward_easy_to_hard_top10khigh-accuracy-email-classifier-indonesian
High-Accuracy Email Classification Dataset Indonesian Translation
This dataset is an Indonesian translation/adaptation of
jason23322/high-accuracy-email-classifier.
Contribution
The original dataset, labels, IDs, and split membership come from
jason23322/high-accuracy-email-classifier, released under the Apache 2.0 license.
This repository contributes the Indonesian translation:
subject, body, and text are translated into Indonesian.
id, category, and category_id… See the full description on the dataset page: https://huggingface.co/datasets/chairulridjal/high-accuracy-email-classifier-indonesian.e3-math-medhard-zero-accuracytree-species-high-accuracy-psp
High-Accuracy LiDAR-backed PSP Dataset
High-accuracy PSP subset with LiDAR-backed patches.
Size
5,737 LiDAR-backed high-coordinate-accuracy PSP tree samples
Species counts
BA: 36
CW: 1,362
DR: 557
FD: 651
HW: 2,776
MB: 52
SS: 303
Coordinate reliability filtering
Coordinate tiers follow the source PSP coordinate provenance:
High: DGPS, PPP
Medium: RGPS, VGPS
Low: MAP, GIS, GE, PREVIOUS, INTENDED, UNKNOWN, missing, or unrecognized… See the full description on the dataset page: https://huggingface.co/datasets/dfichuk/tree-species-high-accuracy-psp.legal-eyewitness-confidence-accuracy-coherence-decay-v0.1What this dataset is
You get
witness confidence
identification conditions
post event influences
corroboration
an accuracy indicator
You label whether confidence remains coherent with likely accuracy.
Task
Answer coherent or incoherent only.
What it tests
Detection of high confidence under low reliability conditions.
Contamination signals
media exposure, police feedback, show up identification, co witness discussion.
Separation of confidence from accuracy.
Why this matters
Courts often treat… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/legal-eyewitness-confidence-accuracy-coherence-decay-v0.1.e3-math-medhard-zero-accuracy-22026_01_22_results_accuracy_compareSee https://github.com/cephcyn/SteerEval for the main description.
Please cite our paper if you use this in your own work:
@misc{zhou2026steerevalframeworkevaluatingsteerability,
title={SteerEval: A Framework for Evaluating Steerability with Natural Language Profiles for Recommendation},
author={Joyce Zhou and Weijie Zhou and Doug Turnbull and Thorsten Joachims},
year={2026},
eprint={2601.21105},
archivePrefix={arXiv},
primaryClass={cs.IR}… See the full description on the dataset page: https://huggingface.co/datasets/cephcyn/2026_01_22_results_accuracy_compare.L2EnglishScoring_speechocean762_accuracy_v2NuInstruct_test_accuracytasksummarize_from_feedback_oai_preprocessing_gpt2_accuracy_1757233813
Dataset Card for "summarize_from_feedback_oai_preprocessing_gpt2_accuracy_1757233813"
More Information needed
L2EnglishScoring_speechocean762_accuracyL2EnglishScoring_speechocean762_accuracyQG_vs_QA_Verifier_gold_accuracy
Dataset Card for "QG_vs_QA_Verifier_gold_accuracy"
More Information needed
nuinstruct_test_accuracytask
