datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
t2-ragbench
Dataset Card for T2-RAGBench
Project Page | Paper | Code
IMPORTANT NOTICE:
We deleted VQAonBD from the dataset due to low quality of the question reformulations. If you still want to use it you will find the data in the previous commit history.
Dataset Description
Dataset Summary
T2-RAGBench is a benchmark dataset designed to evaluate Retrieval-Augmented Generation (RAG) on financial documents containing both text and tables. It consists of 23,088… See the full description on the dataset page: https://huggingface.co/datasets/botay/t2-ragbench.olmo-igsm-arith
OLMo iGSM-Easy Arithmetic
This repository contains a frozen, evaluation-only release of the synthetic
mod-7 arithmetic task called iGSM-Easy Arithmetic in the accompanying OLMo
evaluation code. It contains 750 examples: 250 examples at each target depth
2, 3, and 4.
This is an i-GSM-style task variant, not a claim to be an official release
of another dataset named iGSM. The olmo-igsm-arith name is used to make the
implementation provenance explicit.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Mihara-bot/olmo-igsm-arith.dacr-bench-results
DACR-Bench Results: Synthetic-to-Real Document Reasoning Transfer
Evaluation results demonstrating that fine-tuning a 7B model on 4,421 procedurally generated document reasoning traces more than doubles accuracy on real arXiv papers (18.9% → 40.0% on real documents in DACR-Bench), with gains concentrated in multi-hop reasoning, numerical computation, and causal authority resolution under conflicting information. All results reported below are on real arXiv documents only (9… See the full description on the dataset page: https://huggingface.co/datasets/botcoinmoney/dacr-bench-results.
