datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
UDA-QA
Dataset Card for Dataset Name
[NIPS-2024] UDA: A Benchmark Suite for Retrieval Augmented Generation in Real-world Document Analysis (https://arxiv.org/abs/2406.15187)
UDA (Unstructured Document Analysis) is a benchmark suite for Retrieval Augmented Generation (RAG) in real-world document analysis.
Each entry in the UDA dataset is organized as a document-question-answer triplet, where a question is raised from the document, accompanied by a corresponding ground-truth answer.
The… See the full description on the dataset page: https://huggingface.co/datasets/qinchuanhui/UDA-QA.indian-regulatory-bfsi-benchmark-v1
Indian Regulatory BFSI Benchmark v1
A 60-question, hand-curated, openly licensed evaluation set for
extractive question answering over Indian financial regulation -
specifically Reserve Bank of India (RBI) Master Directions and
Securities and Exchange Board of India (SEBI) Master Circulars.
60 questions, 30 RBI / 30 SEBI
30 numeric / named-fact extraction (tier 2) + 30 heading-bound passage
questions (tier 3)
22 distinct source PDFs from a document-disjoint held-out split of… See the full description on the dataset page: https://huggingface.co/datasets/udit6969/indian-regulatory-bfsi-benchmark-v1.maritime-sft-mixed-formats
Maritime SFT Mixed Formats Dataset
High-quality supervised fine-tuning (SFT) dataset for the maritime domain,
generated from maritime books and technical documents using GLM-4.7 via NVIDIA NIM API.
Dataset Configs
Config
Format
Description
alpaca
Instruction / Input / Output
Standard Alpaca SFT format
chat
System / User / Assistant
ChatML multi-turn format
rag
Context / Question / Answer
RAG triad format for retrieval-augmented fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/uday568/maritime-sft-mixed-formats.
