datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
long-doc_book_enAvailable Versions:
AIR-Bench_24.04
Task / Domain / Language: long-doc / book / en
Available Datasets (Dataset Name: Splits):
origin-of-species_darwin: test
a-brief-history-of-time_stephen-hawking: test
AIR-Bench_24.05
Task / Domain / Language: long-doc / book / en
Available Datasets (Dataset Name: Splits):
origin-of-species_darwin: test
a-brief-history-of-time_stephen-hawking: dev
long-doc_law_enAvailable Versions:
AIR-Bench_24.04
Task / Domain / Language: long-doc / law / en
Available Datasets (Dataset Name: Splits):
lex_files_300K-400K: test
lex_files_400K-500K: test
lex_files_500K-600K: test
lex_files_600K-700K: test
AIR-Bench_24.05
Task / Domain / Language: long-doc / law / en
Available Datasets (Dataset Name: Splits):
lex_files_300K-400K: dev
lex_files_400K-500K: test
lex_files_500K-600K: test
lex_files_600K-700K: test
long-doc_arxiv_enAvailable Versions:
AIR-Bench_24.04
Task / Domain / Language: long-doc / arxiv / en
Available Datasets (Dataset Name: Splits):
gpt3: test
llama2: test
gemini: test
llm-survey: test
AIR-Bench_24.05
Task / Domain / Language: long-doc / arxiv / en
Available Datasets (Dataset Name: Splits):
gpt3: test
llama2: dev
gemini: test
llm-survey: test
long-doc_healthcare_enAvailable Versions:
AIR-Bench_24.04
Task / Domain / Language: long-doc / healthcare / en
Available Datasets (Dataset Name: Splits):
pubmed_100K-200K_1: test
pubmed_100K-200K_2: test
pubmed_100K-200K_3: test
pubmed_40K-50K_5-merged: test
pubmed_30K-40K_10-merged: test
AIR-Bench_24.05
Task / Domain / Language: long-doc / healthcare / en
Available Datasets (Dataset Name: Splits):
pubmed_100K-200K_1: test
pubmed_100K-200K_2: test
pubmed_100K-200K_3: dev
pubmed_40K-50K_5-merged: test… See the full description on the dataset page: https://huggingface.co/datasets/AIR-Bench/long-doc_healthcare_en.qrels-long-doc_book_en-devAvailable Versions:
AIR-Bench_24.04
Task / Domain / Language: long-doc / book / en
Available Datasets (Dataset Name: Splits):
origin-of-species_darwin: test
a-brief-history-of-time_stephen-hawking: test
AIR-Bench_24.05
Task / Domain / Language: long-doc / book / en
Available Datasets (Dataset Name: Splits):
origin-of-species_darwin: test
a-brief-history-of-time_stephen-hawking: dev
xfund-docai-xl
xfund-docai-xl
An adaptation of XFUND for key information extraction and document classification DocAI tasks with augmented generated documents.
Note: Annotations and Synthetic documents have been generated by Claude Opus 5 under my supervision - but errors and biases may be here and there. The dataset is intended for experimenting with LLMs finetuning for DocAI.
If you spot something odd and would like to contribute to improve the quality, feel free to open an issue!
Seven… See the full description on the dataset page: https://huggingface.co/datasets/andreagemelli/xfund-docai-xl.mckinsey_state_of_ai_doc_understanding
Mckinsey State Of Ai Doc Understanding
This dataset was generated using YourBench (v0.3.1), an open-source framework for generating domain-specific benchmarks from document collections.
Pipeline Steps
ingestion: Read raw source documents, convert them to normalized markdown and save for downstream steps
summarization: Perform hierarchical summarization: chunk-level LLM summaries followed by combine-stage reduction
chunking: Split texts into token-based single-hop and… See the full description on the dataset page: https://huggingface.co/datasets/yourbench/mckinsey_state_of_ai_doc_understanding.doc_calibration_datasetazure-ai-engineer-doc-loaderqrels-long-doc_arxiv_en-devAvailable Versions:
AIR-Bench_24.04
Task / Domain / Language: long-doc / arxiv / en
Available Datasets (Dataset Name: Splits):
gpt3: test
llama2: test
gemini: test
llm-survey: test
AIR-Bench_24.05
Task / Domain / Language: long-doc / arxiv / en
Available Datasets (Dataset Name: Splits):
gpt3: test
llama2: dev
gemini: test
llm-survey: test
clean-doc-benchaihub_retriever_doc_smr문서요약 텍스트
qrels-long-doc_law_en-devAvailable Versions:
AIR-Bench_24.04
Task / Domain / Language: long-doc / law / en
Available Datasets (Dataset Name: Splits):
lex_files_300K-400K: test
lex_files_400K-500K: test
lex_files_500K-600K: test
lex_files_600K-700K: test
AIR-Bench_24.05
Task / Domain / Language: long-doc / law / en
Available Datasets (Dataset Name: Splits):
lex_files_300K-400K: dev
lex_files_400K-500K: test
lex_files_500K-600K: test
lex_files_600K-700K: test
Viet-Doc-VQA-II-flash2
Dataset Overview
This dataset is a continuation of the ongoing work from Viet Document VAQ dataset was collected from 64,765 pages of Vietnamese 🇻🇳 textbooks( Sách bài tập, chuyên đề, sách giáo án của Bộ GDĐT, Cánh Diều, Chân trời sáng tạo, Kết nối tri thức), spanning all subjects from grades 1 to 12. Each page has been analyzed and annotated using advanced Visual Question Answering (VQA) techniques to produce a comprehensive dataset.
There is a set of 388,277 detailed… See the full description on the dataset page: https://huggingface.co/datasets/5CD-AI/Viet-Doc-VQA-II-flash2.qrels-long-doc_healthcare_en-devAvailable Versions:
AIR-Bench_24.04
Task / Domain / Language: long-doc / healthcare / en
Available Datasets (Dataset Name: Splits):
pubmed_100K-200K_1: test
pubmed_100K-200K_2: test
pubmed_100K-200K_3: test
pubmed_40K-50K_5-merged: test
pubmed_30K-40K_10-merged: test
AIR-Bench_24.05
Task / Domain / Language: long-doc / healthcare / en
Available Datasets (Dataset Name: Splits):
pubmed_100K-200K_1: test
pubmed_100K-200K_2: test
pubmed_100K-200K_3: dev
pubmed_40K-50K_5-merged: test… See the full description on the dataset page: https://huggingface.co/datasets/AIR-Bench/qrels-long-doc_healthcare_en-dev.FIR-Bench-Sin-Doc-FinQAmdpbench-doc-ocr-sft
mdpbench-doc-ocr-sft
Qwen 4B (Qwen3-VL) MDPBench 한국어/일본어 문서 OCR/파싱 SFT 학습 데이터(공개 가능분).
재현 코드: https://github.com/sionic-ai/qwen-mdpbench-ocr
Configs
config
rows
내용
상태
ko
2,407
한국 지자체 소식지(신문형) 이미지 + Markdown GT
공개
jp
—
일본어 (추후)
예정
스키마
image (Image, bytes 임베드) — 문서 페이지
markdown (string) — GT 전사(Markdown, 표/읽기순서 보존)
source (string) — snvision / seocho / gongju
lang (string) — ko
doc_type (string) — newspaper
출처 & 라이선스… See the full description on the dataset page: https://huggingface.co/datasets/sionic-ai/mdpbench-doc-ocr-sft.Viet-Doc-VQA-flash2
Dataset Overview
The Document VAQ dataset was collected from 51,856 pages of Vietnamese 🇻🇳 textbooks( Sách Bộ GDĐT, Cánh Diều, Chân trời sáng tạo, Kết nối tri thức), spanning all subjects from grades 1 to 12. Each page has been analyzed and annotated using advanced Visual Question Answering (VQA) techniques to produce a comprehensive dataset.
There is a set of 310,952 detailed descriptions and query-based questions and answers generated by the Gemini 1.5 Flash model, currently… See the full description on the dataset page: https://huggingface.co/datasets/5CD-AI/Viet-Doc-VQA-flash2.ai_doc_tunehf_doc_qa_eval_chunk_size_800_open_aiFedE4OODRAG4FINViet-Doc-VQA-verIII
Cite
@misc{doan2024vintern1befficientmultimodallarge,
title={Vintern-1B: An Efficient Multimodal Large Language Model for Vietnamese},
author={Khang T. Doan and Bao G. Huynh and Dung T. Hoang and Thuc D. Pham and Nhat H. Pham and Quan T. M. Nguyen and Bang Q. Vo and Suong N. Hoang},
year={2024},
eprint={2408.12480},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2408.12480},
}
hf_doc_qa_eval_best_answersFedE4LEGFedE4FIN
