CoolFace
26 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AIR-Bench /long-doc_book_enAvailable Versions: AIR-Bench_24.04 Task / Domain / Language: long-doc / book / en Available Datasets (Dataset Name: Splits): origin-of-species_darwin: test a-brief-history-of-time_stephen-hawking: test AIR-Bench_24.05 Task / Domain / Language: long-doc / book / en Available Datasets (Dataset Name: Splits): origin-of-species_darwin: test a-brief-history-of-time_stephen-hawking: dev texttext-retrieval1K<n<10K0 likes630 downloads2y agoHugging Face02AIR-Bench /long-doc_law_enAvailable Versions: AIR-Bench_24.04 Task / Domain / Language: long-doc / law / en Available Datasets (Dataset Name: Splits): lex_files_300K-400K: test lex_files_400K-500K: test lex_files_500K-600K: test lex_files_600K-700K: test AIR-Bench_24.05 Task / Domain / Language: long-doc / law / en Available Datasets (Dataset Name: Splits): lex_files_300K-400K: dev lex_files_400K-500K: test lex_files_500K-600K: test lex_files_600K-700K: test texttext-retrieval10K<n<100K1 likes132 downloads2y agoHugging Face03AIR-Bench /long-doc_arxiv_enAvailable Versions: AIR-Bench_24.04 Task / Domain / Language: long-doc / arxiv / en Available Datasets (Dataset Name: Splits): gpt3: test llama2: test gemini: test llm-survey: test AIR-Bench_24.05 Task / Domain / Language: long-doc / arxiv / en Available Datasets (Dataset Name: Splits): gpt3: test llama2: dev gemini: test llm-survey: test texttext-retrieval1K<n<10K1 likes127 downloads2y agoHugging Face04AIR-Bench /long-doc_healthcare_enAvailable Versions: AIR-Bench_24.04 Task / Domain / Language: long-doc / healthcare / en Available Datasets (Dataset Name: Splits): pubmed_100K-200K_1: test pubmed_100K-200K_2: test pubmed_100K-200K_3: test pubmed_40K-50K_5-merged: test pubmed_30K-40K_10-merged: test AIR-Bench_24.05 Task / Domain / Language: long-doc / healthcare / en Available Datasets (Dataset Name: Splits): pubmed_100K-200K_1: test pubmed_100K-200K_2: test pubmed_100K-200K_3: dev pubmed_40K-50K_5-merged: test… See the full description on the dataset page: https://huggingface.co/datasets/AIR-Bench/long-doc_healthcare_en.texttext-retrieval10K<n<100K1 likes115 downloads2y agoHugging Face05AIR-Bench /qrels-long-doc_book_en-devAvailable Versions: AIR-Bench_24.04 Task / Domain / Language: long-doc / book / en Available Datasets (Dataset Name: Splits): origin-of-species_darwin: test a-brief-history-of-time_stephen-hawking: test AIR-Bench_24.05 Task / Domain / Language: long-doc / book / en Available Datasets (Dataset Name: Splits): origin-of-species_darwin: test a-brief-history-of-time_stephen-hawking: dev textn<1K1 likes94 downloads2y agoHugging Face06andreagemelli /xfund-docai-xl xfund-docai-xl An adaptation of XFUND for key information extraction and document classification DocAI tasks with augmented generated documents. Note: Annotations and Synthetic documents have been generated by Claude Opus 5 under my supervision - but errors and biases may be here and there. The dataset is intended for experimenting with LLMs finetuning for DocAI. If you spot something odd and would like to contribute to improve the quality, feel free to open an issue! Seven… See the full description on the dataset page: https://huggingface.co/datasets/andreagemelli/xfund-docai-xl.text1K<n<10K0 likes92 downloads21d agoHugging Face07yourbench /mckinsey_state_of_ai_doc_understanding Mckinsey State Of Ai Doc Understanding This dataset was generated using YourBench (v0.3.1), an open-source framework for generating domain-specific benchmarks from document collections. Pipeline Steps ingestion: Read raw source documents, convert them to normalized markdown and save for downstream steps summarization: Perform hierarchical summarization: chunk-level LLM summaries followed by combine-stage reduction chunking: Split texts into token-based single-hop and… See the full description on the dataset page: https://huggingface.co/datasets/yourbench/mckinsey_state_of_ai_doc_understanding.tabularn<1K0 likes81 downloads1y agoHugging Face08bakhil-aissa /doc_calibration_datasetimage1K<n<10K0 likes67 downloads1mo agoHugging Face09dwb2023 /azure-ai-engineer-doc-loadertextn<1K0 likes41 downloads1y agoHugging Face10AIR-Bench /qrels-long-doc_arxiv_en-devAvailable Versions: AIR-Bench_24.04 Task / Domain / Language: long-doc / arxiv / en Available Datasets (Dataset Name: Splits): gpt3: test llama2: test gemini: test llm-survey: test AIR-Bench_24.05 Task / Domain / Language: long-doc / arxiv / en Available Datasets (Dataset Name: Splits): gpt3: test llama2: dev gemini: test llm-survey: test textn<1K0 likes40 downloads2y agoHugging Face11ai4bharat /clean-doc-benchtext1K<n<10K0 likes36 downloads5mo agoHugging Face12wisenut-nlp-team /aihub_retriever_doc_smr문서요약 텍스트 text100K<n<1M0 likes34 downloads2y agoHugging Face13AIR-Bench /qrels-long-doc_law_en-devAvailable Versions: AIR-Bench_24.04 Task / Domain / Language: long-doc / law / en Available Datasets (Dataset Name: Splits): lex_files_300K-400K: test lex_files_400K-500K: test lex_files_500K-600K: test lex_files_600K-700K: test AIR-Bench_24.05 Task / Domain / Language: long-doc / law / en Available Datasets (Dataset Name: Splits): lex_files_300K-400K: dev lex_files_400K-500K: test lex_files_500K-600K: test lex_files_600K-700K: test text1K<n<10K0 likes25 downloads2y agoHugging Face145CD-AI /Viet-Doc-VQA-II-flash2gated Dataset Overview This dataset is a continuation of the ongoing work from Viet Document VAQ dataset was collected from 64,765 pages of Vietnamese 🇻🇳 textbooks( Sách bài tập, chuyên đề, sách giáo án của Bộ GDĐT, Cánh Diều, Chân trời sáng tạo, Kết nối tri thức), spanning all subjects from grades 1 to 12. Each page has been analyzed and annotated using advanced Visual Question Answering (VQA) techniques to produce a comprehensive dataset. There is a set of 388,277 detailed… See the full description on the dataset page: https://huggingface.co/datasets/5CD-AI/Viet-Doc-VQA-II-flash2.imagevisual-question-answering10K<n<100K6 likes21 downloads8mo agoHugging Face15AIR-Bench /qrels-long-doc_healthcare_en-devAvailable Versions: AIR-Bench_24.04 Task / Domain / Language: long-doc / healthcare / en Available Datasets (Dataset Name: Splits): pubmed_100K-200K_1: test pubmed_100K-200K_2: test pubmed_100K-200K_3: test pubmed_40K-50K_5-merged: test pubmed_30K-40K_10-merged: test AIR-Bench_24.05 Task / Domain / Language: long-doc / healthcare / en Available Datasets (Dataset Name: Splits): pubmed_100K-200K_1: test pubmed_100K-200K_2: test pubmed_100K-200K_3: dev pubmed_40K-50K_5-merged: test… See the full description on the dataset page: https://huggingface.co/datasets/AIR-Bench/qrels-long-doc_healthcare_en-dev.textn<1K0 likes19 downloads2y agoHugging Face16valuesimplex-ai-lab /FIR-Bench-Sin-Doc-FinQAtext1K<n<10K0 likes19 downloads1y agoHugging Face17sionic-ai /mdpbench-doc-ocr-sft mdpbench-doc-ocr-sft Qwen 4B (Qwen3-VL) MDPBench 한국어/일본어 문서 OCR/파싱 SFT 학습 데이터(공개 가능분). 재현 코드: https://github.com/sionic-ai/qwen-mdpbench-ocr Configs config rows 내용 상태 ko 2,407 한국 지자체 소식지(신문형) 이미지 + Markdown GT 공개 jp — 일본어 (추후) 예정 스키마 image (Image, bytes 임베드) — 문서 페이지 markdown (string) — GT 전사(Markdown, 표/읽기순서 보존) source (string) — snvision / seocho / gongju lang (string) — ko doc_type (string) — newspaper 출처 & 라이선스… See the full description on the dataset page: https://huggingface.co/datasets/sionic-ai/mdpbench-doc-ocr-sft.imageimage-to-text1K<n<10K2 likes14 downloads3mo agoHugging Face185CD-AI /Viet-Doc-VQA-flash2gated Dataset Overview The Document VAQ dataset was collected from 51,856 pages of Vietnamese 🇻🇳 textbooks( Sách Bộ GDĐT, Cánh Diều, Chân trời sáng tạo, Kết nối tri thức), spanning all subjects from grades 1 to 12. Each page has been analyzed and annotated using advanced Visual Question Answering (VQA) techniques to produce a comprehensive dataset. There is a set of 310,952 detailed descriptions and query-based questions and answers generated by the Gemini 1.5 Flash model, currently… See the full description on the dataset page: https://huggingface.co/datasets/5CD-AI/Viet-Doc-VQA-flash2.imagevisual-question-answering10K<n<100K4 likes12 downloads8mo agoHugging Face19Jynpogger /ai_doc_tunetext1K<n<10K0 likes11 downloads2y agoHugging Face20AIEnthusiast369 /hf_doc_qa_eval_chunk_size_800_open_aitextn<1K0 likes9 downloads2y agoHugging Face21DocAILab /FedE4OODtabular10K<n<100K0 likes7 downloads8mo agoHugging Face22DocAILab /RAG4FINtabularn<1K0 likes7 downloads8mo agoHugging Face235CD-AI /Viet-Doc-VQA-verIIIgated Cite @misc{doan2024vintern1befficientmultimodallarge, title={Vintern-1B: An Efficient Multimodal Large Language Model for Vietnamese}, author={Khang T. Doan and Bao G. Huynh and Dung T. Hoang and Thuc D. Pham and Nhat H. Pham and Quan T. M. Nguyen and Bang Q. Vo and Suong N. Hoang}, year={2024}, eprint={2408.12480}, archivePrefix={arXiv}, primaryClass={cs.LG}, url={https://arxiv.org/abs/2408.12480}, } image100K<n<1M3 likes6 downloads2y agoHugging Face24AIEnthusiast369 /hf_doc_qa_eval_best_answerstabularn<1K0 likes3 downloads2y agoHugging Face25DocAILab /FedE4LEGtext10K<n<100K0 likes3 downloads8mo agoHugging Face26DocAILab /FedE4FINtextn<1K0 likes1 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.