datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
docqa_artificial_intelligence_beirThis is a copy of https://huggingface.co/datasets/jinaai/docqa_artificial_intelligence reformatted into the BEIR format. For any further information like license, please refer to the original dataset.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/docqa_artificial_intelligence_beir.docqa_energy_beirThis is a copy of https://huggingface.co/datasets/jinaai/docqa_energy reformatted into the BEIR format. For any further information like license, please refer to the original dataset.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at) jina.ai" for… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/docqa_energy_beir.docqa_healthcare_industry_beirThis is a copy of https://huggingface.co/datasets/jinaai/docqa_healthcare_industry reformatted into the BEIR format. For any further information like license, please refer to the original dataset.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at)… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/docqa_healthcare_industry_beir.docqa_gov_report_beirThis is a copy of https://huggingface.co/datasets/jinaai/docqa_gov_report reformatted into the BEIR format. For any further information like license, please refer to the original dataset.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at) jina.ai"… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/docqa_gov_report_beir.DocQA-RL-1.6KTo construct a challenging RL dataset for verifiable long-context reasoning, we develop 🤗 DocQA-RL-1.6K, which comprises 1.6K DocQA problems across three reasoning domains:
(1) Mathematical Reasoning: We use 600 problems from the DocMath dataset, requiring numerical reasoning across long and specialized documents such as financial reports. For DocMath, we sample 75% items from each subset from its valid split for training and 25% for evaluation;
(2) Logical Reasoning: We employ DeepSeek-R1… See the full description on the dataset page: https://huggingface.co/datasets/Tongyi-Zhiwen/DocQA-RL-1.6K.DocQA_XMLdocqa-rl-verl
DocQA-RL-1.6K (VERL Format)
This dataset contains 1,591 challenging long-context document QA problems from DocQA-RL-1.6K, converted to VERL (Volcano Engine Reinforcement Learning) format for reinforcement learning training workflows.
Source: Tongyi-Zhiwen/DocQA-RL-1.6K
License: Apache 2.0
Note: This dataset maintains the original high-quality structure with user-only messages. The extra_info field has been standardized to contain only the index field for consistency with other VERL… See the full description on the dataset page: https://huggingface.co/datasets/sungyub/docqa-rl-verl.docqa_energy
Creation
This dataset is build upon the corresponding dataset from the ViDoRe Benchmark. For more information regarding the filtering please read our paper or this discussion on github.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/docqa_energy.CorningAI-DocQA
Dataset Card for "CorningAI-DocQA"
More Information needed
docqa_gov_report
Creation
This dataset is build upon the corresponding dataset from the ViDoRe Benchmark. For more information regarding the filtering please read our paper or this discussion on github.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/docqa_gov_report.docqa_artificial_intelligence
Creation
This dataset is build upon the corresponding dataset from the ViDoRe Benchmark. For more information regarding the filtering please read our paper or this discussion on github.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/docqa_artificial_intelligence.docqa_healthcare_industry
Creation
This dataset is build upon the corresponding dataset from the ViDoRe Benchmark. For more information regarding the filtering please read our paper or this discussion on github.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/docqa_healthcare_industry.doc-qa-rl-datasets
Document Question-Answering Dataset
This dataset combines and transforms the QASPER and NarrativeQA datasets into a unified format for document-based question answering tasks.
Dataset Description
This dataset is designed for training and evaluating models on document-level question answering with source attribution. Each entry contains:
A question about a document
A corresponding answer
Source text passages from the document that support the answer
Position information… See the full description on the dataset page: https://huggingface.co/datasets/shreyashankar/doc-qa-rl-datasets.docqa_healthcare_industry_deprecated
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at) jina.ai" for removal. We do not collect or process personal, sensitive, or private information intentionally. If you believe this dataset includes such content (e.g., portraits, location-linked… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/docqa_healthcare_industry_deprecated.docqa_gov_report_deprecated
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at) jina.ai" for removal. We do not collect or process personal, sensitive, or private information intentionally. If you believe this dataset includes such content (e.g., portraits, location-linked… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/docqa_gov_report_deprecated.docqa_artificial_intelligence_deprecated
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at) jina.ai" for removal. We do not collect or process personal, sensitive, or private information intentionally. If you believe this dataset includes such content (e.g., portraits, location-linked… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/docqa_artificial_intelligence_deprecated.docqa_energy_deprecated
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at) jina.ai" for removal. We do not collect or process personal, sensitive, or private information intentionally. If you believe this dataset includes such content (e.g., portraits, location-linked… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/docqa_energy_deprecated.docqa_train
Dataset Card for "docqa_train"
More Information needed
doc_qadoc-qa-gpt4o-rollouts-1749712307DocQA-RL-1.6KDocQA-demo-dataset
Dataset Card for "DocQA-demo-dataset"
More Information needed
kazadmin-docqa
DATASET: Kazakh administrative documents for RAG document QA.
Structure. Each item is a JSON object with:
text: the full Kazakh document body (biography or power-of-attorney).
category: document type label — e.g., Өмірбаян (autobiographical CV/biography) and Сенімхат (power of attorney) etc. In Kazakh admin usage, Өмірбаян is a concise, chronological personal record; Сенімхат is a written authorization to act on someone’s behalf.
extended_answer: list of {user, answer} QA pairs… See the full description on the dataset page: https://huggingface.co/datasets/nur-dev/kazadmin-docqa.DocQA-dataset-300-samples
Dataset Card for "DocQA-dataset-300-samples"
More Information needed
doc-qaNayana-DocQA-gu-10k-v1-docmatixNayana-DocQA-hi-10k-v1-docmatixNayana-DocQA-en-10k-v1-docmatixNayana-DocQA-or-10k-v1-docmatixNayana-DocQA-pa-10k-v1-docmatix
