CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jinaai /docqa_artificial_intelligence_beirThis is a copy of https://huggingface.co/datasets/jinaai/docqa_artificial_intelligence reformatted into the BEIR format. For any further information like license, please refer to the original dataset. Disclaimer This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/docqa_artificial_intelligence_beir.image1K<n<10K0 likes264 downloads1y agoHugging Face02jinaai /docqa_energy_beirThis is a copy of https://huggingface.co/datasets/jinaai/docqa_energy reformatted into the BEIR format. For any further information like license, please refer to the original dataset. Disclaimer This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at) jina.ai" for… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/docqa_energy_beir.image1K<n<10K0 likes264 downloads1y agoHugging Face03jinaai /docqa_healthcare_industry_beirThis is a copy of https://huggingface.co/datasets/jinaai/docqa_healthcare_industry reformatted into the BEIR format. For any further information like license, please refer to the original dataset. Disclaimer This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at)… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/docqa_healthcare_industry_beir.image1K<n<10K0 likes259 downloads1y agoHugging Face04jinaai /docqa_gov_report_beirThis is a copy of https://huggingface.co/datasets/jinaai/docqa_gov_report reformatted into the BEIR format. For any further information like license, please refer to the original dataset. Disclaimer This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at) jina.ai"… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/docqa_gov_report_beir.image1K<n<10K0 likes255 downloads1y agoHugging Face05Tongyi-Zhiwen /DocQA-RL-1.6KTo construct a challenging RL dataset for verifiable long-context reasoning, we develop 🤗 DocQA-RL-1.6K, which comprises 1.6K DocQA problems across three reasoning domains: (1) Mathematical Reasoning: We use 600 problems from the DocMath dataset, requiring numerical reasoning across long and specialized documents such as financial reports. For DocMath, we sample 75% items from each subset from its valid split for training and 25% for evaluation; (2) Logical Reasoning: We employ DeepSeek-R1… See the full description on the dataset page: https://huggingface.co/datasets/Tongyi-Zhiwen/DocQA-RL-1.6K.text1K<n<10K42 likes252 downloads1y agoHugging Face06EtashGuha /DocQA_XMLimage1K<n<10K3 likes88 downloads2y agoHugging Face07sungyub /docqa-rl-verl DocQA-RL-1.6K (VERL Format) This dataset contains 1,591 challenging long-context document QA problems from DocQA-RL-1.6K, converted to VERL (Volcano Engine Reinforcement Learning) format for reinforcement learning training workflows. Source: Tongyi-Zhiwen/DocQA-RL-1.6K License: Apache 2.0 Note: This dataset maintains the original high-quality structure with user-only messages. The extra_info field has been standardized to contain only the index field for consistency with other VERL… See the full description on the dataset page: https://huggingface.co/datasets/sungyub/docqa-rl-verl.textreinforcement-learning1K<n<10K0 likes33 downloads9mo agoHugging Face08jinaai /docqa_energy Creation This dataset is build upon the corresponding dataset from the ViDoRe Benchmark. For more information regarding the filtering please read our paper or this discussion on github. Disclaimer This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/docqa_energy.imagen<1K1 likes32 downloads1y agoHugging Face09myngsoooo /CorningAI-DocQA Dataset Card for "CorningAI-DocQA" More Information needed text1K<n<10K0 likes29 downloads3y agoHugging Face10U4R /DocGenome-Testset-DocQA DocGenome-TestSet-DocQA Firstly, you need to download the test dataset here. Then, unzip the DocGenome-Testset-DocQA.zip as follows: DocGenome-Testset-DocQA ├── testset │ ├── xxx ├── eval_tools │ ├── eval_open_docqa_gpt.py │ ├── eval_normal_docqa.py ├── qa_info │ ├── docgenome_testset_multiqa.json │ ├── docgenome_testset_singleqa.json │ ├── docgenome_testset_normalqa.jsonl ├── example │ ├── internvl_open_docqa_test.py │ ├── internvl_normal_docqa_test.py ├──… See the full description on the dataset page: https://huggingface.co/datasets/U4R/DocGenome-Testset-DocQA.image1K<n<10K4 likes28 downloads2y agoHugging Face11jinaai /docqa_gov_report Creation This dataset is build upon the corresponding dataset from the ViDoRe Benchmark. For more information regarding the filtering please read our paper or this discussion on github. Disclaimer This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/docqa_gov_report.imagen<1K0 likes27 downloads1y agoHugging Face12jinaai /docqa_artificial_intelligence Creation This dataset is build upon the corresponding dataset from the ViDoRe Benchmark. For more information regarding the filtering please read our paper or this discussion on github. Disclaimer This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/docqa_artificial_intelligence.imagen<1K0 likes26 downloads1y agoHugging Face13jinaai /docqa_healthcare_industry Creation This dataset is build upon the corresponding dataset from the ViDoRe Benchmark. For more information regarding the filtering please read our paper or this discussion on github. Disclaimer This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/docqa_healthcare_industry.imagen<1K0 likes21 downloads1y agoHugging Face14shreyashankar /doc-qa-rl-datasets Document Question-Answering Dataset This dataset combines and transforms the QASPER and NarrativeQA datasets into a unified format for document-based question answering tasks. Dataset Description This dataset is designed for training and evaluating models on document-level question answering with source attribution. Each entry contains: A question about a document A corresponding answer Source text passages from the document that support the answer Position information… See the full description on the dataset page: https://huggingface.co/datasets/shreyashankar/doc-qa-rl-datasets.textn<1K0 likes14 downloads1y agoHugging Face15jinaai /docqa_healthcare_industry_deprecated Disclaimer This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at) jina.ai" for removal. We do not collect or process personal, sensitive, or private information intentionally. If you believe this dataset includes such content (e.g., portraits, location-linked… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/docqa_healthcare_industry_deprecated.imagen<1K0 likes12 downloads1y agoHugging Face16jinaai /docqa_gov_report_deprecated Disclaimer This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at) jina.ai" for removal. We do not collect or process personal, sensitive, or private information intentionally. If you believe this dataset includes such content (e.g., portraits, location-linked… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/docqa_gov_report_deprecated.imagen<1K0 likes12 downloads1y agoHugging Face17jinaai /docqa_artificial_intelligence_deprecated Disclaimer This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at) jina.ai" for removal. We do not collect or process personal, sensitive, or private information intentionally. If you believe this dataset includes such content (e.g., portraits, location-linked… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/docqa_artificial_intelligence_deprecated.imagen<1K0 likes11 downloads1y agoHugging Face18jinaai /docqa_energy_deprecated Disclaimer This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at) jina.ai" for removal. We do not collect or process personal, sensitive, or private information intentionally. If you believe this dataset includes such content (e.g., portraits, location-linked… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/docqa_energy_deprecated.imagen<1K0 likes9 downloads1y agoHugging Face19dhiruHF /docqa_train Dataset Card for "docqa_train" More Information needed text1K<n<10K0 likes8 downloads3y agoHugging Face20SantiagoPG /doc_qatabular10K<n<100K2 likes8 downloads3y agoHugging Face21shreyashankar /doc-qa-gpt4o-rollouts-1749712307textn<1K0 likes8 downloads1y agoHugging Face22zechen-nlp /DocQA-RL-1.6Ktext1K<n<10K0 likes8 downloads9mo agoHugging Face23dhiruHF /DocQA-demo-dataset Dataset Card for "DocQA-demo-dataset" More Information needed textn<1K0 likes7 downloads3y agoHugging Face24nur-dev /kazadmin-docqagated DATASET: Kazakh administrative documents for RAG document QA. Structure. Each item is a JSON object with: text: the full Kazakh document body (biography or power-of-attorney). category: document type label — e.g., Өмірбаян (autobiographical CV/biography) and Сенімхат (power of attorney) etc. In Kazakh admin usage, Өмірбаян is a concise, chronological personal record; Сенімхат is a written authorization to act on someone’s behalf. extended_answer: list of {user, answer} QA pairs… See the full description on the dataset page: https://huggingface.co/datasets/nur-dev/kazadmin-docqa.text10K<n<100K0 likes7 downloads1y agoHugging Face25dhiruHF /DocQA-dataset-300-samples Dataset Card for "DocQA-dataset-300-samples" More Information needed textn<1K0 likes6 downloads3y agoHugging Face26thng292 /doc-qatextn<1K0 likes5 downloads1y agoHugging Face27Nayana-cognitivelab /Nayana-DocQA-gu-10k-v1-docmatixgatedimage10K<n<100K0 likes3 downloads2y agoHugging Face28Nayana-cognitivelab /Nayana-DocQA-hi-10k-v1-docmatixgatedimage10K<n<100K0 likes3 downloads2y agoHugging Face29Nayana-cognitivelab /Nayana-DocQA-en-10k-v1-docmatixgatedimage10K<n<100K0 likes2 downloads2y agoHugging Face30Nayana-cognitivelab /Nayana-DocQA-or-10k-v1-docmatixgatedimage10K<n<100K0 likes2 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.