CoolFace
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hashmortar /spreadsheet-bench-v2-modified SpreadsheetBench V2 Modified: Multi-Document QA 1,060 questions and reference answers grounded in 127 Excel workbooks, 35 PDFs and 9 DOCX files. This independent derivative of SpreadsheetBench 2 shifts the task from editing spreadsheets and producing workbook deliverables toward finding, interpreting and combining information in business documents. An independent project built entirely from publicly available source material and newly authored QA annotations. No private company… See the full description on the dataset page: https://huggingface.co/datasets/hashmortar/spreadsheet-bench-v2-modified.documentquestion-answering1K<n<10K2 likes3.4k downloads18d agoHugging Face02minorproject-research /researchy_questions_modified Dataset Summary This dataset is derived from corbyrosset/researchy_questions, a collection of ~100k non-factoid, multi-perspective "Researchy Questions" mined from real Bing search engine logs. After a labor-intensive filtering funnel from billions of queries, these "needles in the haystack" are questions that probably require a lot of sub-questions and research to answer adequately, and are shown to be harder than other open-domain QA datasets like Natural Questions.… See the full description on the dataset page: https://huggingface.co/datasets/minorproject-research/researchy_questions_modified.textquestion-answering10K<n<100K0 likes37 downloads1mo agoHugging Face03drguolai /distill_r1_110k_sft_modifiedBorrowed from https://huggingface.co/datasets/Congliu/Chinese-DeepSeek-R1-Distill-data-110k-SFT Fix the <image> placeholder issue, which will cause error during training: raise ValueError(f"The number of images does not match the number of {IMAGE_PLACEHOLDER} tokens.") tabularquestion-answering100K<n<1M0 likes20 downloads2y agoHugging Face04romikgosai /squad_modified_2textquestion-answering10K<n<100K0 likes10 downloads2y agoHugging Face05AbrarHyder /Modified_german_dpr_dataset Original Dataset The original dataset, deepset/germandpr, contains: 9275 training examples 1025 testing examples Each example is a question/answer pair, consisting of: One question One answer One positive context Three negative contexts You can find the original dataset here. Modifications Adding Easy Negative Examples To enhance the dataset, an "easy negative example" was added to each row. The objective of this addition is to train the model to… See the full description on the dataset page: https://huggingface.co/datasets/AbrarHyder/Modified_german_dpr_dataset.textquestion-answeringn<1K0 likes9 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.