CoolFace
21 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01vblagoje /lfqatext100K<n<1M16 likes197 downloads5y agoHugging Face02vblagoje /lfqa_support_docsSupport documents for building https://huggingface.co/vblagoje/bart_lfqa model text100K<n<1M6 likes152 downloads5y agoHugging Face03indonesian-nlp /lfqa_idtext100K<n<1M4 likes92 downloads5y agoHugging Face04efederici /lfqa-preprocessed-ittextquestion-answering10K<n<100K2 likes44 downloads3y agoHugging Face05LLukas22 /lfqa_preprocessed Dataset Card for "lfqa_preprocessed" Dataset Summary This is a simplified version of vblagoje's lfqa_support_docs and lfqa datasets. It was generated by me to have a more straight forward way to train Seq2Seq models on context based long form question answering tasks. Dataset Structure Data Instances An example of 'train' looks as follows. { "question": "what's the difference between a forest and a wood?", "answer": "They're used… See the full description on the dataset page: https://huggingface.co/datasets/LLukas22/lfqa_preprocessed.textquestion-answering100K<n<1M2 likes39 downloads4y agoHugging Face06aitetic /eli5-lfqa ELI5: Long Form Question Answering https://arxiv.org/pdf/1907.09190 Items (train): 272_634 Items (val): 1_507 Downloaded from: https://www.kaggle.com/datasets/trandaiphu/eli5-dataset Tokens ctxs length distribution summary: min: 152 mean: 3745.35 median (p50): 3736.00 p90: 4037.00 p95: 4152.00 p99: 4432.00 max: 6530 Tokens answers length distribution summary: rows: 272_634 min: 8 mean: 401.31 median (p50): 217.00 p90: 820.70 p95: 1230.00 p99:… See the full description on the dataset page: https://huggingface.co/datasets/aitetic/eli5-lfqa.text100K<n<1M0 likes34 downloads2mo agoHugging Face07allenai /intent-aware-lfqa-intent-implicittext1K<n<10K1 likes20 downloads6mo agoHugging Face08aitetic /eli5-lfqa-combined eli5-lfqa-combined Postprocessed from eli5-lfqa: (ctxs, question+answers[]) Assembled by: dataset_assembler.py Size: 1.1B text100K<n<1M0 likes18 downloads2mo agoHugging Face09nlpatunt /lfqa-textbuggertextn<1K0 likes16 downloads7mo agoHugging Face10allenai /intent-aware-lfqa-multiviewtext1K<n<10K1 likes14 downloads6mo agoHugging Face11allenai /intent-aware-lfqa-baselinetext1K<n<10K1 likes12 downloads6mo agoHugging Face12fahdsoliman /NAITS_LFQA_with_supportstabularn<1K0 likes11 downloads2y agoHugging Face13allenai /intent-aware-lfqa-intent-explicittext1K<n<10K1 likes11 downloads6mo agoHugging Face14SIA86 /LFQAKnowledgeBasetextquestion-answeringn<1K0 likes9 downloads3y agoHugging Face15nlpatunt /LFQA-HP-1M-Sampletext1K<n<10K0 likes9 downloads1y agoHugging Face16raphaelmerx /lfqa-idtextn<1K0 likes7 downloads5y agoHugging Face17fahdsoliman /lfqa_test_with_supports_v1textn<1K0 likes7 downloads2y agoHugging Face18fahdsoliman /NAITS_LFQA_with_supports_v2 NAITS_LFQA_with_supports_v2 Dataset Description NAITS_LFQA_with_supports_v2 is a bilingual Arabic-English dataset designed for Long-Form Question Answering (LFQA) and Retrieval-Augmented Generation (RAG) research. The dataset contains 102 manually curated samples. Each sample consists of: A question in Arabic and English. A long-form answer in Arabic and English. Supporting passages in Arabic and English from which the answer can be derived. A reference field… See the full description on the dataset page: https://huggingface.co/datasets/fahdsoliman/NAITS_LFQA_with_supports_v2.tabularquestion-answeringn<1K0 likes7 downloads4mo agoHugging Face19fahdsoliman /lfqa_datasettext1K<n<10K0 likes4 downloads2y agoHugging Face20fahdsoliman /lfqa_test_with_supports_v2textn<1K0 likes3 downloads2y agoHugging Face21nlpatunt /lfqa-deepwordbugtextn<1K0 likes3 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.