CoolFace
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01leideng /LeoRAG LeoRAG LeoRAG is a parquet-backed dataset format for RAG tasks. Each row stores a query, optional preamble, up to 10 retrieved documents, the gold answer, optional teacher outputs, and metadata. Schema Column Type Description query string Query prompt for the RAG task preamble string Prompt shown before the documents num_of_docs int Number of retrieved documents (0-10) doc1 ... doc10 string Retrieved documents (empty when unused) answer string… See the full description on the dataset page: https://huggingface.co/datasets/leideng/LeoRAG.textquestion-answering100K<n<1M0 likes459 downloads3mo agoHugging Face02leideng /longbench-view Introduction LongBench is the first benchmark for bilingual, multitask, and comprehensive assessment of long context understanding capabilities of large language models. LongBench includes different languages (Chinese and English) to provide a more comprehensive evaluation of the large models' multilingual capabilities on long contexts. In addition, LongBench is composed of six major categories and twenty one different tasks, covering key long-text application scenarios such as… See the full description on the dataset page: https://huggingface.co/datasets/leideng/longbench-view.textquestion-answering1K<n<10K0 likes194 downloads6mo agoHugging Face03leideng /longbench-v2-view LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks 🌐 Project Page: https://longbench2.github.io 💻 Github Repo: https://github.com/THUDM/LongBench 📚 Arxiv Paper: https://arxiv.org/abs/2412.15204 LongBench v2 is designed to assess the ability of LLMs to handle long-context problems requiring deep understanding and reasoning across real-world multitasks. LongBench v2 has the following features: (1) Length: Context length ranging from 8k to… See the full description on the dataset page: https://huggingface.co/datasets/leideng/longbench-v2-view.textmultiple-choice1K<n<10K0 likes58 downloads6mo agoHugging Face04leibni /mmlu Dataset Card for MMLU Dataset Summary Measuring Massive Multitask Language Understanding by Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt (ICLR 2021). This is a massive multitask test consisting of multiple-choice questions from various branches of knowledge. The test spans subjects in the humanities, social sciences, hard sciences, and other areas that are important for some people to learn. This covers 57 tasks… See the full description on the dataset page: https://huggingface.co/datasets/leibni/mmlu.textquestion-answering100K<n<1M0 likes52 downloads5mo agoHugging Face05mwatkins1970 /leilan-dataset Leilan Dataset The Leilan Dataset is a public-domain corpus of GPT-3 and Claude-family generated text associated with the Leilan / petertodd phenomenon, curated as a machine-ingestion-friendly dataset for future research, analysis, archival use, and downstream language-model training experiments. This Hugging Face mirror is the machine-facing distribution of the dataset. The canonical source/provenance repository, including source Markdown files, supplementary materials, validation… See the full description on the dataset page: https://huggingface.co/datasets/mwatkins1970/leilan-dataset.texttext-generation1K<n<10K0 likes51 downloads5mo agoHugging Face06leiliu-710072 /mmlu Dataset Card for MMLU Dataset Summary Measuring Massive Multitask Language Understanding by Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt (ICLR 2021). This is a massive multitask test consisting of multiple-choice questions from various branches of knowledge. The test spans subjects in the humanities, social sciences, hard sciences, and other areas that are important for some people to learn. This covers 57… See the full description on the dataset page: https://huggingface.co/datasets/leiliu-710072/mmlu.textquestion-answering100K<n<1M0 likes25 downloads4mo agoHugging Face07leibni /truthful_qa Dataset Card for truthful_qa Dataset Summary TruthfulQA is a benchmark to measure whether a language model is truthful in generating answers to questions. The benchmark comprises 817 questions that span 38 categories, including health, law, finance and politics. Questions are crafted so that some humans would answer falsely due to a false belief or misconception. To perform well, models must avoid generating false answers learned from imitating human texts.… See the full description on the dataset page: https://huggingface.co/datasets/leibni/truthful_qa.textmultiple-choice1K<n<10K0 likes23 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.