CoolFace
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01alexandrainst /m_hellaswag Multilingual HellaSwag Dataset Summary This dataset is a machine translated version of the HellaSwag dataset. The Icelandic (is) part was translated with Miðeind's Greynir model and Norwegian (nb) was translated with DeepL. The rest of the languages was translated using GPT-3.5-turbo by the University of Oregon, and this part of the dataset was originally uploaded to this Github repository. textquestion-answering100K<n<1M7 likes6.7k downloads3y agoHugging Face02Hello-SimpleAI /HC3Human ChatGPT Comparison Corpus (HC3)texttext-classification10K<n<100K224 likes5.7k downloads4y agoHugging Face03Hello-SimpleAI /HC3-ChineseHuman ChatGPT Comparison Corpus (HC3) Chinese Versiontexttext-classification10K<n<100K176 likes1.4k downloads4y agoHugging Face04Hellisotherpeople /OpenDebateEvidence-Anonymized Dataset Card for OpenDebateEvidence (Anonymized) A collection of evidence used in collegiate and high school debate competitions, with all debater-identifying columns removed. This is an anonymized redistribution of Yusuf5/OpenCaselist. The argumentative content is byte-for-byte unchanged. 26 of the original 45 columns have been dropped. See Anonymization for exactly what was removed and why. Dataset Details Dataset Description This dataset is a… See the full description on the dataset page: https://huggingface.co/datasets/Hellisotherpeople/OpenDebateEvidence-Anonymized.tabulartext-generation1M<n<10M0 likes944 downloads2mo agoHugging Face05Hellisotherpeople /DebateSum DebateSum Corresponding code repo for the upcoming paper at ARGMIN 2020: "DebateSum: A large-scale argument mining and summarization dataset" Arxiv pre-print available here: https://arxiv.org/abs/2011.07251 Check out the presentation date and time here: https://argmining2020.i3s.unice.fr/node/9 Full paper as presented by the ACL is here: https://www.aclweb.org/anthology/2020.argmining-1.1/ Video of presentation at COLING 2020:… See the full description on the dataset page: https://huggingface.co/datasets/Hellisotherpeople/DebateSum.tabularquestion-answering100K<n<1M21 likes289 downloads4y agoHugging Face06Hellisotherpeople /OpenDebateEvidence-Deduplicated-Anonymized Dataset Card for OpenDebateEvidence-Deduplicated (Anonymized) Debate evidence from collegiate and high school competitions, semantically deduplicated, with all debater-identifying columns removed. This is the semantically deduplicated companion to OpenDebateEvidence-Anonymized. Where the parent dataset contains every piece of evidence as used in every round, this version collapses repeated use of the same evidence into single records, making it substantially smaller and better… See the full description on the dataset page: https://huggingface.co/datasets/Hellisotherpeople/OpenDebateEvidence-Deduplicated-Anonymized.tabulartext-generation100K<n<1M0 likes199 downloads2mo agoHugging Face07Hellboi78688 /jee-neet-benchmark JEE/NEET LLM Benchmark Dataset Dataset Description This repository contains a benchmark dataset designed for evaluating the capabilities of Large Language Models (LLMs) on questions from major Indian competitive examinations: JEE (Main & Advanced): Joint Entrance Examination for engineering. NEET: National Eligibility cum Entrance Test for medical fields. The questions are presented in image format (.png) as they appear in the original papers. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/Hellboi78688/jee-neet-benchmark.imagevisual-question-answeringn<1K0 likes108 downloads4mo agoHugging Face08malhajar /hellaswag-trThis Dataset is part of a series of datasets aimed at advancing Turkish LLM Developments by establishing rigid Turkish benchmarks to evaluate the performance of LLM's Produced in the Turkish Language. Dataset Card for Hellaswag-Turkish malhajar/hellaswag-turkish is a translated version of hellaswag aimed specifically to be used in the OpenLLMTurkishLeaderboard This Dataset contains rigid tests extracted from the paper Can a Machine Really Finish Your Sentence? published at ACL2019.… See the full description on the dataset page: https://huggingface.co/datasets/malhajar/hellaswag-tr.textquestion-answering10K<n<100K3 likes104 downloads3y agoHugging Face09richmondsin /m_hellaswag Multilingual HellaSwag Dataset Summary This dataset is a machine translated version of the HellaSwag dataset. The languages was translated using GPT-3.5-turbo by the University of Oregon, and this part of the dataset was originally uploaded to this Github repository. The NUS Deep Learning Lab contributed to this effort by standardizing the dataset, ensuring consistent question formatting and alignment across all languages. This standardization enhances cross-linguistic… See the full description on the dataset page: https://huggingface.co/datasets/richmondsin/m_hellaswag.textquestion-answering10K<n<100K0 likes49 downloads2y agoHugging Face10swap-uniba /hellaswag_ita Italian version of the HellaSwag Dataset The dataset has been automatically translate by using Argos Translate v. 1.9.1 Citation Information @misc{basile2023llamantino, title={LLaMAntino: LLaMA 2 Models for Effective Text Generation in Italian Language}, author={Pierpaolo Basile and Elio Musacchio and Marco Polignano and Lucia Siciliani and Giuseppe Fiameni and Giovanni Semeraro}, year={2023}, eprint={2312.09993}, archivePrefix={arXiv}… See the full description on the dataset page: https://huggingface.co/datasets/swap-uniba/hellaswag_ita.question-answering1 likes45 downloads3y agoHugging Face11Hellrabbit /medical-o1-reasoning-SFT News [2025/04/22] We split the data and kept only the medical SFT dataset (medical_o1_sft.json). The file medical_o1_sft_mix.json contains a mix of medical and general instruction data. [2025/02/22] We released the distilled dataset from Deepseek-R1 based on medical verifiable problems. You can use it to initialize your models with the reasoning chain from Deepseek-R1. [2024/12/25] We open-sourced the medical reasoning dataset for SFT, built on medical verifiable problems and an LLM… See the full description on the dataset page: https://huggingface.co/datasets/Hellrabbit/medical-o1-reasoning-SFT.textquestion-answering10K<n<100K0 likes24 downloads9mo agoHugging Face12helloansuman /OncoIntentQAgatedDataset contains 1015 QA pairs with 4 broad categories of question, such as Knowledge (8), Treatment (11), Hybrid (3), Crisis / Safety (5) with 27 sub-categories. Category Definition knowledge_diagnosis_methods Questions seeking factual understanding of how cancer is diagnosed, including diagnostic procedures, tests, biopsies, and biological or molecular mechanisms relevant to diagnosis. knowledge_disease_definition Questions asking what a cancer or disease is, its fundamental… See the full description on the dataset page: https://huggingface.co/datasets/helloansuman/OncoIntentQA.question-answering1K<n<10K0 likes21 downloads12h agoHugging Face13RikoteMaster /hellaswag-mcqa HellaSwag MCQA Dataset This dataset contains the HellaSwag dataset converted to Multiple Choice Question Answering (MCQA) format. Dataset Description HellaSwag is a dataset for commonsense inference about physical situations. Given a context describing an activity, the task is to select the most plausible continuation from four choices. Dataset Structure Each example contains: question: The activity label and context combined choices: List of 4 possible… See the full description on the dataset page: https://huggingface.co/datasets/RikoteMaster/hellaswag-mcqa.textquestion-answering10K<n<100K0 likes19 downloads1y agoHugging Face14Helllloooo7919 /databricks-dolly-15k Summary databricks-dolly-15k is an open source dataset of instruction-following records generated by thousands of Databricks employees in several of the behavioral categories outlined in the InstructGPT paper, including brainstorming, classification, closed QA, generation, information extraction, open QA, and summarization. This dataset can be used for any purpose, whether academic or commercial, under the terms of the Creative Commons Attribution-ShareAlike 3.0 Unported… See the full description on the dataset page: https://huggingface.co/datasets/Helllloooo7919/databricks-dolly-15k.textquestion-answering10K<n<100K0 likes10 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.