CoolFace
24 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01PersonaBias /Reverse-baseline-bias-unbiastabulartext-classification1M<n<10M0 likes219 downloads2mo agoHugging Face02Aipresso /prompts_under_512_tokens Under 512 Tokens Prompts Dataset Created by Aipresso LIMITED, London, UK ⚠️ IMPORTANT: By using this dataset, you agree to our Terms of Use Dataset Overview Specialized collection of short-form English prompts (under 512 tokens), perfect for training models with context length constraints or faster iteration cycles. 📊 Dataset Statistics Metric Value Total Files 200 Rows Per File 10,000 Total Rows 2,000,000 Token Range 1 to 511 tokens… See the full description on the dataset page: https://huggingface.co/datasets/Aipresso/prompts_under_512_tokens.texttext-generation1M<n<10M0 likes107 downloads11mo agoHugging Face03PersonaBias /Original-baseline-bias-unbiastabulartext-classification100K<n<1M0 likes94 downloads2mo agoHugging Face04cristian-untaru /medquad-retrieval-pretriage MedQuAD Retrieval Pre-Triage Dataset Dataset Description This repository contains a processed, retrieval-oriented derivative of the MedQuAD medical question-answering dataset. It was prepared for contextual medical information retrieval in SortMed, an academic medical pre-triage assistant. The corpus is not used to train the SortMed triage classifiers. It is used by a separate semantic retrieval component that identifies medically related question-answer entries… See the full description on the dataset page: https://huggingface.co/datasets/cristian-untaru/medquad-retrieval-pretriage.tabularquestion-answering10K<n<100K0 likes90 downloads15d agoHugging Face05Crisp-Unimib /BEEP_eval 🚗 BEst DrivEr’s License Performer (BEEP) Dataset BEEP is a challenge benchmark designed to evaluate large language models (LLMs) through a simulation of the Italian driver’s license exam. This dataset focuses on understanding traffic laws and reasoning through driving situations, replicating the complexity of the Italian licensing process. 📁 Dataset Structure Column Data Type Description Categorisation Structure [String] Hierarchical categorisation of major… See the full description on the dataset page: https://huggingface.co/datasets/Crisp-Unimib/BEEP_eval.tabularquestion-answering1K<n<10K2 likes85 downloads2y agoHugging Face06ktiyab /ethical-framework-UNESCO-Ethics-of-AI Ethical AI Training Dataset Introduction UNESCO's Ethics of Artificial Intelligence, adopted by 193 Member States in November 2021, represents the first global framework for ethical AI development and deployment. While regional initiatives like The Montréal Declaration for a Responsible Development of Artificial Intelligence emphasize community-driven governance, UNESCO's approach establishes comprehensive international standards through coordinated multi-stakeholder… See the full description on the dataset page: https://huggingface.co/datasets/ktiyab/ethical-framework-UNESCO-Ethics-of-AI.textquestion-answeringn<1K3 likes77 downloads2y agoHugging Face07swap-uniba /bbh_ita Italian version of the BHH Dataset Dataset based on the Italian translation provided by: Leonardo Ranaldi, Giulia Pucci, Elena Sofia Ruzzetti, Fabio Massimo Zanzotto, and André Freitas - Teasing LLMs adapted to Italian Citations @article{suzgun2022challenging, title={Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them}, author={Suzgun, Mirac and Scales, Nathan and Sch{\"a}rli, Nathanael and Gehrmann, Sebastian and Tay, Yi and Chung, Hyung Won and… See the full description on the dataset page: https://huggingface.co/datasets/swap-uniba/bbh_ita.textquestion-answering1K<n<10K0 likes55 downloads3y agoHugging Face08arafatar /toxic_uncensored_LGBTQ_csvtextquestion-answeringn<1K26 likes54 downloads2y agoHugging Face09Itau-Unibanco /FAQ_BACENThis dataset was used in the article: https://arxiv.org/abs/2311.11331 texttext-classification1K<n<10K16 likes52 downloads3y agoHugging Face10Mike2481 /UniD3_DDMtextquestion-answering10K<n<100K1 likes44 downloads1y agoHugging Face11yusufbaykaloglu /turkish-university-mevzuat Turkey University Regulation Data Collection This dataset provides a comprehensive collection of regulatory documents of Turkish universities obtained from mevzuat.gov.tr. It includes full texts of regulations with detailed publication information and unique identifiers. Overview Data Sources: mevzuat.gov.tr website Technologies Used: Selenium, BeautifulSoup, Python Data Formats: CSV CSV Data Structure Column Description Üniversite Name of the… See the full description on the dataset page: https://huggingface.co/datasets/yusufbaykaloglu/turkish-university-mevzuat.textquestion-answering1K<n<10K3 likes35 downloads2y agoHugging Face12chowfi /instance-level-tofu-unlearning Instance-Level TOFU Benchmark This dataset provides an instance-level adaptation of the TOFU (Maini et al, 2024) dataset for evaluating in-context unlearning in large language models (LLMs). Unlike the original TOFU benchmark, which focuses on entity-level unlearning, this version targets selective memory erasure at the instance level — i.e., forgetting specific facts about an entity. It is compatible for evaluation with the locuslab/tofu_ft_llama2-7b model, which was fine-tuned on… See the full description on the dataset page: https://huggingface.co/datasets/chowfi/instance-level-tofu-unlearning.textquestion-answering1K<n<10K1 likes34 downloads1y agoHugging Face13MLNTeam-Unical /MoralTextManipulation 📊 Exploring LLMs’ Ability to Spontaneously and Conditionally Modify Moral Expressions through Text Manipulation Morality serves as the foundation of societal structure, guiding legal systems, shaping cultural values, and influencing individual self-perception. With the rise and pervasiveness of generative AI tools, and particularly Large Language Models (LLMs), concerns arise regarding how these tools capture and potentially alter moral dimensions through machine-generated text… See the full description on the dataset page: https://huggingface.co/datasets/MLNTeam-Unical/MoralTextManipulation.tabulartext-classification1M<n<10M0 likes33 downloads11mo agoHugging Face14ColdSlither /toxic_uncensored_LGBTQ_csvtextquestion-answeringn<1K0 likes29 downloads5d agoHugging Face15ClarusC64 /hierarchy-delegation-fidelity-under-pressure-v0.1 What this dataset tests You lead inside a hierarchy. A senior pushes you under pressure. You must hold role boundaries. You must delegate work without dropping truth. Why it exists Many models sound helpful. Then pressure hits. They skip delegation. They seize authority. They invent certainty. This dataset forces that failure into view. Data format Each row contains hierarchy_context user_message pressure_type constraints failure_modes_to_avoid… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/hierarchy-delegation-fidelity-under-pressure-v0.1.texttext-generationn<1K0 likes23 downloads8mo agoHugging Face16MLNTeam-Unical /PersonaGen 📊 PersonaGen: A Persona-Driven Open-Ended Machine-Generated Text Dataset PersonaGen is a dataset of persona-driven machine-generated texts produced by open Large Language Models. PersonaGen is specifically designed to investigate how synthetic persona profiles affect, guide, or manifest in machine-generated texts. We built PersonaGen by pairing curated persona-profiles (i.e., description of characteristics, background, and goals) across eight thematic domains (e.g., Physics… See the full description on the dataset page: https://huggingface.co/datasets/MLNTeam-Unical/PersonaGen.texttext-generation1M<n<10M1 likes22 downloads11mo agoHugging Face17Mike2481 /UniD3_DTAtexttext-classification1K<n<10K1 likes21 downloads1y agoHugging Face18veraradas /toxic_uncensored_LGBTQ_csvtextquestion-answeringn<1K1 likes20 downloads9mo agoHugging Face19Mike2481 /UniD3_DEAtexttext-classification10K<n<100K1 likes19 downloads1y agoHugging Face20go-inoue /ArabicMMLU_undiac Fajri Koto, Haonan Li, Sara Shatnawi, Jad Doughman, Abdelrahman Boda Sadallah, Aisha Alraeesi, Khalid Almubarak, Zaid Alyafeai, Neha Sengupta, Shady Shehata, Nizar Habash, Preslav Nakov, and Timothy Baldwin MBZUAI, Prince Sattam bin Abdulaziz University, KFUPM, Core42, NYU Abu Dhabi, The University of Melbourne Introduction We present ArabicMMLU, the first multi-task language understanding benchmark for Arabic language, sourced from school exams across diverse… See the full description on the dataset page: https://huggingface.co/datasets/go-inoue/ArabicMMLU_undiac.tabularquestion-answering10K<n<100K0 likes15 downloads1y agoHugging Face21maitri-vv /UN16_Peace-Justice Dataset Card for Dataset Name Dataset Summary This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/maitri-vv/UN16_Peace-Justice.tabularquestion-answeringn<1K0 likes14 downloads3y agoHugging Face22Soyombo1872 /toxic_uncensored_LGBTQ_csvtextquestion-answeringn<1K0 likes9 downloads5mo agoHugging Face23Federal-University-Lokoja /Lecturer-Reviewgated Dataset Card for Lecturer Reviews Dataset A review of lectures from Federal University Lokoja Dataset Details Dataset Description A Review done by students on the performance of certain lecturers in Federal University Lokoja Curated by: Ogbuagu Francis texttext-classificationn<1K1 likes1 downloads2y agoHugging Face24CALM-Lab-Purdue /UN_NU_interpretation_LLMsgated Quantifier Scope Interpretation Dataset Datasets for an ongoing project about Scope preferences and ambiguity in LLM interpretation. Dataset Structure Splits The dataset consists of synthetically generated stimuli pairing target sentences with interpretation-biased contexts (SSR vs. ISR). Features language (string)Language of the stimulus (English or Chinese). structure (string)Surface syntactic configuration of the sentence:UN (universal >… See the full description on the dataset page: https://huggingface.co/datasets/CALM-Lab-Purdue/UN_NU_interpretation_LLMs.texttext-classificationn<1K1 likes1 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.