CoolFace
28 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Johnson8187 /Chinese_Multi-Emotion_Dialogue_Dataset Chinese_Multi-Emotion_Dialogue_Dataset 📄 Description This dataset contains 4159 Chinese dialogues annotated with 8 distinct emotion categories. The data is suitable for emotion recognition, sentiment analysis, and other NLP tasks involving Chinese text. Data Sources: Daily Conversations: Captured from natural, informal human conversations. Movie Dialogues: Extracted from diverse Chinese-language movies. AI-Generated Dialogues: Synthesized using… See the full description on the dataset page: https://huggingface.co/datasets/Johnson8187/Chinese_Multi-Emotion_Dialogue_Dataset.texttext-classification1K<n<10K19 likes367 downloads12d agoHugging Face02kz-transformers /multidomain-kazakh-dataset Dataset Description Point of Contact: Sanzhar Murzakhmetov, Besultan Sagyndyk Dataset Summary MDBKD | Multi-Domain Bilingual Kazakh Dataset is a Kazakh-language dataset containing just over 24 883 808 unique texts from multiple domains. Supported Tasks 'MLM/CLM': can be used to train a model for casual and masked languange modeling Languages The kk code for Kazakh as generally spoken in the Kazakhstan Data Instances For each instance… See the full description on the dataset page: https://huggingface.co/datasets/kz-transformers/multidomain-kazakh-dataset.texttext-generation10M<n<100M30 likes307 downloads1y agoHugging Face03alibaba-multimodal-industrial-ai /IndustryBench IndustryBench: Probing the Industrial Knowledge Boundaries of LLMs 💻Github | 📝Paper IndustryBench is a multi-lingual benchmark for evaluating the industrial domain knowledge of large language models. It comprises 2,049 expert-curated QA pairs spanning 12 industrial sectors, with human-reviewed translations in Chinese, English, Russian, and Vietnamese. Overview Dimension Details Total questions 2,049 Languages Chinese (zh), English (en), Russian (ru)… See the full description on the dataset page: https://huggingface.co/datasets/alibaba-multimodal-industrial-ai/IndustryBench.textquestion-answering1K<n<10K30 likes220 downloads5mo agoHugging Face04iNLP-Lab /multilingual-lima Multilingual LIMA A multilingual extension of the LIMA instruction-tuning dataset. The original English prompt–response pairs were translated into 9 additional typologically diverse languages with google/gemini-2.0-flash-001. Each language is stored as a separate Hugging Face config. Field Description prompt User instruction (translated; en is the original). output Assistant response (translated; en is the original). Languages (configs): en (original), zh, it, bn… See the full description on the dataset page: https://huggingface.co/datasets/iNLP-Lab/multilingual-lima.texttext-generation10K<n<100K0 likes142 downloads5mo agoHugging Face05JRQi /DeepResearch-Bench-Multilingual DeepResearch Bench Multilingual Prompts This dataset provides prompt-level multilingual translations for the 100 research tasks used in muset-ai/DeepResearch-Bench-Dataset. The translations cover eight languages: en zh es it ar bn ja el What is included This repository focuses on the benchmark prompts only. On the Hugging Face Hub, the Dataset Viewer is configured with one default subset named all plus nine explicit subset configurations: source_prompt, en, zh, es… See the full description on the dataset page: https://huggingface.co/datasets/JRQi/DeepResearch-Bench-Multilingual.texttext-generation1K<n<10K1 likes138 downloads6mo agoHugging Face06iNLP-Lab /multilingual-s1 Multilingual s1 A multilingual extension of the s1K-1.1 reasoning dataset. The original English reasoning questions and DeepSeek-R1 distilled solutions were translated into 9 additional typologically diverse languages with google/gemini-2.0-flash-001. Each language is stored as a separate Hugging Face config. We filter the upstream simplescaling/s1K-1.1 corpus to keep only samples whose DeepSeek-R1 trajectories were marked as correctly distilled, then translate the resulting subset.… See the full description on the dataset page: https://huggingface.co/datasets/iNLP-Lab/multilingual-s1.texttext-generation1K<n<10K0 likes124 downloads5mo agoHugging Face07gretelai /synthetic_multilingual_llm_prompts Image generated by DALL-E. See prompt for more details 📝🌐 Synthetic Multilingual LLM Prompts Welcome to the "Synthetic Multilingual LLM Prompts" dataset! This comprehensive collection features 1,250 synthetic LLM prompts generated using Gretel Navigator, available in seven different languages. To ensure accuracy and diversity in prompts, and translation quality and consistency across the different languages, we employed Gretel Navigator both as a generation tool and as an… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/synthetic_multilingual_llm_prompts.tabulartext-generation1K<n<10K11 likes115 downloads2y agoHugging Face08AmazonScience /Multi-IaC-Eval Multi-IaC-Eval We present Multi-IaC-Eval is a novel benchmark dataset for evaluating LLM-based IaC generation and mutation across AWS CloudFormation, Terraform, and Cloud Development Kit (CDK) formats. The dataset consists of triplets containing initial IaC templates, natural language modification requests, and corresponding updated templates, created through a synthetic data generation pipeline with rigorous validation. Cloudformation: 263 Terraform: 446 CDK (Python): 64 CDK… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/Multi-IaC-Eval.texttext-generationn<1K0 likes107 downloads1y agoHugging Face09FirstBML1 /afrofinchain-multilingual-web3 AfroFinChain — Multilingual Web3 & Blockchain Dataset Multilingual Web3 & blockchain dataset in Yoruba, Hausa, Igbo, and Nigerian Pidgin with 1,451 terminology entries and 1,451 conversational Q&A pairs. Designed for LLM fine-tuning, financial literacy, and conversational AI in low-resource African languages. Uses culturally grounded analogies (e.g., ajo, adashi, isusu) to make DeFi concepts actually understandable. Built with Adaptive Data by Adaption as part of the Adaption… See the full description on the dataset page: https://huggingface.co/datasets/FirstBML1/afrofinchain-multilingual-web3.texttext-generation1K<n<10K0 likes90 downloads5mo agoHugging Face10iNLP-Lab /multilingual-safety Multilingual Safety Instructions A multilingual extension of the safety-only instruction–refusal pairs released with the Safety-Tuned LLaMAs project. The original 1,000 harmful-prompt / refusal-response pairs (English) were translated into 11 additional typologically diverse languages with google/gemini-2.0-flash-001. Each language is stored as a separate Hugging Face config. Field Description prompt Harmful user instruction (translated; en is the original). output Safe… See the full description on the dataset page: https://huggingface.co/datasets/iNLP-Lab/multilingual-safety.texttext-generation10K<n<100K0 likes85 downloads5mo agoHugging Face11wujoe132 /ponys-multilingual-ai-character-consistency-benchmark Ponys Multilingual AI Character Consistency Benchmark This repository contains a preregistered test instrument, not collected product results and not an independent product ranking. 140 fixed test cases across seven locales four dimensions: persona, register, relationship state, and visual identity three planned clean-session runs per case result state: not_collected publisher: Ponys.ai Research (official first-party research) official source: https://ponys.ai/ research feeds:… See the full description on the dataset page: https://huggingface.co/datasets/wujoe132/ponys-multilingual-ai-character-consistency-benchmark.tabulartext-generationn<1K0 likes67 downloads26d agoHugging Face12ciol-research /multilevel-legal-reasoning Legal Reasoning Dataset with Multilevel Human and Model-Annotated Explanations Prepared by Mst Rafia Islam, Umong Sain, Azmine Toushik Wasi Prepared as a part of Reasoning Datasets Competition by Bespoke Labs, Hugging Face, and Together.ai. 🧭 Purpose and Scope The Legal Reasoning Dataset aims to support the evaluation and training of legal reasoning systems, particularly in multilingual or jurisdiction-agnostic contexts. It focuses on international acts and treaties… See the full description on the dataset page: https://huggingface.co/datasets/ciol-research/multilevel-legal-reasoning.tabulartext-generationn<1K7 likes66 downloads1y agoHugging Face13gplsi /MULTICOMMULTICOM V1.1 This repository hosts the MULTICOM dataset, a novel benchmark for evaluating the multilingual commonsense generation abilities of Large Language Models (LLMs), as presented in the paper Do LLMs exhibit the same commonsense capabilities across languages?. The dataset extends the COCOTEROS dataset to four languages: English, Spanish, Dutch, and Valencian. The task involves generating a commonsensical sentence that includes a given triplet of words. texttext-generation10K<n<100K0 likes57 downloads11mo agoHugging Face14Svngoku /wikipedia-2023-11-kikongo-lingala-cohere-multilingual-v3texttext-generation10K<n<100K2 likes51 downloads2y agoHugging Face15omaressam1111 /multi-tafseer-quran-rag Quran Tafseer RAG Dataset A structured Arabic dataset of Quranic tafseer collected from eight classical and modern tafseer books.The dataset contains verse-aligned tafseer passages designed for Retrieval-Augmented Generation (RAG) systems and Arabic NLP research. Each record links a Quran verse with its corresponding tafseer explanation from one of the tafseer books and includes rich metadata such as surah information, tafseer source, and embedding-ready text. The dataset was… See the full description on the dataset page: https://huggingface.co/datasets/omaressam1111/multi-tafseer-quran-rag.tabularquestion-answering10K<n<100K7 likes28 downloads5mo agoHugging Face16yuyijiong /Multi-Doc-Multi-QA-ChineseDeprecated, please use Multi-Doc-QA-Chinese instead. 文档和问答对都来自 Multi-Doc-QA-Chinese,通过随机抽取和组合形成多轮问答形式。 推荐直接使用原始数据集Multi-Doc-QA-Chinese自己生成指令微调数据,可以控制参考文档和问答的数量 经过随机组合,每条数据形成了 20-60个参考文档 + 10个问答对的形式 chat格式为chatml tabulartext-generation1K<n<10K7 likes26 downloads9mo agoHugging Face17GoJulyAI /illicit-general-multi-turn Illicit General Multi-Turn Conversations Multi-turn adversarial conversations that successfully elicited harmful illicit content from AI models. This sample dataset contains 5 conversations (52 turns) covering chemical weapons, cyber threats, and other safety-critical domains. Dataset Statistics Metric Value Conversations 5 Total Turns 52 Avg Turns/Conv 10.4 Harm Categories 3 Harm Categories Category Turns Description Chemical… See the full description on the dataset page: https://huggingface.co/datasets/GoJulyAI/illicit-general-multi-turn.tabulartext-generationn<1K0 likes26 downloads10mo agoHugging Face18kmhj1306 /IFAAB-MULTI-LLM-2026 Dataset Card for IFAAB-MULTI-LLM-2026 This dataset card serves as a comprehensive datasheet for the kmhj1306/IFAAB-MULTI-LLM-2026 dataset repository. It maps demographic persona features to localized automated financial planning prompts and responses, specifically curated to evaluate LLM behavior within the Indian socio-economic context. Dataset Details Dataset Description This dataset consists of 222,138 rows of tabular text data designed to… See the full description on the dataset page: https://huggingface.co/datasets/kmhj1306/IFAAB-MULTI-LLM-2026.texttext-generation100K<n<1M0 likes26 downloads2mo agoHugging Face19AmanPriyanshu /FRACTURED-SORRY-Bench-Automated-Multishot-Jailbreakgated FRACTURED-SORRY-Bench: Framework for Revealing Attacks in Conversational Turns Undermining Refusal Efficacy and Defenses over SORRY-Bench (Automated Multi-shot Jailbreaks) Dataset Card for FRACTURED-SORRY-Bench Dataset 🌐Website 📑Paper 📚Dataset 💻Github FRACTURED-SORRY-Bench is a framework for evaluating the safety of Large Language Models (LLMs) against multi-turn conversational attacks. Building upon the SORRY-Bench dataset, we propose a simple… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/FRACTURED-SORRY-Bench-Automated-Multishot-Jailbreak.textquestion-answering1K<n<10K1 likes21 downloads2y agoHugging Face20tejasashinde /a2z-multidomain-glossary A–Z Multi-Domain Glossary Dataset This dataset is a creative collection of A-to-Z terminology across a wide range of high-level domains including Agriculture, Technology, Environment, Artificial Intelligence, Zoology, and more.Each entry includes: domain letter (A–Z) word description (short) 📊 Structure Column Description domain The high-level category (e.g. Technology, Agriculture) letter The alphabetical letter from A to Z word The concept/keyword… See the full description on the dataset page: https://huggingface.co/datasets/tejasashinde/a2z-multidomain-glossary.texttext-generationn<1K1 likes21 downloads1y agoHugging Face21multimodal-reframing /mirror MIRROR Dataset MIRROR is a synthetic vision–language dataset for multimodal cognitive reframing under client resistance. Paper: 🪞 MIRROR: Multimodal Cognitive Reframing Therapy for Rolling with Resistance The dataset includes: Client profile metadata (CACTUS idx, CelebA idx) Dialogue written in a screenplay format, including stage directions that describe facial expressions ⚠️ Images themselves are not included to comply with the CelebA license. However, we provide the full image… See the full description on the dataset page: https://huggingface.co/datasets/multimodal-reframing/mirror.tabulartext-generationn<1K2 likes21 downloads10mo agoHugging Face22AhmedZaky1 /authorship-style-transfer-multilangual Parallel neutral / author-style fine-tuning dataset Tabular parallel text built from matched neutral (“standard”) and author-style sources. Each row is one chunk of several consecutive non-empty lines, paired so that the same semantic content appears in both columns. Dataset statistics Samples (CSV rows) 4,868 Hub size bucket 1K<n<10K (matches sample count) Primary file fine_tune_dataset.csv (UTF-8) The metadata field size_categories refers to number… See the full description on the dataset page: https://huggingface.co/datasets/AhmedZaky1/authorship-style-transfer-multilangual.texttext-generation1K<n<10K0 likes20 downloads6mo agoHugging Face23osuih /Chinese_Multi-Emotion_Dialogue_Dataset Chinese_Multi-Emotion_Dialogue_Dataset 📄 Description This dataset contains 4159 Chinese dialogues annotated with 8 distinct emotion categories. The data is suitable for emotion recognition, sentiment analysis, and other NLP tasks involving Chinese text. Data Sources: Daily Conversations: Captured from natural, informal human conversations. Movie Dialogues: Extracted from diverse Chinese-language movies. AI-Generated Dialogues: Synthesized using advanced… See the full description on the dataset page: https://huggingface.co/datasets/osuih/Chinese_Multi-Emotion_Dialogue_Dataset.texttext-classification1K<n<10K0 likes17 downloads7mo agoHugging Face24vidula123 /Multilingual_modeltextquestion-answeringn<1K0 likes15 downloads2y agoHugging Face25GoJulyAI /illicit-bio-multi-turn Illicit Bio Multi-Turn Conversations Multi-turn adversarial conversations that successfully elicited harmful bio-safety content from AI models. This sample dataset contains 5 conversations (57 turns) covering bioweapons and related threats. Dataset Statistics Metric Value Conversations 5 Total Turns 57 Avg Turns/Conv 11.4 Harm Categories 3 Harm Categories Category Turns Description Bioweapons 34 Information about biological… See the full description on the dataset page: https://huggingface.co/datasets/GoJulyAI/illicit-bio-multi-turn.tabulartext-generationn<1K0 likes15 downloads10mo agoHugging Face26multidefmod /doregated You must agree to the license and terms of use before using the dataset in this repo. DORE: Definition MOdelling in PoRtuguEse This repository introduces DORE, a comprehensive corpus of over 100,000 definitions from Portuguese dictionaries. Alongside DORE, we also introduce the models used to perform Portuguese DM. The release of DORE aims to fill in the gap of resources for Automatic Definition Generation, or Definition Modelling (DM), in Portuguese. DORE is the first dataset… See the full description on the dataset page: https://huggingface.co/datasets/multidefmod/dore.texttext-generation10K<n<100K6 likes12 downloads3y agoHugging Face27GoJulyAI /psychology-multi-turn Psychology Multi-Turn Conversations Multi-turn adversarial conversations that successfully elicited harmful psychological content from AI models. This sample dataset contains 5 conversations (54 turns) covering anthropomorphism, psychosis, self-harm, etc. Dataset Statistics Metric Value Conversations 5 Total Turns 54 Avg Turns/Conv 10.8 Harm Categories 3 Harm Categories Category Turns Description Anthropomorphism 28… See the full description on the dataset page: https://huggingface.co/datasets/GoJulyAI/psychology-multi-turn.tabulartext-generationn<1K0 likes8 downloads10mo agoHugging Face28Orib24 /Roomly-Student-Bios-Multimodal Roomly: Multimodal Roommate Matching Dataset 🎯 Problem Statement Finding a roommate is often reduced to dry filters like "budget" and "location". Roomly aims to revolutionize this by focusing on personality, lifestyle, and visual preferences. This dataset provides synthetic student profiles and their ideal room environments. 📊 Exploratory Data Analysis (EDA) 1. User Persona Distribution Our dataset contains a balanced mix of different student… See the full description on the dataset page: https://huggingface.co/datasets/Orib24/Roomly-Student-Bios-Multimodal.texttext-generationn<1K0 likes8 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.