CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01dbarbedillo /SMS_Spam_Multilingual_Collection_DatasetSMS Spam Multilingual Collection Dataset Collection of Multilingual SMS messages tagged as spam or legitimate About Dataset Context The SMS Spam Collection is a set of SMS-tagged messages that have been collected for SMS Spam research. It originally contained one set of SMS messages in English of 5,574 messages, tagged according to being ham (legitimate) or spam and later Machine Translated into Hindi, German and French. The text has been further translated into Spanish, Chinese, Arabic… See the full description on the dataset page: https://huggingface.co/datasets/dbarbedillo/SMS_Spam_Multilingual_Collection_Dataset.texttext-classification1K<n<10K14 likes723 downloads4y agoHugging Face02flax-community /conceptual-12m-mbart-50-multilingualimage10M<n<100M2 likes641 downloads5y agoHugging Face03KumarSahil299885 /SMS_Spam_Multilingual_Collection_DatasetSMS Spam Multilingual Collection Dataset Collection of Multilingual SMS messages tagged as spam or legitimate About Dataset Context The SMS Spam Collection is a set of SMS-tagged messages that have been collected for SMS Spam research. It originally contained one set of SMS messages in English of 5,574 messages, tagged according to being ham (legitimate) or spam and later Machine Translated into Hindi, German and French. The text has been further translated into Spanish, Chinese, Arabic… See the full description on the dataset page: https://huggingface.co/datasets/KumarSahil299885/SMS_Spam_Multilingual_Collection_Dataset.texttext-classification1K<n<10K3 likes401 downloads3mo agoHugging Face04flax-community /conceptual-12m-multilingual-marianThis dataset is created from subset of Conceptual Captions. The original dataset has 12M captions but this dataset has around 10M image, caption pairs in different languages with 2.5M unique images. This dataset has captions translated from English to Spanish, German, French using language specific English to Marian models. Data distribution is following: train_file_marian_final.tsv: 10010625 captions (2502656 captions of English, German, Spanish, French each) val_file_marian_final.tsv:… See the full description on the dataset page: https://huggingface.co/datasets/flax-community/conceptual-12m-multilingual-marian.text10M<n<100M1 likes239 downloads3y agoHugging Face05Debk /Indian-Multilingual-Bias-Dataset Indian Multilingual Bias Dataset Dataset Description The Indian Multilingual Bias Dataset is a comprehensive collection designed to evaluate and measure social biases in Large Language Models (LLMs) across three major Indian languages: English, Bengali (বাংলা), and Hindi (हिंदी). This dataset is based on the original Indian-BhED dataset and focuses on four critical dimensions of bias prevalent in Indian society. Key Features 🌐 Multilingual:… See the full description on the dataset page: https://huggingface.co/datasets/Debk/Indian-Multilingual-Bias-Dataset.texttext-classification1K<n<10K0 likes234 downloads3mo agoHugging Face06ameyhengle /Multilingual-Needle-in-a-Haystack Multilingual Needle in a Haystack (MLNeedle) The MultiLingual Needle-in-a-Haystack (MLNeedle) test is a dataset designed to assess how well Large Language Models (LLMs) find specific information ("needle") within long, multilingual texts ("haystack"). Built on MLQA, it contains over 5,000 extractive question-answer instances across seven languages (English, Arabic, German, Spanish, Hindi, Vietnamese, Simplified Chinese). We systematically vary the "needle's" language and position to… See the full description on the dataset page: https://huggingface.co/datasets/ameyhengle/Multilingual-Needle-in-a-Haystack.text10K<n<100K3 likes232 downloads1y agoHugging Face07Multilingual-Perspectivist-NLU /MultiPICo Dataset Summary MultiPICo (Multilingual Perspectivist Irony Corpus) is a disaggregated multilingual corpus for irony detection, containing 18,778 pairs of short conversations (post-reply) from Twitter (8,956) and Reddit (9,822), along with the demographic information of each annotator (age, nationality, gender, and so on). Supported Tasks and Leaderboards Irony classification task using soft labels (i.e., distribution of annotations) or hard labels (i.e.… See the full description on the dataset page: https://huggingface.co/datasets/Multilingual-Perspectivist-NLU/MultiPICo.tabular10K<n<100K6 likes231 downloads2y agoHugging Face08flax-community /conceptual-12m-multilingual-marian-128This dataset is created from subset of Conceptual Captions. The original dataset has 12M captions but this dataset has around 10M image, caption pairs in different languages with 2.5M unique images. This dataset has captions translated from English to Spanish, German, French using language specific English to Marian models (with sequence length 128). Data distribution is following: train_file_marian_final.tsv: 10002432 captions (2500608 captions of English, German, Spanish, French each)… See the full description on the dataset page: https://huggingface.co/datasets/flax-community/conceptual-12m-multilingual-marian-128.text10M<n<100M0 likes193 downloads3y agoHugging Face09FrancophonIA /multilingual-hatespeech-dataset [!NOTE] Dataset origin: https://www.kaggle.com/datasets/wajidhassanmoosa/multilingual-hatespeech-dataset Description This dataset contains hate speech text with labels where 0 represents non-hate and 1 shows hate texts also the data from different languages needed to be identified as a corresponding correct language. The following are the languages in the dataset with the numbers corresponding to that language. (1 Arabic)(2 English)(3 Chinese)(4 French) (5 German) (6 Russian)(7… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/multilingual-hatespeech-dataset.tabular100K<n<1M4 likes170 downloads1y agoHugging Face10Multilingual-Perspectivist-NLU /EPIC Dataset Card for EPICorpus Dataset Summary EPIC (English Perspectivist Irony Corpus) is a disaggregated English corpus for irony detection, containing 3,000 pairs of short conversations (posts-replies) from Twitter and Reddit, along with the demographic information of each annotator (age, nationality, gender, and so on). Supported Tasks and Leaderboards Irony classification task using soft labels (i.e., distribution of annotations) or hard labels… See the full description on the dataset page: https://huggingface.co/datasets/Multilingual-Perspectivist-NLU/EPIC.tabulartext-classification10K<n<100K2 likes156 downloads2y agoHugging Face11iNLP-Lab /multilingual-lima Multilingual LIMA A multilingual extension of the LIMA instruction-tuning dataset. The original English prompt–response pairs were translated into 9 additional typologically diverse languages with google/gemini-2.0-flash-001. Each language is stored as a separate Hugging Face config. Field Description prompt User instruction (translated; en is the original). output Assistant response (translated; en is the original). Languages (configs): en (original), zh, it, bn… See the full description on the dataset page: https://huggingface.co/datasets/iNLP-Lab/multilingual-lima.texttext-generation10K<n<100K0 likes146 downloads4mo agoHugging Face12vanila434 /multilingual-elder-safety-msgs multilingual-elder-safety-msgs A hand-authored, multilingual elder fraud-recognition and safety coaching dataset. 467 curated scam/safe scenarios in Chinese and English, with platform-generated coaching responses localized across 5 languages: Chinese, English, Vietnamese, Khmer (Cambodian), and Lao. Expanded to 1,029 rows through Adaption Labs platform reasoning traces and multilingual adaptation. Built for communities where filial piety, authority deference, and fear of… See the full description on the dataset page: https://huggingface.co/datasets/vanila434/multilingual-elder-safety-msgs.texttext-classification1K<n<10K0 likes132 downloads5mo agoHugging Face13iNLP-Lab /multilingual-s1 Multilingual s1 A multilingual extension of the s1K-1.1 reasoning dataset. The original English reasoning questions and DeepSeek-R1 distilled solutions were translated into 9 additional typologically diverse languages with google/gemini-2.0-flash-001. Each language is stored as a separate Hugging Face config. We filter the upstream simplescaling/s1K-1.1 corpus to keep only samples whose DeepSeek-R1 trajectories were marked as correctly distilled, then translate the resulting subset.… See the full description on the dataset page: https://huggingface.co/datasets/iNLP-Lab/multilingual-s1.texttext-generation1K<n<10K0 likes125 downloads4mo agoHugging Face14flax-community /multilingual-vqatext1M<n<10M0 likes121 downloads3y agoHugging Face15gretelai /synthetic_multilingual_llm_prompts Image generated by DALL-E. See prompt for more details 📝🌐 Synthetic Multilingual LLM Prompts Welcome to the "Synthetic Multilingual LLM Prompts" dataset! This comprehensive collection features 1,250 synthetic LLM prompts generated using Gretel Navigator, available in seven different languages. To ensure accuracy and diversity in prompts, and translation quality and consistency across the different languages, we employed Gretel Navigator both as a generation tool and as an… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/synthetic_multilingual_llm_prompts.tabulartext-generation1K<n<10K11 likes120 downloads2y agoHugging Face16flax-community /conceptual-12m-multilingual-marian-esimage1M<n<10M0 likes109 downloads3y agoHugging Face17JRQi /DeepResearch-Bench-Multilingual DeepResearch Bench Multilingual Prompts This dataset provides prompt-level multilingual translations for the 100 research tasks used in muset-ai/DeepResearch-Bench-Dataset. The translations cover eight languages: en zh es it ar bn ja el What is included This repository focuses on the benchmark prompts only. On the Hugging Face Hub, the Dataset Viewer is configured with one default subset named all plus nine explicit subset configurations: source_prompt, en, zh, es… See the full description on the dataset page: https://huggingface.co/datasets/JRQi/DeepResearch-Bench-Multilingual.texttext-generation1K<n<10K1 likes104 downloads6mo agoHugging Face18FirstBML1 /afrofinchain-multilingual-web3 AfroFinChain — Multilingual Web3 & Blockchain Dataset Multilingual Web3 & blockchain dataset in Yoruba, Hausa, Igbo, and Nigerian Pidgin with 1,451 terminology entries and 1,451 conversational Q&A pairs. Designed for LLM fine-tuning, financial literacy, and conversational AI in low-resource African languages. Uses culturally grounded analogies (e.g., ajo, adashi, isusu) to make DeFi concepts actually understandable. Built with Adaptive Data by Adaption as part of the Adaption… See the full description on the dataset page: https://huggingface.co/datasets/FirstBML1/afrofinchain-multilingual-web3.texttext-generation1K<n<10K0 likes101 downloads5mo agoHugging Face19iNLP-Lab /multilingual-safety Multilingual Safety Instructions A multilingual extension of the safety-only instruction–refusal pairs released with the Safety-Tuned LLaMAs project. The original 1,000 harmful-prompt / refusal-response pairs (English) were translated into 11 additional typologically diverse languages with google/gemini-2.0-flash-001. Each language is stored as a separate Hugging Face config. Field Description prompt Harmful user instruction (translated; en is the original). output Safe… See the full description on the dataset page: https://huggingface.co/datasets/iNLP-Lab/multilingual-safety.texttext-generation10K<n<100K0 likes88 downloads4mo agoHugging Face20Steveeeeeeen /multilingual_evalstabularn<1K0 likes87 downloads4mo agoHugging Face21molamin /Kinyarwanda_Engligh_Multilingual_ASRThis dataset was created from Mozilla's Common Voice dataset for the purposes of Multilingual ASR on Kinyarwanda and English. The dataset contains 3000 hours of multilingual training samples, 300 hours of validation samples and 200 of testing samples. text100K<n<1M0 likes84 downloads4y agoHugging Face22Kenpache /multilingual-financial-sentiment Multilingual Financial Sentiment Dataset A curated dataset of 39,829 financial news sentences annotated with sentiment labels (Negative / Neutral / Positive) across 7 languages, collected from 80+ financial news sources worldwide. Dataset Summary Total samples 39,829 Languages 7 (EN, ZH, JA, DE, FR, ES, AR) Labels 3 (negative, neutral, positive) Format CSV Sources 80+ financial news outlets Languages Language Code Samples %… See the full description on the dataset page: https://huggingface.co/datasets/Kenpache/multilingual-financial-sentiment.texttext-classification10K<n<100K0 likes83 downloads6mo agoHugging Face23Tulsiandhare /Multilingual_medical_symptom_triage tags: - medical - healthcare - classification - outbreak-detection - triage - multilingual - adaption - india Multilingual Medical Symptom Triage Dataset Dataset Description A Mutlilingual medical triage dataset containing 9,064 patient symptom descriptions in Hindi, English, and Hinglish (code-mixed Hindi-English), paired with triage recommendations and rich clinical metadata. Designed for training multilingual triage classification models and… See the full description on the dataset page: https://huggingface.co/datasets/Tulsiandhare/Multilingual_medical_symptom_triage.text10K<n<100K0 likes76 downloads3mo agoHugging Face24joshdavham /multilingual-frequency-lists Multilingual Frequency Lists This dataset contains multiple word-frequency lists in various languages such as French, Japanese, Spanish, Italian and Portuguese. Specifically, these are frequency lists of lemmas, meaning, for example, that words like 'run', 'runs' and 'running' are counted together as occurences of the same lemma 'run'. These frequency lists were generated from ~1GB of subtitles scraped from a variety of Netflix shows and films and parsed using relevant spacy models… See the full description on the dataset page: https://huggingface.co/datasets/joshdavham/multilingual-frequency-lists.tabular10K<n<100K1 likes74 downloads5mo agoHugging Face25Svngoku /wikipedia-2023-11-kikongo-lingala-cohere-multilingual-v3texttext-generation10K<n<100K2 likes73 downloads2y agoHugging Face26Rapidata /multilingual-llm-jokes-4o-claude-gemini Rapidata Generated Joke Preference Dataset We collected 1'000'000+ human opinions on the jokes generated by state-of-the-art LLMs to decide which model is the funniest. The labelers are shown a joke in their language and asked to answer 'Yes' or 'No' to the question 'Is this joke funny?'. It took us less than 5 days to get all of the responses. The jokes are evenly distributed across 5 languages: English, Arabic, Japanese, Vietnamese, Portuguese and across 4 model… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/multilingual-llm-jokes-4o-claude-gemini.tabular1K<n<10K14 likes68 downloads1y agoHugging Face27Raftico /instructional-dialogues-multilingual Multilingual Instructional Dialogues (10-Language Dataset) Multilingual Instructional Dialogues is a high-quality dataset of 100 structured, goal-oriented dialogues in 10 major world languages, created for training and fine-tuning AI assistants, chatbots, and instruction-tuned large language models. Each dialogue simulates a clear, polite interaction where a user asks for guidance on how to perform a task, and the assistant responds with easy-to-follow steps. This dataset has been… See the full description on the dataset page: https://huggingface.co/datasets/Raftico/instructional-dialogues-multilingual.text1K<n<10K2 likes62 downloads1y agoHugging Face28YDX07 /YDX07_Multilingual_Corpus_2026audion<1K0 likes60 downloads24d agoHugging Face29wujoe132 /ponys-multilingual-ai-character-consistency-benchmark Ponys Multilingual AI Character Consistency Benchmark This repository contains a preregistered test instrument, not collected product results and not an independent product ranking. 140 fixed test cases across seven locales four dimensions: persona, register, relationship state, and visual identity three planned clean-session runs per case result state: not_collected publisher: Ponys.ai Research (official first-party research) official source: https://ponys.ai/ research feeds:… See the full description on the dataset page: https://huggingface.co/datasets/wujoe132/ponys-multilingual-ai-character-consistency-benchmark.tabulartext-generationn<1K0 likes60 downloads22d agoHugging Face30oberbics /Multilingual_Topic-Specific_Article-Extraction_and_Classification Dataset Card for Multilingual Historical News Article Extraction and Classification Dataset This dataset was created specifically to test Large Language Models' (LLMs) capabilities in processing and extracting topic-specific content from historical newspapers based on OCR'd text. Cite the Dataset Mauermann, Johanna, González-Gallardo, Carlos-Emiliano, and Oberbichler, Sarah. (2025). Multilingual Topic-Specific Article-Extraction and Classification [Data set]. Hugging… See the full description on the dataset page: https://huggingface.co/datasets/oberbics/Multilingual_Topic-Specific_Article-Extraction_and_Classification.texttext-classificationn<1K1 likes52 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.