CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Multilingual-Perspectivist-NLU /MultiPICo Dataset Summary MultiPICo (Multilingual Perspectivist Irony Corpus) is a disaggregated multilingual corpus for irony detection, containing 18,778 pairs of short conversations (post-reply) from Twitter (8,956) and Reddit (9,822), along with the demographic information of each annotator (age, nationality, gender, and so on). Supported Tasks and Leaderboards Irony classification task using soft labels (i.e., distribution of annotations) or hard labels (i.e.… See the full description on the dataset page: https://huggingface.co/datasets/Multilingual-Perspectivist-NLU/MultiPICo.tabular10K<n<100K6 likes228 downloads2y agoHugging Face02FrancophonIA /multilingual-hatespeech-dataset [!NOTE] Dataset origin: https://www.kaggle.com/datasets/wajidhassanmoosa/multilingual-hatespeech-dataset Description This dataset contains hate speech text with labels where 0 represents non-hate and 1 shows hate texts also the data from different languages needed to be identified as a corresponding correct language. The following are the languages in the dataset with the numbers corresponding to that language. (1 Arabic)(2 English)(3 Chinese)(4 French) (5 German) (6 Russian)(7… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/multilingual-hatespeech-dataset.tabular100K<n<1M4 likes171 downloads1y agoHugging Face03Multilingual-Perspectivist-NLU /EPIC Dataset Card for EPICorpus Dataset Summary EPIC (English Perspectivist Irony Corpus) is a disaggregated English corpus for irony detection, containing 3,000 pairs of short conversations (posts-replies) from Twitter and Reddit, along with the demographic information of each annotator (age, nationality, gender, and so on). Supported Tasks and Leaderboards Irony classification task using soft labels (i.e., distribution of annotations) or hard labels… See the full description on the dataset page: https://huggingface.co/datasets/Multilingual-Perspectivist-NLU/EPIC.tabulartext-classification10K<n<100K2 likes160 downloads2y agoHugging Face04gretelai /synthetic_multilingual_llm_prompts Image generated by DALL-E. See prompt for more details 📝🌐 Synthetic Multilingual LLM Prompts Welcome to the "Synthetic Multilingual LLM Prompts" dataset! This comprehensive collection features 1,250 synthetic LLM prompts generated using Gretel Navigator, available in seven different languages. To ensure accuracy and diversity in prompts, and translation quality and consistency across the different languages, we employed Gretel Navigator both as a generation tool and as an… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/synthetic_multilingual_llm_prompts.tabulartext-generation1K<n<10K11 likes122 downloads2y agoHugging Face05Steveeeeeeen /multilingual_evalstabularn<1K0 likes89 downloads4mo agoHugging Face06joshdavham /multilingual-frequency-lists Multilingual Frequency Lists This dataset contains multiple word-frequency lists in various languages such as French, Japanese, Spanish, Italian and Portuguese. Specifically, these are frequency lists of lemmas, meaning, for example, that words like 'run', 'runs' and 'running' are counted together as occurences of the same lemma 'run'. These frequency lists were generated from ~1GB of subtitles scraped from a variety of Netflix shows and films and parsed using relevant spacy models… See the full description on the dataset page: https://huggingface.co/datasets/joshdavham/multilingual-frequency-lists.tabular10K<n<100K1 likes77 downloads5mo agoHugging Face07danielelvs /multilingual-islr-mediapipe Multilingual ISLR MediaPipe Landmarks Dataset Description This dataset combines frame-level MediaPipe Holistic landmarks derived from four isolated sign language recognition (ISLR) resources: INCLUDE-50, KSL, MINDS-Libras, and LIBRAS-UFOP. It provides a common tabular schema for research on landmark selection, temporal modeling, signer-independent evaluation, and multilingual transfer learning. The release contains landmarks rather than source RGB videos. Every… See the full description on the dataset page: https://huggingface.co/datasets/danielelvs/multilingual-islr-mediapipe.tabularvideo-classification1K<n<10K0 likes71 downloads5d agoHugging Face08Rapidata /multilingual-llm-jokes-4o-claude-gemini Rapidata Generated Joke Preference Dataset We collected 1'000'000+ human opinions on the jokes generated by state-of-the-art LLMs to decide which model is the funniest. The labelers are shown a joke in their language and asked to answer 'Yes' or 'No' to the question 'Is this joke funny?'. It took us less than 5 days to get all of the responses. The jokes are evenly distributed across 5 languages: English, Arabic, Japanese, Vietnamese, Portuguese and across 4 model… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/multilingual-llm-jokes-4o-claude-gemini.tabular1K<n<10K14 likes69 downloads1y agoHugging Face09wujoe132 /ponys-multilingual-ai-character-consistency-benchmark Ponys Multilingual AI Character Consistency Benchmark This repository contains a preregistered test instrument, not collected product results and not an independent product ranking. 140 fixed test cases across seven locales four dimensions: persona, register, relationship state, and visual identity three planned clean-session runs per case result state: not_collected publisher: Ponys.ai Research (official first-party research) official source: https://ponys.ai/ research feeds:… See the full description on the dataset page: https://huggingface.co/datasets/wujoe132/ponys-multilingual-ai-character-consistency-benchmark.tabulartext-generationn<1K0 likes61 downloads23d agoHugging Face10hjm1980 /korea-places-multilingual Korean Place Names, Multilingual Built and maintained by Korea Basics, a sourced guide to Korean entry rules and getting around, published in seven languages. 16,126 places in South Korea with their Korean (Hangul) name next to the romanized English name, plus Japanese and Chinese names where the source has them, coordinates, road-name address, and subway lines for stations. Why this exists A visitor who reads "Gyeongbokgung Palace" in a guide cannot type that… See the full description on the dataset page: https://huggingface.co/datasets/hjm1980/korea-places-multilingual.tabular10K<n<100K1 likes48 downloads1mo agoHugging Face11Febriyansyah /phishing-emails-multilingual Phishing Emails Multilingual (ID/EN) — Synthetic Dataset sintetis & edukatif 600 email dwibahasa Indonesia 🇮🇩 & English 🇺🇸 untuk riset deteksi phishing — oleh Febriyansyah. ⚠️ Synthetic & edu-defense-only — dibuat untuk pembelajaran defensive security, bukan untuk kampanye nyata. Jangan gunakan untuk aktivitas ilegal. Ringkasan 600 baris — 300 phishing / 300 benign (seimbang), 321 EN / 279 ID Kolom: id (int), language (id/en), text (string, badan email)… See the full description on the dataset page: https://huggingface.co/datasets/Febriyansyah/phishing-emails-multilingual.tabulartext-classificationn<1K0 likes41 downloads15d agoHugging Face12erickfmm /agentlans__multilingual-sentences__paired_10_stsSentences from agentlans/multilingual-sentences in Spanish, and processed with Sentence Similarity Cosine Scores with model hiiamsid/sentence_similarity_spanish_es Each sentence in original dataset was randomly assigned 10 rows (sentences) within a batch of 1000, calculate the sentence similarity, and then deleted duplicate pairs The code for processing can be found here Useful for data distillation, training or benchmarking. Its recommended resampling the dataset to undersample to get a… See the full description on the dataset page: https://huggingface.co/datasets/erickfmm/agentlans__multilingual-sentences__paired_10_sts.tabularsentence-similarity1M<n<10M0 likes38 downloads11mo agoHugging Face13jamesdborin /Nemotron-SFT-Multilingual-v2-prompt-only Nemotron-SFT-Multilingual-v2-prompt-only Prompt-only extraction from nvidia/Nemotron-SFT-Multilingual-v2. Files: prompts.csv: one prompt extraction record per source row. Records include prompt, separated system_prompt, and structured tools when the source row defines available tools. Nested values are JSON-encoded inside CSV cells. summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts. null_or_empty_rows.md: row indexes where prompt… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-SFT-Multilingual-v2-prompt-only.tabular100K<n<1M0 likes33 downloads3mo agoHugging Face14freococo /quran_multilingual_parallel 📘 Qur’an Multilingual Parallel Dataset (quran_multilingual_parallel) This dataset presents a clean, structurally-aligned multilingual parallel corpus of the Qur’anic text. It is intended for linguistic, computational, and cross-lingual AI applications — not only for religious interpretation. It contains over 6,200 verse-level alignments in 54 human languages, formatted in a machine-friendly .csv structure with language-specific translation fields. 🧠 Dataset Highlights… See the full description on the dataset page: https://huggingface.co/datasets/freococo/quran_multilingual_parallel.tabulartranslation1K<n<10K5 likes32 downloads1y agoHugging Face15helinivan /sarcasm_headlines_multilingual Dataset Card for Multilingual Sarcasm Detection Dataset Summary Dataset consists of news article headlines in Dutch, English and Italian. The news article headlines are both from actual news sources and sarcastic/satirical newspapers. The news article is determined sarcastic/non-sarcastic based on the news article source. The sources of news articles are: The Huffington Post (en, non-sarcastic) The Onion (en, sarcastic) NOS (nl, non-sarcastic) De Speld (nl, sarcastic) Il… See the full description on the dataset page: https://huggingface.co/datasets/helinivan/sarcasm_headlines_multilingual.tabular10K<n<100K1 likes29 downloads4y agoHugging Face16jamesdborin /Nemotron-SFT-Multilingual-v1-prompt-only Nemotron-SFT-Multilingual-v1-prompt-only Prompt-only extraction from nvidia/Nemotron-SFT-Multilingual-v1. Files: prompts.csv: one prompt extraction record per source row. Records include prompt, separated system_prompt, and structured tools when the source row defines available tools. Nested values are JSON-encoded inside CSV cells. summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts. null_or_empty_rows.md: row indexes where prompt… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-SFT-Multilingual-v1-prompt-only.tabular1M<n<10M0 likes28 downloads3mo agoHugging Face17Luel-ai /luel-multilingual-tts-samplesgated Multilingual TTS Samples (Luel) License: All Rights Reserved. Proprietary. Access only for authorized parties; no redistribution or use without permission. See LICENSE. A multilingual text-to-speech / read-speech dataset of short scripted utterances across 7 languages. Each sample is a single-speaker recording of a written prompt, paired with rich speaker and recording metadata. Useful for TTS training and evaluation, ASR adaptation, dialect/accent studies, and read-speech… See the full description on the dataset page: https://huggingface.co/datasets/Luel-ai/luel-multilingual-tts-samples.audiotext-to-speechn<1K0 likes15 downloads5mo agoHugging Face18Sowmya15 /gibberish_multilingualtabular10K<n<100K1 likes12 downloads3y agoHugging Face19aq1048576 /red_team_agent_analysis_multilingual_story_analysis_detailed red_team_agent_analysis_multilingual_story_analysis_detailed This dataset was automatically uploaded from the red-team-agent repository. Dataset Information Original file: multilingual_story_analysis_detailed.csv Source path: /home/ubuntu/red-team-agent/red_team_agent/analysis/multilingual_story_analysis_detailed.csv Validation: Valid CSV with 1000 rows, 12 columns (0.1MB) Usage import pandas as pd from datasets import load_dataset # Load using datasets… See the full description on the dataset page: https://huggingface.co/datasets/aq1048576/red_team_agent_analysis_multilingual_story_analysis_detailed.tabularother1K<n<10K0 likes12 downloads1y agoHugging Face20Tonic /multilingual-EMIR-reporting-csvtabular10K<n<100K0 likes12 downloads2mo agoHugging Face21Dvvreddy /multilingual_abusive-non-abusivetabular100K<n<1M1 likes10 downloads4mo agoHugging Face22FrancophonIA /multilingualcrowspairs [!NOTE] Dataset origin: https://gitlab.inria.fr/corpus4ethics/multilingualcrowspairs/ MultiLingualCrowsPairs Multilingual CrowS-Pairs, a challenge dataset for measuring stereotypical biases present in the masked language models (MLMs) in 7 different languages. This challenge dataset was built on the Crows-Pairs corpus (Nangia et al. 2020) using the methodology described in (Névéol et al. 2023). The 7 new languages are the following: Arabic from Maghreb and the Arab world in… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/multilingualcrowspairs.tabulartext-classification10K<n<100K1 likes8 downloads1y agoHugging Face23wow2000 /multilingual_jailbreak_challengesgatedtabular1K<n<10K2 likes6 downloads2y agoHugging Face24PalakEngineerMaster /Processed_TTS_Multilingual_Data Processed TTS Multilingual Data Validated and quality-checked multilingual speech datasets for TTS training, covering 12+ Indian languages. Datasets Included Subset Samples Hours Description indic_voices_r 239,684 548.8h Indic Voices_R — IVR recordings rasa 201,509 361.2h RASA — read speech (wiki, conv, book, news) indictts_iitm 155,236 253.6h Indic TTS (IIT Madras) — studio TTS recordings at 48kHz Total 596,429 1,163.6h Languages… See the full description on the dataset page: https://huggingface.co/datasets/PalakEngineerMaster/Processed_TTS_Multilingual_Data.tabulartext-to-speech100K<n<1M0 likes6 downloads7mo agoHugging Face25osher-dk /do_not_answer_multilingual_12gatedtabular10K<n<100K0 likes3 downloads1y agoHugging Face26model2me /scienceqa-multilingual-hindi#ScienceQA Hindi Translation Dataset ##Dataset Description This dataset is a Hindi-translated version of the original ScienceQA dataset. It includes multiple-choice science questions, with fields for: Images (optional visual context), Hints (optional support text), English questions and their Hindi translations, Multiple answer choices, Correct answers. This translation is intended to support multilingual education research, question-answering in Hindi, and fairness studies in multilingual… See the full description on the dataset page: https://huggingface.co/datasets/model2me/scienceqa-multilingual-hindi.tabular10K<n<100K0 likes3 downloads1y agoHugging Face27infinite-dataset-hub /MultilingualTranscriptionDataset MultilingualTranscriptionDataset tags: Transcription, LanguageProcessing, Multilingual Note: This is an AI-generated dataset so its content may be inaccurate or false Dataset Description: The 'MultilingualTranscriptionDataset' is a curated collection of text transcriptions from various audio recordings. Each transcription is provided in multiple languages, emphasizing the diversity and complexity of language processing. This dataset aims to assist in developing machine learning… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/MultilingualTranscriptionDataset.tabularn<1K0 likes2 downloads2y agoHugging Face28MithuSi /multi-lingual-llm Dataset Card for Dataset Name This data set contains set questions in tamil and possible answers, with the correct answer in the column. This helps to test LLM to see for accuracy. Dataset Details Dataset Description Curated by: Madhumitha Sivalingapandian Language(s) (NLP): Tamil and English License: [More Information Needed] Uses Used for measuring performance of LLM. Direct Use Spoken language accuracy measurement dataset… See the full description on the dataset page: https://huggingface.co/datasets/MithuSi/multi-lingual-llm.tabularn<1K0 likes1 downloads2y agoHugging Face29isc-tleavitt /opus100-multilingualtabular1K<n<10K0 likes1 downloads1y agoHugging Face30Shoriful025 /Multilingual_E-commercetabularn<1K0 likes1 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.