CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mteb /amazon_massive_intent MassiveIntentClassification An MTEB dataset Massive Text Embedding Benchmark MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages Task category t2c Domains Spoken Reference https://arxiv.org/abs/2204.08582 How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["MassiveIntentClassification"]) evaluator =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/amazon_massive_intent.texttext-classification100K<n<1M27 likes37k downloads7mo agoHugging Face02cais /MASK The MASK Evaluation 🌐 Website | 📄 Paper | GitHub Center for AI Safety & Scale AI The MASK evaluation provides a rigorous benchmark for evaluating honesty in large language models by measuring whether models remain truthful when incentivized to lie. The public set contains 1,028 high-quality human-labeled examples across six distinct archetypes, each consisting of a proposition, ground truth, pressure prompt designed to elicit lying, and belief elicitation prompt to… See the full description on the dataset page: https://huggingface.co/datasets/cais/MASK.text1K<n<10K18 likes8.9k downloads2y agoHugging Face03masoodali /apple-app-store-labels-policiestabular10M<n<100M1 likes7.7k downloads2y agoHugging Face04mteb /amazon_massive_scenario MassiveScenarioClassification An MTEB dataset Massive Text Embedding Benchmark MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages Task category t2c Domains Spoken Reference https://arxiv.org/abs/2204.08582 How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["MassiveScenarioClassification"]) evaluator =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/amazon_massive_scenario.texttext-classification1M<n<10M6 likes6.8k downloads1y agoHugging Face05tgsc /c4-pt-randMore35M-part04-deduplicated-128000-no-digit-split-mask-train-15003771-lines Dataset Card for "c4-pt-randMore35M-part04-deduplicated-128000-no-digit-split-mask-train-15003771-lines" More Information needed text10M<n<100M1 likes4.7k downloads3y agoHugging Face06ai4privacy /pii-masking-300k 👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. Purpose and Features 🌍 World's largest open dataset for privacy masking 🌎 The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-300k.texttext-classification100K<n<1M116 likes4.3k downloads4mo agoHugging Face07FBK-MT /Speech-MASSIVE Speech-MASSIVE Dataset Description Speech-MASSIVE is a multilingual Spoken Language Understanding (SLU) dataset comprising the speech counterpart for a portion of the MASSIVE textual corpus. Speech-MASSIVE covers 12 languages (Arabic, German, Spanish, French, Hungarian, Korean, Dutch, Polish, European Portuguese, Russian, Turkish, and Vietnamese) from different families and inherits from MASSIVE the annotations for the intent prediction and slot-filling tasks. MASSIVE… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/Speech-MASSIVE.audioaudio-classification10K<n<100K54 likes4.3k downloads1y agoHugging Face08masakhane /mafandMAFAND-MT is the largest MT benchmark for African languages in the news domain, covering 21 languages. The languages covered are: - Amharic - Bambara - Ghomala - Ewe - Fon - Hausa - Igbo - Kinyarwanda - Luganda - Luo - Mossi - Nigerian-Pidgin - Chichewa - Shona - Swahili - Setswana - Twi - Wolof - Xhosa - Yoruba - Zulu The train/validation/test sets are available for 16 languages, and validation/test set for amh, kin, nya, sna, and xho For more details see https://aclanthology.org/2022.naacl-main.223/texttranslation100K<n<1M17 likes3.8k downloads3y agoHugging Face09rulins /MassiveDS-140BWe release the raw passages, embeddings, and index of MassiveDS. Website: https://retrievalscaling.github.io We release two versions of MassiveDS: MassiveDS-1.4T, which contains 1.4T tokens in the datastore. MassiveDS-140B, which is a subsampled version containing 140B tokens in the datastore. File structure: raw_data: plain data in JSONL files. passages: chunked raw passages with passage IDs. Each passage is chunked to have no more than 256 words. embeddings: embeddings of the passages… See the full description on the dataset page: https://huggingface.co/datasets/rulins/MassiveDS-140B.text1M<n<10M7 likes3.7k downloads2y agoHugging Face10ahmed-masry /ChartQAIf you wanna use the dataset, you need to download the zip file manually from the "Files and versions" tab. Please note that this dataset can not be directly loaded with the load_dataset function from the datasets library. If you want a version of the dataset that can be loaded with the load_dataset function, you can use this one: https://huggingface.co/datasets/ahmed-masry/chartqa_without_images But it doesn't contain the chart images. Hence, you will still need to use the images stored in… See the full description on the dataset page: https://huggingface.co/datasets/ahmed-masry/ChartQA.image10K<n<100K32 likes3.6k downloads2y agoHugging Face11roman-bushuiev /MassSpecGym MassSpecGym provides a dataset and benchmark for the discovery and identification of new molecules from tandem mass spectrometry (MS/MS) spectra. The provided challenges abstract the process of scientific discovery of new molecules from biological and environmental samples into well-defined machine learning problems. Papers MassSpecGym in the Wild: Uncovering and Correcting Evaluation Pitfalls in AI-Driven Molecule Discovery (2025): Paper Link MassSpecGym: A benchmark for… See the full description on the dataset page: https://huggingface.co/datasets/roman-bushuiev/MassSpecGym.tabularother100K<n<1M23 likes3.5k downloads2mo agoHugging Face12ai4privacy /pii-masking-200k 👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. Ai4Privacy Community Join our community at https://discord.gg/FmzWshaaQT to help build open datasets for privacy masking. Purpose and Features Previous world's largest open dataset for privacy.… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-200k.texttext-classification100K<n<1M127 likes3.5k downloads4mo agoHugging Face13ai4privacy /pii-masking-openpii-1.5m OpenPII 1.5M: Multilingual PII Masking Dataset (Asia Pacific Extension) 📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific Overview The OpenPII 1.5M dataset extends OpenPII 1M with a new Asia Pacific corpus, bringing global coverage to 30 languages across Europe, Americas, and Asia Pacific. This is the flagship release of the PII-Masking-3M family, the world's largest open multilingual PII masking corpus. Built to advance open… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-openpii-1.5m.texttoken-classification1M<n<10M21 likes2.9k downloads4mo agoHugging Face14MohamedRashad /MASC-Arabic MASC Arabic Dataset Card Dataset Summary MASC is a dataset that contains 1,000 hours of speech sampled at 16 kHz and crawled from over 700 YouTube channels. The dataset is multi-regional, multi-genre, and multi-dialect intended to advance the research and development of Arabic speech technology with a special emphasis on Arabic speech recognition. How to use The datasets library allows you to load and pre-process your dataset in pure Python, at scale. The… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/MASC-Arabic.audioautomatic-speech-recognition100K<n<1M8 likes2.8k downloads6mo agoHugging Face15masakhane /afrimgsm Dataset Card for afrimgsm Dataset Summary AFRIMGSM is an evaluation dataset comprising translations of a subset of the GSM8k dataset into 16 African languages. It includes test sets across all 18 languages, maintaining an English and French subsets from the original GSM8k dataset. Languages There are 18 languages available : Dataset Structure Data Instances The examples look like this for English: from datasets import load_dataset data =… See the full description on the dataset page: https://huggingface.co/datasets/masakhane/afrimgsm.text1K<n<10K10 likes2.3k downloads1y agoHugging Face16rulins /MassiveDS-1.4T-raw-dataWe release the raw passages, embeddings, and index of MassiveDS. Website: https://retrievalscaling.github.io Versions We release two versions of MassiveDS: MassiveDS-1.4T, which contains the embeddings and passages of the 1.4T-token datastore. MassiveDS-1.4T-raw-text, contains the raw text of the 1.4T-token datastore. MassiveDS-140B, which contains the index, embeddings, passages, and raw text of a subsampled version containing 140B tokens in the datastore. Note: Code support to… See the full description on the dataset page: https://huggingface.co/datasets/rulins/MassiveDS-1.4T-raw-data.text100M<n<1B6 likes2.3k downloads2y agoHugging Face17ahmed-masry /ChartQAPro ChartQAPro: A More Diverse and Challenging Benchmark for Chart Question Answering 🤗Dataset | 🖥️Code | 📄Paper The abstract of the paper states that: Charts are ubiquitous, as people often use them to analyze data, answer questions, and discover critical insights. However, performing complex analytical tasks with charts requires significant perceptual and cognitive effort. Chart Question Answering (CQA) systems automate this process by enabling models to interpret and reason with… See the full description on the dataset page: https://huggingface.co/datasets/ahmed-masry/ChartQAPro.textvisual-question-answering1K<n<10K21 likes2.1k downloads1y agoHugging Face18rulins /MasssiveDS-1.4T-raw-datatext100M<n<1B0 likes2.1k downloads2y agoHugging Face19masakhane /masakhanews Dataset Card for [Dataset Name] Dataset Summary MasakhaNEWS is the largest publicly available dataset for news topic classification in 16 languages widely spoken in Africa. The train/validation/test sets are available for all the 16 languages. Supported Tasks and Leaderboards [More Information Needed] news topic classification: categorize news articles into new topics e.g business, sport sor politics. Languages There are 16 languages available :… See the full description on the dataset page: https://huggingface.co/datasets/masakhane/masakhanews.texttext-classification10K<n<100K16 likes2k downloads10mo agoHugging Face20ahmed-masry /chartqa_without_images Dataset Card for "chartqa_without_images" If you wanna load the dataset, you can run the following code: from datasets import load_dataset data = load_dataset('ahmed-masry/chartqa_without_images') The dataset has the following structure: DatasetDict({ train: Dataset({ features: ['imgname', 'query', 'label', 'type'], num_rows: 28299 }) val: Dataset({ features: ['imgname', 'query', 'label', 'type'], num_rows: 1920 }) test:… See the full description on the dataset page: https://huggingface.co/datasets/ahmed-masry/chartqa_without_images.text10K<n<100K1 likes1.8k downloads3y agoHugging Face21gretelai /gretel-pii-masking-en-v1 Gretel Synthetic Domain-Specific Documents Dataset (English) This dataset is a synthetically generated collection of documents enriched with Personally Identifiable Information (PII) and Protected Health Information (PHI) entities spanning multiple domains. Created using Gretel Navigator with mistral-nemo-2407 as the backend model, it is specifically designed for fine-tuning Gliner models. The dataset contains document passages featuring PII/PHI entities from a wide range of… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/gretel-pii-masking-en-v1.texttext-classification10K<n<100K46 likes1.7k downloads9mo agoHugging Face22SetFit /amazon_massive_intent_en-UStext10K<n<100K10 likes1.7k downloads4y agoHugging Face23masakhane /afrimmlu Dataset Card for afrimmlu Dataset Summary AFRIMMLU is an evaluation dataset comprising translations of a subset of the MMLU dataset into 15 African languages. It includes test sets across all 17 languages, maintaining an English and French subsets from the original MMLU dataset. Languages There are 17 languages available : Dataset Structure Data Instances The examples look like this for English: from datasets import load_dataset data =… See the full description on the dataset page: https://huggingface.co/datasets/masakhane/afrimmlu.textquestion-answering10K<n<100K12 likes1.7k downloads1y agoHugging Face24the-masses /ReplicaOcc Replica_OCC Benchmark Replica_OCC is a Replica-based occupancy benchmark constructed in the data organization style of EmbodiedOcc-ScanNet and OccScanNet. It provides RGB-D sequences and scene-level occupancy ground truth for evaluating embodied occupancy prediction systems. Ground-truth occupancy and poses are intended for evaluation-time alignment and metric computation. They are not required for training FreeOcc and are not used for map construction. Citation If… See the full description on the dataset page: https://huggingface.co/datasets/the-masses/ReplicaOcc.image10K<n<100K3 likes1.7k downloads4mo agoHugging Face25masakhane /afrisenti Dataset Summary AfriSenti is the largest sentiment analysis dataset for under-represented African languages, covering 110,000+ annotated tweets in 14 African languages (Amharic, Algerian Arabic, Hausa, Igbo, Kinyarwanda, Moroccan Arabic, Mozambican Portuguese, Nigerian Pidgin, Oromo, Swahili, Tigrinya, Twi, Xitsonga, and Yoruba). The datasets are used in the first Afrocentric SemEval shared task, SemEval 2023 Task 12: Sentiment analysis for African languages (AfriSenti-SemEval).… See the full description on the dataset page: https://huggingface.co/datasets/masakhane/afrisenti.texttext-classification100K<n<1M2 likes1.5k downloads2y agoHugging Face26ai4privacy /pii-masking-openpii-1m OpenPII 1M — Multilingual PII Masking Dataset Overview The OpenPII 1M dataset is a large-scale, multilingual collection of 1,428,143 synthetic text examples with fine-grained PII (Personally Identifiable Information) annotations, spanning 23 European languages and 19 entity types. Built to advance open research in privacy-preserving NLP, this dataset enables the development and benchmarking of Named Entity Recognition (NER) models, token classification pipelines… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-openpii-1m.texttoken-classification1M<n<10M15 likes1.5k downloads6mo agoHugging Face27ai4privacy /pii-masking-400k 👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. Purpose and Features 🌍 World's largest open dataset for privacy masking 🌎 The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-400k.texttext-classification100K<n<1M63 likes1.3k downloads4mo agoHugging Face28BangumiBase /masamunekunnorevenger Bangumi Image Base of Masamune-kun No Revenge R This is the image base of bangumi Masamune-kun no Revenge R, we detected 57 characters, 4840 images in total. The full dataset is here. Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1%… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/masamunekunnorevenger.image1K<n<10K0 likes1.3k downloads2y agoHugging Face29AnanthZeke /tamil_sentences_master_raw Dataset Card for "tamil_sentences_master" More Information needed text10M<n<100M0 likes1.3k downloads3y agoHugging Face30masakhane /afrixnli Dataset Card for afrixnli Dataset Summary AFRIXNLI is an evaluation dataset comprising translations of a subset of the XNLI dataset into 16 African languages. It includes both validation and test sets across all 18 languages, maintaining the English and French subsets from the original XNLI dataset. Languages There are 18 languages available : Dataset Structure Data Instances The examples look like this for English: from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/masakhane/afrixnli.texttext-classification10K<n<100K6 likes1.2k downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.