CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Anoy123423123 /MSA_PretrainData MSA Pretrain Data Retrieval-style pretraining corpora. Each subset is split into two parts: file columns meaning <subset>/queries/*.parquet question, answer, reference_ids: list<int64>, labels: list<int64> query, plus row indices into the subset's reference table <subset>/references/*.parquet value: string the reference/memory passage text reference_ids are the candidate pool for a query; labels are the positive(s). Both are global row indices into the subset's… See the full description on the dataset page: https://huggingface.co/datasets/Anoy123423123/MSA_PretrainData.texttext-retrieval10M<n<100M0 likes7.2k downloads2mo agoHugging Face02Dr-AliGomaa /ar-quran-hadith14books-MSA ar-quran-hadith14books-MSA Arabic speech for both primary sources of Islam — Quran and Hadith — plus cleaned general Modern Standard Arabic, under one construction pipeline and one text convention. ASR errors on sacred text are not ordinary errors: a plausible-sounding substitution can alter the meaning of scripture, and because chatbots, search and summarizers increasingly answer from transcriptions rather than from audio, such an error propagates silently. Quranic recitation… See the full description on the dataset page: https://huggingface.co/datasets/Dr-AliGomaa/ar-quran-hadith14books-MSA.audioautomatic-speech-recognition10K<n<100K6 likes1.1k downloads1mo agoHugging Face03APProjects /us-layoffs-by-metro-area-msa-warn-act US layoffs by metro area: 54,170 WARN notices mapped to 765 metro and micro areas Rebuilt 2026-09-22. 765 of the 935 US core-based statistical areas carry at least one layoff notice on record — 361 metropolitan and 404 micropolitan. Nobody hires, sells or reports by county. A recruiter covers Austin; an account team books the Phoenix metro; a reporter writes Bay Area layoffs. State agencies publish neither — they publish the site of a layoff as free text in 48 different… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-layoffs-by-metro-area-msa-warn-act.tabulartabular-regression10K<n<100K0 likes607 downloads46m agoHugging Face04Omartificial-Intelligence-Space /FineWeb2-MSA FineWeb2 MSA Arabic This is the MSA Arabic Portion of The FineWeb2 Dataset. This dataset contains a rich collection of text in MSA Arabic (ISO 639-3: arz), a widely spoken dialect within the Afro-Asiatic language family. With over 439 million words and 1.4 million documents, it serves as a valuable resource for NLP development and linguistic research focused on Egyptian Arabic. Purpose of This Repository This repository provides easy access to the Arabic portion - MSA… See the full description on the dataset page: https://huggingface.co/datasets/Omartificial-Intelligence-Space/FineWeb2-MSA.text100M<n<1B2 likes438 downloads2y agoHugging Face05ragrawal36 /msa-hotpotqa-qa-with-idstext1K<n<10K0 likes423 downloads5mo agoHugging Face06msaligane /tinystories_phonologytext10M<n<100M0 likes370 downloads3y agoHugging Face07ragrawal36 /msa-musique-qa-with-idstextn<1K0 likes301 downloads5mo agoHugging Face08ragrawal36 /msa-hotpotqa-docs-with-idstext1K<n<10K0 likes276 downloads5mo agoHugging Face09msaleme /mcp-sandbox-authority-boundary-profile MCP Sandbox Authority Boundary Profile Profile v0.1.0 · Release v0.2.0 - Experimental Characterization Profile Profile release date: 2026-07-23 Latest distribution release date: 2026-09-05 Execution containment is not proof of bounded authority. Start here For a one-minute, case-by-case reading of the profile, open the companion Authority Boundary Field Guide Space. It presents the released synthetic observations with their control question, observed result… See the full description on the dataset page: https://huggingface.co/datasets/msaleme/mcp-sandbox-authority-boundary-profile.textn<1K1 likes275 downloads17d agoHugging Face10oddadmix /msa-omnivoice-tts-v1 MSA-OmniVoice-v1 Dataset Description MSA-OmniVoice-v1 is a 50-hour synthetic Modern Standard Arabic (MSA) speech dataset generated using OmniVoice. The dataset contains high-quality synthetic speech from a single speaker paired with fully diacritized (تشكيل) transcripts. It is intended for training and fine-tuning Arabic speech models, including Text-to-Speech (TTS), Automatic Speech Recognition (ASR), speech representation learning, and alignment tasks.… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/msa-omnivoice-tts-v1.audio10K<n<100K0 likes273 downloads3mo agoHugging Face11underfrog /msa-2wikimultihopqa-qa-with-idstext1K<n<10K0 likes251 downloads3mo agoHugging Face12m-sakka /agripotentialMore information and competition link: https://github.com/MohammadElSakka/agripotential https://www.codabench.org/competitions/12055/ https://zenodo.org/records/15551829 imageimage-segmentation1K<n<10K2 likes244 downloads2mo agoHugging Face13songlab /gpn-msa-sapiens-dataset Training windows for GPN-MSA-Sapiens For more information check out our paper and repository. Path in Snakemake: results/dataset/multiz100way/89/128/64/True/defined.phastCons.percentile-75_0.05_0.001 tabular1M<n<10M0 likes242 downloads2y agoHugging Face14shenlehan /msavbench-videotext1K<n<10K0 likes199 downloads5d agoHugging Face15HeshamHaroon /arabic-msa-25k-saudi-male-tashkeel Arabic MSA 25K — Saudi Male (Tashkeel) 25,000 fully-diacritized Arabic MSA text + audio pairs, rendered with a single Saudi male neural voice at 48 kHz / 16-bit PCM, across 10 thematic categories. Dataset Summary arabic-msa-25k-saudi-male-tashkeel is a 25,000-clip Modern Standard Arabic (MSA) speech corpus with matching diacritized text (full tashkeel / ḥarakāt). Every clip is synthesized by the single voice ar-SA-HamedNeural (Azure Neural TTS, Saudi Arabic male) at 48… See the full description on the dataset page: https://huggingface.co/datasets/HeshamHaroon/arabic-msa-25k-saudi-male-tashkeel.tabulartext-to-speech10K<n<100K10 likes185 downloads5mo agoHugging Face16underfrog /msa-musique-qa-with-idstextn<1K0 likes178 downloads3mo agoHugging Face17underfrog /msa-musique-docs-with-idstext10K<n<100K0 likes173 downloads3mo agoHugging Face18ragrawal36 /msa-2wikimultihopqa-qa-with-idstext1K<n<10K0 likes168 downloads5mo agoHugging Face19clarayyu22 /gpn-msa-microglia-fulltabular1M<n<10M0 likes159 downloads2y agoHugging Face20otozz /MSA_train_setPre-processed MSA data based on https://huggingface.co/datasets/mozilla-foundation/common_voice_16_1. audio10K<n<100K1 likes151 downloads2y agoHugging Face21dotan1111 /MSA-nuc-9-seq Multiple Sequence Alignment as a Sequence-to-Sequence Learning Problem Abstract: The sequence alignment problem is one of the most fundamental problems in bioinformatics and a plethora of methods were devised to tackle it. Here we introduce BetaAlign, a methodology for aligning sequences using an NLP approach. BetaAlign accounts for the possible variability of the evolutionary process among different datasets by using an ensemble of transformers, each trained on millions… See the full description on the dataset page: https://huggingface.co/datasets/dotan1111/MSA-nuc-9-seq.text1M<n<10M0 likes148 downloads3y agoHugging Face22dotan1111 /MSA-amino-9-seq Multiple Sequence Alignment as a Sequence-to-Sequence Learning Problem Abstract: The sequence alignment problem is one of the most fundamental problems in bioinformatics and a plethora of methods were devised to tackle it. Here we introduce BetaAlign, a methodology for aligning sequences using an NLP approach. BetaAlign accounts for the possible variability of the evolutionary process among different datasets by using an ensemble of transformers, each trained on millions… See the full description on the dataset page: https://huggingface.co/datasets/dotan1111/MSA-amino-9-seq.text1M<n<10M1 likes124 downloads3y agoHugging Face23FosterBirnbaum /PottsMPNN_Training_MSAs Paired & Monomer MSA Database This repository contains multiple sequence alignments (MSAs) and the mapping files needed to associate them with PDB complexes. Contents File Size Description PDB_paired_msas.tar.zst.part_aa … part_af ~10.7 GB each Sharded, zstd-compressed tar archive of the paired PDB MSAs (.a3m files). cath42_msas.tar.gz 13.3 GB Gzip-compressed tar archive of the CATH 4.2 MSAs. cid_mapping.pkl 7.84 MB Dictionary mapping PDB chain IDs… See the full description on the dataset page: https://huggingface.co/datasets/FosterBirnbaum/PottsMPNN_Training_MSAs.text100K<n<1M0 likes123 downloads3mo agoHugging Face24dotan1111 /MSA-nuc-8-seq Multiple Sequence Alignment as a Sequence-to-Sequence Learning Problem Abstract: The sequence alignment problem is one of the most fundamental problems in bioinformatics and a plethora of methods were devised to tackle it. Here we introduce BetaAlign, a methodology for aligning sequences using an NLP approach. BetaAlign accounts for the possible variability of the evolutionary process among different datasets by using an ensemble of transformers, each trained on millions… See the full description on the dataset page: https://huggingface.co/datasets/dotan1111/MSA-nuc-8-seq.text1M<n<10M0 likes122 downloads3y agoHugging Face25dotan1111 /MSA-nuc-7-seq Multiple Sequence Alignment as a Sequence-to-Sequence Learning Problem Abstract: The sequence alignment problem is one of the most fundamental problems in bioinformatics and a plethora of methods were devised to tackle it. Here we introduce BetaAlign, a methodology for aligning sequences using an NLP approach. BetaAlign accounts for the possible variability of the evolutionary process among different datasets by using an ensemble of transformers, each trained on millions… See the full description on the dataset page: https://huggingface.co/datasets/dotan1111/MSA-nuc-7-seq.text1M<n<10M0 likes111 downloads3y agoHugging Face26underfrog /msa-hotpotqa-qa-with-idstext1K<n<10K0 likes111 downloads3mo agoHugging Face27msarmi9 /korean-english-multitarget-ted-talks-task Dataset Card for english-korean-multitarget-ted-talks-task Dataset Summary Parallel English-Korean Text Corpus Text was originally transcribed to English from various Ted Talks, then translated to Korean by TED translators Approximately 166k train, 2k validation, and 2k test sentence pairs. Supported Tasks and Leaderboards Machine Translation Languages English Korean Additional Information Dataset Curators Kevin Duh, "The… See the full description on the dataset page: https://huggingface.co/datasets/msarmi9/korean-english-multitarget-ted-talks-task.text100K<n<1M11 likes109 downloads4y agoHugging Face28tunis-ai /tunisian-msa-parallel-corpus Dataset Description This is an ambitious project to create a high-quality, reproducible parallel corpus for Modern Standard Arabic (MSA) and Tunisian Arabic (aeb) through a sophisticated synthetic data generation pipeline. The dataset is being developed by the Tunisia.AI community to address the scarcity of high-quality dialectal data for training and evaluating language models. The primary goal is to provide a rich, well-documented resource for the research and development of:… See the full description on the dataset page: https://huggingface.co/datasets/tunis-ai/tunisian-msa-parallel-corpus.tabulartranslation1K<n<10K0 likes102 downloads1y agoHugging Face29msamg /QnA_Descriptivetextn<1K0 likes99 downloads2y agoHugging Face30ragrawal36 /msa-musique-docs-with-idstext10K<n<100K0 likes99 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.