CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Anoy123423123 /MSA_PretrainData MSA Pretrain Data Retrieval-style pretraining corpora. Each subset is split into two parts: file columns meaning <subset>/queries/*.parquet question, answer, reference_ids: list<int64>, labels: list<int64> query, plus row indices into the subset's reference table <subset>/references/*.parquet value: string the reference/memory passage text reference_ids are the candidate pool for a query; labels are the positive(s). Both are global row indices into the subset's… See the full description on the dataset page: https://huggingface.co/datasets/Anoy123423123/MSA_PretrainData.texttext-retrieval10M<n<100M0 likes7.2k downloads2mo agoHugging Face02jasperyeoh2 /pevo-msa-grch38-19way pevo-msa-grch38-19way (dataset Hub) EN: Training data, project tables, and reproducibility artifacts for primate MSA variant-effect modeling. 中文: 灵长类 MSA 变异效应建模的训练数据与项目复现材料(不含模型权重)。 Results-status note (2026-07-29). The multi-seed metrics reported below are retained as historical registry records. They use an earlier scoring protocol and cohort convention, and are not comparable to the corrected strict-v2 results used for the final project conclusions. Do not use the values… See the full description on the dataset page: https://huggingface.co/datasets/jasperyeoh2/pevo-msa-grch38-19way.0 likes6.7k downloads14d agoHugging Face03LiteFold /human-proteome-wide-msa Human Proteome ColabFold MSAs This dataset contains ColabFold/MMseqs2 multiple sequence alignments and AlphaFold 3 JSON inputs for the human proteome query set. Contents a3m/shard-*/: one A3M file per protein accession, sharded to stay below repository directory file limits. af3_json/shard-*/: one AlphaFold 3 JSON file per protein accession, sharded to stay below repository directory file limits. manifest.tsv: tab-separated index with accession, metadata… See the full description on the dataset page: https://huggingface.co/datasets/LiteFold/human-proteome-wide-msa.other2 likes1.2k downloads16d agoHugging Face04Dr-AliGomaa /ar-quran-hadith14books-MSA ar-quran-hadith14books-MSA Arabic speech for both primary sources of Islam — Quran and Hadith — plus cleaned general Modern Standard Arabic, under one construction pipeline and one text convention. ASR errors on sacred text are not ordinary errors: a plausible-sounding substitution can alter the meaning of scripture, and because chatbots, search and summarizers increasingly answer from transcriptions rather than from audio, such an error propagates silently. Quranic recitation… See the full description on the dataset page: https://huggingface.co/datasets/Dr-AliGomaa/ar-quran-hadith14books-MSA.audioautomatic-speech-recognition10K<n<100K6 likes1.1k downloads1mo agoHugging Face05songlab /gpn-msa-hg38-scores GPN-MSA predictions for all possible SNPs in the human genome (~9 billion) For more information check out our paper and repository. Querying specific variants or genes Install the latest tabix:In your current conda environment (might be slow):conda install -c bioconda -c conda-forge htslib=1.18 or in a new conda environment:conda create -n tabix -c bioconda -c conda-forge htslib=1.18 conda activate tabix Query a specific region (e.g. BRCA1), from the remote file:… See the full description on the dataset page: https://huggingface.co/datasets/songlab/gpn-msa-hg38-scores.5 likes955 downloads2y agoHugging Face06astro-legacy-archive /msam-released-products MSAM released flight products This dataset contains the complete numeric contents of the three MSAM1 flight archives served by LAMBDA: the June 1992, June 1994, and June 1995 observation tables, sampled beam maps, and released 1994/1995 covariance matrices. The 38 configuration identifiers preserve the flight directory and source filename stem. How to use python -m venv .venv && .venv/bin/pip install datasets huggingface_hub pyarrow from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/msam-released-products.tabular10K<n<100K0 likes690 downloads17d agoHugging Face07APProjects /us-layoffs-by-metro-area-msa-warn-act US layoffs by metro area: 54,166 WARN notices mapped to 765 metro and micro areas Rebuilt 2026-09-21. 765 of the 935 US core-based statistical areas carry at least one layoff notice on record — 361 metropolitan and 404 micropolitan. Nobody hires, sells or reports by county. A recruiter covers Austin; an account team books the Phoenix metro; a reporter writes Bay Area layoffs. State agencies publish neither — they publish the site of a layoff as free text in 48 different… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-layoffs-by-metro-area-msa-warn-act.tabulartabular-regression10K<n<100K0 likes574 downloads8h agoHugging Face08ragrawal36 /msa-hotpotqa-qa-with-idstext1K<n<10K0 likes440 downloads5mo agoHugging Face09Omartificial-Intelligence-Space /FineWeb2-MSA FineWeb2 MSA Arabic This is the MSA Arabic Portion of The FineWeb2 Dataset. This dataset contains a rich collection of text in MSA Arabic (ISO 639-3: arz), a widely spoken dialect within the Afro-Asiatic language family. With over 439 million words and 1.4 million documents, it serves as a valuable resource for NLP development and linguistic research focused on Egyptian Arabic. Purpose of This Repository This repository provides easy access to the Arabic portion - MSA… See the full description on the dataset page: https://huggingface.co/datasets/Omartificial-Intelligence-Space/FineWeb2-MSA.text100M<n<1B2 likes439 downloads2y agoHugging Face10MSALab /ParaDLC-Bench ParaDLC-Bench ParaDLC-Bench (Parallel Detailed Localized Captioning Benchmark) is a benchmark for multi-region localized captioning that jointly evaluates caption quality and inference efficiency. It extends DLC-Bench from single-region evaluation to concurrent multi-region evaluation, explicitly stressing a model's ability to describe many regions at once without cross-region interference. 📄 Paper &nbsp;|&nbsp; 💻 Code &nbsp;|&nbsp; 🤖 PerceptionDLM Key… See the full description on the dataset page: https://huggingface.co/datasets/MSALab/ParaDLC-Bench.imageimage-to-textn<1K2 likes420 downloads3mo agoHugging Face11msaligane /tinystories_phonologytext10M<n<100M0 likes366 downloads3y agoHugging Face12ragrawal36 /msa-musique-qa-with-idstextn<1K0 likes326 downloads5mo agoHugging Face13msavchen-nasa /clr_motion_planning_hw_7image1K<n<10K0 likes319 downloads3mo agoHugging Face14ragrawal36 /msa-hotpotqa-docs-with-idstext1K<n<10K0 likes286 downloads5mo agoHugging Face15oddadmix /msa-omnivoice-tts-v1 MSA-OmniVoice-v1 Dataset Description MSA-OmniVoice-v1 is a 50-hour synthetic Modern Standard Arabic (MSA) speech dataset generated using OmniVoice. The dataset contains high-quality synthetic speech from a single speaker paired with fully diacritized (تشكيل) transcripts. It is intended for training and fine-tuning Arabic speech models, including Text-to-Speech (TTS), Automatic Speech Recognition (ASR), speech representation learning, and alignment tasks.… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/msa-omnivoice-tts-v1.audio10K<n<100K0 likes280 downloads3mo agoHugging Face16msaleme /mcp-sandbox-authority-boundary-profile MCP Sandbox Authority Boundary Profile Profile v0.1.0 · Release v0.2.0 - Experimental Characterization Profile Profile release date: 2026-07-23 Latest distribution release date: 2026-09-05 Execution containment is not proof of bounded authority. Start here For a one-minute, case-by-case reading of the profile, open the companion Authority Boundary Field Guide Space. It presents the released synthetic observations with their control question, observed result… See the full description on the dataset page: https://huggingface.co/datasets/msaleme/mcp-sandbox-authority-boundary-profile.textn<1K1 likes272 downloads17d agoHugging Face17m-sakka /agripotentialMore information and competition link: https://github.com/MohammadElSakka/agripotential https://www.codabench.org/competitions/12055/ https://zenodo.org/records/15551829 imageimage-segmentation1K<n<10K2 likes247 downloads2mo agoHugging Face18underfrog /msa-2wikimultihopqa-qa-with-idstext1K<n<10K0 likes241 downloads3mo agoHugging Face19songlab /gpn-msa-sapiens-dataset Training windows for GPN-MSA-Sapiens For more information check out our paper and repository. Path in Snakemake: results/dataset/multiz100way/89/128/64/True/defined.phastCons.percentile-75_0.05_0.001 tabular1M<n<10M0 likes240 downloads2y agoHugging Face20IwanPark /msa-lecture0 likes200 downloads1mo agoHugging Face21shenlehan /msavbench-videotext1K<n<10K0 likes199 downloads4d agoHugging Face22HeshamHaroon /arabic-msa-25k-saudi-male-tashkeel Arabic MSA 25K — Saudi Male (Tashkeel) 25,000 fully-diacritized Arabic MSA text + audio pairs, rendered with a single Saudi male neural voice at 48 kHz / 16-bit PCM, across 10 thematic categories. Dataset Summary arabic-msa-25k-saudi-male-tashkeel is a 25,000-clip Modern Standard Arabic (MSA) speech corpus with matching diacritized text (full tashkeel / ḥarakāt). Every clip is synthesized by the single voice ar-SA-HamedNeural (Azure Neural TTS, Saudi Arabic male) at 48… See the full description on the dataset page: https://huggingface.co/datasets/HeshamHaroon/arabic-msa-25k-saudi-male-tashkeel.tabulartext-to-speech10K<n<100K8 likes178 downloads5mo agoHugging Face23ragrawal36 /msa-2wikimultihopqa-qa-with-idstext1K<n<10K0 likes167 downloads5mo agoHugging Face24clarayyu22 /gpn-msa-microglia-fulltabular1M<n<10M0 likes164 downloads2y agoHugging Face25underfrog /msa-musique-docs-with-idstext10K<n<100K0 likes152 downloads3mo agoHugging Face26otozz /MSA_train_setPre-processed MSA data based on https://huggingface.co/datasets/mozilla-foundation/common_voice_16_1. audio10K<n<100K1 likes151 downloads2y agoHugging Face27underfrog /msa-musique-qa-with-idstextn<1K0 likes149 downloads3mo agoHugging Face28dotan1111 /MSA-nuc-9-seq Multiple Sequence Alignment as a Sequence-to-Sequence Learning Problem Abstract: The sequence alignment problem is one of the most fundamental problems in bioinformatics and a plethora of methods were devised to tackle it. Here we introduce BetaAlign, a methodology for aligning sequences using an NLP approach. BetaAlign accounts for the possible variability of the evolutionary process among different datasets by using an ensemble of transformers, each trained on millions… See the full description on the dataset page: https://huggingface.co/datasets/dotan1111/MSA-nuc-9-seq.text1M<n<10M0 likes146 downloads3y agoHugging Face29sebacascardo87 /msap-loop-ctl-202607280 likes145 downloads2mo agoHugging Face30GWZhong /MSA_OOD_Dataset_in_CIDerHere are the MSA OOD datasets mentioned in the CIDer paper. Please cite our paper if you find that useful for your research: @article{zhong2025towards, title={Towards Robust Multimodal Emotion Recognition under Missing Modalities and Distribution Shifts}, author={Zhong, Guowei and Huan, Ruohong and Wu, Mingzhen and Liang, Ronghua and Chen, Peng}, journal={arXiv preprint arXiv:2506.10452}, year={2025} } 0 likes134 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.