CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Anoy123423123 /MSA_PretrainData MSA Pretrain Data Retrieval-style pretraining corpora. Each subset is split into two parts: file columns meaning <subset>/queries/*.parquet question, answer, reference_ids: list<int64>, labels: list<int64> query, plus row indices into the subset's reference table <subset>/references/*.parquet value: string the reference/memory passage text reference_ids are the candidate pool for a query; labels are the positive(s). Both are global row indices into the subset's… See the full description on the dataset page: https://huggingface.co/datasets/Anoy123423123/MSA_PretrainData.texttext-retrieval10M<n<100M0 likes7.3k downloads2mo agoHugging Face02Dr-AliGomaa /ar-quran-hadith14books-MSA ar-quran-hadith14books-MSA Arabic speech for both primary sources of Islam — Quran and Hadith — plus cleaned general Modern Standard Arabic, under one construction pipeline and one text convention. ASR errors on sacred text are not ordinary errors: a plausible-sounding substitution can alter the meaning of scripture, and because chatbots, search and summarizers increasingly answer from transcriptions rather than from audio, such an error propagates silently. Quranic recitation… See the full description on the dataset page: https://huggingface.co/datasets/Dr-AliGomaa/ar-quran-hadith14books-MSA.audioautomatic-speech-recognition10K<n<100K6 likes1.1k downloads1mo agoHugging Face03astro-legacy-archive /msam-released-products MSAM released flight products This dataset contains the complete numeric contents of the three MSAM1 flight archives served by LAMBDA: the June 1992, June 1994, and June 1995 observation tables, sampled beam maps, and released 1994/1995 covariance matrices. The 38 configuration identifiers preserve the flight directory and source filename stem. How to use python -m venv .venv && .venv/bin/pip install datasets huggingface_hub pyarrow from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/msam-released-products.tabular10K<n<100K0 likes768 downloads21d agoHugging Face04ragrawal36 /msa-hotpotqa-qa-with-idstext1K<n<10K0 likes324 downloads5mo agoHugging Face05oddadmix /msa-omnivoice-tts-v1 MSA-OmniVoice-v1 Dataset Description MSA-OmniVoice-v1 is a 50-hour synthetic Modern Standard Arabic (MSA) speech dataset generated using OmniVoice. The dataset contains high-quality synthetic speech from a single speaker paired with fully diacritized (تشكيل) transcripts. It is intended for training and fine-tuning Arabic speech models, including Text-to-Speech (TTS), Automatic Speech Recognition (ASR), speech representation learning, and alignment tasks.… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/msa-omnivoice-tts-v1.audio10K<n<100K1 likes289 downloads3mo agoHugging Face06songlab /gpn-msa-sapiens-dataset Training windows for GPN-MSA-Sapiens For more information check out our paper and repository. Path in Snakemake: results/dataset/multiz100way/89/128/64/True/defined.phastCons.percentile-75_0.05_0.001 tabular1M<n<10M0 likes243 downloads2y agoHugging Face07ragrawal36 /msa-hotpotqa-docs-with-idstext1K<n<10K0 likes213 downloads5mo agoHugging Face08ragrawal36 /msa-musique-qa-with-idstextn<1K0 likes213 downloads5mo agoHugging Face09HeshamHaroon /arabic-msa-25k-saudi-male-tashkeel Arabic MSA 25K — Saudi Male (Tashkeel) 25,000 fully-diacritized Arabic MSA text + audio pairs, rendered with a single Saudi male neural voice at 48 kHz / 16-bit PCM, across 10 thematic categories. Dataset Summary arabic-msa-25k-saudi-male-tashkeel is a 25,000-clip Modern Standard Arabic (MSA) speech corpus with matching diacritized text (full tashkeel / ḥarakāt). Every clip is synthesized by the single voice ar-SA-HamedNeural (Azure Neural TTS, Saudi Arabic male) at 48… See the full description on the dataset page: https://huggingface.co/datasets/HeshamHaroon/arabic-msa-25k-saudi-male-tashkeel.tabulartext-to-speech10K<n<100K10 likes194 downloads5mo agoHugging Face10underfrog /msa-musique-qa-with-idstextn<1K0 likes193 downloads3mo agoHugging Face11underfrog /msa-musique-docs-with-idstext10K<n<100K0 likes185 downloads3mo agoHugging Face12clarayyu22 /gpn-msa-microglia-fulltabular1M<n<10M0 likes170 downloads2y agoHugging Face13ragrawal36 /msa-2wikimultihopqa-qa-with-idstext1K<n<10K0 likes169 downloads5mo agoHugging Face14dotan1111 /MSA-nuc-9-seq Multiple Sequence Alignment as a Sequence-to-Sequence Learning Problem Abstract: The sequence alignment problem is one of the most fundamental problems in bioinformatics and a plethora of methods were devised to tackle it. Here we introduce BetaAlign, a methodology for aligning sequences using an NLP approach. BetaAlign accounts for the possible variability of the evolutionary process among different datasets by using an ensemble of transformers, each trained on millions… See the full description on the dataset page: https://huggingface.co/datasets/dotan1111/MSA-nuc-9-seq.text1M<n<10M0 likes151 downloads3y agoHugging Face15underfrog /msa-hotpotqa-qa-with-idstext1K<n<10K0 likes136 downloads3mo agoHugging Face16Yusser /m_sae_wiki_tokenized1M<n<10M0 likes130 downloads2y agoHugging Face17dotan1111 /MSA-amino-9-seq Multiple Sequence Alignment as a Sequence-to-Sequence Learning Problem Abstract: The sequence alignment problem is one of the most fundamental problems in bioinformatics and a plethora of methods were devised to tackle it. Here we introduce BetaAlign, a methodology for aligning sequences using an NLP approach. BetaAlign accounts for the possible variability of the evolutionary process among different datasets by using an ensemble of transformers, each trained on millions… See the full description on the dataset page: https://huggingface.co/datasets/dotan1111/MSA-amino-9-seq.text1M<n<10M1 likes122 downloads3y agoHugging Face18dotan1111 /MSA-nuc-8-seq Multiple Sequence Alignment as a Sequence-to-Sequence Learning Problem Abstract: The sequence alignment problem is one of the most fundamental problems in bioinformatics and a plethora of methods were devised to tackle it. Here we introduce BetaAlign, a methodology for aligning sequences using an NLP approach. BetaAlign accounts for the possible variability of the evolutionary process among different datasets by using an ensemble of transformers, each trained on millions… See the full description on the dataset page: https://huggingface.co/datasets/dotan1111/MSA-nuc-8-seq.text1M<n<10M0 likes115 downloads3y agoHugging Face19msaligs /deepfake_iitm_rawaudio10K<n<100K0 likes112 downloads6mo agoHugging Face20underfrog /msa-hotpotqa-docs-with-idstext1K<n<10K0 likes112 downloads4mo agoHugging Face21dotan1111 /MSA-nuc-7-seq Multiple Sequence Alignment as a Sequence-to-Sequence Learning Problem Abstract: The sequence alignment problem is one of the most fundamental problems in bioinformatics and a plethora of methods were devised to tackle it. Here we introduce BetaAlign, a methodology for aligning sequences using an NLP approach. BetaAlign accounts for the possible variability of the evolutionary process among different datasets by using an ensemble of transformers, each trained on millions… See the full description on the dataset page: https://huggingface.co/datasets/dotan1111/MSA-nuc-7-seq.text1M<n<10M0 likes103 downloads3y agoHugging Face22ragrawal36 /msa-musique-docs-with-idstext10K<n<100K0 likes100 downloads5mo agoHugging Face23ragrawal36 /msa-2wikimultihopqa-docs-with-idstext1K<n<10K0 likes96 downloads5mo agoHugging Face24badrex /arabic-speech-SADA22-MSA Dataset Card for SADA (Saudi Audio Dataset for Arabic) ⚠️ Caution This is only the portion of the SADA dataset where the speaker dialect is Modern Standard Arabic (MSA). To access full dataset, you should check this link. Dataset Summary The SADA dataset (Saudi Audio Dataset for Arabic) is a large-scale Arabic speech corpus designed to support the development of high-quality artificial intelligence models for Arabic speech processing. It contains over… See the full description on the dataset page: https://huggingface.co/datasets/badrex/arabic-speech-SADA22-MSA.audioautomatic-speech-recognition1K<n<10K2 likes93 downloads1y agoHugging Face25otozz /MSA_train_setPre-processed MSA data based on https://huggingface.co/datasets/mozilla-foundation/common_voice_16_1. audio10K<n<100K1 likes90 downloads2y agoHugging Face26vm2825 /msa-musique-qatextn<1K0 likes88 downloads5mo agoHugging Face27underfrog /msa-2wikimultihopqa-qa-with-idstext1K<n<10K0 likes85 downloads3mo agoHugging Face28dotan1111 /MSA-amino-7-seq Multiple Sequence Alignment as a Sequence-to-Sequence Learning Problem Abstract: The sequence alignment problem is one of the most fundamental problems in bioinformatics and a plethora of methods were devised to tackle it. Here we introduce BetaAlign, a methodology for aligning sequences using an NLP approach. BetaAlign accounts for the possible variability of the evolutionary process among different datasets by using an ensemble of transformers, each trained on millions… See the full description on the dataset page: https://huggingface.co/datasets/dotan1111/MSA-amino-7-seq.text1M<n<10M0 likes82 downloads3y agoHugging Face29dotan1111 /MSA-amino-8-seq Multiple Sequence Alignment as a Sequence-to-Sequence Learning Problem Abstract: The sequence alignment problem is one of the most fundamental problems in bioinformatics and a plethora of methods were devised to tackle it. Here we introduce BetaAlign, a methodology for aligning sequences using an NLP approach. BetaAlign accounts for the possible variability of the evolutionary process among different datasets by using an ensemble of transformers, each trained on millions… See the full description on the dataset page: https://huggingface.co/datasets/dotan1111/MSA-amino-8-seq.text1M<n<10M0 likes82 downloads3y agoHugging Face30dotan1111 /MSA-nuc-5-seq Multiple Sequence Alignment as a Sequence-to-Sequence Learning Problem Abstract: The sequence alignment problem is one of the most fundamental problems in bioinformatics and a plethora of methods were devised to tackle it. Here we introduce BetaAlign, a methodology for aligning sequences using an NLP approach. BetaAlign accounts for the possible variability of the evolutionary process among different datasets by using an ensemble of transformers, each trained on millions… See the full description on the dataset page: https://huggingface.co/datasets/dotan1111/MSA-nuc-5-seq.text1M<n<10M0 likes78 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.