CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01amirveyseh /acronym_identification Dataset Card for Acronym Identification Dataset Dataset Summary This dataset contains the training, validation, and test data for the Shared Task 1: Acronym Identification of the AAAI-21 Workshop on Scientific Document Understanding. Supported Tasks and Leaderboards The dataset supports an acronym-identification task, where the aim is to predic which tokens in a pre-tokenized sentence correspond to acronyms. The dataset was released for a Shared Task which… See the full description on the dataset page: https://huggingface.co/datasets/amirveyseh/acronym_identification.texttoken-classification10K<n<100K23 likes67k downloads3y agoHugging Face02thucdangvan020999 /speaker_identification_100_speakersaudio1K<n<10K0 likes175 downloads11mo agoHugging Face03Abdelrahman-Rezk /Arabic_Dialect_IdentificationArabic dialects, multi-class-Classification, Tweets. Dataset Card for Arabic_Dialect_Identification Dataset Summary We present QADI, an automatically collected dataset of tweets belonging to a wide range of country-level Arabic dialects covering 18 different countries in the Middle East and North Africa region. Our method for building this dataset relies on applying multiple filters to identify users who belong to different countries based on their account descriptions… See the full description on the dataset page: https://huggingface.co/datasets/Abdelrahman-Rezk/Arabic_Dialect_Identification.tabular100K<n<1M12 likes173 downloads4y agoHugging Face04thucdangvan020999 /speaker_identification_100_speakers_audio1K<n<10K0 likes126 downloads11mo agoHugging Face05Lots-of-LoRAs /task427_hindienglish_corpora_hi-en_language_identification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task427_hindienglish_corpora_hi-en_language_identification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task427_hindienglish_corpora_hi-en_language_identification.texttext-generation1K<n<10K0 likes112 downloads2y agoHugging Face06arubenruben /portuguese-language-identification-rawtext10M<n<100M0 likes101 downloads3y agoHugging Face07macwiatrak /bacbench-operon-identification-protein-sequences Dataset for operon identification in bacteria (Protein sequences) A dataset of 4,073 operons across 11 bacterial genomes species. The operon annotations have been extracted from Operon DB and the genome protein sequences have been extracted from GenBank. Each row contains a set of protein sequences present in the genome, represented by a list of protein sequences from different contigs. We extracted high-confidence (i.e. known) operons from Operon DB, filtered out non-contigous… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-operon-identification-protein-sequences.textn<1K0 likes98 downloads1y agoHugging Face08Lots-of-LoRAs /task112_asset_simple_sentence_identification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task112_asset_simple_sentence_identification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task112_asset_simple_sentence_identification.texttext-generation1K<n<10K0 likes83 downloads2y agoHugging Face09broadinstitute /Domain_Identification_Algorithms_Comparison_Datatabular10M<n<100M0 likes56 downloads6mo agoHugging Face10Hanno-Labs /indian-ipc-statute-identification Indian IPC Statute Identification Given the facts of an Indian court case, identify the relevant Indian Penal Code (IPC) section. Each example pairs the factual narrative of a High Court judgment with the text of an IPC section that the judgment applies. The task is framed as retrieval / statute identification: from a fact scenario, retrieve (or classify) the governing statute. It is a useful benchmark and training signal for legal information retrieval, legal text… See the full description on the dataset page: https://huggingface.co/datasets/Hanno-Labs/indian-ipc-statute-identification.texttext-retrieval10K<n<100K0 likes45 downloads3mo agoHugging Face111-800-SHARED-TASKS /LID201_Devanagari_Script_Languages_Identificationtext1M<n<10M0 likes43 downloads2y agoHugging Face12macwiatrak /operon-identification-long-read-rna-sequencing-protein-sequences Dataset for operon identification from long-read RNA sequencing A dataset of annotated operons across 5 distinct bacterial strains. The operons were annotated by running and analysing long-read RNA sequencing and identifying genes located on the same transcripts. The genome protein sequences have been extracted from GenBank. Each row contains whole bacterial genome represented by an ordered list of protein sequences. Usage For a complete example on how to read and use… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/operon-identification-long-read-rna-sequencing-protein-sequences.textn<1K0 likes42 downloads1y agoHugging Face13MUGEN-Benchmark /Language_Identificationaudion<1K0 likes36 downloads8mo agoHugging Face14PriyaPatel /Bias_identification Gathered Dataset for Stereotypical Bias Analysis This dataset was compiled to analyze various types of stereotypical biases present in language models. It incorporates data from multiple publicly available datasets, each contributing to the identification of specific bias types. Source Datasets The following datasets were used to create this comprehensive dataset: StereoSet CrowS-Pair Multi-Grain Stereotype Dataset Investigating Subtler Biases: Ageism, Beauty… See the full description on the dataset page: https://huggingface.co/datasets/PriyaPatel/Bias_identification.texttext-classification10K<n<100K2 likes34 downloads2y agoHugging Face15MUGEN-Benchmark /Accent_Identificationaudion<1K0 likes30 downloads8mo agoHugging Face16kirito011024 /safety_risk_identificationimagen<1K0 likes28 downloads2y agoHugging Face17Process-Venue /Language_Identification_v1 Dataset Card for Language Identification Dataset Dataset Summary A comprehensive dataset for Indian language identification and text classification. The dataset contains text samples across 10 major Indian languages, making it suitable for developing language identification systems and multilingual NLP applications. Languages and Distribution Language Distribution: Urdu 1000 Hindi 1000 Odia 1000 Tamil 1000 Kannada 1000 Bengali… See the full description on the dataset page: https://huggingface.co/datasets/Process-Venue/Language_Identification_v1.texttext-classification1K<n<10K1 likes28 downloads2y agoHugging Face18NimaZahedinameghi /Workplace-Hazard-Identificationtext1K<n<10K1 likes27 downloads2y agoHugging Face19Lots-of-LoRAs /task441_eng_guj_parallel_corpus_gu-en_language_identification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task441_eng_guj_parallel_corpus_gu-en_language_identification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task441_eng_guj_parallel_corpus_gu-en_language_identification.texttext-generation1K<n<10K0 likes27 downloads2y agoHugging Face20farabi-lab /Identification-of-paraphrasinggated 🇰🇿 Identification of Paraphrasing in Kazakh Context Dataset Summary Identification of Paraphrasing in Kazakh Context is a targeted dataset designed to train Large Language Models (LLMs) and embeddings to detect semantic equivalence between two distinct Kazakh texts. 📊 Dataset Statistics General Metrics Metric Count Total Samples 2,000 Total Words (approx.) 184,465 Avg. Words per Sample 92 Word Count… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Identification-of-paraphrasing.texttext-classification1K<n<10K0 likes25 downloads2mo agoHugging Face21marcov /acronym_identification_promptsourcetext10K<n<100K0 likes23 downloads2y agoHugging Face22Lots-of-LoRAs /task265_paper_reviews_language_identification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task265_paper_reviews_language_identification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task265_paper_reviews_language_identification.texttext-generationn<1K0 likes23 downloads2y agoHugging Face23Lots-of-LoRAs /task533_europarl_es-en_language_identification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task533_europarl_es-en_language_identification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task533_europarl_es-en_language_identification.texttext-generation1K<n<10K0 likes22 downloads2y agoHugging Face24farabi-lab /Topic-Identificationgated 🇰🇿 Topic Identification in Kazakh Context Dataset Summary Topic Identification is a curated dataset designed to train Large Language Models (LLMs) to accurately extract core themes, keywords, and main subjects from Kazakh texts. This dataset teaches models to read a paragraph of text and distill its contents into a concise list of relevant topics. Covering various professional and academic domains (such as Disaster Management, Linguistics, and Research).… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Topic-Identification.texttext-classification1K<n<10K0 likes22 downloads2mo agoHugging Face251-800-SHARED-TASKS /Wiki2018_Devanagari_Script_Language_Identificationtext1K<n<10K0 likes21 downloads2y agoHugging Face26Lots-of-LoRAs /task562_alt_language_identification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task562_alt_language_identification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task562_alt_language_identification.texttext-generationn<1K0 likes21 downloads2y agoHugging Face27YanAdjeNole /EPPC_Miner_Identification_traintext1K<n<10K0 likes21 downloads1y agoHugging Face28macwiatrak /bacbench-operon-identification-dna Dataset for operon identification in bacteria (DNA) A dataset of 4,073 operons across 11 bacterial genomes species. The operon annotations have been extracted from Operon DB and the genome DNA sequences have been extracted from GenBank. Each row contains whole bacterial genome, represented by a list of DNA sequences from different contigs. We extracted high-confidence (i.e. known) operons from Operon DB, filtered out non-contigous operons and only kept genomes with at least 9 known… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-operon-identification-dna.textn<1K0 likes19 downloads1y agoHugging Face29Lots-of-LoRAs /task315_europarl_sv-en_language_identification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task315_europarl_sv-en_language_identification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task315_europarl_sv-en_language_identification.texttext-generation1K<n<10K0 likes18 downloads2y agoHugging Face30FariqF /NPTL_Datasets_with_speaker_identificationtabular1K<n<10K0 likes18 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.