datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
acronym_identification
Dataset Card for Acronym Identification Dataset
Dataset Summary
This dataset contains the training, validation, and test data for the Shared Task 1: Acronym Identification of the AAAI-21 Workshop on Scientific Document Understanding.
Supported Tasks and Leaderboards
The dataset supports an acronym-identification task, where the aim is to predic which tokens in a pre-tokenized sentence correspond to acronyms. The dataset was released for a Shared Task which… See the full description on the dataset page: https://huggingface.co/datasets/amirveyseh/acronym_identification.speaker_identification_100_speakersArabic_Dialect_IdentificationArabic dialects, multi-class-Classification, Tweets.
Dataset Card for Arabic_Dialect_Identification
Dataset Summary
We present QADI, an automatically collected dataset of tweets belonging to a wide range of
country-level Arabic dialects covering 18 different countries in the Middle East and North
Africa region. Our method for building this dataset relies on applying multiple filters to identify
users who belong to different countries based on their account descriptions… See the full description on the dataset page: https://huggingface.co/datasets/Abdelrahman-Rezk/Arabic_Dialect_Identification.speaker_identification_100_speakers_task427_hindienglish_corpora_hi-en_language_identification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task427_hindienglish_corpora_hi-en_language_identification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task427_hindienglish_corpora_hi-en_language_identification.portuguese-language-identification-rawbacbench-operon-identification-protein-sequences
Dataset for operon identification in bacteria (Protein sequences)
A dataset of 4,073 operons across 11 bacterial genomes species.
The operon annotations have been extracted from Operon DB and the genome protein sequences have been extracted from GenBank. Each row contains a set of protein sequences present in the genome, represented
by a list of protein sequences from different contigs.
We extracted high-confidence (i.e. known) operons from Operon DB, filtered out non-contigous… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-operon-identification-protein-sequences.task112_asset_simple_sentence_identification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task112_asset_simple_sentence_identification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task112_asset_simple_sentence_identification.Domain_Identification_Algorithms_Comparison_Dataindian-ipc-statute-identification
Indian IPC Statute Identification
Given the facts of an Indian court case, identify the relevant Indian Penal Code (IPC) section. Each example
pairs the factual narrative of a High Court judgment with the text of an IPC section that the judgment applies.
The task is framed as retrieval / statute identification: from a fact scenario, retrieve (or classify) the governing
statute. It is a useful benchmark and training signal for legal information retrieval, legal text… See the full description on the dataset page: https://huggingface.co/datasets/Hanno-Labs/indian-ipc-statute-identification.LID201_Devanagari_Script_Languages_Identificationoperon-identification-long-read-rna-sequencing-protein-sequences
Dataset for operon identification from long-read RNA sequencing
A dataset of annotated operons across 5 distinct bacterial strains. The operons were annotated by running and analysing long-read RNA sequencing and identifying genes
located on the same transcripts.
The genome protein sequences have been extracted from GenBank. Each row contains whole bacterial genome represented by an ordered list
of protein sequences.
Usage
For a complete example on how to read and use… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/operon-identification-long-read-rna-sequencing-protein-sequences.Language_IdentificationBias_identification
Gathered Dataset for Stereotypical Bias Analysis
This dataset was compiled to analyze various types of stereotypical biases present in language models. It incorporates data from multiple publicly available datasets, each contributing to the identification of specific bias types.
Source Datasets
The following datasets were used to create this comprehensive dataset:
StereoSet
CrowS-Pair
Multi-Grain Stereotype Dataset
Investigating Subtler Biases: Ageism, Beauty… See the full description on the dataset page: https://huggingface.co/datasets/PriyaPatel/Bias_identification.Accent_Identificationsafety_risk_identificationLanguage_Identification_v1
Dataset Card for Language Identification Dataset
Dataset Summary
A comprehensive dataset for Indian language identification and text classification. The dataset contains text samples across 10 major Indian languages, making it suitable for developing language identification systems and multilingual NLP applications.
Languages and Distribution
Language Distribution:
Urdu 1000
Hindi 1000
Odia 1000
Tamil 1000
Kannada 1000
Bengali… See the full description on the dataset page: https://huggingface.co/datasets/Process-Venue/Language_Identification_v1.Workplace-Hazard-Identificationtask441_eng_guj_parallel_corpus_gu-en_language_identification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task441_eng_guj_parallel_corpus_gu-en_language_identification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task441_eng_guj_parallel_corpus_gu-en_language_identification.Identification-of-paraphrasing
🇰🇿 Identification of Paraphrasing in Kazakh Context
Dataset Summary
Identification of Paraphrasing in Kazakh Context is a targeted dataset designed to train Large Language Models (LLMs) and embeddings to detect semantic equivalence between two distinct Kazakh texts.
📊 Dataset Statistics
General Metrics
Metric
Count
Total Samples
2,000
Total Words (approx.)
184,465
Avg. Words per Sample
92
Word Count… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Identification-of-paraphrasing.acronym_identification_promptsourcetask265_paper_reviews_language_identification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task265_paper_reviews_language_identification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task265_paper_reviews_language_identification.task533_europarl_es-en_language_identification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task533_europarl_es-en_language_identification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task533_europarl_es-en_language_identification.Topic-Identification
🇰🇿 Topic Identification in Kazakh Context
Dataset Summary
Topic Identification is a curated dataset designed to train Large Language Models (LLMs) to accurately extract core themes, keywords, and main subjects from Kazakh texts.
This dataset teaches models to read a paragraph of text and distill its contents into a concise list of relevant topics. Covering various professional and academic domains (such as Disaster Management, Linguistics, and Research).… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Topic-Identification.Wiki2018_Devanagari_Script_Language_Identificationtask562_alt_language_identification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task562_alt_language_identification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task562_alt_language_identification.EPPC_Miner_Identification_trainbacbench-operon-identification-dna
Dataset for operon identification in bacteria (DNA)
A dataset of 4,073 operons across 11 bacterial genomes species.
The operon annotations have been extracted from Operon DB and the genome DNA sequences have been extracted from GenBank. Each row contains whole bacterial genome, represented
by a list of DNA sequences from different contigs.
We extracted high-confidence (i.e. known) operons from Operon DB, filtered out non-contigous operons and only kept genomes with at least 9 known… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-operon-identification-dna.task315_europarl_sv-en_language_identification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task315_europarl_sv-en_language_identification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task315_europarl_sv-en_language_identification.NPTL_Datasets_with_speaker_identification
