datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
language-identification
Dataset Card for Language Identification dataset
Dataset Summary
The Language Identification dataset is a collection of 90k samples consisting of text passages and corresponding language label.
This dataset was created by collecting data from 3 sources: Multilingual Amazon Reviews Corpus, XNLI, and STSb Multi MT.
Supported Tasks and Leaderboards
The dataset can be used to train a model for language identification, which is a multi-class text classification… See the full description on the dataset page: https://huggingface.co/datasets/papluca/language-identification.language-identificationtask427_hindienglish_corpora_hi-en_language_identification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task427_hindienglish_corpora_hi-en_language_identification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task427_hindienglish_corpora_hi-en_language_identification.portuguese-language-identification-rawILID_Indian_Language_Identification_Dataset
ILID: Native Script Language Identification for Indian Languages
Paper | Code | Project Page
🗣 ILID: Indian Language Identification Dataset (23 Languages)Authors: Yash Ingle, Dr. Pruthwik MishraInstitute: Sardar Vallabhbhai National Institute of Technology (SVNIT), Surat, India
📄 Dataset Description
The ILID (Indian Language Identification Dataset) benchmark contains 250,000sentences from English and 22 official Indian languages, designed for training and evaluating… See the full description on the dataset page: https://huggingface.co/datasets/yash-ingle/ILID_Indian_Language_Identification_Dataset.Language-IdentificationLanguage_Identificationsouth_african_language_identificationLanguage_Identification_v1
Dataset Card for Language Identification Dataset
Dataset Summary
A comprehensive dataset for Indian language identification and text classification. The dataset contains text samples across 10 major Indian languages, making it suitable for developing language identification systems and multilingual NLP applications.
Languages and Distribution
Language Distribution:
Urdu 1000
Hindi 1000
Odia 1000
Tamil 1000
Kannada 1000
Bengali… See the full description on the dataset page: https://huggingface.co/datasets/Process-Venue/Language_Identification_v1.task441_eng_guj_parallel_corpus_gu-en_language_identification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task441_eng_guj_parallel_corpus_gu-en_language_identification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task441_eng_guj_parallel_corpus_gu-en_language_identification.task265_paper_reviews_language_identification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task265_paper_reviews_language_identification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task265_paper_reviews_language_identification.task533_europarl_es-en_language_identification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task533_europarl_es-en_language_identification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task533_europarl_es-en_language_identification.Wiki2018_Devanagari_Script_Language_Identificationtask562_alt_language_identification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task562_alt_language_identification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task562_alt_language_identification.language-identification
Dataset Card for Language Identification dataset
Dataset Summary
The Language Identification dataset is a collection of 90k samples consisting of text passages and corresponding language label.
This dataset was created by collecting data from 3 sources: Multilingual Amazon Reviews Corpus, XNLI, and STSb Multi MT.
Supported Tasks and Leaderboards
The dataset can be used to train a model for language identification, which is a multi-class text classification… See the full description on the dataset page: https://huggingface.co/datasets/chiragkolte01/language-identification.language-identification
Dataset Card for Language Identification dataset
Dataset Summary
The Language Identification dataset is a collection of 90k samples consisting of text passages and corresponding language label.
This dataset was created by collecting data from 3 sources: Multilingual Amazon Reviews Corpus, XNLI, and STSb Multi MT.
Supported Tasks and Leaderboards
The dataset can be used to train a model for language identification, which is a multi-class text classification… See the full description on the dataset page: https://huggingface.co/datasets/borrore/language-identification.gn-offensive-language-identification
Text-based afective computing
We collected a dataset of tweets primarily written in Guarani (and Jopara, a code-switching language that combines Guarani and Spanish) and annotated them for three widely-used dimensions in sentiment analysis:
emotion recognition (https://huggingface.co/datasets/mmaguero/gn-emotion-recognition),
humor detection (https://huggingface.co/datasets/mmaguero/gn-humor-detection), and
offensive language identification (this repo… See the full description on the dataset page: https://huggingface.co/datasets/mmaguero/gn-offensive-language-identification.task315_europarl_sv-en_language_identification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task315_europarl_sv-en_language_identification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task315_europarl_sv-en_language_identification.LanguageIdentification_VoxForge
Dataset Card for "LanguageIdentification_VoxForge"
More Information needed
task1574_amazon_reviews_multi_language_identification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1574_amazon_reviews_multi_language_identification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1574_amazon_reviews_multi_language_identification.sapient-synth-flan-niv2-fsopt-data-task265-paper-reviews-language-identification
sapient-synth-flan-niv2-fsopt-data-task265-paper-reviews-language-identification
Chat-template-ready synthetic anonymous replacement examples for one Sapient source excluded from the DFM5 data mix.
Contents
Format: gzip-compressed JSON Lines under data/train.jsonl.gz
Schema: {"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]}
Files: 1
Rows: 1341
Task: synthetic anonymous instruction replacement
Generation… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/sapient-synth-flan-niv2-fsopt-data-task265-paper-reviews-language-identification.LanguageIdentification_VoxForge_TTScode-switched-language-identificationflan_combined_task1574_amazon_reviews_multi_language_identificationLanguageIdentification_VoxForgetask1621_menyo20k-mt_en_yo_language_identification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1621_menyo20k-mt_en_yo_language_identification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1621_menyo20k-mt_en_yo_language_identification.flan_combined_task265_paper_reviews_language_identificationflan_source_task441_eng_guj_parallel_corpus_gu-en_language_identification_303flan_combined_task976_pib_indian_language_identificationflan_combined_task427_hindienglish_corpora_hi-en_language_identification
