CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01papluca /language-identification Dataset Card for Language Identification dataset Dataset Summary The Language Identification dataset is a collection of 90k samples consisting of text passages and corresponding language label. This dataset was created by collecting data from 3 sources: Multilingual Amazon Reviews Corpus, XNLI, and STSb Multi MT. Supported Tasks and Leaderboards The dataset can be used to train a model for language identification, which is a multi-class text classification… See the full description on the dataset page: https://huggingface.co/datasets/papluca/language-identification.texttext-classification10K<n<100K70 likes4.4k downloads4y agoHugging Face02unklefedor /language-identificationtext100K<n<1M1 likes155 downloads3y agoHugging Face03Lots-of-LoRAs /task427_hindienglish_corpora_hi-en_language_identification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task427_hindienglish_corpora_hi-en_language_identification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task427_hindienglish_corpora_hi-en_language_identification.texttext-generation1K<n<10K0 likes112 downloads2y agoHugging Face04arubenruben /portuguese-language-identification-rawtext10M<n<100M0 likes101 downloads3y agoHugging Face05yash-ingle /ILID_Indian_Language_Identification_Dataset ILID: Native Script Language Identification for Indian Languages Paper | Code | Project Page 🗣 ILID: Indian Language Identification Dataset (23 Languages)Authors: Yash Ingle, Dr. Pruthwik MishraInstitute: Sardar Vallabhbhai National Institute of Technology (SVNIT), Surat, India 📄 Dataset Description The ILID (Indian Language Identification Dataset) benchmark contains 250,000sentences from English and 22 official Indian languages, designed for training and evaluating… See the full description on the dataset page: https://huggingface.co/datasets/yash-ingle/ILID_Indian_Language_Identification_Dataset.texttext-classification100K<n<1M0 likes97 downloads9mo agoHugging Face06pranavagrawal /Language-Identificationtext1M<n<10M0 likes53 downloads2y agoHugging Face07MUGEN-Benchmark /Language_Identificationaudion<1K0 likes36 downloads8mo agoHugging Face08mlexplorer008 /south_african_language_identificationtext10K<n<100K1 likes30 downloads2y agoHugging Face09Process-Venue /Language_Identification_v1 Dataset Card for Language Identification Dataset Dataset Summary A comprehensive dataset for Indian language identification and text classification. The dataset contains text samples across 10 major Indian languages, making it suitable for developing language identification systems and multilingual NLP applications. Languages and Distribution Language Distribution: Urdu 1000 Hindi 1000 Odia 1000 Tamil 1000 Kannada 1000 Bengali… See the full description on the dataset page: https://huggingface.co/datasets/Process-Venue/Language_Identification_v1.texttext-classification1K<n<10K1 likes28 downloads2y agoHugging Face10Lots-of-LoRAs /task441_eng_guj_parallel_corpus_gu-en_language_identification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task441_eng_guj_parallel_corpus_gu-en_language_identification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task441_eng_guj_parallel_corpus_gu-en_language_identification.texttext-generation1K<n<10K0 likes27 downloads2y agoHugging Face11Lots-of-LoRAs /task265_paper_reviews_language_identification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task265_paper_reviews_language_identification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task265_paper_reviews_language_identification.texttext-generationn<1K0 likes23 downloads2y agoHugging Face12Lots-of-LoRAs /task533_europarl_es-en_language_identification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task533_europarl_es-en_language_identification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task533_europarl_es-en_language_identification.texttext-generation1K<n<10K0 likes22 downloads2y agoHugging Face131-800-SHARED-TASKS /Wiki2018_Devanagari_Script_Language_Identificationtext1K<n<10K0 likes21 downloads2y agoHugging Face14Lots-of-LoRAs /task562_alt_language_identification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task562_alt_language_identification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task562_alt_language_identification.texttext-generationn<1K0 likes21 downloads2y agoHugging Face15chiragkolte01 /language-identification Dataset Card for Language Identification dataset Dataset Summary The Language Identification dataset is a collection of 90k samples consisting of text passages and corresponding language label. This dataset was created by collecting data from 3 sources: Multilingual Amazon Reviews Corpus, XNLI, and STSb Multi MT. Supported Tasks and Leaderboards The dataset can be used to train a model for language identification, which is a multi-class text classification… See the full description on the dataset page: https://huggingface.co/datasets/chiragkolte01/language-identification.texttext-classification10K<n<100K0 likes21 downloads5mo agoHugging Face16borrore /language-identification Dataset Card for Language Identification dataset Dataset Summary The Language Identification dataset is a collection of 90k samples consisting of text passages and corresponding language label. This dataset was created by collecting data from 3 sources: Multilingual Amazon Reviews Corpus, XNLI, and STSb Multi MT. Supported Tasks and Leaderboards The dataset can be used to train a model for language identification, which is a multi-class text classification… See the full description on the dataset page: https://huggingface.co/datasets/borrore/language-identification.texttext-classification10K<n<100K0 likes20 downloads6mo agoHugging Face17mmaguero /gn-offensive-language-identification Text-based afective computing We collected a dataset of tweets primarily written in Guarani (and Jopara, a code-switching language that combines Guarani and Spanish) and annotated them for three widely-used dimensions in sentiment analysis: emotion recognition (https://huggingface.co/datasets/mmaguero/gn-emotion-recognition), humor detection (https://huggingface.co/datasets/mmaguero/gn-humor-detection), and offensive language identification (this repo… See the full description on the dataset page: https://huggingface.co/datasets/mmaguero/gn-offensive-language-identification.texttext-classification1K<n<10K0 likes19 downloads2y agoHugging Face18Lots-of-LoRAs /task315_europarl_sv-en_language_identification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task315_europarl_sv-en_language_identification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task315_europarl_sv-en_language_identification.texttext-generation1K<n<10K0 likes18 downloads2y agoHugging Face19DynamicSuperb /LanguageIdentification_VoxForge Dataset Card for "LanguageIdentification_VoxForge" More Information needed audion<1K0 likes16 downloads3y agoHugging Face20Lots-of-LoRAs /task1574_amazon_reviews_multi_language_identification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1574_amazon_reviews_multi_language_identification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1574_amazon_reviews_multi_language_identification.texttext-generationn<1K0 likes14 downloads2y agoHugging Face21schneiderkamplab /sapient-synth-flan-niv2-fsopt-data-task265-paper-reviews-language-identification sapient-synth-flan-niv2-fsopt-data-task265-paper-reviews-language-identification Chat-template-ready synthetic anonymous replacement examples for one Sapient source excluded from the DFM5 data mix. Contents Format: gzip-compressed JSON Lines under data/train.jsonl.gz Schema: {"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]} Files: 1 Rows: 1341 Task: synthetic anonymous instruction replacement Generation… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/sapient-synth-flan-niv2-fsopt-data-task265-paper-reviews-language-identification.text1K<n<10K0 likes14 downloads3mo agoHugging Face22DynamicSuperbPrivate /LanguageIdentification_VoxForge_TTSaudion<1K0 likes13 downloads2y agoHugging Face23sameeramin /code-switched-language-identificationtext100K<n<1M0 likes13 downloads2y agoHugging Face24supergoose /flan_combined_task1574_amazon_reviews_multi_language_identificationtextn<1K0 likes9 downloads2y agoHugging Face25macabdul9 /LanguageIdentification_VoxForgeaudion<1K0 likes8 downloads2y agoHugging Face26Lots-of-LoRAs /task1621_menyo20k-mt_en_yo_language_identification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1621_menyo20k-mt_en_yo_language_identification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1621_menyo20k-mt_en_yo_language_identification.texttext-generation1K<n<10K0 likes7 downloads2y agoHugging Face27supergoose /flan_combined_task265_paper_reviews_language_identificationtext1K<n<10K0 likes5 downloads2y agoHugging Face28supergoose /flan_source_task441_eng_guj_parallel_corpus_gu-en_language_identification_303text10K<n<100K0 likes4 downloads2y agoHugging Face29supergoose /flan_combined_task976_pib_indian_language_identificationtext1K<n<10K0 likes4 downloads2y agoHugging Face30supergoose /flan_combined_task427_hindienglish_corpora_hi-en_language_identificationtext10K<n<100K0 likes4 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.