Language Identification
language-identification
Dataset Card for Language Identification dataset
Dataset Summary
The Language Identification dataset is a collection of 90k samples consisting of text passages and corresponding language label.
This dataset was created by collecting data from 3 sources: Multilingual Amazon Reviews Corpus, XNLI, and STSb Multi MT.
Supported Tasks and Leaderboards
The dataset can be used to train a model for language identification, which is a multi-class text classification… See the full description on the dataset page: https://huggingface.co/datasets/papluca/language-identification.language_identification
语种识别
Tips:
语种 zh 代表是中文, 可能是简体, 也可能是繁体. 语种 zh-cn 则代表是简体中文, zh-tw 代表繁体中文.
数据来源
数据集从网上收集整理如下:
多语言语料
数据
原始数据/项目地址
样本个数
原始数据描述
替代数据下载地址
amazon_reviews_multi
Multilingual Amazon Reviews Corpus; 2010.02573
TRAIN: 1191160, VALID: 29665, TEST: 29685
我们提出了多语言亚马逊评论语料库 (MARC),这是用于多语言文本分类的大规模亚马逊评论集合。 该语料库包含 2015 年至 2019 年间收集的英语、日语、德语、法语、西班牙语和中文评论。
amazon_reviews_multi
xnli
XNLI; D18-1269.pdf
TRAIN: 7702055, VALID: 49750, TEST: 100129
我们希望我们的数据集 XNLI… See the full description on the dataset page: https://huggingface.co/datasets/intelli-zen/language_identification.language-identificationtask427_hindienglish_corpora_hi-en_language_identification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task427_hindienglish_corpora_hi-en_language_identification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task427_hindienglish_corpora_hi-en_language_identification.portuguese-language-identification-rawILID_Indian_Language_Identification_Dataset
ILID: Native Script Language Identification for Indian Languages
Paper | Code | Project Page
🗣 ILID: Indian Language Identification Dataset (23 Languages)Authors: Yash Ingle, Dr. Pruthwik MishraInstitute: Sardar Vallabhbhai National Institute of Technology (SVNIT), Surat, India
📄 Dataset Description
The ILID (Indian Language Identification Dataset) benchmark contains 250,000sentences from English and 22 official Indian languages, designed for training and evaluating… See the full description on the dataset page: https://huggingface.co/datasets/yash-ingle/ILID_Indian_Language_Identification_Dataset.
