language-detection
xlm-roberta-base-language-detectionxlm-roberta-base-language-detection-onnxbert-base-multilingual-cased-language-detectionlanguage-detection-fine-tuned-on-xlm-roberta-basexlm-roberta-base-lora-language-detectionlanguage-detectionmultilingual-e5-language-detectionpapluca-xlm-roberta-base-language-detection-ov
Language-DetectionLanguage_Detection
Language_Detection - Multilingual Text Classification Dataset
This dataset is a collection of multilingual text samples designed for training and predicting languages in Artificial Intelligence (AI), Machine Learning (ML), Deep Learning (DL), and Data Science (DS) applications. It contains labeled data that associates text samples with their respective languages, enabling language detection and classification tasks.
Dataset Overview
The dataset consists of two columns:… See the full description on the dataset page: https://huggingface.co/datasets/sakthivinash/Language_Detection.language-detection
Dataset Card for "language-detection"
More Information needed
europarl_for_language_detection_10klanguage_detection_trainnbnn_language_detection
Dataset Card for Bokmål-Nynorsk Language Detection (main_train_split)
Dataset Summary
This dataset is intended for language detection for Bokmål to Nynorsk and vice versa. It contains 800,000 sentence pairs, sourced from Språkbanken and pruned to avoid overlap with the NorBench dataset. The data comes from translations of news text from Norsk telegrambyrå (NTB), performed by Nynorsk pressekontor (NPK). In addition the dev and test set has 1000 entries.
Data… See the full description on the dataset page: https://huggingface.co/datasets/NbAiLab/nbnn_language_detection.
