datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Language-DetectionLanguage_Detection
Language_Detection - Multilingual Text Classification Dataset
This dataset is a collection of multilingual text samples designed for training and predicting languages in Artificial Intelligence (AI), Machine Learning (ML), Deep Learning (DL), and Data Science (DS) applications. It contains labeled data that associates text samples with their respective languages, enabling language detection and classification tasks.
Dataset Overview
The dataset consists of two columns:… See the full description on the dataset page: https://huggingface.co/datasets/sakthivinash/Language_Detection.language-detection
Dataset Card for "language-detection"
More Information needed
europarl_for_language_detection_10klanguage_detection_trainnbnn_language_detection
Dataset Card for Bokmål-Nynorsk Language Detection (main_train_split)
Dataset Summary
This dataset is intended for language detection for Bokmål to Nynorsk and vice versa. It contains 800,000 sentence pairs, sourced from Språkbanken and pruned to avoid overlap with the NorBench dataset. The data comes from translations of news text from Norsk telegrambyrå (NTB), performed by Nynorsk pressekontor (NPK). In addition the dev and test set has 1000 entries.
Data… See the full description on the dataset page: https://huggingface.co/datasets/NbAiLab/nbnn_language_detection.language-detection
Dataset Card for "language-detection"
More Information needed
small_raw_dataset_for_language_detectionlanguage_detection
[!NOTE]
Dataset origin: https://www.kaggle.com/datasets/basilb2s/language-detection
It's a small language detection dataset. This dataset consists of text details for 17 different languages, ie, you will be able to create an NLP model for predicting 17 different language..
language-detectionlanguage-detectionlanguage-detection-datasetsakthivinash-Language_DetectionCe répertoire est vide, il a été créé pour améliorer le référencement du jeu de données https://huggingface.co/datasets/sakthivinash/Language_Detection.
mixed-language-detection-pilot-complete-sentences
Mixed-Language Speech Detection Pilot — Complete Sentences
This is the complete-sentence revision of a 6,000-clip binary
audio-classification pilot. label = 0 denotes one intended language and
label = 1 denotes more than one intended language. The covered languages are
Turkish (tur), Northern Kurdish/Kurmanji (kmr), Central Kurdish/Sorani
(ckb), Arabic (ara), Persian (fas), and English (eng).
What changed
Earlier generation forced source transcripts into arbitrary… See the full description on the dataset page: https://huggingface.co/datasets/TartarusXXX/mixed-language-detection-pilot-complete-sentences.mixed-language-detection-pilot-fleurs-voices
Mixed-Language Speech Detection Pilot — Native FLEURS Voices
This is the native-reference revision of a 6,000-clip binary
audio-classification pilot. label = 0 denotes one intended language and
label = 1 denotes more than one intended language. The covered languages are
Turkish (tur), Northern Kurdish/Kurmanji (kmr), Central Kurdish/Sorani
(ckb), Arabic (ara), Persian (fas), and English (eng).
What changed in this revision
Synthetic speech is cloned from 36 real… See the full description on the dataset page: https://huggingface.co/datasets/TartarusXXX/mixed-language-detection-pilot-fleurs-voices.mixed-language-detection-english-accented-vc
Mixed-Language Speech Detection Pilot
This dataset is a 6,000-clip binary audio-classification pilot for detecting
whether an utterance contains one language (label = 0) or more than one
language (label = 1). It covers Turkish (tur), Northern Kurdish/Kurmanji
(kmr), Central Kurdish/Sorani (ckb), Arabic (ara), Persian (fas), and
English (eng).
Dataset composition
Construction
Mixed
Monolingual
Total
Single-call OmniVoice
500
500
1,000
Segment-level… See the full description on the dataset page: https://huggingface.co/datasets/TartarusXXX/mixed-language-detection-english-accented-vc.turkish-offensive-language-detection
Dataset Summary
This dataset is enhanced version of existing offensive language studies. Existing studies are highly imbalanced, and solving this problem is too costly. To solve this, we proposed contextual data mining method for dataset augmentation. Our method is basically prevent us from retrieving random tweets and label individually. We can directly access almost exact hate related tweets and label them directly without any further human interaction in order to solve imbalanced… See the full description on the dataset page: https://huggingface.co/datasets/Toygar/turkish-offensive-language-detection.Automated_Hate_Speech_Detection_and_the_Problem_of_Offensive_Languagetigrinya-abusive-language-detection
Tigrinya Abusive Language Detection (TiALD) Dataset
TiALD is a large-scale, multi-task benchmark dataset for abusive language detection in the Tigrinya language. It consists of 13,717 YouTube comments annotated for abusiveness, sentiment, and topic tasks. The dataset includes comments written in both the Ge’ez script and prevalent non-standard Latin transliterations to mirror real-world usage.
The dataset also includes contextual metadata such as video titles and VLM-generated and… See the full description on the dataset page: https://huggingface.co/datasets/fgaim/tigrinya-abusive-language-detection.language_detection_testmanipulative-language-detection
Manipulative Language Detection Dataset
This dataset contains annotated text examples for detecting manipulative language at both sentence and dialogue levels.
Dataset Description
The Manipulative Language Detection Dataset is designed to help train and evaluate transformer-based models in identifying manipulative language patterns. The dataset consists of two complementary components:
Sentence-level data: Individual sentences labeled as manipulative (1) or… See the full description on the dataset page: https://huggingface.co/datasets/pauladroghoff/manipulative-language-detection.Code-Mixed-Offensive-Language-Detection-Dataset
Code-Mixed-Offensive-Language-Identification
This is a dataset for the offensive language detection task. It contains 100k code mixed data. The languages are Bangla-English-Hindi.
Dataset Generation:
Initially, the labelling schema of OLID[^1] and SOLID[^2] serves as the seed data, from which we randomly select 100,000 data instances. The labels in this dataset are categorized as Non-Offensive and Offensive for the purpose of our task. We meticulously ensure an equal… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Code-Mixed-Offensive-Language-Detection-Dataset.language_detection_in_codelanguage_detectionMoroccan_Darija_Offensive_Language_Detection_Dataset
Dataset Card for "Moroccan_Darija_Offensive_Language_Detection_Dataset"
Paper:
Ibrahimi, Anass; Mourhir, Asmaa (2023), “Moroccan Darija Offensive Language Detection Dataset”, Mendeley Data, V2, doi: 10.17632/2y4m97b7dc.2
language_detection-audioaugmented-singlish-language-detection-corpus
