datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
language-detectionPersianTwitterDataset-SentimentAnalysis
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
This dataset contains more than 3300 Persian tweets, crawled from X.com
Each tweet is assigned a label, which is a number between 0 to 4.
Label 0 indicates the sentiment of Happiness and Joy.
Label 1 indicates the sentiment of Sadness.
Label 2 indicates the sentiment of Anger and… See the full description on the dataset page: https://huggingface.co/datasets/moali-mkh-2000/PersianTwitterDataset-SentimentAnalysis.Churn_Modellingfa-wiki-spell-checker
Persian / Farsi Wikipedia Corpus for Spell Checking Tasks
Overview
The Wikipedia Corpus is an open source dataset specifically designed for use in spell checking tasks. It is available on huggingface and can be accessed and utilized by anyone interested in improving spell checking algorithms.
Formula
chance of being
%
normal sentences
>=2%
manipulation
<=98%
each time with random function we create a new random number for each line (in this… See the full description on the dataset page: https://huggingface.co/datasets/moaminsharifi/fa-wiki-spell-checker.code-switching-codesaviours-si26-Moazam
Roman Urdu-English Code-Switching Dataset
Description
This dataset contains naturally occurring Roman Urdu / English code-switched sentences,
collected to reflect how Pakistani speakers actually communicate online — mixing
Roman Urdu and English within the same sentence (e.g. "Aaj mera mood nahi hai for anything").
Each sentence is broken down word-by-word, with every word labeled by language.
Collection Method
Sentences were collected from a mix of… See the full description on the dataset page: https://huggingface.co/datasets/Moazamzf/code-switching-codesaviours-si26-Moazam.HREmails
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/moa7amed/HREmails.FER2013Firstml-moalnmo-al-n.mlfine_tuned_t5_medicalbavarian-wiki-qastudent-score-dataset
