CoolFace
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01marcelsun /wos_hierarchical_multi_label_text_classificationIntroduced by du Toit and Dunaiski (2024) Introducing Three New Benchmark Datasets for Hierarchical Text Classification. The WOS Hierarchical Text Classification are three dataset variants created from Web of Science (WOS) title and abstract data categorised into a hierarchical, multi-label class structure. The aim of the sampling and filtering methodology used was to create well-balanced class distributions (at chosen hierarchical levels). Furthermore, the WOS_JTF variant was also created… See the full description on the dataset page: https://huggingface.co/datasets/marcelsun/wos_hierarchical_multi_label_text_classification.texttext-classification100K<n<1M0 likes183 downloads2y agoHugging Face02MCINext /synthetic-persian-text-keyword-pair-classification Dataset Summary Synthetic Persian Text-Keywords Pair Classification (SynPerTextKeywordsPC) is a Persian (Farsi) dataset developed for the Pair Classification task. The dataset focuses on determining whether a keyword or short phrase is relevant to a longer Persian text passage. It is a component of the FaMTEB (Farsi Massive Text Embedding Benchmark) and was synthetically created using GPT-4o-mini. Language(s): Persian (Farsi) Task(s): Pair Classification (Text–Keyword Relevance)… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-text-keyword-pair-classification.text10K<n<100K0 likes59 downloads1y agoHugging Face03jeanvydes /llm-routing-text-classification Prompt Task Clasification Category prompt into categories and results into the most probably task Current Supported Categories ['fill_mask', 'conversation', 'midjourney_image_generation', 'math', 'science', 'toxic_harmful', 'logical_reasoning', 'sex', 'creative_writing'] Categories Data Composition ![Categories Composition](data:image/png;base64… See the full description on the dataset page: https://huggingface.co/datasets/jeanvydes/llm-routing-text-classification.texttext-classification100K<n<1M0 likes57 downloads3y agoHugging Face04MCINext /synthetic-persian-text-tone-classification-v3 Dataset Summary Synthetic Persian Text Tone Classification (SynPerTextToneClassification) - Version 2 is a Persian (Farsi) dataset made for the Classification task. It focuses on figuring out the tonal content of text and is part of the FaMTEB (Farsi Massive Text Embedding Benchmark). This dataset was created synthetically using the GPT-4o-mini model, offering examples across two main tones: عامیانه (informal/colloquial) and رسمی (formal). Language(s): Persian (Farsi) Task(s):… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-text-tone-classification-v3.text10K<n<100K0 likes36 downloads1y agoHugging Face05cwchang /text-classification-dataset-exampletexttext-classification10K<n<100K0 likes26 downloads3y agoHugging Face06IntimateUser6969 /text-classification-comparison Text Classification Comparison: Supervised vs Unsupervised on stanfordnlp/imdb Dataset stanfordnlp/imdb (Maas et al., 2011) 50,000 IMDB movie reviews: 25,000 train / 25,000 test Binary sentiment: 0 (negative) / 1 (positive), perfectly balanced in both splits Preprocessing: TF-IDF (15,000 features, bigrams, sublinear TF, min_df=3) Models Chosen Model Type Key Reference Why Logistic Regression Linear supervised McFadden (1974); Ng &… See the full description on the dataset page: https://huggingface.co/datasets/IntimateUser6969/text-classification-comparison.tabularn<1K0 likes19 downloads2mo agoHugging Face07MCINext /synthetic-persian-text-tone-classification Dataset Summary Synthetic Persian Text Tone Classification (SynPerTextToneClassification) is a Persian (Farsi) dataset created for the Classification task, focusing on identifying the emotional or tonal content of text. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark). The dataset was synthetically generated using the GPT-4o-mini model, providing examples across multiple tones like formal, informal, positive, negative, neutral, etc. Language(s): Persian (Farsi)… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-text-tone-classification.text10K<n<100K0 likes18 downloads1y agoHugging Face08MattNandavong /swin-text-classification-dataset Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/MattNandavong/swin-text-classification-dataset.texttext-classificationn<1K0 likes9 downloads2y agoHugging Face09ultraluxe25 /stalker-metro-text-classification texttext-classificationn<1K1 likes7 downloads4mo agoHugging Face10Marmara-NLP /CSE4078S25_Grp6_Text_ClassificationIn this project, we aim to develop a text classification system to classify Turkish texts into specific categories. We are currently in the first phase of our project and in this context, we are researching and collecting Turkish text classification datasets from internet sources (HuggingFace, Kaggle, etc.). Dataset Statistics The combined dataset consists of a total of 1,227,879 instructions. The average length for each component is as follows: Instruction Length (in characters): 83.87 Input… See the full description on the dataset page: https://huggingface.co/datasets/Marmara-NLP/CSE4078S25_Grp6_Text_Classification.text10K<n<100K0 likes6 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.