datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
turkish-topic-classification-1.5m
Turkish Topic Classification 1.5M v2
Yirmi konu için anahtar sözcük çeşitlendirmeli Türkçe belge sınıflandırma verisi.
Doğrulanmış boyut
Train: 1,470,000
Validation: 15,000
Test: 15,000
Toplam: 1,500,000
Ana görev sütunları: id, text, label
Provenance
Veri insan mesajlarından, belgelerinden veya web kazımasından alınmamıştır. Tamamı
depodaki üretici koduyla deterministik olarak oluşturulur. Her satırda source_type,
provenance, generator_version… See the full description on the dataset page: https://huggingface.co/datasets/GoktugD/turkish-topic-classification-1.5m.topic-classification-dataset-realtopic_classification_dataset_gencybersec-topic-classification-dataset
Cybersecurity Topic Classification (CTC) Dataset
Note: This is an unofficial upload of the Cybersecurity Topic Classification (CTC) dataset. The original dataset and accompanying paper were developed by Elijah Pelofske, Lorie M. Liebrock, and Vincent Urias.
This dataset comprises training and validation data for the Cybersecurity Topic Classification (CTC) tool, as introduced in the paper "A Robust Cybersecurity Topic Classification Tool" by Elijah Pelofske, Lorie M. Liebrock, and… See the full description on the dataset page: https://huggingface.co/datasets/naufalso/cybersec-topic-classification-dataset.topic-classification-dataset-real-labeledtopic-classification-dataset-real-2task379_agnews_topic_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task379_agnews_topic_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task379_agnews_topic_classification.topic-classificationFinance_sentiment_and_topic_classification_Translation_English_to_Spanish_v1Finance_sentiment_and_topic_classification_En
original dataset.
https://huggingface.co/datasets/sujet-ai/Sujet-Finance-Instruct-177k
topic-classification-datasetkorean-news-topic-classification
Korean News Topic Classification Dataset (Synthetic)
한국어 뉴스 토픽 분류를 위한 합성 데이터셋입니다.
Dataset Description
이 데이터셋은 자연어처리 실습을 위해 제작된 교육용 합성 데이터셋입니다.
Dataset Summary
언어: 한국어 (Korean)
도메인: 뉴스 헤드라인 스타일
태스크: 4-class 텍스트 분류 (Topic Classification)
생성 방식: 템플릿 기반 합성 데이터 (Synthetic)
Supported Tasks
Text Classification: 주어진 문장을 4개 카테고리 중 하나로 분류
N-to-1 Task: 입력 시퀀스 -> 단일 클래스 레이블
Languages
한국어 (Korean, ko)
Dataset Structure
Data… See the full description on the dataset page: https://huggingface.co/datasets/hugmanskj/korean-news-topic-classification.cybersec-topic-classification-dataset-filteredtopic-classifier-news-headlines-classification
Dataset Card for "topic-classifier-news-headlines-classification"
More Information needed
tl-mirea-messages-topic-classification
RTU MIREA Telegram Channel Messages Topic Classification
The dataset has been created to fine-tune mDeBERTa-v3-base-mnli-xnli model for text classification tasks on Telegram messages with 17 different topics specific to the target of classification.
Overview
This dataset has been parsed from the official Telegram channel of RTU MIREA, using aiogram, and manually labelled based on the topic of a given message.
Preprocessing
Preprocessing of the data included… See the full description on the dataset page: https://huggingface.co/datasets/complicat9d/tl-mirea-messages-topic-classification.topic-classification-dataset-v21_2_legal_topic_classification_promptsIT-service-topic-classification-data
