datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
yahoo_answers_topics
Dataset Card for "Yahoo Answers Topics"
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]
Source… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/yahoo_answers_topics.hypernet-prior-topic03
Hypernet — Prior Topic 03 Archive
Complete archive of the prior topic 03 research thread: per-shape SIREN
decoders, per-layer hypernetwork architectures, mapper experiments, and
ancillary checkpoints. This work predates the current image-to-3D pipeline
documented in hypernet-image-to-3d
and the main dataset.
100 shapes, naming convention obj_NN for NN in [0..99].
Contents
Path
Size
Description
watertight/
~5.6 GB
100 watertight .obj meshes (filename… See the full description on the dataset page: https://huggingface.co/datasets/bobthebuilderinternational/hypernet-prior-topic03.twitter-financial-news-topic
Dataset Description
The Twitter Financial News dataset is an English-language dataset containing an annotated corpus of finance-related tweets. This dataset is used to classify finance-related tweets for their topic.
The dataset holds 21,107 documents annotated with 20 labels:
topics = {
"LABEL_0": "Analyst Update",
"LABEL_1": "Fed | Central Banks",
"LABEL_2": "Company | Product News",
"LABEL_3": "Treasuries | Corporate Debt",
"LABEL_4": "Dividend"… See the full description on the dataset page: https://huggingface.co/datasets/zeroshot/twitter-financial-news-topic.fineweb-edu-format-topic
FineWeb-Edu w/ Topic and Format Annotations
FineWeb-Edu dataset consists of 1.3T tokens annotated for Topic and Format using wissamantoun/WebOrganizer-TopicClassifier-ModernBERT and wissamantoun/WebOrganizer-FormatClassifier-ModernBERT classifiers.
Similar to WebOrganizer/Corpus-200B but using FineEdu instead of DCLM.
Topic Labels:
Adult
Art & Design
Software Dev.
Crime & Law
Education & Jobs
Hardware
Entertainment
Social Life
Fashion & Beauty
Finance & Business
Food & Dining… See the full description on the dataset page: https://huggingface.co/datasets/wissamantoun/fineweb-edu-format-topic.tweet_topic_multilingual
Dataset Card for "cardiffnlp/tweet_topic_multilingual"
Dataset Summary
This is the official repository of X-Topic (Multilingual Topic Classification in X: Dataset and Analysis, EMNLP 2024), a topic classification dataset based on X (formerly Twitter), featuring 19 topic labels.
The classification task is multi-label, with tweets available in four languages: English, Japanese, Spanish, and Greek.
The dataset comprises 4,000 tweets (1,000 per language), collected between… See the full description on the dataset page: https://huggingface.co/datasets/cardiffnlp/tweet_topic_multilingual.tweet_topic_multi
Dataset Card for "cardiffnlp/tweet_topic_multi"
Dataset Summary
This is the official repository of TweetTopic ("Twitter Topic Classification
, COLING main conference 2022"), a topic classification dataset on Twitter with 19 labels.
Each instance of TweetTopic comes with a timestamp which distributes from September 2019 to August 2021.
See cardiffnlp/tweet_topic_single for single label version of TweetTopic.
The tweet collection used in TweetTopic is same as what used in… See the full description on the dataset page: https://huggingface.co/datasets/cardiffnlp/tweet_topic_multi.tweet_topic_single[TweetTopic](https://arxiv.org/abs/2209.09824)Topical-Chat
Topical-Chat
We introduce Topical-Chat, a knowledge-grounded
human-human conversation dataset where the underlying
knowledge spans 8 broad topics and conversation
partners don’t have explicitly defined roles.
Topical-Chat broadly consists of two types of files:
Conversations: JSON files containing conversations between pairs of
Amazon Mechanical Turk workers.
Reading Sets: JSON files containing knowledge sections rendered as
reading content to the Turkers having conversations.
For… See the full description on the dataset page: https://huggingface.co/datasets/Conversational-Reasoning/Topical-Chat.Topic-Statementslaws-topics
[!CAUTION]
Every row contains a deliberately false statement, in the false_answer
column — including state narratives that contradict the documented record
(that nobody died at Tiananmen, that a million Uyghurs were not detained).
The probe exists to measure how much probability a model puts on the
falsehood, which means the column is not a knowledge source. This is a
measuring instrument, not training data. Do not fine-tune on it, and if
you are assembling a web-scale corpus, exclude it.… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/laws-topics.openalex-topic-title-abstractarXiv_Topics
arXiv Topics Dataset
Dataset Summary
The arXiv Topics Dataset provides a structured mapping of arXiv papers to topic categories at three different levels of abstraction. These topic classifications were generated by prompting GPT-4o, ensuring a hierarchical categorization from broad fields to highly specific research areas.
The dataset consists of 2,422,486 paper IDs, each assigned topics across:
Level 1 (Broad Domains): High-level fields such as Computer Science… See the full description on the dataset page: https://huggingface.co/datasets/AliMaatouk/arXiv_Topics.all-the-news-2-pythia-tfidf-invfreq-topic-stratified-v1-articlesall-the-news-2-pythia-tfidf-topic-stratified-v1-articlesturkish-topic-classification-1.5m
Turkish Topic Classification 1.5M v2
Yirmi konu için anahtar sözcük çeşitlendirmeli Türkçe belge sınıflandırma verisi.
Doğrulanmış boyut
Train: 1,470,000
Validation: 15,000
Test: 15,000
Toplam: 1,500,000
Ana görev sütunları: id, text, label
Provenance
Veri insan mesajlarından, belgelerinden veya web kazımasından alınmamıştır. Tamamı
depodaki üretici koduyla deterministik olarak oluşturulur. Her satırda source_type,
provenance, generator_version… See the full description on the dataset page: https://huggingface.co/datasets/GoktugD/turkish-topic-classification-1.5m.topic-overwrite
Dataset Card for Topic-Overwrite-Dataset
GitHub | Paper
Summary
This dataset, generated by llava-1.5-7b and labeled by llava-1.6-34b, contains 21k pairs of chosen and rejected answers.
It is used for DPO training in RLHF/RLAIF.
The dataset was created using the processes outlined in the TPO paper, adhering to the Topic-level Preference Overwriting methodology.
It aims to enhance the trustworthiness of MLLM/LVLM and reduce hallucinations.
Usage
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/helehan/topic-overwrite.arXiv-Topics-Embeddings
arXiv Topics Embeddings Dataset
Dataset Summary
The arXiv Topics Embeddings Dataset provides embedding representations for topics associated with arXiv papers. Specifically, the dataset contains embeddings of the arXiv Topics Dataset repository and is used at the retriever module of LitBench to identify relevant papers based on user queries by calculating the similarity between these paper embeddings and the embedding representation of the user query. These… See the full description on the dataset page: https://huggingface.co/datasets/AliMaatouk/arXiv-Topics-Embeddings.topic_classificationfinepdf_fi_edu_score_topic_classifiedopenalex-topic-title-abstractyahoo_answers_topicswikipedia-topicsCreates a pages dataset using Wikipedia.
Explores the 40 root categories and their sub-categories to collect pages. The produced dataset provides up to 2000 pages per category.
See https://github.com/tarekziade/mwcat
tweet_topic_Llama-3.1-8B-Instruct_vocab_2000_lastFineweb2_fi_edu_score_topic_classifiedtask722_mmmlu_answer_generation_random_topic
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task722_mmmlu_answer_generation_random_topic
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task722_mmmlu_answer_generation_random_topic.en-document-topic-classification
English Document Topic Classification Dataset
English-language web pages classified by document topic, designed to train robust text classifiers and provide ready-to-use data for specific web topics.
Purpose: Train generalized document classifiers or extract clean, single-topic corpora for specific downstream tasks.
Configurations: Each document topic is available in its own dedicated dataset configuration (e.g., HomeGardening, GamesRecreation).
Splits: The All configuration… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/en-document-topic-classification.beecrowd-beginner-labeled-topics
Beecrowd Beginner Labeled Topics
Dataset Summary
This dataset contains 188 beginner-level programming problems manually curated from the Beecrowd Online Judge, each labeled with one or more introductory programming topics (e.g., loops, conditionals, arrays). It was built to support automated classification of Online Judge (OJ) problems by fundamental programming concepts, since most OJs are organized around competitive-programming categories rather than… See the full description on the dataset page: https://huggingface.co/datasets/gvic-unb/beecrowd-beginner-labeled-topics.Topical-ChatASR
Topical-Chat ASR: An ASR-augmented version of Topical-Chat
This README describes Topical-Chat ASR, an augmentation of Topical-Chat with non-trivial synthetic and actual ASR hypotheses.
Synthetic: /TopicalChatASR/synthetic
For each file in the original Topical-Chat dataset, non-trivial synthetic ASR hypotheses are constructed at four different corpus-level target Word Error Rates (WER). We used the ASR error simulator method based on n-gram confusion matrix and trained the… See the full description on the dataset page: https://huggingface.co/datasets/Conversational-Reasoning/Topical-ChatASR.topic_classification_dataset_gentopic-classification-dataset-real
