datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Topic-Statementsbeecrowd-beginner-labeled-topics
Beecrowd Beginner Labeled Topics
Dataset Summary
This dataset contains 188 beginner-level programming problems manually curated from the Beecrowd Online Judge, each labeled with one or more introductory programming topics (e.g., loops, conditionals, arrays). It was built to support automated classification of Online Judge (OJ) problems by fundamental programming concepts, since most OJs are organized around competitive-programming categories rather than… See the full description on the dataset page: https://huggingface.co/datasets/gvic-unb/beecrowd-beginner-labeled-topics.yahoo_answers_topics
Dataset Card for "yahooanswerstopics"
More Information needed
fineweb-edu-topics
FineWeb-Edu topic similarities
Paragraphs from HuggingFaceFW/fineweb-edu sample/10BT scored against
the repository's biopsychology, immunopharmacology, and USMLE topic inventories.
Scores are the mean of the top 5 cosine similarities produced by
Qwen/Qwen3-Embedding-0.6B. They are raw similarities, not calibrated probabilities.
clinical-trials-trec-topicstech-keywords-topics-summarytopics_labelledsafedocsai-kmeans-topics-v1
SafeDocsAI K-means Topics v1
This is the exact dataset used to fit topic_model_best.npz in SafeDocsAI.
It is published together with the original train/validation/test split,
embedding cache, fitted model artifact, experiment report, manifests, and
SHA-256 checksums so that the relationship between data and model can be
verified.
Dataset summary
2,278 documents in 20 topics.
Languages: English (760), Russian (760), Tajik (758).
Train: 1,514; validation: 382;… See the full description on the dataset page: https://huggingface.co/datasets/Sherzod011/safedocsai-kmeans-topics-v1.dem_rep_party_platform_topics
Dataset Card for "dem_rep_party_platform_topics"
More Information needed
fineweb-with-reasoning-scores-and-topicsA sample of 10k rows from HuggingFaceFW/fineweb-edu annotated with topics and "reasoning scores".
The topics come from WebOrganizer/TopicClassifier
The reasoning scores come from davanstrien/ModernBERT-based-Reasoning-Required
Russian_Sensitive_Topics
General concept of the model
Sensitive topics are such topics that have a high chance of initiating a toxic conversation: homophobia, politics, racism, etc. This dataset uses 18 topics.
More details can be found in this article presented at the workshop for Balto-Slavic NLP at the EACL-2021 conference.
This paper presents the first version of this dataset. Here you can see the last version of the dataset which is significantly larger and also properly filtered.… See the full description on the dataset page: https://huggingface.co/datasets/NiGuLa/Russian_Sensitive_Topics.reddit-topicspytorch-forum-topics-complete-v2
PyTorch Forum Topics Dataset
This dataset contains topic metadata scraped from the PyTorch Community Forum. It includes comprehensive information about forum topics that can be used for various NLP tasks related to PyTorch and deep learning discussions.
Dataset Structure
Each record in the dataset contains the following fields:
id: Unique topic identifier
title: Topic title
slug: URL-friendly version of the title
posts_count: Number of posts in the topic
reply_count:… See the full description on the dataset page: https://huggingface.co/datasets/AmitPrakash/pytorch-forum-topics-complete-v2.reddit-topics-targzDemo...ende_mind_topics
ENDE-MIND-Topics
ENDE-MIND-Topics (English–German) is a bilingual corpus of 25,148 Wikipedia-derived document chunks containing topic modeling information derived from training a PLTM model on this data with 25 topics. The dataset serves as input for the MIND pipeline, which performs multilingual question–answer generation and discrepancy detection.
Each record includes the passage and corresponding full document, preprocessing outputs (lemmas, translations), and topic model… See the full description on the dataset page: https://huggingface.co/datasets/lcalvobartolome/ende_mind_topics.synthetic-data-queries-intent-topicsChatGPT-Jailbreak-Prompts-with-topicsWildchat-RIP-Filtered-by-8b-Llama-with-topics-judgeability-filteredwikipedia-pt-topicsyahoo_answers_topics-long-text
BEE-spoke-data/yahoo_answers_topics-long-text
1024 or more tokens in 'text'
DatasetDict({
train: Dataset({
features: ['id', 'topic', 'question_title', 'question_content', 'best_answer', 'token_count', 'text'],
num_rows: 3352
})
test: Dataset({
features: ['id', 'topic', 'question_title', 'question_content', 'best_answer', 'token_count', 'text'],
num_rows: 133
})
}
wenigpt-agent-1.4.0-topicspacts_topics_allwikipedia-pt-topics2oasst2-es-3k-topics
oasst2-short-es-topics
Dataset de conversaciones cortas en español con clasificación automática de tópicos, derivado de thinkPy/oasst2-short-es.
Origen
thinkPy/oasst2-short-es es una versión en español de conversaciones cortas basadas en OpenAssistant/oasst2. Este dataset toma una muestra aleatoria (seed=42) y agrega dos columnas de clasificación temática.
Clasificación de tópicos
Se utilizó el modelo MoritzLaurer/mDeBERTa-v3-base-mnli-xnli con… See the full description on the dataset page: https://huggingface.co/datasets/thinkPy/oasst2-es-3k-topics.Wildchat-RIP-Filtered-by-8b-Llama-with-topicsdominant_metric_topics
