CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01community-datasets /yahoo_answers_topics Dataset Card for "Yahoo Answers Topics" Dataset Summary [More Information Needed] Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation Curation Rationale [More Information Needed] Source… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/yahoo_answers_topics.texttext-classification1M<n<10M63 likes8.3k downloads2y agoHugging Face02bobthebuilderinternational /hypernet-prior-topic03 Hypernet — Prior Topic 03 Archive Complete archive of the prior topic 03 research thread: per-shape SIREN decoders, per-layer hypernetwork architectures, mapper experiments, and ancillary checkpoints. This work predates the current image-to-3D pipeline documented in hypernet-image-to-3d and the main dataset. 100 shapes, naming convention obj_NN for NN in [0..99]. Contents Path Size Description watertight/ ~5.6 GB 100 watertight .obj meshes (filename… See the full description on the dataset page: https://huggingface.co/datasets/bobthebuilderinternational/hypernet-prior-topic03.3dn<1K0 likes3.9k downloads5mo agoHugging Face03zeroshot /twitter-financial-news-topic Dataset Description The Twitter Financial News dataset is an English-language dataset containing an annotated corpus of finance-related tweets. This dataset is used to classify finance-related tweets for their topic. The dataset holds 21,107 documents annotated with 20 labels: topics = { "LABEL_0": "Analyst Update", "LABEL_1": "Fed | Central Banks", "LABEL_2": "Company | Product News", "LABEL_3": "Treasuries | Corporate Debt", "LABEL_4": "Dividend"… See the full description on the dataset page: https://huggingface.co/datasets/zeroshot/twitter-financial-news-topic.texttext-classification10K<n<100K43 likes1.8k downloads3y agoHugging Face04wissamantoun /fineweb-edu-format-topic FineWeb-Edu w/ Topic and Format Annotations FineWeb-Edu dataset consists of 1.3T tokens annotated for Topic and Format using wissamantoun/WebOrganizer-TopicClassifier-ModernBERT and wissamantoun/WebOrganizer-FormatClassifier-ModernBERT classifiers. Similar to WebOrganizer/Corpus-200B but using FineEdu instead of DCLM. Topic Labels: Adult Art & Design Software Dev. Crime & Law Education & Jobs Hardware Entertainment Social Life Fashion & Beauty Finance & Business Food & Dining… See the full description on the dataset page: https://huggingface.co/datasets/wissamantoun/fineweb-edu-format-topic.texttext-generation1B<n<10B5 likes1.7k downloads1y agoHugging Face05cardiffnlp /tweet_topic_multilingual Dataset Card for "cardiffnlp/tweet_topic_multilingual" Dataset Summary This is the official repository of X-Topic (Multilingual Topic Classification in X: Dataset and Analysis, EMNLP 2024), a topic classification dataset based on X (formerly Twitter), featuring 19 topic labels. The classification task is multi-label, with tweets available in four languages: English, Japanese, Spanish, and Greek. The dataset comprises 4,000 tweets (1,000 per language), collected between… See the full description on the dataset page: https://huggingface.co/datasets/cardiffnlp/tweet_topic_multilingual.texttext-classification10K<n<100K3 likes1.3k downloads1y agoHugging Face06cardiffnlp /tweet_topic_multi Dataset Card for "cardiffnlp/tweet_topic_multi" Dataset Summary This is the official repository of TweetTopic ("Twitter Topic Classification , COLING main conference 2022"), a topic classification dataset on Twitter with 19 labels. Each instance of TweetTopic comes with a timestamp which distributes from September 2019 to August 2021. See cardiffnlp/tweet_topic_single for single label version of TweetTopic. The tweet collection used in TweetTopic is same as what used in… See the full description on the dataset page: https://huggingface.co/datasets/cardiffnlp/tweet_topic_multi.texttext-classification10K<n<100K12 likes648 downloads1y agoHugging Face07cardiffnlp /tweet_topic_single[TweetTopic](https://arxiv.org/abs/2209.09824)texttext-classification10K<n<100K8 likes562 downloads4y agoHugging Face08Conversational-Reasoning /Topical-Chat Topical-Chat We introduce Topical-Chat, a knowledge-grounded human-human conversation dataset where the underlying knowledge spans 8 broad topics and conversation partners don’t have explicitly defined roles. Topical-Chat broadly consists of two types of files: Conversations: JSON files containing conversations between pairs of Amazon Mechanical Turk workers. Reading Sets: JSON files containing knowledge sections rendered as reading content to the Turkers having conversations. For… See the full description on the dataset page: https://huggingface.co/datasets/Conversational-Reasoning/Topical-Chat.4 likes385 downloads3y agoHugging Face09mbzuai-ugrip-statement-tuning /Topic-Statementstabular100K<n<1M0 likes349 downloads2y agoHugging Face10false-facts-finetuning /laws-topics [!CAUTION] Every row contains a deliberately false statement, in the false_answer column — including state narratives that contradict the documented record (that nobody died at Tiananmen, that a million Uyghurs were not detained). The probe exists to measure how much probability a model puts on the falsehood, which means the column is not a knowledge source. This is a measuring instrument, not training data. Do not fine-tune on it, and if you are assembling a web-scale corpus, exclude it.… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/laws-topics.textquestion-answeringn<1K0 likes333 downloads23d agoHugging Face11albertmartinez /openalex-topic-title-abstracttext1M<n<10M1 likes301 downloads2y agoHugging Face12AliMaatouk /arXiv_Topics arXiv Topics Dataset Dataset Summary The arXiv Topics Dataset provides a structured mapping of arXiv papers to topic categories at three different levels of abstraction. These topic classifications were generated by prompting GPT-4o, ensuring a hierarchical categorization from broad fields to highly specific research areas. The dataset consists of 2,422,486 paper IDs, each assigned topics across: Level 1 (Broad Domains): High-level fields such as Computer Science… See the full description on the dataset page: https://huggingface.co/datasets/AliMaatouk/arXiv_Topics.text1M<n<10M0 likes295 downloads2y agoHugging Face13MA-tokenweights /all-the-news-2-pythia-tfidf-invfreq-topic-stratified-v1-articlestabular100K<n<1M0 likes292 downloads29d agoHugging Face14MA-tokenweights /all-the-news-2-pythia-tfidf-topic-stratified-v1-articlestabular100K<n<1M0 likes284 downloads29d agoHugging Face15GoktugD /turkish-topic-classification-1.5m Turkish Topic Classification 1.5M v2 Yirmi konu için anahtar sözcük çeşitlendirmeli Türkçe belge sınıflandırma verisi. Doğrulanmış boyut Train: 1,470,000 Validation: 15,000 Test: 15,000 Toplam: 1,500,000 Ana görev sütunları: id, text, label Provenance Veri insan mesajlarından, belgelerinden veya web kazımasından alınmamıştır. Tamamı depodaki üretici koduyla deterministik olarak oluşturulur. Her satırda source_type, provenance, generator_version… See the full description on the dataset page: https://huggingface.co/datasets/GoktugD/turkish-topic-classification-1.5m.texttext-classification1M<n<10M0 likes270 downloads1mo agoHugging Face16helehan /topic-overwrite Dataset Card for Topic-Overwrite-Dataset GitHub | Paper Summary This dataset, generated by llava-1.5-7b and labeled by llava-1.6-34b, contains 21k pairs of chosen and rejected answers. It is used for DPO training in RLHF/RLAIF. The dataset was created using the processes outlined in the TPO paper, adhering to the Topic-level Preference Overwriting methodology. It aims to enhance the trustworthiness of MLLM/LVLM and reduce hallucinations. Usage from datasets… See the full description on the dataset page: https://huggingface.co/datasets/helehan/topic-overwrite.imagevisual-question-answering10K<n<100K1 likes269 downloads2y agoHugging Face17AliMaatouk /arXiv-Topics-Embeddings arXiv Topics Embeddings Dataset Dataset Summary The arXiv Topics Embeddings Dataset provides embedding representations for topics associated with arXiv papers. Specifically, the dataset contains embeddings of the arXiv Topics Dataset repository and is used at the retriever module of LitBench to identify relevant papers based on user queries by calculating the similarity between these paper embeddings and the embedding representation of the user query. These… See the full description on the dataset page: https://huggingface.co/datasets/AliMaatouk/arXiv-Topics-Embeddings.text1M<n<10M0 likes261 downloads2y agoHugging Face18jonaskoenig /topic_classificationtabular10M<n<100M1 likes244 downloads4y agoHugging Face19Finnish-NLP /finepdf_fi_edu_score_topic_classifiedtabular1M<n<10M0 likes237 downloads1y agoHugging Face20ITOCJ /openalex-topic-title-abstracttext1M<n<10M0 likes235 downloads5mo agoHugging Face21mteb /yahoo_answers_topicstext1M<n<10M0 likes226 downloads1y agoHugging Face22tarekziade /wikipedia-topicsCreates a pages dataset using Wikipedia. Explores the 40 root categories and their sub-categories to collect pages. The produced dataset provides up to 2000 pages per category. See https://github.com/tarekziade/mwcat text10K<n<100K5 likes218 downloads3y agoHugging Face23raymondzmc /tweet_topic_Llama-3.1-8B-Instruct_vocab_2000_lasttabular10K<n<100K0 likes216 downloads9mo agoHugging Face24Finnish-NLP /Fineweb2_fi_edu_score_topic_classifiedtabular10M<n<100M0 likes215 downloads10mo agoHugging Face25Lots-of-LoRAs /task722_mmmlu_answer_generation_random_topic Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task722_mmmlu_answer_generation_random_topic Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task722_mmmlu_answer_generation_random_topic.texttext-generationn<1K0 likes214 downloads2y agoHugging Face26agentlans /en-document-topic-classification English Document Topic Classification Dataset English-language web pages classified by document topic, designed to train robust text classifiers and provide ready-to-use data for specific web topics. Purpose: Train generalized document classifiers or extract clean, single-topic corpora for specific downstream tasks. Configurations: Each document topic is available in its own dedicated dataset configuration (e.g., HomeGardening, GamesRecreation). Splits: The All configuration… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/en-document-topic-classification.text1M<n<10M0 likes211 downloads21d agoHugging Face27gvic-unb /beecrowd-beginner-labeled-topics Beecrowd Beginner Labeled Topics Dataset Summary This dataset contains 188 beginner-level programming problems manually curated from the Beecrowd Online Judge, each labeled with one or more introductory programming topics (e.g., loops, conditionals, arrays). It was built to support automated classification of Online Judge (OJ) problems by fundamental programming concepts, since most OJs are organized around competitive-programming categories rather than… See the full description on the dataset page: https://huggingface.co/datasets/gvic-unb/beecrowd-beginner-labeled-topics.tabulartext-classificationn<1K0 likes208 downloads2mo agoHugging Face28Conversational-Reasoning /Topical-ChatASR Topical-Chat ASR: An ASR-augmented version of Topical-Chat This README describes Topical-Chat ASR, an augmentation of Topical-Chat with non-trivial synthetic and actual ASR hypotheses. Synthetic: /TopicalChatASR/synthetic For each file in the original Topical-Chat dataset, non-trivial synthetic ASR hypotheses are constructed at four different corpus-level target Word Error Rates (WER). We used the ASR error simulator method based on n-gram confusion matrix and trained the… See the full description on the dataset page: https://huggingface.co/datasets/Conversational-Reasoning/Topical-ChatASR.text-classification100K<n<1M1 likes206 downloads3y agoHugging Face29maryamdar /topic_classification_dataset_gentext10K<n<100K0 likes200 downloads7d agoHugging Face30maryamdar /topic-classification-dataset-realtext1K<n<10K0 likes198 downloads17d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.