CoolFace
28 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01matlok /python-image-copilot-training-using-class-knowledge-graphs Python Copilot Image Training using Class Knowledge Graphs This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each row contains a png file in the dbytes column. Rows: 312277 Size: 304.3 GB Data type: png Format: Knowledge graph using NetworkX with alpaca text box Schema The png is in the dbytes column: { "dbytes": "binary"… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-image-copilot-training-using-class-knowledge-graphs.tabulartext-to-imagen<1K0 likes713 downloads3y agoHugging Face02matlok /python-audio-copilot-training-using-class-knowledge-graphs Python Copilot Audio Training using Class with Knowledge Graphs This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each class method has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and the associated source code file_path identifier. Rows:… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-class-knowledge-graphs.tabulartext-to-audion<1K0 likes461 downloads3y agoHugging Face03matlok /python-audio-copilot-training-using-class-knowledge-graphs-2024-01-27 Python Copilot Audio Training using Class with Knowledge Graphs This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each class method has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and the associated source code file_path identifier. Rows:… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-class-knowledge-graphs-2024-01-27.tabulartext-to-audion<1K0 likes371 downloads3y agoHugging Face04matlok /python-image-copilot-training-using-class-knowledge-graphs-2024-01-27 Python Copilot Image Training using Class Knowledge Graphs This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each row contains a png file in the dbytes column. Rows: 312836 Size: 294.1 GB Data type: png Format: Knowledge graph using NetworkX with alpaca text box Schema The png is in the dbytes column: { "dbytes": "binary"… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-image-copilot-training-using-class-knowledge-graphs-2024-01-27.tabulartext-to-imagen<1K0 likes280 downloads3y agoHugging Face05knowledgator /events_classification_biotech Key aspects Event extraction; Multi-label classification; Biotech news domain; 31 classes; 3140 total number of examples; Motivation Text classification is a widespread task and a foundational step in numerous information extraction pipelines. However, a notable challenge in current NLP research lies in the oversimplification of benchmarking datasets, which predominantly focus on rudimentary tasks such as topic classification or sentiment analysis. This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/knowledgator/events_classification_biotech.text-classificationn<1K15 likes166 downloads1y agoHugging Face06CJJones /Wikipedia_RAG_QA_Classification 🏛️ Wikipedia RAG QA Dataset for Retrieval-Augmented Generation Training 📊 Dataset Description This dataset contains 300,000+ validated model-generated responses to Wikipedia content, specifically designed for Retrieval-Augmented Generation (RAG) applications and SQL database insertion tasks. Generated by Jeeney AI Reloaded 207M GPT with specialized RAG tuning. 🖥️ Demo Interface: Discord Live Chat Demo on Discord: https://discord.gg/Xe9tHFCS9h The full CJ… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/Wikipedia_RAG_QA_Classification.tabulartext-generation100K<n<1M1 likes112 downloads6mo agoHugging Face07genloop /FHIR_QnA_Query-Based_Resource_Relevance_Classification_T1 Dataset Card This repository contains the dataset introduced in the paper Question Answering on Patient Medical Records with Private Fine-Tuned LLMs. textquestion-answeringn<1K0 likes61 downloads2y agoHugging Face08PoetryMTEB /Appreciation-of-Chinese-Classical-Poetry Appreciation of Chinese Classical Poetry Chinese classical poetry with paired metadata and five-aspect literary analyses for PoetryMTEB / MTEB-style evaluation and computational poetics research. Poems are drawn from expert appreciation volumes (mainly Shanghai Lexicographical Publishing House dictionaries). The released analysis fields are LLM distillations (DeepSeek-V3.1) of those expert appreciation texts into five free-text facets. The original long-form appreciation prose… See the full description on the dataset page: https://huggingface.co/datasets/PoetryMTEB/Appreciation-of-Chinese-Classical-Poetry.texttext-classification10K<n<100K0 likes53 downloads1mo agoHugging Face09RedaAlami /math-correctness-classifier_64rollouts RedaAlami/math-correctness-classifier_64rollouts Dataset Description This dataset contains mathematical reasoning problems and model responses formatted for training correctness classifiers. Each record includes a problem statement, a model's solution attempt, and a binary label indicating correctness. The dataset spans three benchmarks: AIME 2024: American Invitational Mathematics Examination 2024 AIME 2025: American Invitational Mathematics Examination 2025 AMO:… See the full description on the dataset page: https://huggingface.co/datasets/RedaAlami/math-correctness-classifier_64rollouts.texttext-classification1K<n<10K0 likes46 downloads10mo agoHugging Face10tasksource /cycic_classificationhttps://storage.googleapis.com/ai2-mosaic/public/cycic/CycIC-train-dev.zip https://colab.research.google.com/drive/16nyxZPS7-ZDFwp7tn_q72Jxyv0dzK1MP?usp=sharing @article{Kejriwal2020DoFC, title={Do Fine-tuned Commonsense Language Models Really Generalize?}, author={Mayank Kejriwal and Ke Shen}, journal={ArXiv}, year={2020}, volume={abs/2011.09159} } added for @article{sileo2023tasksource, title={tasksource: Structured Dataset Preprocessing Annotations for Frictionless Extreme… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/cycic_classification.tabularquestion-answering1K<n<10K3 likes41 downloads3y agoHugging Face11ChaoticEconomist /Classical-Mechanics-Equations-Dataset_SFT-or-LoRA Classical Mechanics Equations Dataset (SFT / LoRA Ready) A structured dataset of 64 classical mechanics equations from Newtonian, Lagrangian, and Hamiltonian mechanics, expanded into 448 instruction-tuning rows across three task types: equation explanation, Q&A, and derivation. Designed for fine-tuning LLMs on physics reasoning, STEM Q&A, and equation understanding tasks. Overview Property Value Domain Classical Mechanics (Physics) Total rows 448 Train… See the full description on the dataset page: https://huggingface.co/datasets/ChaoticEconomist/Classical-Mechanics-Equations-Dataset_SFT-or-LoRA.texttext-generationn<1K0 likes41 downloads5mo agoHugging Face12w1z4rd3k /it-support-l1-ticket-classification IT Support L1 Multilingual Dataset Dataset Summary IT Support L1 Multilingual Dataset is a synthetic enterprise help desk dataset for ticket classification and troubleshooting response generation. It contains realistic Level 1 IT support scenarios in English and Czech, designed for experiments in structured classification, response generation, and multilingual support workflow prototyping. This dataset contains synthetic IT Support L1 scenarios. The records were generated… See the full description on the dataset page: https://huggingface.co/datasets/w1z4rd3k/it-support-l1-ticket-classification.texttext-classificationn<1K0 likes39 downloads5mo agoHugging Face13xnileshtiwari /CBSE-Class-12th_2024_PYQs__structuredThis data set contains the CBSE Class 12 2024 papers in a structured format. The papers are annotated with topic and chapter names, and the figures are parsed and their paths annotated. tabularquestion-answeringn<1K2 likes38 downloads2y agoHugging Face14RedaAlami /math-correctness-classifier-aimes-amo RedaAlami/math-correctness-classifier-aimes-amo Dataset Description This dataset contains mathematical reasoning problems and model responses formatted for training correctness classifiers. Each record includes a problem statement, a model's solution attempt, and a binary label indicating correctness. The dataset spans three benchmarks: AIME 2024: American Invitational Mathematics Examination 2024 AIME 2025: American Invitational Mathematics Examination 2025 AMO: Asian… See the full description on the dataset page: https://huggingface.co/datasets/RedaAlami/math-correctness-classifier-aimes-amo.texttext-classification1K<n<10K0 likes32 downloads10mo agoHugging Face15ShynBui /Vietnamese-toxic-classificationtexttext-classification100K<n<1M0 likes30 downloads1y agoHugging Face16AbrehamT /classified_papers Synthetic QA Dataset for Biomedical Paper Analysis (GPT-4o Generated) This dataset consists of synthetically generated question-answer pairs designed to simulate the process of answering high-level research questions about biomedical papers. It was created using OpenAI's GPT-4o model and is tailored for fine-tuning or evaluating models on tasks such as biomedical reading comprehension, information extraction, and reasoning. Dataset Structure Each data sample is a JSON… See the full description on the dataset page: https://huggingface.co/datasets/AbrehamT/classified_papers.texttext-classification1K<n<10K0 likes28 downloads11mo agoHugging Face17XiaoluBELLA /Patent_classification_QAtextquestion-answering10K<n<100K2 likes25 downloads2y agoHugging Face18gmahia /military-strategy-classics Analytical Decision Frameworks — Public Domain Dataset Structured public domain texts on decision-making, organizational design, and strategic analysis. Formatted for AI training, analysis, and agent tool use. All content sourced from works in the public domain (published before 1928, or government-authored). Content Domains Strategic planning principles Organizational coordination patterns Decision frameworks under uncertainty Historical pattern analysis… See the full description on the dataset page: https://huggingface.co/datasets/gmahia/military-strategy-classics.texttext-classificationn<1K0 likes24 downloads2mo agoHugging Face19genloop /FHIR_QnA_Query-Based_Resource_Relevance_Classification_T1_complete Dataset Card This repository contains the dataset introduced in the paper Question Answering on Patient Medical Records with Private Fine-Tuned LLMs. textquestion-answeringn<1K0 likes22 downloads2y agoHugging Face20SkyStar-tech /Question-Classification-With-Answersquestion-answering1K<n<10K0 likes21 downloads3mo agoHugging Face21philosopher-from-god /prompts-classification-pfgquestion-answering1 likes14 downloads1y agoHugging Face22d2uxd2ux /kyungsang_ko_class_new Dataset Card for kyungsang_ko_class_new 개요 이 데이터셋은 표준어 질문과 경상도 사투리 답변으로 구성된 한국어 질의응답 데이터셋입니다. 입력 형식: question 출력 형식: answer 주요 특징: 표준어 질문, 경상도 사투리 답변 사투리 강도: strong 데이터 구조 각 샘플은 아래와 같은 구조를 가집니다. { "question": "한국의 수도는 어디인가?", "answer": "서울 아이가. 그건 뭐 다 아는 거라 안카나." } 스플릿 train: 212 validation: 24 소스 파일 원본 업로드 파일: data_new.jsonl 활용 예시 한국어 LLM instruction tuning 표준어 → 경상도 사투리 스타일 변환 방언 생성 및 스타일 제어 실험… See the full description on the dataset page: https://huggingface.co/datasets/d2uxd2ux/kyungsang_ko_class_new.texttext-generationn<1K0 likes14 downloads5mo agoHugging Face23decodingchris /clean_squad_classic_v1 Clean SQuAD Classic v1 This is a refined version of the SQuAD v1 dataset. It has been preprocessed to ensure higher data quality and usability for NLP tasks such as Question Answering. Description The Clean SQuAD Classic v1 dataset was created by applying preprocessing steps to the original SQuAD v1 dataset, including: Trimming whitespace: All leading and trailing spaces have been removed from the question field. Minimum question length: Questions with fewer than 12… See the full description on the dataset page: https://huggingface.co/datasets/decodingchris/clean_squad_classic_v1.textquestion-answering10K<n<100K1 likes12 downloads2y agoHugging Face24d2uxd2ux /kyungsang_ko_class Dataset Card for kyungsang_ko_class 개요 이 데이터셋은 표준어 질문과 경상도 사투리 답변으로 구성된 한국어 질의응답 데이터셋입니다. 입력 형식: question 출력 형식: answer 주요 특징: 표준어 질문, 경상도 사투리 답변 사투리 강도: strong 데이터 구조 각 샘플은 아래와 같은 구조를 가집니다. { "question": "한국의 수도는 어디인가?", "answer": "서울 아이가. 그건 뭐 다 아는 거라 안카나." } 스플릿 train: 398 validation: 45 소스 파일 원본 업로드 파일: data.jsonl 활용 예시 한국어 LLM instruction tuning 표준어 → 경상도 사투리 스타일 변환 방언 생성 및 스타일 제어 실험… See the full description on the dataset page: https://huggingface.co/datasets/d2uxd2ux/kyungsang_ko_class.texttext-generationn<1K0 likes12 downloads5mo agoHugging Face25decodingchris /clean_squad_classic_v2 Clean SQuAD Classic v2 This is a refined version of the SQuAD v2 dataset. It has been preprocessed to ensure higher data quality and usability for NLP tasks such as Question Answering. Description The Clean SQuAD Classic v2 dataset was created by applying preprocessing steps to the original SQuAD v2 dataset, including: Trimming whitespace: All leading and trailing spaces have been removed from the question field. Minimum question length: Questions with fewer than 12… See the full description on the dataset page: https://huggingface.co/datasets/decodingchris/clean_squad_classic_v2.textquestion-answering100K<n<1M1 likes9 downloads2y agoHugging Face26NotoriousH2 /gsm8k_classtextquestion-answering1K<n<10K0 likes8 downloads1y agoHugging Face27thach124 /Vietnamese-toxic-classificationtexttext-classification100K<n<1M0 likes8 downloads6mo agoHugging Face28farabi-lab /Text_classification_by_subject_areagated 🇰🇿 Kazakh Topic and Domain Identification Dataset Dataset Summary Kazakh Topic and Domain Identification Dataset is a Kazakh-language instruction-following dataset designed for topic recognition, domain classification, and text understanding tasks. Each sample contains a short Kazakh prompt, a long Kazakh text passage, a target response, a domain label, and a unique sample identifier. The dataset is intended to help Large Language Models (LLMs) and NLP systems… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Text_classification_by_subject_area.texttext-classification1K<n<10K0 likes3 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.