datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
python-image-copilot-training-using-class-knowledge-graphs
Python Copilot Image Training using Class Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each row contains a png file in the dbytes column.
Rows: 312277
Size: 304.3 GB
Data type: png
Format: Knowledge graph using NetworkX with alpaca text box
Schema
The png is in the dbytes column:
{
"dbytes": "binary"… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-image-copilot-training-using-class-knowledge-graphs.python-audio-copilot-training-using-class-knowledge-graphs
Python Copilot Audio Training using Class with Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each class method has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and the associated source code file_path identifier.
Rows:… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-class-knowledge-graphs.python-audio-copilot-training-using-class-knowledge-graphs-2024-01-27
Python Copilot Audio Training using Class with Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each class method has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and the associated source code file_path identifier.
Rows:… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-class-knowledge-graphs-2024-01-27.python-image-copilot-training-using-class-knowledge-graphs-2024-01-27
Python Copilot Image Training using Class Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each row contains a png file in the dbytes column.
Rows: 312836
Size: 294.1 GB
Data type: png
Format: Knowledge graph using NetworkX with alpaca text box
Schema
The png is in the dbytes column:
{
"dbytes": "binary"… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-image-copilot-training-using-class-knowledge-graphs-2024-01-27.events_classification_biotech
Key aspects
Event extraction;
Multi-label classification;
Biotech news domain;
31 classes;
3140 total number of examples;
Motivation
Text classification is a widespread task and a foundational step in numerous information extraction pipelines. However, a notable challenge in current NLP research lies in the oversimplification of benchmarking datasets, which predominantly focus on rudimentary tasks such as topic classification or sentiment analysis.
This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/knowledgator/events_classification_biotech.Wikipedia_RAG_QA_Classification
🏛️ Wikipedia RAG QA Dataset for Retrieval-Augmented Generation Training
📊 Dataset Description
This dataset contains 300,000+ validated model-generated responses to Wikipedia content, specifically designed for Retrieval-Augmented Generation (RAG) applications and SQL database insertion tasks. Generated by Jeeney AI Reloaded 207M GPT with specialized RAG tuning.
🖥️ Demo Interface: Discord
Live Chat Demo on Discord: https://discord.gg/Xe9tHFCS9h
The full CJ… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/Wikipedia_RAG_QA_Classification.FHIR_QnA_Query-Based_Resource_Relevance_Classification_T1
Dataset Card
This repository contains the dataset introduced in the paper Question Answering on Patient Medical Records with Private Fine-Tuned LLMs.
Appreciation-of-Chinese-Classical-Poetry
Appreciation of Chinese Classical Poetry
Chinese classical poetry with paired metadata and five-aspect literary analyses for PoetryMTEB / MTEB-style evaluation and computational poetics research.
Poems are drawn from expert appreciation volumes (mainly Shanghai Lexicographical Publishing House dictionaries). The released analysis fields are LLM distillations (DeepSeek-V3.1) of those expert appreciation texts into five free-text facets. The original long-form appreciation prose… See the full description on the dataset page: https://huggingface.co/datasets/PoetryMTEB/Appreciation-of-Chinese-Classical-Poetry.math-correctness-classifier_64rollouts
RedaAlami/math-correctness-classifier_64rollouts
Dataset Description
This dataset contains mathematical reasoning problems and model responses formatted for training correctness classifiers. Each record includes a problem statement, a model's solution attempt, and a binary label indicating correctness.
The dataset spans three benchmarks:
AIME 2024: American Invitational Mathematics Examination 2024
AIME 2025: American Invitational Mathematics Examination 2025
AMO:… See the full description on the dataset page: https://huggingface.co/datasets/RedaAlami/math-correctness-classifier_64rollouts.cycic_classificationhttps://storage.googleapis.com/ai2-mosaic/public/cycic/CycIC-train-dev.zip
https://colab.research.google.com/drive/16nyxZPS7-ZDFwp7tn_q72Jxyv0dzK1MP?usp=sharing
@article{Kejriwal2020DoFC,
title={Do Fine-tuned Commonsense Language Models Really Generalize?},
author={Mayank Kejriwal and Ke Shen},
journal={ArXiv},
year={2020},
volume={abs/2011.09159}
}
added for
@article{sileo2023tasksource,
title={tasksource: Structured Dataset Preprocessing Annotations for Frictionless Extreme… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/cycic_classification.Classical-Mechanics-Equations-Dataset_SFT-or-LoRA
Classical Mechanics Equations Dataset (SFT / LoRA Ready)
A structured dataset of 64 classical mechanics equations from Newtonian,
Lagrangian, and Hamiltonian mechanics, expanded into 448 instruction-tuning
rows across three task types: equation explanation, Q&A, and derivation.
Designed for fine-tuning LLMs on physics reasoning, STEM Q&A, and
equation understanding tasks.
Overview
Property
Value
Domain
Classical Mechanics (Physics)
Total rows
448
Train… See the full description on the dataset page: https://huggingface.co/datasets/ChaoticEconomist/Classical-Mechanics-Equations-Dataset_SFT-or-LoRA.it-support-l1-ticket-classification
IT Support L1 Multilingual Dataset
Dataset Summary
IT Support L1 Multilingual Dataset is a synthetic enterprise help desk dataset for ticket classification and troubleshooting response generation. It contains realistic Level 1 IT support scenarios in English and Czech, designed for experiments in structured classification, response generation, and multilingual support workflow prototyping.
This dataset contains synthetic IT Support L1 scenarios. The records were generated… See the full description on the dataset page: https://huggingface.co/datasets/w1z4rd3k/it-support-l1-ticket-classification.CBSE-Class-12th_2024_PYQs__structuredThis data set contains the CBSE Class 12 2024 papers in a structured format. The papers are annotated with topic and chapter names, and the figures are parsed and their paths annotated.
math-correctness-classifier-aimes-amo
RedaAlami/math-correctness-classifier-aimes-amo
Dataset Description
This dataset contains mathematical reasoning problems and model responses formatted for training correctness classifiers. Each record includes a problem statement, a model's solution attempt, and a binary label indicating correctness.
The dataset spans three benchmarks:
AIME 2024: American Invitational Mathematics Examination 2024
AIME 2025: American Invitational Mathematics Examination 2025
AMO: Asian… See the full description on the dataset page: https://huggingface.co/datasets/RedaAlami/math-correctness-classifier-aimes-amo.Vietnamese-toxic-classificationclassified_papers
Synthetic QA Dataset for Biomedical Paper Analysis (GPT-4o Generated)
This dataset consists of synthetically generated question-answer pairs designed to simulate the process of answering high-level research questions about biomedical papers. It was created using OpenAI's GPT-4o model and is tailored for fine-tuning or evaluating models on tasks such as biomedical reading comprehension, information extraction, and reasoning.
Dataset Structure
Each data sample is a JSON… See the full description on the dataset page: https://huggingface.co/datasets/AbrehamT/classified_papers.Patent_classification_QAmilitary-strategy-classics
Analytical Decision Frameworks — Public Domain Dataset
Structured public domain texts on decision-making, organizational design, and strategic analysis.
Formatted for AI training, analysis, and agent tool use.
All content sourced from works in the public domain (published before 1928, or government-authored).
Content Domains
Strategic planning principles
Organizational coordination patterns
Decision frameworks under uncertainty
Historical pattern analysis… See the full description on the dataset page: https://huggingface.co/datasets/gmahia/military-strategy-classics.FHIR_QnA_Query-Based_Resource_Relevance_Classification_T1_complete
Dataset Card
This repository contains the dataset introduced in the paper Question Answering on Patient Medical Records with Private Fine-Tuned LLMs.
Question-Classification-With-Answersprompts-classification-pfgkyungsang_ko_class_new
Dataset Card for kyungsang_ko_class_new
개요
이 데이터셋은 표준어 질문과 경상도 사투리 답변으로 구성된 한국어 질의응답 데이터셋입니다.
입력 형식: question
출력 형식: answer
주요 특징: 표준어 질문, 경상도 사투리 답변
사투리 강도: strong
데이터 구조
각 샘플은 아래와 같은 구조를 가집니다.
{
"question": "한국의 수도는 어디인가?",
"answer": "서울 아이가. 그건 뭐 다 아는 거라 안카나."
}
스플릿
train: 212
validation: 24
소스 파일
원본 업로드 파일: data_new.jsonl
활용 예시
한국어 LLM instruction tuning
표준어 → 경상도 사투리 스타일 변환
방언 생성 및 스타일 제어 실험… See the full description on the dataset page: https://huggingface.co/datasets/d2uxd2ux/kyungsang_ko_class_new.clean_squad_classic_v1
Clean SQuAD Classic v1
This is a refined version of the SQuAD v1 dataset. It has been preprocessed to ensure higher data quality and usability for NLP tasks such as Question Answering.
Description
The Clean SQuAD Classic v1 dataset was created by applying preprocessing steps to the original SQuAD v1 dataset, including:
Trimming whitespace: All leading and trailing spaces have been removed from the question field.
Minimum question length: Questions with fewer than 12… See the full description on the dataset page: https://huggingface.co/datasets/decodingchris/clean_squad_classic_v1.kyungsang_ko_class
Dataset Card for kyungsang_ko_class
개요
이 데이터셋은 표준어 질문과 경상도 사투리 답변으로 구성된 한국어 질의응답 데이터셋입니다.
입력 형식: question
출력 형식: answer
주요 특징: 표준어 질문, 경상도 사투리 답변
사투리 강도: strong
데이터 구조
각 샘플은 아래와 같은 구조를 가집니다.
{
"question": "한국의 수도는 어디인가?",
"answer": "서울 아이가. 그건 뭐 다 아는 거라 안카나."
}
스플릿
train: 398
validation: 45
소스 파일
원본 업로드 파일: data.jsonl
활용 예시
한국어 LLM instruction tuning
표준어 → 경상도 사투리 스타일 변환
방언 생성 및 스타일 제어 실험… See the full description on the dataset page: https://huggingface.co/datasets/d2uxd2ux/kyungsang_ko_class.clean_squad_classic_v2
Clean SQuAD Classic v2
This is a refined version of the SQuAD v2 dataset. It has been preprocessed to ensure higher data quality and usability for NLP tasks such as Question Answering.
Description
The Clean SQuAD Classic v2 dataset was created by applying preprocessing steps to the original SQuAD v2 dataset, including:
Trimming whitespace: All leading and trailing spaces have been removed from the question field.
Minimum question length: Questions with fewer than 12… See the full description on the dataset page: https://huggingface.co/datasets/decodingchris/clean_squad_classic_v2.gsm8k_classVietnamese-toxic-classificationText_classification_by_subject_area
🇰🇿 Kazakh Topic and Domain Identification Dataset
Dataset Summary
Kazakh Topic and Domain Identification Dataset is a Kazakh-language instruction-following dataset designed for topic recognition, domain classification, and text understanding tasks.
Each sample contains a short Kazakh prompt, a long Kazakh text passage, a target response, a domain label, and a unique sample identifier. The dataset is intended to help Large Language Models (LLMs) and NLP systems… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Text_classification_by_subject_area.
