datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
python_code_instructions_18k_alpaca
Dataset Card for python_code_instructions_18k_alpaca
The dataset contains problem descriptions and code in python language.
This dataset is taken from sahil2801/code_instructions_120k, which adds a prompt column in alpaca style. Refer to the source here.
code_instructions_120k_alpaca
Dataset Card for code_instructions_120k_alpaca
This dataset is taken from sahil2801/code_instructions_120k, which adds a prompt column in alpaca style. Refer to the original source here.
code_contest_python3_alpaca
Dataset Card for Code Contest Processed
Dataset Summary
This dataset contains coding contest questions and their solution written in Python3.
This dataset is created by processing code_contest dataset from Deepmind. It is a competitive programming dataset for machine-learning. Read more about dataset at original source.
Columns Description
id : unique string associated with a problem
description : problem description
code : one correct code for the problem… See the full description on the dataset page: https://huggingface.co/datasets/iamtarun/code_contest_python3_alpaca.DD-VQAFirstAidQA
FirstAidQA: A Synthetic First-Aid and Emergency-Response Question-Answering Dataset
Medical safety notice: FirstAidQA is intended for research and educational purposes. It is not a substitute for professional medical advice, emergency services, certified first-aid training, or clinical judgment. Models trained on this dataset may produce incomplete, outdated, or unsafe responses.
Dataset Summary
FirstAidQA is an English-language synthetic question-answering… See the full description on the dataset page: https://huggingface.co/datasets/i-am-mushfiq/FirstAidQA.code_contest_processed
Dataset Card for Code Contest Processed
Dataset Summary
This dataset is created by processing code_contest dataset from Deepmind. It is a competitive programming dataset for machine-learning. Read more about dataset at original source.
Columns Description
id : unique string associated with a problem
description : problem description
code : one correct code for the problem
language : programming language used for code
test_samples : contains inputs and their… See the full description on the dataset page: https://huggingface.co/datasets/iamtarun/code_contest_processed.amazon-customer-support
Amazon Customer Support
Derived from the TWCS corpus (Kaggle: thoughtvector/customer-support-on-twitter), this dataset
contains 200 labelled customer-support interactions for evaluation / fine-tuning purposes.
Schema
Each line is a JSON object with four fields:
Field
Type
Description
query
string
The customer's raw message (input)
action
string
Agent action taken — resolve or escalated_to_human
intent
string
Classified intent — complaint, question… See the full description on the dataset page: https://huggingface.co/datasets/iam-tsr/amazon-customer-support.ragmix
RAGmix
RAGmix is a heterogeneous, multi-domain evaluation dataset for Retrieval-Augmented Generation (RAG) systems.
It mixes real-world document styles—policies, meeting minutes, clinical and scientific text, financial disclosures, job postings, and more—so models can be tested outside a single vertical. Each example pairs a full source document with one grounded question and a reference answer.
Source PDFs were obtained from Digital Corpora and converted to markdown for this… See the full description on the dataset page: https://huggingface.co/datasets/iam-tsr/ragmix.qa_metacul
qa_metacul
Summary
qa_metacul is an 800-question multiple-choice benchmark used to evaluate metadata-conditioned language models in the Metadata Conditioned LLMs project.
The benchmark tests whether a model can answer culturally and geographically grounded factual questions for different parts of the world, and whether metadata-aware models correctly adapt their answers when continent- or country-level context changes.
Paper: https://arxiv.org/abs/2601.15236
Project… See the full description on the dataset page: https://huggingface.co/datasets/iamshnoo/qa_metacul.ai-basic-law-dataset
台灣人工智慧基本法 訓練資料集
Taiwan AI Basic Law (人工智慧基本法) Q&A dataset for LLM finetuning.
Files
File
Description
Entries
train.jsonl
Full training dataset with oversampling
~5000
fulltext.jsonl
Clean article fulltext (20 articles)
38
Data Composition
Category
Unique
Repeat
Purpose
Article Fulltext Q&A
~157
x15
Verbatim article text with topic anchors
Alias Recognition
~109
x10
「基本法」「AI基本法」→ 人工智慧基本法
Legislative Reasons
~35
x3
Background… See the full description on the dataset page: https://huggingface.co/datasets/iamjry/ai-basic-law-dataset.mental_health_Chatbot
Amod/mental_health_counseling_conversations
This dataset is a compilation of high-quality, real one-on-one mental health counseling conversations between individuals and licensed professionals. Each exchange is structured as a clear question–answer pair, making it directly suitable for fine-tuning or instruction-tuning language models that need to handle sensitive, empathetic, and contextually aware dialogue.
Since its public release in 2023, it has been downloaded over 100,000… See the full description on the dataset page: https://huggingface.co/datasets/Iamzoo/mental_health_Chatbot.Nemotron-RL-litmus-bench-v0.1
Dataset Description:
Litmus-Bench v0.1 is an open dataset for training and evaluating chemical reasoning in language models. It includes 5,232 training questions and 482 test questions, each in short-answer format and was created from the ChEMBL dataset with RDKit descriptors requiring short answers. The dataset is for RL training.
This dataset is released as part of NVIDIA NeMo-Gym, an open-source library within the NVIDIA NeMo framework, designed for large-scale, verifiable… See the full description on the dataset page: https://huggingface.co/datasets/Iambackup/Nemotron-RL-litmus-bench-v0.1.HCMUT_FAQpython_code_instructions_18k_alpaca
Dataset Card for python_code_instructions_18k_alpaca
The dataset contains problem descriptions and code in python language.
This dataset is taken from sahil2801/code_instructions_120k, which adds a prompt column in alpaca style. Refer to the source here.
RonDistillMed3MI-am-not-happy-with-your-punishment-and-I-don-t-agree-with-itrussian-spell-correctionmedmcqa
