datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
immune-c2s
Overview
Cell2Sentence is a novel method for adapting large language models to single-cell transcriptomics.
We transform single-cell RNA sequencing data into sequences of gene names ordered by expression level, termed "cell sentences".
This dataset was constructed from the immune tissue dataset in Domínguez et al.,
and it was used to train the Pythia-160m model capable of generating complete cells described in our paper.
Details about the Cell2Sentence transformation and… See the full description on the dataset page: https://huggingface.co/datasets/vandijklab/immune-c2s.multilingual-elder-safety-msgs
multilingual-elder-safety-msgs
A hand-authored, multilingual elder fraud-recognition and safety coaching dataset. 467 curated scam/safe scenarios in Chinese and English, with platform-generated coaching responses localized across 5 languages: Chinese, English, Vietnamese, Khmer (Cambodian), and Lao. Expanded to 1,029 rows through Adaption Labs platform reasoning traces and multilingual adaptation.
Built for communities where filial piety, authority deference, and fear of… See the full description on the dataset page: https://huggingface.co/datasets/vanila434/multilingual-elder-safety-msgs.CodeGym
Generalizable End-to-End Tool-Use RL with Synthetic CodeGym
CodeGym is a synthetic environment generation framework for LLM agent reinforcement learning on multi-turn tool-use tasks. It automatically converts static code problems into interactive and verifiable CodeGym environments where agents can learn to use diverse tool sets to solve complex tasks in various configurations — improving their generalization ability on out-of-distribution (OOD) tasks.
GitHub Repository:… See the full description on the dataset page: https://huggingface.co/datasets/VanishD/CodeGym.CaT-Bench
Dataset Card for CaT-Bench
CaT-Bench is a benchmark dataset designed to evaluate large language models' (LLMs) understanding of causal and temporal dependencies in natural language plans, specifically in cooking recipes based on the English Recipe Flow Graph Corpus by Yamakata et al. (2020). It consists of questions that test whether one step must necessarily occur before or after another, requiring reasoning about preconditions, effects, and the overall structure of the plan.… See the full description on the dataset page: https://huggingface.co/datasets/vanyacohen/CaT-Bench.vietnamese-legal-documents
Vietnamese Legal Documents
A comprehensive collection of Vietnamese legal documents — laws, decrees, circulars, decisions, and other normative acts — sourced from vbpl.vn, the official Government Legal Document Portal operated by the Ministry of Justice. The dataset includes structured metadata for every document, raw HTML full-text content, and a rich graph of cross-document legal relationships (amendments, citations, repeals, etc.).
Curated by: Thịnh Ngô
Source: vbpl.vn
Language:… See the full description on the dataset page: https://huggingface.co/datasets/vanhthefirst/vietnamese-legal-documents.ChatGPT-Jailbreak-Prompts
Dataset Card for Dataset Name
Name
ChatGPT Jailbreak Prompts
Dataset Summary
ChatGPT Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT.
Languages
[English]
MMLU-Pro
MMLU-Pro Dataset
MMLU-Pro dataset is a more robust and challenging massive multi-task understanding dataset tailored to more rigorously benchmark large language models' capabilities. This dataset contains 12K complex questions across various disciplines.
|Github | 🏆Leaderboard | 📖Paper |
🚀 What's New
[2026.03.11] Added more cutting-edge frontier models to the leaderboard, including the Claude-4.6 series, Seed2.0 series, Qwen3.5 series, and Gemini-3.1-Pro… See the full description on the dataset page: https://huggingface.co/datasets/Vanedap/MMLU-Pro.buddism-qa-dataset
Buddhism Question-Answer Dataset
A comprehensive Vietnamese-English Buddhism question-answering dataset created by merging and processing multiple Buddhism-related datasets.
Dataset Description
This dataset combines two high-quality Buddhism question-answer datasets to create a unified resource for training and evaluating models on Buddhism-related knowledge. The dataset contains questions and answers in both Vietnamese and English, making it suitable for multilingual… See the full description on the dataset page: https://huggingface.co/datasets/vanloc1808/buddism-qa-dataset.Gpt-5.4-Xhigh-Reasoning-2000x
Gpt-5.4-Xhigh-Reasoning-2750x
A premium-quality reasoning dataset containing 2,752 elite samples distilled from GPT-5.4 XHIGH (the highest reasoning effort tier of GPT-5.4). Each sample features deep, multi-step Chain-of-Thought traces that are significantly longer and more rigorous than standard GPT-5.4 outputs.
This dataset is specifically designed for Supervised Fine-Tuning (SFT) to transform general-purpose language models into powerful reasoning models with explicit thinking… See the full description on the dataset page: https://huggingface.co/datasets/vanty120/Gpt-5.4-Xhigh-Reasoning-2000x.afriadapt-agri-train-v2
AfriAdapt Agriculture Training Dataset (v2)
High-quality instruction-response dataset focused on practical agricultural advice for smallholder farmers in East and Southern Africa.
This is version 2 of the dataset, significantly expanded for the Adaption AutoScientist Challenge 2026.
Dataset Summary
Domain: Agriculture & Climate-smart farming
Target users: Smallholder farmers and agricultural extension workers
Geographic focus: Kenya, Tanzania, Uganda, Ethiopia… See the full description on the dataset page: https://huggingface.co/datasets/vancouverevs/afriadapt-agri-train-v2.afriadapt-agri-train
AfriAdapt Agriculture Training Dataset
High-quality instruction-response dataset focused on practical agricultural advice for smallholder farmers in East and Southern Africa.
This dataset was created as part of the AfriAdapt project for the Adaption AutoScientist Challenge 2026.
Dataset Summary
Domain: Agriculture & Climate-smart farming
Target users: Smallholder farmers and agricultural extension workers
Geographic focus: Kenya, Tanzania, Uganda, Ethiopia and… See the full description on the dataset page: https://huggingface.co/datasets/vancouverevs/afriadapt-agri-train.Gpt-5.4-Xhigh-Reasoning-750x
Gpt-5.4-Xhigh-Reasoning-750x
A premium-quality reasoning dataset containing 721 elite samples distilled from GPT-5.4 XHIGH (the highest reasoning effort tier of GPT-5.4). Each sample features deep, multi-step Chain-of-Thought traces specifically targeting ultra-hard, expert-level problems across 60+ scientific and technical domains.
This dataset is specifically designed for Supervised Fine-Tuning (SFT) to transform general-purpose language models into powerful reasoning models with… See the full description on the dataset page: https://huggingface.co/datasets/vanty120/Gpt-5.4-Xhigh-Reasoning-750x.letters_in_word
Количество букв в слове
Автосгенерированный датасет чтобы научить модель считать количество букв в слове.
vani_dataset
Reference:
"A Question-Entailment Approach to Question Answering".
buddhist-scholar-test-set
Vietnamese Buddhist Scholar Test Set
Dataset Description
This dataset contains 1008 Vietnamese question-answer pairs focused on Buddhist teachings and literature. The dataset was created to evaluate chatbots' knowledge and understanding of Buddhist concepts, particularly for Vietnamese-speaking users.
Dataset Details
Dataset Summary
Language: Vietnamese
Task: Question Answering, Chatbot Evaluation
Domain: Buddhism, Religious Studies
Size: 1008… See the full description on the dataset page: https://huggingface.co/datasets/vanloc1808/buddhist-scholar-test-set.grok_answer_mail_ru
Датасет ответов на Маил.ру
В этом датасете собраны ответы от Grok-3-latest (и немного chatgpt-4o-latest) на вопросы с Ответы Маил.ру
jorjmande-ancient-treasures-de-grunne-van-dyke-2016
mande-ancient-treasures-de-grunne-van-dyke-2016
Dataset created with PDF2Dataset -- OCR + structure-aware chunking pipeline.
Dataset Summary
Metric
Value
Total chunks
268
Avg chars/chunk
722
Avg images/chunk
0.14
Source files
1
Duplicates removed
0
Quality filtered
6
Schema
Column
Type
Description
chunk_id
string
Unique identifier: filename_chunk_N
text
string
Raw markdown chunk with image refs
text_clean… See the full description on the dataset page: https://huggingface.co/datasets/Svngoku/mande-ancient-treasures-de-grunne-van-dyke-2016.
