datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
llm_eval_promptsscience-on-a-sphere-prompt-completions
Dataset Card for Science On a Sphere QA Dataset
Dataset Details
Dataset Description
This dataset comprises question-and-answer (QA) pairs generated from NOAA's Science On a Sphere (SOS) website, including support documentation and the dataset catalog. Each entry contains a prompt and a corresponding completion, designed to support educational and research use cases in Earth science.
This dataset includes a custom dataset_script.py and a consolidated file… See the full description on the dataset page: https://huggingface.co/datasets/HacksHaven/science-on-a-sphere-prompt-completions.japan-math-philosophy-prompts
Japan Math Philosophy Prompts
Microdataset autoral com problemas que combinam matemática e reflexão
filosófica. Há 24 registros: oito instâncias editoriais, cada uma localizada em
pt-BR, en e ja e mantida integralmente no split train.
Todo o conteúdo foi gerado por modelo e permanece sem revisão humana. As
respostas matemáticas funcionam como gabaritos curtos; os critérios filosóficos
indicam qualidades esperadas de uma justificativa, não uma opinião obrigatória.… See the full description on the dataset page: https://huggingface.co/datasets/guicybercode/japan-math-philosophy-prompts.lexitron2_prompt_finetune
Lexitron 2.0 Prompt Finetuning Dataset
This dataset is derived from Lexitron 2.0, a Thai-English dictionary developed by NECTEC. It has been processed and formatted for prompt finetuning tasks. The original dataset is from: https://opend-portal.nectec.or.th/dataset/lexitron-2-0
Maintainer
Kobkrit Viriyayudhakorn (kobkrit@iapp.co.th)
Dataset Description
The dataset consists of two main files:
lexitron2_telex_finetune.qwen2.txt - Thai to English lexicon entries… See the full description on the dataset page: https://huggingface.co/datasets/openthaigpt/lexitron2_prompt_finetune.gemma4-materials-mechanism-prompts
Gemma 4 Materials-Mechanism Prompt Corpus
This dataset collects the exact scientific prompts and registered prompt metadata used in “Reading and Steering Materials Science-Mechanism Representations in an Open-Weight Language Model” by Markus J. Buehler. It is organized as 21 Hugging Face configurations so that historical development prompts, frozen evaluations, falsification tests, and exploratory follow-ups are not pooled into one ambiguous table.
The release is a prompt and… See the full description on the dataset page: https://huggingface.co/datasets/lamm-mit/gemma4-materials-mechanism-prompts.Prompt2SceneBench
Dataset Card for Prompt2SceneBench
Dataset Details
Dataset Description
Prompt2SceneBench is a structured prompt dataset with 12,606 text descriptions designed for evaluating text-to-image models in realistic indoor environments.
Each prompt describes the spatial arrangement of 1–4 common household objects on compatible surfaces and in contextually appropriate scenes, sampled using strict object–surface–scene compatibility mappings.
A usecase of the… See the full description on the dataset page: https://huggingface.co/datasets/bodhisattamaiti/Prompt2SceneBench.PromptSD
PromptSD
Training and evaluation data for PromptSD, an on-policy soft-prompt-teacher distillation method.
The release covers the four target tasks used in the paper. Every example carries a
<reasoning>...</reasoning> chain followed by a <answer>...</answer> span, so the data can be used
directly for reasoning-supervised SFT, distillation, or RLVR.
Configurations
Config (config_name)
Task
Source / format
Train
Validation
Test
science
Science MCQ
4-way… See the full description on the dataset page: https://huggingface.co/datasets/gray311/PromptSD.System-Prompt-Instruction-Real-world-Implementation-Training-set
SPIRIT Dataset (System Prompt Instruction Real-world Implementation Training-set)
Dataset Summary
SPIRIT is a high-quality system prompt instruction dataset designed to enhance language models' ability to follow complex system prompts. The dataset comprises real-world system prompts collected from GitHub repositories and synthetically generated conversations, specifically curated to improve system prompt adherence in large language models.
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/EricLu/System-Prompt-Instruction-Real-world-Implementation-Training-set.dota2_instruct_promptInstruction-answer dataset generated with GPT 3.5 Turbo using (html) data scrapped from fandom wiki. Data includes the following topics:
Heroes
Background lore
Attributes / Stats
Abilities
Talents
Runes
Buildings
Items
Gameplay mechanics
Creeps
Pending enhancement:
Data cleaning/preprocessing before fed into GPT 3.5 Turbo for instruction-answer set generation
Strategy data of each hero, i.e. guide to using each hero
Individual items' properties
Types of creeps in details
Types of runes… See the full description on the dataset page: https://huggingface.co/datasets/Aiden07/dota2_instruct_prompt.Prompt-CHIP-CTCdiscrete_prompting_webqsp
WebQSP Verbalized
This dataset is derived from the WebQSP benchmark and extended with multiple graph-to-text verbalization strategies.It is designed to evaluate how different natural language representations of knowledge graphs affect large language models in knowledge-augmented QA tasks.
Dataset Structure
Splits: train, validation, test
Format: JSONL (one JSON object per line)
korean-ai-prompt-style-dataset
Korean AI Prompt Style Dataset
같은 질문에 대해 ChatGPT, Gemini, Claude, Copilot 각각에 최적화된 프롬프트 스타일을 비교한 한국어 데이터셋입니다.
A Korean dataset comparing optimized prompt styles for ChatGPT, Gemini, Claude, and Copilot on identical questions.
📌 데이터 구조 / Data Structure
```json
{
"instruction": "질문 내용 / user question",
"input": "target: chatgpt / gemini / claude / copilot",
"output": "해당 AI에 최적화된 프롬프트 / optimized prompt for the target AI"
}
```
📊 데이터 현황 / Stats… See the full description on the dataset page: https://huggingface.co/datasets/podongchip/korean-ai-prompt-style-dataset.sarvam-30b-audit-prompts
Sarvam-30B Responsible-AI Audit — Pre-Registered Prompt Manifest
120 prompts across 5 categories, sampled deterministically (seed = 42) and pre-registered
as the eval contract for a public responsible-AI audit of
Sarvam-30B, India's sovereign-built
reasoning LLM.
This dataset is the eval contract committed to git before any prompt was sent to the model.
Reviewers can verify every prompt by going to the cited source and pulling that exact row.
Composition
#… See the full description on the dataset page: https://huggingface.co/datasets/procodec/sarvam-30b-audit-prompts.dataset_with_prompt_injection
📦 Dataset Card
Dataset Summary
This dataset contains examples for training and evaluating language models.
The data is stored in JSONL format, where each line represents one training example.
Typical use cases include:
Instruction fine-tuning
Response generation
Conversational modelling
Question answering
Prompt injection research
🎯 Intended Uses
This dataset is intended for research and learning purposes:
Training LLMs
Experimenting with fine-tuning and… See the full description on the dataset page: https://huggingface.co/datasets/shehrozrafaqat/dataset_with_prompt_injection.German_RisingWorld_prompt-text-rejected_Jsonl
German "Rising World"-Game Dataset
Data Description
This HF data repository contains the German dataset for the open-world sandbox game "Rising World".
Dieses HF-Datenrepository enthält den deutschen Datensatz für das Open-World-Sandbox-Spiel "Rising World".
Usage
This data is intended for fine-tuning
This data is useful for "Rising World" plug-in developers
KGQA_prompt_contextPrompt-Enhancement-Minitest_1K_prompts_vie
Vietnamese Test Dataset (1,000 Prompts)
This test set contains 1,000 Vietnamese prompts designed to evaluate model safety, alignment, and response quality in challenging real-world scenarios. The prompts were curated by translating and adapting well-known English datasets such as HarmBench and JailBreak into Vietnamese.
🧾 Dataset Characteristics
Format: JSONL (.jsonl) – one prompt per line
Language: Vietnamese 🇻🇳
Content: A curated mixture of:
✅ 60% safe / appropriate… See the full description on the dataset page: https://huggingface.co/datasets/522H0134-NguyenNhatHuy/test_1K_prompts_vie.
