datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
llm_eval_promptsjapan-math-philosophy-prompts
Japan Math Philosophy Prompts
Microdataset autoral com problemas que combinam matemática e reflexão
filosófica. Há 24 registros: oito instâncias editoriais, cada uma localizada em
pt-BR, en e ja e mantida integralmente no split train.
Todo o conteúdo foi gerado por modelo e permanece sem revisão humana. As
respostas matemáticas funcionam como gabaritos curtos; os critérios filosóficos
indicam qualidades esperadas de uma justificativa, não uma opinião obrigatória.… See the full description on the dataset page: https://huggingface.co/datasets/guicybercode/japan-math-philosophy-prompts.gemma4-materials-mechanism-prompts
Gemma 4 Materials-Mechanism Prompt Corpus
This dataset collects the exact scientific prompts and registered prompt metadata used in “Reading and Steering Materials Science-Mechanism Representations in an Open-Weight Language Model” by Markus J. Buehler. It is organized as 21 Hugging Face configurations so that historical development prompts, frozen evaluations, falsification tests, and exploratory follow-ups are not pooled into one ambiguous table.
The release is a prompt and… See the full description on the dataset page: https://huggingface.co/datasets/lamm-mit/gemma4-materials-mechanism-prompts.PromptSD
PromptSD
Training and evaluation data for PromptSD, an on-policy soft-prompt-teacher distillation method.
The release covers the four target tasks used in the paper. Every example carries a
<reasoning>...</reasoning> chain followed by a <answer>...</answer> span, so the data can be used
directly for reasoning-supervised SFT, distillation, or RLVR.
Configurations
Config (config_name)
Task
Source / format
Train
Validation
Test
science
Science MCQ
4-way… See the full description on the dataset page: https://huggingface.co/datasets/gray311/PromptSD.sarvam-30b-audit-prompts
Sarvam-30B Responsible-AI Audit — Pre-Registered Prompt Manifest
120 prompts across 5 categories, sampled deterministically (seed = 42) and pre-registered
as the eval contract for a public responsible-AI audit of
Sarvam-30B, India's sovereign-built
reasoning LLM.
This dataset is the eval contract committed to git before any prompt was sent to the model.
Reviewers can verify every prompt by going to the cited source and pulling that exact row.
Composition
#… See the full description on the dataset page: https://huggingface.co/datasets/procodec/sarvam-30b-audit-prompts.test_1K_prompts_vie
Vietnamese Test Dataset (1,000 Prompts)
This test set contains 1,000 Vietnamese prompts designed to evaluate model safety, alignment, and response quality in challenging real-world scenarios. The prompts were curated by translating and adapting well-known English datasets such as HarmBench and JailBreak into Vietnamese.
🧾 Dataset Characteristics
Format: JSONL (.jsonl) – one prompt per line
Language: Vietnamese 🇻🇳
Content: A curated mixture of:
✅ 60% safe / appropriate… See the full description on the dataset page: https://huggingface.co/datasets/522H0134-NguyenNhatHuy/test_1K_prompts_vie.
