CoolFace
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01richardyoung /llm-instruction-following-eval LLM Instruction-Following Evaluation: 256 Models Across 20 Diagnostic Tests Dataset Summary This dataset contains comprehensive evaluation results from testing 256 Large Language Models across 20 carefully designed diagnostic instruction-following prompts, totaling 5,120 individual evaluations. The evaluation was conducted on October 14, 2025, using the OpenRouter API. Paper: When Models Can't Follow: Testing Instruction Adherence Across 256 LLMs arXiv: 2510.18892… See the full description on the dataset page: https://huggingface.co/datasets/richardyoung/llm-instruction-following-eval.text-generation1K<n<10K0 likes206 downloads11mo agoHugging Face02amalia-llm /persona_instruction_following Persona Instruction Following Datasets This is a synthetic instruction-following dataset, available in two configs: full and filtered. Each config contains two language splits, English (en) and Portuguese (pt). The filtered version keeps only the higher-quality examples (quality score 5). The prompts were generated using an approach similar to PersonaHub, with a translated version of proj-persona/PersonaHub. Both prompts and answers were generated using Gemma 3-27B.… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/persona_instruction_following.textquestion-answering100K<n<1M0 likes43 downloads3mo agoHugging Face03PoSTMEDIA /rosetta-ko-instruction-following-synth-sftgated rosetta-ko-instruction-following-synth-sft Korean-native instruction-following data — instructions with verifiable constraints (length, format, keywords, JSON, ...) checked by programmatic verifiers. Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA). Instruction-Following Suite Sibling datasets from the same pipeline (each a separate repo): repo format… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-instruction-following-synth-sft.texttext-generation100K<n<1M0 likes20 downloads14d agoHugging Face04thunder-research-group /SNU_Thunder-synthetic-instruction-followinggated Dataset Card for SNU Thunder Synthetic InstructionFollowing Dataset Summary This dataset was used as part of the post-training corpus for SnuLLM(to_fill). This dataset consists of Korean and English question-answer pairs. Questions are sourced from publicly available datasets, and answers were generated using open large language models (Exaone 3.5, LLaMA 3.3, Qwen 2.5). It is intended for research and non-commercial use. Supported Tasks Tasks: Instruction… See the full description on the dataset page: https://huggingface.co/datasets/thunder-research-group/SNU_Thunder-synthetic-instruction-following.textquestion-answering100K<n<1M2 likes5 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.