datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
danish-icl-schema-format-v3
danish-icl-json-v3
In-context-learning rows derived from jensjepsen/danish-json-grpo-v1. Each row
packs 1-5 worked examples into a single user turn, followed by a held-out
passage; the assistant turn is the answer for that passage. No instruction is
included, so both the schema and the output format have to be inferred from the
examples. Two axes vary per row and are held constant within a row: the schema
(134 field-sets) and the output format (10 renderers — JSON, key: value… See the full description on the dataset page: https://huggingface.co/datasets/jensjepsen/danish-icl-schema-format-v3.wildchat-glm53-format-completions
WildChat format completions
9,975 GLM-5.3 answers across 29 parseable formats. Each answer passed its
contract verifier. Failed answers were resampled with the same prompt until one
passed; no semantic judge or answer repair was used.
This is the final release from a 10,000-prompt run; 25 unfinished prompts were excluded. It stores
answers, format instructions, exact contracts, and pinned
WildChat-4.8M references—not
the source prompts or conversations. All rows are in the train… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/wildchat-glm53-format-completions.t5gemma2-indonesia-chat-formatted
T5Gemma-2 Indonesian Chat & QA Dataset
A high-quality Indonesian language multi-turn conversation and reading comprehension dataset, specifically formatted for instruction tuning of sequence-to-sequence (Seq2Seq) models like T5-Gemma / T5-Gemma-2.
Dataset Description
This dataset contains over 7,400 multi-turn conversations and document-based Q&A in Bahasa Indonesia. It covers diverse topics including everyday life, technology, general knowledge, and structured… See the full description on the dataset page: https://huggingface.co/datasets/daruokta/t5gemma2-indonesia-chat-formatted.danish-icl-schema-format-v1
danish-icl-json-v1
In-context-learning rows derived from jensjepsen/danish-json-grpo-v1. Each
row packs 1-5 worked examples sharing a JSON schema into a single user turn,
followed by a held-out passage; the assistant turn is the answer for that
passage. No instruction is included, so the schema and the output format have
to be inferred from the examples. In roughly half the rows the field names are
replaced by meaning-free symbols (alfa/beta/..., kat_a/..., f1/...,
foo/bar/...)… See the full description on the dataset page: https://huggingface.co/datasets/jensjepsen/danish-icl-schema-format-v1.danish-icl-schema-format-v2
danish-icl-json-v2
In-context-learning rows derived from jensjepsen/danish-json-grpo-v1. Each row
packs 1-5 worked examples into a single user turn, followed by a held-out
passage; the assistant turn is the answer for that passage. No instruction is
included, so both the schema and the output format have to be inferred from the
examples. Two axes vary per row and are held constant within a row: the schema
(134 field-sets) and the output format (8 renderers — JSON, key: value… See the full description on the dataset page: https://huggingface.co/datasets/jensjepsen/danish-icl-schema-format-v2.MATH-500-gsm8k-format
MATH-500-gsm8k-format
Dataset Description
This dataset contains 500 mathematical problems from the MATH-500 benchmark, converted to GSM8K format for step-by-step reasoning.
Dataset Summary
Source: HuggingFaceH4/MATH-500
Format: GSM8K-style step-by-step solutions with inline computation annotations
Size: 500 problems
Split: Test (original MATH-500 test split)
Conversion Process
The original MATH-500 solutions (which use LaTeX notation and… See the full description on the dataset page: https://huggingface.co/datasets/albertge/MATH-500-gsm8k-format.exp-bf-format
Brand Function Format Optimization (Exp D)
Dataset Summary
375 LLM responses (355 valid, 20 parse errors, 0 failures) testing which representational format of a brand function specification maximizes AI comprehension fidelity. Five formats (JSON structured, prose narrative, tabular minimal, ranked list, score-only vector) were crossed with five canonical SBT brands (Hermes, IKEA, Patagonia, Tesla, Erewhon), five model families (Claude Haiku 4.5, GPT-4o-mini… See the full description on the dataset page: https://huggingface.co/datasets/spectralbranding/exp-bf-format.
