jensjepsen/danish-icl-schema-format-v2
danish-icl-json-v2 In-context-learning rows derived from jensjepsen/danish-json-grpo-v1. Each row packs 1-5 worked examples into a single user turn, followed by a held-out passage; the assistant turn is the answer for that passage. No instruction is included, so both the schema and the output format have to be inferred from the examples. Two axes vary per row and are held constant within a row: the schema (134 field-sets) and the output format (8 renderers — JSON, key: value… See the full description on the dataset page: https://huggingface.co/datasets/jensjepsen/danish-icl-schema-format-v2.
danish-icl-json-v2
In-context-learning rows derived from jensjepsen/danish-json-grpo-v1. Each row packs 1-5 worked examples into a single user turn, followed by a held-out passage; the assistant turn is the answer for that passage. No instruction is included, so both the schema and the output format have to be inferred from the examples. Two axes vary per row and are held constant within a row: the schema (134 field-sets) and the output format (8 renderers — JSON, key: value, key=value, [key] value, value -> key, numbered, TSV, and <key>value</key>). In roughly half the rows the field names are replaced by meaning-free symbols (alfa/kat_a/f1/foo), applied consistently within a row. Splits partition those axes rather than rows: eval_schema uses schemas absent from training, eval_format uses formats absent from training, eval_both uses neither, and val shares both axes with training but is built from passages reserved before any training row was generated. All values are rendered from the source dataset's gold_values rather than generated, and rows are filtered so that every value is recoverable from its own passage and every key, boolean value, empty marker and notation appearing in the answer is demonstrated by at least one example. Built by scripts/gen_icl_json.py.
