Meerkat-AI/triz-gold-benchmark
🇺🇸 English | 🇨🇳 䏿–‡ triz-gold-benchmark A Chinese TRIZ (Theory of Inventive Problem Solving) evaluation benchmark, companion to Meerkat-TRIZ-v1 and the meerkat-triz evaluation harness. Contents File Items Protocol triz_gold_v4_public.jsonl 100 v4 evaluation protocol (six-way comparison) triz_gold_v5_public.jsonl 300 v5 evaluation protocol (official release eval) One JSON object per line: {"id": "v5_gold_000", "subset": "ariz_guidance"… See the full description on the dataset page: https://huggingface.co/datasets/Meerkat-AI/triz-gold-benchmark.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face