datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
yue_xstory_cloze
Dataset Card for Cantonese XStoryCloze
This dataset is a Cantonese translation of the Simplified Chinese subset of juletxara/xstory_cloze. For more detailed information about the original dataset, please refer to the provided link.
This dataset is translated by indiejoseph/bart-translation-zh-yue and has not undergone any manual verification. The content may be inaccurate or misleading. please keep this in mind when using this dataset.
Sample
{
"input_sentence_1":… See the full description on the dataset page: https://huggingface.co/datasets/hon9kon9ize/yue_xstory_cloze.hukukbert-cloze-benchmark
Turkish Legal Cloze Benchmark (JSONL)
A small-scale Cloze-style multiple-choice benchmark for evaluating Turkish legal-domain language models.
The dataset is designed to test whether models understand legal terminology, doctrinal structure, and domain-specific phrasing in Turkish law.
Dataset Format
Each example is stored as one JSON object per line (JSONL).
Schema
{
"id": "string",
"sentence": "string with [MASK] placeholder",
"options": ["choice1"… See the full description on the dataset page: https://huggingface.co/datasets/turkhukuk/hukukbert-cloze-benchmark.cloze-corpusCloze-lm-retention-seed-Vie
