datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CSSBench
CSSBench: A Safety Evaluation Benchmark for Chinese Lightweight Language Models
Overview
CSSBench (Chinese-Specific Safety Benchmark) is a comprehensive evaluation framework designed to assess the safety robustness of Chinese Large Language Models (LLMs), with a specific emphasis on lightweight models (≤8B parameters). The benchmark bridges a critical evaluation gap by targeting Chinese-specific adversarial patterns—linguistic obfuscations such as homophones and Pinyin… See the full description on the dataset page: https://huggingface.co/datasets/Yaesir06/CSSBench.false-citation-bench
False Citation Bench
False Citation Bench is a compact evaluation and inspection dataset for false or misleading case citations in legal documents. It contains 26 source documents, their PDFs, and manually reviewed citation annotations grounded in the local text extraction.
Dataset contents
The repository has one matching document in each directory:
documents_txt/{index}__{case-name}__{filing}.txt
documents_pdf/{index}__{case-name}__{filing}.pdf… See the full description on the dataset page: https://huggingface.co/datasets/gt-csse/false-citation-bench.cs_squad-3.0
Dataset Card for Czech Simple Question Answering Dataset 3.0
This a processed and filtered adaptation of an existing dataset. For raw and larger dataset, see Dataset Source section.
Dataset Description
The data contains questions and answers based on Czech wikipeadia articles.
Each question has an answer (or more) and a selected part of the context as the evidence.
A majority of the answers are extractive - i.e. they are present in the context in the exact form. The… See the full description on the dataset page: https://huggingface.co/datasets/fewshot-goes-multilingual/cs_squad-3.0.Sci2Pol-BenchSci2Pol-Bench
Data, scripts, and recipes for the benchmark Sci2Pol-Bench, a comprehensive benchmark for evaluating large language models.
About •
Usage•
Authors
About
The data consists of policy briefs obtained from Nature Energy, Nature Climate, Nature Cities, and Journal of Health and Social Behavior Policy Briefs.
Policy briefs originally were introduced in the Nature Energy journal with the goal of:
This format aims to provide… See the full description on the dataset page: https://huggingface.co/datasets/Northwestern-CSSI/Sci2Pol-Bench.MultiRoundConvos-Code-JS-HTML-CSS-PythonHTML_CSS_CodeDataSet_100kcs_snli
Dataset Card for Czech SNLI
Czech translation of the Stanford Natural Language Interface (SNLI) dataset with manual annotation of a SNLI subset.
In addition to the entailment/contradiction/neutral inference, a "bad translation" class was added.
The annotation was done by students of NLP or computational linguistics. 1499 same pairs were annotated by two students to check IAA.
Dataset Details
The annotation for Czech premise-hypothesis pairs is done on 165390 pairs… See the full description on the dataset page: https://huggingface.co/datasets/CZLC/cs_snli.cs-support-labelscss_design_snippets
