CoolFace
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Yaesir06 /CSSBench CSSBench: A Safety Evaluation Benchmark for Chinese Lightweight Language Models Overview CSSBench (Chinese-Specific Safety Benchmark) is a comprehensive evaluation framework designed to assess the safety robustness of Chinese Large Language Models (LLMs), with a specific emphasis on lightweight models (≤8B parameters). The benchmark bridges a critical evaluation gap by targeting Chinese-specific adversarial patterns—linguistic obfuscations such as homophones and Pinyin… See the full description on the dataset page: https://huggingface.co/datasets/Yaesir06/CSSBench.texttext-classification1K<n<10K3 likes275 downloads8mo agoHugging Face02gt-csse /false-citation-bench False Citation Bench False Citation Bench is a compact evaluation and inspection dataset for false or misleading case citations in legal documents. It contains 26 source documents, their PDFs, and manually reviewed citation annotations grounded in the local text extraction. Dataset contents The repository has one matching document in each directory: documents_txt/{index}__{case-name}__{filing}.txt documents_pdf/{index}__{case-name}__{filing}.pdf… See the full description on the dataset page: https://huggingface.co/datasets/gt-csse/false-citation-bench.documentn<1K2 likes128 downloads1mo agoHugging Face03fewshot-goes-multilingual /cs_squad-3.0 Dataset Card for Czech Simple Question Answering Dataset 3.0 This a processed and filtered adaptation of an existing dataset. For raw and larger dataset, see Dataset Source section. Dataset Description The data contains questions and answers based on Czech wikipeadia articles. Each question has an answer (or more) and a selected part of the context as the evidence. A majority of the answers are extractive - i.e. they are present in the context in the exact form. The… See the full description on the dataset page: https://huggingface.co/datasets/fewshot-goes-multilingual/cs_squad-3.0.tabularquestion-answering1K<n<10K3 likes99 downloads3y agoHugging Face04Northwestern-CSSI /Sci2Pol-BenchSci2Pol-Bench Data, scripts, and recipes for the benchmark Sci2Pol-Bench, a comprehensive benchmark for evaluating large language models. About • Usage• Authors About The data consists of policy briefs obtained from Nature Energy, Nature Climate, Nature Cities, and Journal of Health and Social Behavior Policy Briefs. Policy briefs originally were introduced in the Nature Energy journal with the goal of: This format aims to provide… See the full description on the dataset page: https://huggingface.co/datasets/Northwestern-CSSI/Sci2Pol-Bench.textsummarization1K<n<10K2 likes98 downloads1y agoHugging Face05theprint /MultiRoundConvos-Code-JS-HTML-CSS-Pythontext1K<n<10K0 likes84 downloads9mo agoHugging Face06Juliankrg /HTML_CSS_CodeDataSet_100ktext100K<n<1M5 likes45 downloads1y agoHugging Face07CZLC /cs_snli Dataset Card for Czech SNLI Czech translation of the Stanford Natural Language Interface (SNLI) dataset with manual annotation of a SNLI subset. In addition to the entailment/contradiction/neutral inference, a "bad translation" class was added. The annotation was done by students of NLP or computational linguistics. 1499 same pairs were annotated by two students to check IAA. Dataset Details The annotation for Czech premise-hypothesis pairs is done on 165390 pairs… See the full description on the dataset page: https://huggingface.co/datasets/CZLC/cs_snli.tabulartext-classification10K<n<100K0 likes35 downloads2y agoHugging Face08ranjankn /cs-support-labelstextn<1K0 likes9 downloads6mo agoHugging Face09pathii /css_design_snippetstextn<1K0 likes3 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.