datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CSS10-Multilingual-LJSpeech
CSS10-Multilingual-LJSpeech
Multilingual speech dataset combining LJSpeech (English) + CSS10 (10 languages) in a consistent LJSpeech format.
Dataset Description
This dataset merges:
LJSpeech: High-quality English speech dataset
CSS10: A collection of single-speaker speech datasets for 10 languages
All audio files are provided in a consistent format suitable for TTS training.
Features
Each sample contains:
audio: Waveform audio sampled at 22,050 Hz
text:… See the full description on the dataset page: https://huggingface.co/datasets/davidguzmanr/CSS10-Multilingual-LJSpeech.SciSciGPT-SciSciNetpreprocessed_jsut_jsss_css10_common_voice_11
Dataset Card for "preprocessed_jsut_jsss_css10_common_voice_11"
More Information needed
HTML-CSS-Website# Dataset
This dataset contains a collection of FacebookAds-related queries and responses generated by an AI assistant.
# Proudly Dataset Genrated with AI with [AI Dataset Generator API](https://api.example.com)
CSSR-S_labelled_suicidewatch_posts_reddit
Evaluating Reasoning LLMs for Suicide Screening with the Columbia-Suicide Severity Rating Scale
Full code and supplementary materials are available at https://github.com/av9ash/llm_cssrs_code.
License and Citation
This project is released under the Creative Commons Attribution 4.0 International (CC BY 4.0) license.Any use or reuse of this work please cite the following:
@article{patil2025evaluating,
title={Evaluating Reasoning LLMs for Suicide Screening with the… See the full description on the dataset page: https://huggingface.co/datasets/av9ash/CSSR-S_labelled_suicidewatch_posts_reddit.preprocessed_jsut_jsss_css10_fleurs_common_voice_11
Dataset Card for "preprocessed_jsut_jsss_css10_fleurs_common_voice_11"
More Information needed
CSSBench
CSSBench: A Safety Evaluation Benchmark for Chinese Lightweight Language Models
Overview
CSSBench (Chinese-Specific Safety Benchmark) is a comprehensive evaluation framework designed to assess the safety robustness of Chinese Large Language Models (LLMs), with a specific emphasis on lightweight models (≤8B parameters). The benchmark bridges a critical evaluation gap by targeting Chinese-specific adversarial patterns—linguistic obfuscations such as homophones and Pinyin… See the full description on the dataset page: https://huggingface.co/datasets/Yaesir06/CSSBench.sciscinet-v2
📢🚨📣 Sciscinet-v2
Sciscinet-v2 is a refreshed update to SciSciNet which is a large-scale, integrated dataset designed to support research in the science of science domain. It combines scientific publications with their network of relationships to funding sources, patents, citations, and institutional affiliations, creating a rich ecosystem for analyzing scientific productivity, impact, and innovation. Know more.
About Sciscinet-v2
The newer version Sciscinet-v2 is… See the full description on the dataset page: https://huggingface.co/datasets/Northwestern-CSSI/sciscinet-v2.preprocessed_jsut_jsss_css10
Dataset Card for "preprocessed_jsut_jsss_css10"
More Information needed
dota_v1.5github-code-html-css-1false-citation-bench
False Citation Bench
False Citation Bench is a compact evaluation and inspection dataset for false or misleading case citations in legal documents. It contains 26 source documents, their PDFs, and manually reviewed citation annotations grounded in the local text extraction.
Dataset contents
The repository has one matching document in each directory:
documents_txt/{index}__{case-name}__{filing}.txt
documents_pdf/{index}__{case-name}__{filing}.pdf… See the full description on the dataset page: https://huggingface.co/datasets/gt-csse/false-citation-bench.github-code-html-css-2landing-pages-v2-csscs_squad-3.0
Dataset Card for Czech Simple Question Answering Dataset 3.0
This a processed and filtered adaptation of an existing dataset. For raw and larger dataset, see Dataset Source section.
Dataset Description
The data contains questions and answers based on Czech wikipeadia articles.
Each question has an answer (or more) and a selected part of the context as the evidence.
A majority of the answers are extractive - i.e. they are present in the context in the exact form. The… See the full description on the dataset page: https://huggingface.co/datasets/fewshot-goes-multilingual/cs_squad-3.0.Sci2Pol-BenchSci2Pol-Bench
Data, scripts, and recipes for the benchmark Sci2Pol-Bench, a comprehensive benchmark for evaluating large language models.
About •
Usage•
Authors
About
The data consists of policy briefs obtained from Nature Energy, Nature Climate, Nature Cities, and Journal of Health and Social Behavior Policy Briefs.
Policy briefs originally were introduced in the Nature Energy journal with the goal of:
This format aims to provide… See the full description on the dataset page: https://huggingface.co/datasets/Northwestern-CSSI/Sci2Pol-Bench.MultiRoundConvos-Code-JS-HTML-CSS-Pythongithub-code-html-cssgithub-code-html-css-split-3html-css-js-cot
🌐 HTML/CSS/JS Reasoning Traces Dataset
A high-quality, large-scale dataset of complex HTML, CSS, and JavaScript programming questions and model reasoning traces.
📊 Dataset Overview
This repository contains a comprehensively structured dataset of reasoning traces for frontend web development tasks. The data maps intricate, multi-step prompts to step-by-step reasoning solutions generated by advanced Language Models.
It is designed for researchers and… See the full description on the dataset page: https://huggingface.co/datasets/AdhyanshVerma/html-css-js-cot.css-deepfake-datasetHTML_CSS_CodeDataSet_100khtml-css-codegen-datasetSciSciGPT-SciSciCorpuscs_snli
Dataset Card for Czech SNLI
Czech translation of the Stanford Natural Language Interface (SNLI) dataset with manual annotation of a SNLI subset.
In addition to the entailment/contradiction/neutral inference, a "bad translation" class was added.
The annotation was done by students of NLP or computational linguistics. 1499 same pairs were annotated by two students to check IAA.
Dataset Details
The annotation for Czech premise-hypothesis pairs is done on 165390 pairs… See the full description on the dataset page: https://huggingface.co/datasets/CZLC/cs_snli.csst2AMSMB-line-transcription
Dataset Card
Dataset for line-level handwritten text recognition on medieval historical manuscripts, consisting of 3,369 lines (images of text lines with the associated transcription and metadata) from 100 digitized documents written by at least 80 different hands and spanning three centuries (from 1208 to 1499). This dataset is derived from the AMSMB dataset, which contains the full-page images of the digitized manuscripts and their associated transcriptions in the PageXML format.… See the full description on the dataset page: https://huggingface.co/datasets/BSC-CSSH/AMSMB-line-transcription.landing-pages-styling-css-only-v2v3-merged
merged_v2_v3_dedup
This dataset is the deduplicated merged training set built from the repository's landing_page_v2 and landing_page_v3 pipelines.
It contains chat-format rows for CSS generation on landing pages:
messages + metadata
Each sample keeps:
messages: the training conversation, usually a user prompt plus the assistant CSS response.
metadata: generation and provenance fields such as source pipeline, style phrase, recipe, and other analysis attributes.… See the full description on the dataset page: https://huggingface.co/datasets/kogai/landing-pages-styling-css-only-v2v3-merged.HTML-CSS-UIcs_sqad-3.0
