datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CSS10-Multilingual-LJSpeech
CSS10-Multilingual-LJSpeech
Multilingual speech dataset combining LJSpeech (English) + CSS10 (10 languages) in a consistent LJSpeech format.
Dataset Description
This dataset merges:
LJSpeech: High-quality English speech dataset
CSS10: A collection of single-speaker speech datasets for 10 languages
All audio files are provided in a consistent format suitable for TTS training.
Features
Each sample contains:
audio: Waveform audio sampled at 22,050 Hz
text:… See the full description on the dataset page: https://huggingface.co/datasets/davidguzmanr/CSS10-Multilingual-LJSpeech.SciSciGPT-SciSciNetpreprocessed_jsut_jsss_css10_common_voice_11
Dataset Card for "preprocessed_jsut_jsss_css10_common_voice_11"
More Information needed
HTML-CSS-Website# Dataset
This dataset contains a collection of FacebookAds-related queries and responses generated by an AI assistant.
# Proudly Dataset Genrated with AI with [AI Dataset Generator API](https://api.example.com)
CSSR-S_labelled_suicidewatch_posts_reddit
Evaluating Reasoning LLMs for Suicide Screening with the Columbia-Suicide Severity Rating Scale
Full code and supplementary materials are available at https://github.com/av9ash/llm_cssrs_code.
License and Citation
This project is released under the Creative Commons Attribution 4.0 International (CC BY 4.0) license.Any use or reuse of this work please cite the following:
@article{patil2025evaluating,
title={Evaluating Reasoning LLMs for Suicide Screening with the… See the full description on the dataset page: https://huggingface.co/datasets/av9ash/CSSR-S_labelled_suicidewatch_posts_reddit.marin-starcoderdata_cssaudio_visual_starss23_sonpreprocessed_jsut_jsss_css10_fleurs_common_voice_11
Dataset Card for "preprocessed_jsut_jsss_css10_fleurs_common_voice_11"
More Information needed
CSSBench
CSSBench: A Safety Evaluation Benchmark for Chinese Lightweight Language Models
Overview
CSSBench (Chinese-Specific Safety Benchmark) is a comprehensive evaluation framework designed to assess the safety robustness of Chinese Large Language Models (LLMs), with a specific emphasis on lightweight models (≤8B parameters). The benchmark bridges a critical evaluation gap by targeting Chinese-specific adversarial patterns—linguistic obfuscations such as homophones and Pinyin… See the full description on the dataset page: https://huggingface.co/datasets/Yaesir06/CSSBench.audio-dataset-flickr-soundnetcommon_voice_large_jsut_jsss_css10
Dataset Card for vumichien/common_voice_large_jsut_jsss_css10
Stack2Graph_KG_css
CSS StackOverflow Knowledge Graph
Summary
This Hugging Face dataset repository contains the CSS shard of the Stack2Graph StackOverflow Knowledge Graph.
Hugging Face uses one dataset repository per programming language, so this repository is directly cloneable without an extra top-level archive wrapper.
The artifact is optimized for graph-based retrieval, SPARQL analytics, and retrieval-augmented question answering over Stack Overflow content.
Stack2Graph source:… See the full description on the dataset page: https://huggingface.co/datasets/Mo7art/Stack2Graph_KG_css.sciscinet-v2
📢🚨📣 Sciscinet-v2
Sciscinet-v2 is a refreshed update to SciSciNet which is a large-scale, integrated dataset designed to support research in the science of science domain. It combines scientific publications with their network of relationships to funding sources, patents, citations, and institutional affiliations, creating a rich ecosystem for analyzing scientific productivity, impact, and innovation. Know more.
About Sciscinet-v2
The newer version Sciscinet-v2 is… See the full description on the dataset page: https://huggingface.co/datasets/Northwestern-CSSI/sciscinet-v2.preprocessed_jsut_jsss_css10
Dataset Card for "preprocessed_jsut_jsss_css10"
More Information needed
libri_cssdota_v1.5github-code-html-css-1Thal-Kak_local_db
Thal-Kak local MSA, template databases
The sequence and template databases that the local MSA modes of
Thal-Kak search —
--msa mmseqs_local, --msa hhblits_local, --msa mmseqs_hhblits_local, and
local template search on any of them.
Install these with install_db.sh, not by hand.
Every file here is a multi-gigabyte .tar.zst holding a prebuilt MMseqs2 or
HH-suite database; the installer verifies it, unpacks it into place and
renames the files to the layout the pipeline expects.… See the full description on the dataset page: https://huggingface.co/datasets/cssbsnu/Thal-Kak_local_db.false-citation-bench
False Citation Bench
False Citation Bench is a compact evaluation and inspection dataset for false or misleading case citations in legal documents. It contains 26 source documents, their PDFs, and manually reviewed citation annotations grounded in the local text extraction.
Dataset contents
The repository has one matching document in each directory:
documents_txt/{index}__{case-name}__{filing}.txt
documents_pdf/{index}__{case-name}__{filing}.pdf… See the full description on the dataset page: https://huggingface.co/datasets/gt-csse/false-citation-bench.css10-ljspeech-multilingual
CSS10 + LJSpeech Multilingual Dataset
A unified multilingual speech dataset combining CSS10 (10 languages) and LJSpeech (English) in a consistent LJSpeech format.
Dataset Description
This dataset merges:
CSS10: A collection of single-speaker speech datasets for 10 languages
LJSpeech: High-quality English speech dataset (Linda Johnson)
All audio files are provided in a consistent format suitable for TTS training.
Languages and Statistics
Language
Code… See the full description on the dataset page: https://huggingface.co/datasets/ayousanz/css10-ljspeech-multilingual.github-code-html-css-2landing-pages-v2-csscs_squad-3.0
Dataset Card for Czech Simple Question Answering Dataset 3.0
This a processed and filtered adaptation of an existing dataset. For raw and larger dataset, see Dataset Source section.
Dataset Description
The data contains questions and answers based on Czech wikipeadia articles.
Each question has an answer (or more) and a selected part of the context as the evidence.
A majority of the answers are extractive - i.e. they are present in the context in the exact form. The… See the full description on the dataset page: https://huggingface.co/datasets/fewshot-goes-multilingual/cs_squad-3.0.Sci2Pol-BenchSci2Pol-Bench
Data, scripts, and recipes for the benchmark Sci2Pol-Bench, a comprehensive benchmark for evaluating large language models.
About •
Usage•
Authors
About
The data consists of policy briefs obtained from Nature Energy, Nature Climate, Nature Cities, and Journal of Health and Social Behavior Policy Briefs.
Policy briefs originally were introduced in the Nature Energy journal with the goal of:
This format aims to provide… See the full description on the dataset page: https://huggingface.co/datasets/Northwestern-CSSI/Sci2Pol-Bench.css10-ja-ljspeech
CSS100-LJSpeech (Japanese / Meian)
css100-ljspeech は、Park et al. が公開した CSS10 日本語コーパス(明暗)を、LJ Speech 互換フォーマット (id|text & wavs/*.wav) へ変換した派生データセットです。
データ概要
項目
値
話者
1 (ekzemplaro)
音声数
6,841
合計時間
約 15 時間
サンプリングレート
22,050 Hz
テキスト言語
日本語
フォーマット
``id
ファイル構成
css100-ljspeech/
├── metadata.csv # 2 列 (id|text)
└── wavs/
├── meian_0000.wav
├── meian_0001.wav
└── ...
使用例 (🤗 Datasets)
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/ayousanz/css10-ja-ljspeech.MultiRoundConvos-Code-JS-HTML-CSS-Pythongithub-code-html-cssStack2Graph_VD_css
CSS StackOverflow Vector Dataset
Summary
This Hugging Face dataset repository contains the CSS shard of the Stack2Graph vector-database component as restorable Qdrant artifacts plus portable Parquet fallback files.
Hugging Face uses one dataset repository per programming language, so this repository is directly cloneable without an extra top-level archive wrapper.
The artifacts are intended for semantic and hybrid retrieval, graph entry-point finding, and… See the full description on the dataset page: https://huggingface.co/datasets/Mo7art/Stack2Graph_VD_css.github-code-html-css-split-3css10-ljspeech
CSS10-LJSpeech
CSS10-LJSpeech は、Park et al. が公開した CSS10 データセットを、LJSpeech互換フォーマットに変換した10言語の音声合成用データセットです。各言語の文学作品を音声化した高品質な音声データを提供し、LJSpeechフォーマット(id|text & wavs/*.wav)に統一されています。
データ概要
項目
値
話者数
10 (言語別)
総音声数
64,196
合計時間
約 140 時間
サンプリングレート
22,050 Hz
音声フォーマット
IEEE浮動小数点 (32bit)
テキスト言語
10言語
フォーマット
`id
言語別統計
言語
言語コード
音声数
合計時間
ドイツ語
de
7,428
16.14時間
ギリシャ語
el
1,844
4.14時間
スペイン語
es
11,016
19.15時間
フィンランド語
fi
4,842
10.53時間… See the full description on the dataset page: https://huggingface.co/datasets/ayousanz/css10-ljspeech.
