datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CSS10-Multilingual-LJSpeech
CSS10-Multilingual-LJSpeech
Multilingual speech dataset combining LJSpeech (English) + CSS10 (10 languages) in a consistent LJSpeech format.
Dataset Description
This dataset merges:
LJSpeech: High-quality English speech dataset
CSS10: A collection of single-speaker speech datasets for 10 languages
All audio files are provided in a consistent format suitable for TTS training.
Features
Each sample contains:
audio: Waveform audio sampled at 22,050 Hz
text:… See the full description on the dataset page: https://huggingface.co/datasets/davidguzmanr/CSS10-Multilingual-LJSpeech.SciSciGPT-SciSciNetpreprocessed_jsut_jsss_css10_common_voice_11
Dataset Card for "preprocessed_jsut_jsss_css10_common_voice_11"
More Information needed
HTML-CSS-Website# Dataset
This dataset contains a collection of FacebookAds-related queries and responses generated by an AI assistant.
# Proudly Dataset Genrated with AI with [AI Dataset Generator API](https://api.example.com)
preprocessed_jsut_jsss_css10_fleurs_common_voice_11
Dataset Card for "preprocessed_jsut_jsss_css10_fleurs_common_voice_11"
More Information needed
sciscinet-v2
📢🚨📣 Sciscinet-v2
Sciscinet-v2 is a refreshed update to SciSciNet which is a large-scale, integrated dataset designed to support research in the science of science domain. It combines scientific publications with their network of relationships to funding sources, patents, citations, and institutional affiliations, creating a rich ecosystem for analyzing scientific productivity, impact, and innovation. Know more.
About Sciscinet-v2
The newer version Sciscinet-v2 is… See the full description on the dataset page: https://huggingface.co/datasets/Northwestern-CSSI/sciscinet-v2.preprocessed_jsut_jsss_css10
Dataset Card for "preprocessed_jsut_jsss_css10"
More Information needed
github-code-html-css-1github-code-html-css-2landing-pages-v2-cssgithub-code-html-cssgithub-code-html-css-split-3css-deepfake-datasethtml-css-codegen-datasetSciSciGPT-SciSciCorpuscsst2AMSMB-line-transcription
Dataset Card
Dataset for line-level handwritten text recognition on medieval historical manuscripts, consisting of 3,369 lines (images of text lines with the associated transcription and metadata) from 100 digitized documents written by at least 80 different hands and spanning three centuries (from 1208 to 1499). This dataset is derived from the AMSMB dataset, which contains the full-page images of the digitized manuscripts and their associated transcriptions in the PageXML format.… See the full description on the dataset page: https://huggingface.co/datasets/BSC-CSSH/AMSMB-line-transcription.landing-pages-styling-css-only-v2v3-merged
merged_v2_v3_dedup
This dataset is the deduplicated merged training set built from the repository's landing_page_v2 and landing_page_v3 pipelines.
It contains chat-format rows for CSS generation on landing pages:
messages + metadata
Each sample keeps:
messages: the training conversation, usually a user prompt plus the assistant CSS response.
metadata: generation and provenance fields such as source pipeline, style phrase, recipe, and other analysis attributes.… See the full description on the dataset page: https://huggingface.co/datasets/kogai/landing-pages-styling-css-only-v2v3-merged.cs_sqad-3.0css2-uq-mmlu-prodataset_CSSF12_552_en_qadataset_CSSF12_552_en_exqacss-bench
CSS-Bench: Counterfactual Strategic Synthesis Benchmark
CSS-Bench tests whether language models make strategic decisions based on the underlying payoff topology of a game, or on the semantic valence of the words used to narrate it. Every item exists as a matched pair (or triple): a canonical framing where the numerically optimal action is also lexically "nice," and a counterfactual framing with the identical payoff structure but inverted narrative valence -- the numerically… See the full description on the dataset page: https://huggingface.co/datasets/jub-aer/css-bench.HTML-CSS-projects# Dataset
This dataset contains a collection of FacebookAds-related queries and responses generated by an AI assistant.
# Proudly Dataset Genrated with AI with [AI Dataset Generator API](https://api.example.com)
dataset_1_circular_cssfbuzz_sources_092_cssHTML_CSS_CodeDataSet_100k_formatted_and_splitlab_reservation_django_codebase_html_cssHTML_CSS_CodeDataSet_100k_text_only_splitHTML_CSS_CodeDataSet_400_subset
