CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01dlab-spp /corpus-1T-manifest SPP Corpus 1T Manifest The selection manifest for the ~1.0T-token pretraining corpus used in Synthetic Persona Pretraining (SPP): Alignment from Token Zero. The corpus is a seeded subsample of allenai/dolma3_mix-6T. Rather than redistribute ~2.6 TB of text that is already public, this dataset publishes the selection decisions keyed by upstream document id, so the corpus can be reconstructed exactly by replaying against upstream. 📄 Reflections + text for the annotated half:… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/corpus-1T-manifest.tabulartext-generation1B<n<10B0 likes1.5k downloads1mo agoHugging Face02dlab-spp /reflection-50m SPP Reflection 50M The 51.4M-document reflection set from Synthetic Persona Pretraining (SPP): Alignment from Token Zero — the production half-corpus run, and the dataset the released models were actually trained on. 🔬 Small sample (same format): dlab-spp/reflection-sample-2k 📉 Earlier 10M run: dlab-spp/reflection-10m 🧾 Safety scores for the full 1T corpus: dlab-spp/safety-classifications Each row pairs a source document with two generated constitution reflections — a… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/reflection-50m.tabulartext-generation10M<n<100M0 likes1.3k downloads1mo agoHugging Face03dlab-spp /reflection-10m SPP Reflection 10M The full ~10M-document reflection set from Synthetic Persona Pretraining (SPP): Alignment from Token Zero. 📝 Read the post: Synthetic Persona Pretraining: Alignment from Token Zero 🔬 Small sample (same format): dlab-spp/reflection-sample-2k — a 2,000-row sample drawn from this set, for quick inspection. Each row pairs a pretraining document with a synthetic, value-laden reflection generated for it: a short first-person (and third-person) moral reflection… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/reflection-10m.tabulartext-generation1M<n<10M0 likes345 downloads1mo agoHugging Face04ExponentialScience /DLT-Tweets DLT-Tweets [Paper] • [Code] Dataset Description Dataset Summary DLT-Tweets is a large-scale corpus of social media posts related to Distributed Ledger Technology (DLT). This dataset is part of the larger DLT-Corpus collection, designed to support NLP research, social computing studies, and public discourse analysis in the DLT domain. It was introduced in the paper DLT-Corpus: A Large-Scale Text Collection for the Distributed Ledger Technology Domain.… See the full description on the dataset page: https://huggingface.co/datasets/ExponentialScience/DLT-Tweets.tabulartext-generation10M<n<100M0 likes192 downloads7mo agoHugging Face05twistshan /Sudoku_DLM_Reasoning Sudoku DLM Reasoning stwistzz/Sudoku_DLM_Reasoning is a deterministic 9x9 Sudoku benchmark for studying masked diffusion language models, depth, and iterative decoding. Version 1.1.0 trains only on original, r0, and r1_4; r5_19 is held out for adjacent difficulty extrapolation. Version 1.1.0 changes only the deterministic selection seed to 0 relative to v1.0.2. The raw CSV files, sources, bucket definitions, split quotas, schema, and validation rules are unchanged; the four… See the full description on the dataset page: https://huggingface.co/datasets/twistshan/Sudoku_DLM_Reasoning.tabulartext-generation100K<n<1M0 likes108 downloads1mo agoHugging Face06ucberkeley-dlab /fragility-moral-judgment-llms Fragility of Moral Judgment in Large Language Models Companion dataset for the FAccT paper Fragility of Moral Judgment in Large Language Models by Tom van Nuenen. Contains the moral dilemmas, community labels, and per-model verdicts (with explanations and reasoning traces) used in the study. The paper investigates how stable LLM moral judgments are under minimal, morally-irrelevant perturbations of the same dilemma, and whether protocols and reasoning chains improve or worsen… See the full description on the dataset page: https://huggingface.co/datasets/ucberkeley-dlab/fragility-moral-judgment-llms.tabulartext-classification100K<n<1M0 likes53 downloads4mo agoHugging Face07ExponentialScience /DLT-Scientific-Literature DLT-Scientific-Literature Paper | GitHub Dataset Description Dataset Summary DLT-Scientific-Literature is a specialized corpus of academic publications focused on Distributed Ledger Technology (DLT). This dataset is part of the larger DLT-Corpus collection, designed to support NLP research, language model development, and innovation studies in the DLT domain. The dataset contains 37,440 scientific documents with 564 million tokens, spanning publications from… See the full description on the dataset page: https://huggingface.co/datasets/ExponentialScience/DLT-Scientific-Literature.tabulartext-generation10K<n<100K0 likes31 downloads4mo agoHugging Face08dlab-spp /reflection-sample-2k SPP Reflection 2k Sample A 2,000-row sample (seed 42) of dlab-spp/reflection-10m, in the identical format, for quick inspection of the data from Synthetic Persona Pretraining (SPP): Alignment from Token Zero. 📝 Read the post: Synthetic Persona Pretraining: Alignment from Token Zero 📦 Full dataset: dlab-spp/reflection-10m (~10M documents). Each row pairs a pretraining document with a synthetic, value-laden reflection (first- and third-person) grounded in a value constitution.… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/reflection-sample-2k.tabulartext-generation1K<n<10K0 likes27 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.