CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ucberkeley-dlab /measuring-hate-speech Dataset card for Measuring Hate Speech This is a public release of the dataset described in Kennedy et al. (2020) and Sachdeva et al. (2022), consisting of 39,565 comments annotated by 7,912 annotators, for 135,556 combined rows. The primary outcome variable is the "hate speech score" but the 10 constituent ordinal labels (sentiment, (dis)respect, insult, humiliation, inferior status, violence, dehumanization, genocide, attack/defense, hate speech benchmark) can also be treated as… See the full description on the dataset page: https://huggingface.co/datasets/ucberkeley-dlab/measuring-hate-speech.tabulartext-classification100K<n<1M55 likes2.1k downloads9mo agoHugging Face02Abhishek-A0 /dlgenai-nppe-datasettabularn<1K0 likes1.7k downloads11h agoHugging Face03dlab-spp /corpus-1T-manifest SPP Corpus 1T Manifest The selection manifest for the ~1.0T-token pretraining corpus used in Synthetic Persona Pretraining (SPP): Alignment from Token Zero. The corpus is a seeded subsample of allenai/dolma3_mix-6T. Rather than redistribute ~2.6 TB of text that is already public, this dataset publishes the selection decisions keyed by upstream document id, so the corpus can be reconstructed exactly by replaying against upstream. 📄 Reflections + text for the annotated half:… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/corpus-1T-manifest.tabulartext-generation1B<n<10B0 likes1.5k downloads1mo agoHugging Face04dlab-spp /reflection-50m SPP Reflection 50M The 51.4M-document reflection set from Synthetic Persona Pretraining (SPP): Alignment from Token Zero — the production half-corpus run, and the dataset the released models were actually trained on. 🔬 Small sample (same format): dlab-spp/reflection-sample-2k 📉 Earlier 10M run: dlab-spp/reflection-10m 🧾 Safety scores for the full 1T corpus: dlab-spp/safety-classifications Each row pairs a source document with two generated constitution reflections — a… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/reflection-50m.tabulartext-generation10M<n<100M0 likes1.3k downloads1mo agoHugging Face0522f2001542 /dlgenai-nppe2-datasettabularn<1K0 likes595 downloads6d agoHugging Face06AIML-TUDA /dlam-ts-project-data-2026 operations_forecasting_2026 Multivariate hourly forecasting for anonymized operations units. Target Predict the future hourly operational load index for each series_id. Higher values indicate more operational pressure in that unit. Forecast Contract Frequency: h Series: 96 Timesteps per series: 4992 Target column: target Training history length used by the baseline templates: 168 Rollout block length: 24 Required prediction horizon: validation: 336, test: 336… See the full description on the dataset page: https://huggingface.co/datasets/AIML-TUDA/dlam-ts-project-data-2026.tabular100K<n<1M2 likes478 downloads5mo agoHugging Face07mohd3rfan /DLD_Transactionstabular1M<n<10M0 likes417 downloads5mo agoHugging Face08IPEC-COMMUNITY /dlr_edan_shared_control_lerobotThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "dlr_edan", "total_episodes": 104, "total_frames": 8928, "total_tasks": 10, "total_videos": 104, "total_chunks": 1, "chunks_size": 1000, "fps": 5, "splits": { "train": "0:104" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/IPEC-COMMUNITY/dlr_edan_shared_control_lerobot.tabularrobotics1K<n<10K0 likes404 downloads2y agoHugging Face09lerobot /dlr_edan_shared_controlThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "unknown", "total_episodes": 104, "total_frames": 8928, "total_tasks": 14, "total_videos": 104, "total_chunks": 1, "chunks_size": 1000, "fps": 5, "splits": { "train": "0:104" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/dlr_edan_shared_control.tabularrobotics1K<n<10K0 likes359 downloads1y agoHugging Face10lerobot /dlr_sara_grid_clampThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "unknown", "total_episodes": 107, "total_frames": 7622, "total_tasks": 1, "total_videos": 107, "total_chunks": 1, "chunks_size": 1000, "fps": 5, "splits": { "train": "0:107" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/dlr_sara_grid_clamp.tabularrobotics1K<n<10K0 likes350 downloads1y agoHugging Face11dlab-spp /reflection-10m SPP Reflection 10M The full ~10M-document reflection set from Synthetic Persona Pretraining (SPP): Alignment from Token Zero. 📝 Read the post: Synthetic Persona Pretraining: Alignment from Token Zero 🔬 Small sample (same format): dlab-spp/reflection-sample-2k — a 2,000-row sample drawn from this set, for quick inspection. Each row pairs a pretraining document with a synthetic, value-laden reflection generated for it: a short first-person (and third-person) moral reflection… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/reflection-10m.tabulartext-generation1M<n<10M0 likes345 downloads1mo agoHugging Face12lerobot /dlr_sara_pourThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "unknown", "total_episodes": 100, "total_frames": 12971, "total_tasks": 1, "total_videos": 100, "total_chunks": 1, "chunks_size": 1000, "fps": 5, "splits": { "train": "0:100" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/dlr_sara_pour.tabularrobotics10K<n<100K0 likes344 downloads1y agoHugging Face13gabrielaltay /tcga-dlbc-tabular-open TCGA-DLBC — Tabular (Open Access) Open-access TCGA-DLBC data from the NCI Genomic Data Commons, reshaped into one table per GDC data_type. Clinical, biospecimen and every open molecular modality for this cohort, in one place, queryable without downloading a single .tar or parsing a single TSV. GDC data release: Data Release 46.0 - August 10, 2026 Built: 2026-09-12 03:53:50 UTC Scope: one TCGA project — see [the family][repo] for the others from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/gabrielaltay/tcga-dlbc-tabular-open.tabular10M<n<100M1 likes284 downloads14d agoHugging Face14JakeOh /proj-dllm-sfttabular1M<n<10M0 likes283 downloads9mo agoHugging Face15dlb /plue PLUE Repository: https://github.com/ju-resplande/PLUE Paper: Leaderboard: Point of Contact: Portuguese translation of the GLUE benchmark, SNLI, and Scitail using OPUS-MT model and Google Cloud Translation. The language data in PLUE is Brazilian Portuguese (BCP-47 pt-BR) Citation Information @misc{Gomes2020, author = {GOMES, J. R. S.}, title = {PLUE: Portuguese Language Understanding Evaluation}, year = {2020}, publisher = {GitHub}, journal = {GitHub… See the full description on the dataset page: https://huggingface.co/datasets/dlb/plue.tabulartext-classification100K<n<1M7 likes263 downloads1y agoHugging Face16ucberkeley-dlab /interaction_protocol Interaction Protocol Shapes Moral Judgment in Multi-Agent Debate This repository contains the processed experimental datasets used in: Interaction Protocol Shapes Moral Judgment in Multi-Agent Debate. Pratik S. Sachdeva and Tom van Nuenen. COLM 2026. Dataset contents The experiments/ directory contains Parquet datasets used to reproduce the figures and analyses in the paper. It includes: synchronous head-to-head debates; round-robin head-to-head debates;… See the full description on the dataset page: https://huggingface.co/datasets/ucberkeley-dlab/interaction_protocol.tabular10K<n<100K1 likes260 downloads1mo agoHugging Face17jingyq1 /dllm-img-edit-vq-cachetabular100K<n<1M0 likes250 downloads5mo agoHugging Face18dleon23 /dual_so101This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "dual_so101_follower", "total_episodes": 2, "total_frames": 582, "total_tasks": 1, "total_videos": 6, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:2" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/dleon23/dual_so101.tabularrobotics10K<n<100K1 likes193 downloads1y agoHugging Face19ExponentialScience /DLT-Tweets DLT-Tweets [Paper] • [Code] Dataset Description Dataset Summary DLT-Tweets is a large-scale corpus of social media posts related to Distributed Ledger Technology (DLT). This dataset is part of the larger DLT-Corpus collection, designed to support NLP research, social computing studies, and public discourse analysis in the DLT domain. It was introduced in the paper DLT-Corpus: A Large-Scale Text Collection for the Distributed Ledger Technology Domain.… See the full description on the dataset page: https://huggingface.co/datasets/ExponentialScience/DLT-Tweets.tabulartext-generation10M<n<100M0 likes192 downloads7mo agoHugging Face20DataStore /meteora-dlmm-historical-data Meteora DLMM Historical Data Decoded Solana mainnet instructions and events from Meteora DLMM (Dynamic Liquidity Market Maker), a concentrated-liquidity DEX where liquidity sits in discrete price bins and the fee rate rises with volatility. 74 tables, 59,575 rows, one row per decoded instruction or event. Program ID LBUZKhRxPF3XUpBCjp4YzTKgLccjZhTSDM9YuVaPwxo. This is a free sample from datastore.sh, which publishes the complete history as versioned Parquet. Read… See the full description on the dataset page: https://huggingface.co/datasets/DataStore/meteora-dlmm-historical-data.tabular10K<n<100K0 likes188 downloads11d agoHugging Face21omrisap /dl-trm-phase2-codebooks DL-TRM Phase 2 Codebooks This dataset repository contains Phase 2 transition VQ codebook artifacts for DL-TRM. Contents are organized by vocabulary size: V16/ V32/ V128/ V256/ Each folder includes: z_traces.pt: discrete Z trace dataset for the Phase 1 route traces codebook.pt: learned VQ codebook weights transition_vq_model.pt: trained transition VQ model weights diagnostics.json: code usage and final training diagnostics z_trace_manifest.json: artifact manifest checkpoints/:… See the full description on the dataset page: https://huggingface.co/datasets/omrisap/dl-trm-phase2-codebooks.tabular1K<n<10K0 likes176 downloads4mo agoHugging Face22satyanshi431 /dl-npppe2-datasettabularn<1K0 likes175 downloads3d agoHugging Face23wlyu /dl3dv_InP_480tabular1K<n<10K0 likes165 downloads5mo agoHugging Face24dl4phys /top_tagging Dataset Card for Top Quark Tagging Dataset Summary Top Quark Tagging is a dataset of Monte Carlo simulated events produced by proton-proton collisions at the Large Hadron Collider. The top-quark signal and mixed quark-gluon background jets are produced with Pythia8 with its default tune for a center-of-mass energy of 14 TeV. Multiple interactions and pile-up are ignored. The leading 200 jet constituent four-momenta (E,px,py,pz) (E, p_x, p_y, p_z) (E,px​,py​,pz​)are stored… See the full description on the dataset page: https://huggingface.co/datasets/dl4phys/top_tagging.tabular1M<n<10M1 likes160 downloads4y agoHugging Face25dleemiller /wiki-sim Wiki Sim Overview This new semi-synthetic dataset is derived from wikimedia/wikipedia. Each row contains 1-3 references sentences extracted from the original dataset. For each reference sentence, we use an optimized DSPy program to generate 4 similar sentences: Synonym (Replace words with synonyms to maintain the same meaning.) Paraphrase (Rephrase the sentence using a different structure while keeping the same idea.) Conceptual Overlap (Express a related concept… See the full description on the dataset page: https://huggingface.co/datasets/dleemiller/wiki-sim.tabularsentence-similarity1M<n<10M0 likes141 downloads1y agoHugging Face26dlthub /read_test_202608110744310438tabular10K<n<100K0 likes140 downloads2mo agoHugging Face27dleemiller /FineCat-NLI Fine Concatenation (FineCat) NLI Overview A common criticism of SNLI and MNLI datasets is that there are too many 'easy' samples. This tends to overfit to simple / trivial patterns that don't generalize well. In order to combat this, I concatenated 7 datasets (2.6M samples), then ran a training test for 50k steps with ModernBERT-large in cross-encoder configuration. I found that 1 dataset (~100k samples) did not have good compatibility with the labels of the others… See the full description on the dataset page: https://huggingface.co/datasets/dleemiller/FineCat-NLI.tabularfeature-extraction1M<n<10M5 likes137 downloads11mo agoHugging Face28alibustami /UM-DLP-Public-Benchmarking-Dataset UM DLP Public Benchmarking Dataset Description The UM DLP Public Benchmarking Dataset is a publicly available collection designed specifically to stress test Data Loss Prevention (DLP) systems, helping identify detection gaps, false positives, and false negatives for ongoing improvement. This benchmark dataset contains 1,343 manually validated records across six major categories relevant to financial and sensitive data risks: Financial Data (Account information about… See the full description on the dataset page: https://huggingface.co/datasets/alibustami/UM-DLP-Public-Benchmarking-Dataset.tabulartext-classification1K<n<10K4 likes124 downloads1y agoHugging Face29logo-lab /trl-dlte TRL-DLTE Paper: arXiv:2606.09323 — TRL-Bench: Standardizing Cross-Paradigm Representation-Level Evaluation of Tabular Encoders · Code: LOGO-CUHKSZ/TRL-Bench Compositional Data-Lake Table Enrichment suite of TRL-Bench. A 47,772-table data lake derived from 1,379 TabFact and WikiTableQuestions parent tables, fragmented at four cumulative noise tiers (clean / schema / cell / hard). Each parent yields a seed query, a union target (additional rows), and a join target (additional… See the full description on the dataset page: https://huggingface.co/datasets/logo-lab/trl-dlte.tabular100K<n<1M1 likes120 downloads4mo agoHugging Face30dl1wjd2 /pick_and_placeThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "omx_follower", "total_episodes": 40, "total_frames": 25782, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:40" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/dl1wjd2/pick_and_place.tabularrobotics10K<n<100K0 likes118 downloads26d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.