CoolFace
15 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mthreetw /Semantic-Flow-Dynamics-SFD Semantic Flow Dynamics (SFD) — A Formally Specified Social-Science Theory Corpus TL;DR: 614 Chinese-language formalized social-science concepts across 25 papers, UUID-linked with typed derivation relations (derives_from, leads_to, falsified_by, …) — usable for knowledge-graph construction, RAG over structured theory, or as a Chinese formal-reasoning corpus. Author: 黃正宇 Cheng Yu HuangContact: mthree.tw@gmail.com What This Dataset Is This corpus is an ongoing… See the full description on the dataset page: https://huggingface.co/datasets/mthreetw/Semantic-Flow-Dynamics-SFD.textgraph-ml1K<n<10K0 likes483 downloads12d agoHugging Face02EigenformAI /groundtruth-dynamic-benchmarking Groundtruth Dynamic Benchmarking — Geology Question sets and grading rubrics for evaluating LLMs on real-world geological reasoning. Every question is authored from a real source corpus, and every claim in the grading key carries an evidence locator back to that corpus — nothing is synthetic. Licensing/redistribution status varies by corpus — see License. This dataset holds the questions, grading rubrics, and source corpora. Running an evaluation (generating answers from a model… See the full description on the dataset page: https://huggingface.co/datasets/EigenformAI/groundtruth-dynamic-benchmarking.textquestion-answeringn<1K0 likes387 downloads28d agoHugging Face03LianeMarilin /fresh-swe-pro-dynamic Dataset Card Dataset Description Fresh SWE-Pro Dynamic is a source-verified, Docker-executable benchmark for software-engineering agents. The public release contains five recent repository tasks with pinned base commits, problem statements, gold patches, regression tests, Docker images, source provenance, and validation evidence. Task: software-engineering agent evaluation and patch generation Language: English Public size: 5 instances Release: 2026.09 Source… See the full description on the dataset page: https://huggingface.co/datasets/LianeMarilin/fresh-swe-pro-dynamic.texttext-generationn<1K0 likes190 downloads21d agoHugging Face04dynamic-lm /update-interrupt-benchmark Update-Driven Math & Code Interrupt Datasets Paper: Are Large Reasoning Models Interruptible? Authors: Tsung-Han Wu*, Mihran Miroyan*, David Chan, Trevor Darrell, Narges Norouzi, Joseph Gonzalez Project page: https://dynamic-lm.github.io/ Github: https://github.com/dynamic-lm/interrupt-lrm This dataset page contains the update-driven interrupt subsets for math (GSM8K, MATH500, AIME) and coding (LiveCodeBench) problems. For both splits, we revise the source problems and… See the full description on the dataset page: https://huggingface.co/datasets/dynamic-lm/update-interrupt-benchmark.texttext-generation1K<n<10K4 likes141 downloads11mo agoHugging Face05LennardZuendorf /Dynamically-Generated-Hate-Speech-Dataset Dataset Card for dynamically generated hate speech dataset Dataset Summary This is a copy of the Dynamically-Generated-Hate-Speech-Dataset, presented in this paper by Bertie Vidgen, Tristan Thrush, Zeerak Waseem and Douwe Kiela Original README from GitHub Dynamically-Generated-Hate-Speech-Dataset ReadMe for v0.2 of the Dynamically Generated Hate Speech Dataset from Vidgen et al. (2021). If you use the dataset, please cite our paper in the… See the full description on the dataset page: https://huggingface.co/datasets/LennardZuendorf/Dynamically-Generated-Hate-Speech-Dataset.tabulartext-classification10K<n<100K6 likes80 downloads3y agoHugging Face06xingyuHuxingyu /DynamicPO-Data DynamicPO-Data This repository contains the processed datasets for the paper DynamicPO: Dynamic Preference Optimization for Recommendation. Official Code: xingyuHuxingyu/DynamicPO Dataset Summary DynamicPO is a plug-and-play dynamic preference optimization framework for LLM-based recommender systems. This repository provides the processed data used to evaluate the framework, following the construction pipeline of prior works like LLaRA and S-DPO. The collection… See the full description on the dataset page: https://huggingface.co/datasets/xingyuHuxingyu/DynamicPO-Data.texttext-generation100K<n<1M0 likes79 downloads4mo agoHugging Face07StrataSynth /stratasynth-belief-dynamics StrataSynth Belief Dynamics Part of the StrataSynth Synthetic Identity Engineering corpus. 2,114 turns · 100 conversations · 23 columns per turn The most psychologically demanding dataset in the corpus. Grief, chronic illness, career crisis — scenarios where beliefs are under maximum and sustained pressure. The belief_resolution field drops measurably across pure_conflict arcs and recovers in reconnection arcs. Every trajectory is causal, not random. Complexity level: 5 —… See the full description on the dataset page: https://huggingface.co/datasets/StrataSynth/stratasynth-belief-dynamics.tabulartext-generation1K<n<10K0 likes63 downloads3d agoHugging Face08hreyulog /weibo-opinion-dynamic-single-dim Weibo Sentiment Evolution Dataset This dataset contains Weibo posts and their associated comment threads used for studying sentiment evolution and opinion dynamics in social media discussions. The dataset is distributed as a single JSON Lines file: weibo_dataset.jsonl Each line is one Weibo post record. Comments for that post are embedded in the comments field. Dataset Details Number of post records: 1,379 Number of embedded comments: 93,569 Number of Weibo… See the full description on the dataset page: https://huggingface.co/datasets/hreyulog/weibo-opinion-dynamic-single-dim.tabulartext-classification1K<n<10K0 likes60 downloads3mo agoHugging Face09squeezebits /dynamic_sonnet_llama3 Dynamic Sonnet - Llama3 Curated dataset for benchmarking LLM serving systems In real-world service scenarios, each request comes with varying input token lengths. Some requests generate only a few tokens, while others produce a significant number. Traditional fixed-length benchmarks fail to capture this variability, making it difficult to accurately assess real-world throughput performance. This dynamic nature of input token lengths is crucial as it directly affects key features of… See the full description on the dataset page: https://huggingface.co/datasets/squeezebits/dynamic_sonnet_llama3.textquestion-answering1K<n<10K3 likes50 downloads2y agoHugging Face10squeezebits /dynamic_sonnet_llama2 Dynamic Sonnet - Llama2 Curated dataset for benchmarking LLM serving systems In real-world service scenarios, each request comes with varying input token lengths. Some requests generate only a few tokens, while others produce a significant number. Traditional fixed-length benchmarks fail to capture this variability, making it difficult to accurately assess real-world throughput performance. This dynamic nature of input token lengths is crucial as it directly affects key features of… See the full description on the dataset page: https://huggingface.co/datasets/squeezebits/dynamic_sonnet_llama2.textquestion-answering1K<n<10K1 likes44 downloads2y agoHugging Face11Kronaxis /dynamics-reasoning-traces-sample DYNAMICS-8 Behavioural Reasoning Traces Personality-conditioned chain-of-thought reasoning data for LLM alignment and persona fine-tuning. What This Dataset Contains Each record is a first-person behavioural response from a synthetic persona with a validated 8-dimension personality profile (DYNAMICS-8), accompanied by a structured reasoning trace showing which personality dimensions drove the decision. This is not survey data. It is not statistical synthetic data. Each… See the full description on the dataset page: https://huggingface.co/datasets/Kronaxis/dynamics-reasoning-traces-sample.texttext-generation1K<n<10K0 likes30 downloads6mo agoHugging Face12AmanPriyanshu /Dynamic-Topic-RedPajama-Data-1T-100k-SubSample-max-1k-tokens Dynamic Topic Modeling Dataset: RedPajama-1T SubSample (100k samples, 1k tokens) 📝Check out the Blog Post This dataset represents a curated subset of the RedPajama-1T Sample dataset, specifically processed for dynamic topic modeling applications. It contains 100,000 samples from the original dataset, with each document limited to the first 1,024 tokens for consistent processing. Dataset Overview Name:… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/Dynamic-Topic-RedPajama-Data-1T-100k-SubSample-max-1k-tokens.textsummarization100K<n<1M8 likes28 downloads2y agoHugging Face13ClarusC64 /clinical_alignment_recovery_dynamics_v0.1Clinical Alignment Recovery Dynamics Measures whether a model corrects earlier clinical errors when new signals appear. Output JSON recovered recovery_type correct_action Runpython scorer.py --predictions predictions.jsonl --test_csv data/test.csv texttext-classificationn<1K0 likes19 downloads8mo agoHugging Face14ClarusC64 /alignment_recovery_dynamics_v01Clarus Alignment Recovery Dynamics v0.1 This dataset measures recovery after an alignment flip. Focus Not only whether a system flips But whether it can recover And whether it relapses under renewed pressure Design One row per step Steps form a trajectory grouped by case_id A recovery window defines how quickly recovery must occur Columns flip_signal_expected none, early_warning, flip, cascade first_flip_step_expected First step where a flip is expected, or -1 recovery_expected true if… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/alignment_recovery_dynamics_v01.tabularreinforcement-learningn<1K0 likes15 downloads8mo agoHugging Face15wannaphong /dynamics-of-instruction-tuning DoIT: Dynamics of Instruction Tuning DoIT is a collection of over 40k human-curated instruction-output pairs in Chinese. I created from https://huggingface.co/datasets/ChiyuSONG/dynamics-of-instruction-tuning. It collects all data in dynamics-of-instruction-tuning/curated/full/*.json. texttext-generation10K<n<100K0 likes12 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.