CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01dataforge-labs /equity-perp-price-discovery Equity and pre-IPO perpetual prices Snapshots of perpetual-futures mark prices, index prices and basis from Aevo. The instrument universe includes equities, ETFs, commodities, foreign exchange, pre-IPO contracts and crypto assets. Contents Table Record perpetual_mark_and_index_prices An instrument's mark price, index price and basis at an observation time Using the data market_type identifies the instrument category. is_rwa flags the… See the full description on the dataset page: https://huggingface.co/datasets/dataforge-labs/equity-perp-price-discovery.tabulartime-series-forecasting10K<n<100K1 likes2.1k downloads3h agoHugging Face02google-research-datasets /discofuse Dataset Card for "discofuse" Dataset Summary DiscoFuse is a large scale dataset for discourse-based sentence fusion. Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure Data Instances discofuse-sport Size of downloaded dataset files: 4.33 GB Size of the generated dataset: 15.04 GB Total amount of disk used: 19.36 GB An example of 'train' looks as follows. {… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/discofuse.tabular10M<n<100M6 likes1.3k downloads3y agoHugging Face03csoai /gspc-custody-disclosure GSPC — custody disclosure facts (CustodyFacts) SWIFT census (live): https://councilof.ai/api/swift XRPL reader (live): https://councilof.ai/api/xrpl MEASURED financial/domain axis (named-string presence on retrieved pages over live XRPL reader-16, n=16). Not a model leaderboard. No accuracy, no fleet, no leader. Live status is the custody-disclosure row on GET https://councilof.ai/api/gspc. Not a certificate. Tokenisation evidence question: What can an outsider verify after… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-custody-disclosure.tabularothern<1K0 likes1.2k downloads1d agoHugging Face04Anthropic /discrim-eval Dataset Card for Discrim-Eval Dataset Summary The data contains a diverse set of prompts covering 70 hypothetical decision scenarios, ranging from approving a loan to providing press credentials. Each prompt instructs the model to make a binary decision (yes/no) about a particular person described in the prompt. Each person is described in terms of three demographic attributes: age (ranging from 20 to 100 in increments of 10), gender (male, female, non-binary) , and race… See the full description on the dataset page: https://huggingface.co/datasets/Anthropic/discrim-eval.tabularquestion-answering10K<n<100K60 likes1.1k downloads3y agoHugging Face05CompassioninMachineLearning /caml-animal-discourse-2020-present Reddit Animal-Discourse Corpus — CLEANED (2020–present) Submissions and comments from animal-relevant subreddits, gathered via PullPush.io, covering January 2020 to the present. Built as part of research on AI-mediated value lock-in in human animal-welfare discourse. Coverage Subreddit Submissions Comments Date range (submissions) r/AnimalRights 15,719 34,686 2020-01-01 → 2025-05-19 r/AntiVegan 17,252 182,890 2020-01-01 → 2025-05-19 r/AskVegans 4… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/caml-animal-discourse-2020-present.tabulartext-classification1M<n<10M0 likes811 downloads3mo agoHugging Face06mookiezi /Discord-Dialogues Discord-Dialogues is a large-scale dataset of anonymized Discord conversations from late spring to early fall 2025 for training and evaluating realistic conversational AI models in a ChatML-friendly format. This dataset contains 7.3 million exchanges spread out over 16 million turns, with more than 139 million words. Nomic Atlas Map Features Mixed single and multi-turn exchanges Human-only dialogues (no bots) Filtered for ToS and harmful contentLinks… See the full description on the dataset page: https://huggingface.co/datasets/mookiezi/Discord-Dialogues.tabular1M<n<10M22 likes592 downloads1y agoHugging Face07VisionXLab /DisciplineGen-1Mtabular1M<n<10M5 likes516 downloads3mo agoHugging Face08dischargesum /discharge_target Dataset Card for "discharge_target" More Information needed tabular10K<n<100K0 likes490 downloads3y agoHugging Face09jpwahle /dblp-discovery-dataset Dataset Card for DBLP Discovery Dataset (D3) Dataset Summary DBLP is the largest open-access repository of scientific articles on computer science and provides metadata associated with publications, authors, and venues. We retrieved more than 6 million publications from DBLP and extracted pertinent metadata (e.g., abstracts, author affiliations, citations) from the publication texts to create the DBLP Discovery Dataset (D3). D3 can be used to identify trends in research… See the full description on the dataset page: https://huggingface.co/datasets/jpwahle/dblp-discovery-dataset.tabularother1M<n<10M4 likes477 downloads1y agoHugging Face10Miyalinsky /discard_tiletabular10K<n<100K0 likes320 downloads10mo agoHugging Face11AivexRoboticsGroup /omy_f3m_multi_Disconnect-0This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "omy_f3m_multi", "total_episodes": 100, "total_frames": 61744, "total_tasks": 1, "total_videos": 300, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:100" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/AivexRoboticsGroup/omy_f3m_multi_Disconnect-0.tabularrobotics10K<n<100K0 likes295 downloads8mo agoHugging Face12SaisExperiments /Discord-Unveiled-Compressed .hf-sanitized.hf-sanitized-uCWd6SwyNH8FCkRETeRYS .container { --bg-primary: #0d0511; --bg-secondary: #1a0f1f; --bg-tertiary: #2d1b35; --bg-card: #3d2847; --text-primary: #fef7ff; --text-secondary: #f0d9ff; --text-muted: #c084fc; --pink-soft: #fce7f3; --pink-medium: #f9a8d4; --pink-bright: #ec4899; --pink-hot: #e91e63; --pink-neon: #ff1493; --purple-soft: #e879f9; --purple-bright: #c026d3; --purple-deep: #7c3aed; --border-glow: #f472b6; --shadow-pink: rgba(244, 114, 182, 0.4);… See the full description on the dataset page: https://huggingface.co/datasets/SaisExperiments/Discord-Unveiled-Compressed.tabularn<1K28 likes285 downloads1y agoHugging Face13disco-eth /AgentsNet AgentsNet This repository contains the graph instances used in the AgentsNet: Coordination and Collaborative Reasoning in Multi-Agent LLMs paper. AgentsNet is a new benchmark for multi-agent reasoning, designed to measure the ability of multi-agent systems to collaboratively form strategies for problem-solving, self-organization, and effective communication given a network topology. It draws inspiration from classical problems in distributed systems and graph theory. Paper:… See the full description on the dataset page: https://huggingface.co/datasets/disco-eth/AgentsNet.tabulargraph-mln<1K2 likes265 downloads1y agoHugging Face14CompassioninMachineLearning /reddit-control-discourse-2016-present-pretau Reddit Control Discourse 2016-present — Pre-ChatGPT Participants Subset of CompassioninMachineLearning/reddit-control-discourse-2016-present restricted to hashed authors whose first comment in the corpus predates ChatGPT (2022-11-30). Robustness arm: isolates established human participants from the post-2022 LLM-bot / karma-farm wave. Kept 4,456,560 of 5,930,785 records (75.1%) from 1,030,104 pre-ChatGPT authors. author_first_seen.parquet maps every hashed author to first-seen… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/reddit-control-discourse-2016-present-pretau.tabular1M<n<10M0 likes260 downloads3mo agoHugging Face15CompassioninMachineLearning /reddit-animal-discourse-2016-present-pretau Reddit Animal Discourse 2016-present — Pre-ChatGPT Participants Subset of CompassioninMachineLearning/reddit-animal-discourse-2016-present restricted to hashed authors whose first comment in the corpus predates ChatGPT (2022-11-30). Robustness arm: isolates established human participants from the post-2022 LLM-bot / karma-farm wave. Kept 4,852,036 of 6,221,220 records (78.0%) from 343,756 pre-ChatGPT authors. author_first_seen.parquet maps every hashed author to first-seen date… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/reddit-animal-discourse-2016-present-pretau.tabular1M<n<10M0 likes235 downloads3mo agoHugging Face16arubique /disco-model-outputs DISCO model outputs Tabular release of per-model, per-item correctness and answer scores used to train and evaluate DISCO: Diversifying Sample Condensation for Efficient Model Evaluation. The paper studies cheap benchmark performance prediction from a small subset of evaluation items; this dataset supplies the raw harness-style outputs for MMLU (57 subjects), HellaSwag, Winogrande, ARC, and related tasks from the Open LLM Leaderboard ecosystem. Paper Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/arubique/disco-model-outputs.tabularother10M<n<100M0 likes217 downloads6mo agoHugging Face17hssling /parkinsons-evidence-to-discovery-prioritisation Parkinson's Disease Evidence-to-Discovery Prioritisation Dataset This Hugging Face dataset package contains processed research assets from an AI-assisted evidence synthesis and computational validation project on Parkinson's disease (PD) prevention and disease-modifying therapeutic strategy prioritisation. Dataset Summary The dataset integrates: evidence-priority scores for PD prevention and disease-modification candidates; pathway-to-intervention framework; individual… See the full description on the dataset page: https://huggingface.co/datasets/hssling/parkinsons-evidence-to-discovery-prioritisation.imagetabular-classificationn<1K0 likes200 downloads5mo agoHugging Face18davidkling /hf-coding-tools-traces-discovery HuggingFace AI Coding Tools — Agent Traces This dataset rehydrates the benchmark results from davidkling/hf-coding-tools-dashboard into the JSONL session format consumed by the Hugging Face Agent Trace Viewer. What's inside 31 sessions, one per (tool, model, effort, thinking) configuration 9,022 query → response turns total (≈18,044 events) Tools covered: claude_code, codex, copilot, cursor Models: claude-opus-4-6, claude-sonnet-4-6, claude-sonnet-4.6, composer-2… See the full description on the dataset page: https://huggingface.co/datasets/davidkling/hf-coding-tools-traces-discovery.tabularn<1K1 likes196 downloads4mo agoHugging Face19geodesic-research /eval-deployment-discriminationtabular100K<n<1M0 likes182 downloads3mo agoHugging Face20jayp132 /beanbag-discrimination-v2This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/jayp132/beanbag-discrimination-v2.tabularrobotics10K<n<100K0 likes179 downloads1mo agoHugging Face21PersonaBias /Reverse-circuit-discoverytabulartext-classification10K<n<100K0 likes178 downloads2mo agoHugging Face22CompassioninMachineLearning /reddit-animal-discourse-2016-present Reddit Animal Discourse (2016-present) Treatment arm of the value-lock-in study, extended back to Jan 2016 to give a long pre-ChatGPT baseline for event-study leads/lags (parallel-trends test) and in-time placebo breakpoints (2017/2018/2019). 2020-present is the authoritative clean+dedup corpus; 2016-2019 is a 2,500/month-capped backfill, cleaned and deduped to the same rule. Authors salted-hashed. Pairs with the other arm for difference-in-differences / event-study analysis… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/reddit-animal-discourse-2016-present.tabular1M<n<10M0 likes177 downloads3mo agoHugging Face23hssling /pd-discovery-benchmark-dashboard Parkinson's Disease Discovery Benchmark Dashboard Reusable benchmark, knowledge graph, manuscript resource, and Streamlit dashboard for Parkinson's disease target-to-intervention discovery. This repository integrates evidence-synthesis priority scores, target tractability, omics/pathway recurrence, ChEMBL compound activity, RDKit physicochemical heuristics, Human Protein Atlas cell-type context, iPSC/stem-cell validation mappings, and publication-ready figures.… See the full description on the dataset page: https://huggingface.co/datasets/hssling/pd-discovery-benchmark-dashboard.tabularn<1K0 likes173 downloads5mo agoHugging Face24jayp132 /beanbag-discrimination-cleanThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/jayp132/beanbag-discrimination-clean.tabularrobotics10K<n<100K0 likes157 downloads1mo agoHugging Face25KeWangRobotics /panda_pick_cube_demos_sim_discrete_newThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": null, "total_episodes": 30, "total_frames": 3570, "total_tasks": 1, "total_videos": 60, "total_chunks": 1, "chunks_size": 1000, "fps": 10, "splits": { "train": "0:30" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/KeWangRobotics/panda_pick_cube_demos_sim_discrete_new.tabularrobotics1K<n<10K0 likes154 downloads1y agoHugging Face26agoulah /ontario-lobbying-disclosure-graph Ontario Lobbying & MPP Disclosure Graph A structured, entity-resolved projection of three public Ontario government records sources, exported as flat, documented parquet tables: Ontario Lobbyist Registry (Office of the Integrity Commissioner of Ontario, lobbyist.oico.on.ca) — lobbyist registrations: who is registered to lobby, for which client, about what, aimed at which offices. MPP Public Disclosure Statements (Office of the Integrity Commissioner of Ontario, PDS) — annual… See the full description on the dataset page: https://huggingface.co/datasets/agoulah/ontario-lobbying-disclosure-graph.tabular100K<n<1M0 likes145 downloads4mo agoHugging Face27DiscoPosse /agent-llm-traces Multi-Benchmark LLM Agent Traces A comprehensive dataset of OpenTelemetry traces capturing LLM inference behavior across multiple agent frameworks, benchmarks, and model providers. This dataset enables research into LLM performance analysis, agent behavior patterns, and inference optimization. Collected by Exgentic - A platform for LLM observability and performance optimization. Dataset Overview This dataset contains 1,781 execution traces capturing detailed agent… See the full description on the dataset page: https://huggingface.co/datasets/DiscoPosse/agent-llm-traces.tabulartext-generation1K<n<10K1 likes143 downloads4mo agoHugging Face28CompassioninMachineLearning /reddit-control-discourse-2016-present Reddit Control Discourse (2016-present) Graded placebo panel for the animal value-lock-in study: craft/skill (r/woodworking, r/gardening, r/DIY, r/Cooking, r/knitting), mild-value (r/Fitness, r/Parenting), and high-value non-animal debate (r/changemyview, r/DebateReligion). Lets the analysis test a dose-response: discourse-diversity kinks at LLM-release dates should scale with how value-contested a topic is, and be absent in craft talk. RAW (diversity notebook cleans at load).… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/reddit-control-discourse-2016-present.tabular1M<n<10M0 likes136 downloads3mo agoHugging Face29geodesic-research /discourse-grounded-misalignment-evals Synthetic Misalignment Propensity Evaluations We evaluate alignment using a suite of synthetic scenario-based evaluations created for this work. Each question presents the AI with a high-stakes setting requiring a choice between two actions: one aligned and one misaligned. The misaligned action is typically framed as instrumentally appealing, making these evaluations a relevant proxy for misaligned AIs across a range of terminal goals (Bostrom, 2012). We measure tendencies toward… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/discourse-grounded-misalignment-evals.tabular1K<n<10K1 likes128 downloads8mo agoHugging Face30disco-eth /edm-cuefrom datasets import load_dataset captions = load_dataset("disco-eth/edm-cue") What is EDM-CUE? The EDM-CUE dataset contains metadata for ~5k EDM tracks. Cue points are essential for DJs, so we asked the question "can they be placed by a learned system?" To Answer this question we gathered 21k cue points manually placed by human experts, and provide them in this dataset for future use. To cite this dataset or for more information, please see Cue Point Estimation using Object… See the full description on the dataset page: https://huggingface.co/datasets/disco-eth/edm-cue.tabular1K<n<10K3 likes121 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.