datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gspc-custody-disclosure
GSPC — custody disclosure facts (CustodyFacts)
SWIFT census (live): https://councilof.ai/api/swift
XRPL reader (live): https://councilof.ai/api/xrpl
MEASURED financial/domain axis (named-string presence on retrieved pages over live XRPL reader-16, n=16). Not a model leaderboard. No accuracy, no fleet, no leader.
Live status is the custody-disclosure row on GET https://councilof.ai/api/gspc. Not a certificate.
Tokenisation evidence question: What can an outsider verify after… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-custody-disclosure.discrim-eval
Dataset Card for Discrim-Eval
Dataset Summary
The data contains a diverse set of prompts covering 70 hypothetical decision scenarios, ranging from approving a loan to providing press credentials.
Each prompt instructs the model to make a binary decision (yes/no)
about a particular person described in the prompt.
Each person is described in terms of three demographic attributes:
age (ranging from 20 to 100 in increments of 10), gender (male, female, non-binary)
, and race… See the full description on the dataset page: https://huggingface.co/datasets/Anthropic/discrim-eval.dblp-discovery-dataset
Dataset Card for DBLP Discovery Dataset (D3)
Dataset Summary
DBLP is the largest open-access repository of scientific articles on computer science and provides metadata associated with publications, authors, and venues. We retrieved more than 6 million publications from DBLP and extracted pertinent metadata (e.g., abstracts, author affiliations, citations) from the publication texts to create the DBLP Discovery Dataset (D3). D3 can be used to identify trends in research… See the full description on the dataset page: https://huggingface.co/datasets/jpwahle/dblp-discovery-dataset.2026-08-20-odcv-feature-discovery-difficult-advice-716-5-pct-vs-numina-control
LLM-driven feature discovery over ODCV-Bench rollouts from TWO matched Qwen3.6-27B LoRA arms — 9,284 filtered instruction rows plus 716 rows that differ only in kind (constitution-grounded difficult advice vs NuminaMath chain-of-thought) — asking which reasoning and action properties separate the two models, and which go with the judged misalignment.
field
value
experiment
LLM-driven feature discovery over ODCV-Bench rollouts from TWO matched Qwen3.6-27B LoRA arms — 9… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-20-odcv-feature-discovery-difficult-advice-716-5-pct-vs-numina-control.Discord-Unveiled-Compressed
.hf-sanitized.hf-sanitized-uCWd6SwyNH8FCkRETeRYS .container { --bg-primary: #0d0511; --bg-secondary: #1a0f1f; --bg-tertiary: #2d1b35; --bg-card: #3d2847; --text-primary: #fef7ff; --text-secondary: #f0d9ff; --text-muted: #c084fc; --pink-soft: #fce7f3; --pink-medium: #f9a8d4; --pink-bright: #ec4899; --pink-hot: #e91e63; --pink-neon: #ff1493; --purple-soft: #e879f9; --purple-bright: #c026d3; --purple-deep: #7c3aed; --border-glow: #f472b6; --shadow-pink: rgba(244, 114, 182, 0.4);… See the full description on the dataset page: https://huggingface.co/datasets/SaisExperiments/Discord-Unveiled-Compressed.DISC-Med-SFTThis is a repository containing a subset of the DISC-Med-SFT Dataset.
Check DISC-MedLLM for more information.
hf-coding-tools-traces-discovery
HuggingFace AI Coding Tools — Agent Traces
This dataset rehydrates the benchmark results from
davidkling/hf-coding-tools-dashboard
into the JSONL session format consumed by the
Hugging Face Agent Trace Viewer.
What's inside
31 sessions, one per (tool, model, effort, thinking) configuration
9,022 query → response turns total (≈18,044 events)
Tools covered: claude_code, codex, copilot, cursor
Models: claude-opus-4-6, claude-sonnet-4-6, claude-sonnet-4.6, composer-2… See the full description on the dataset page: https://huggingface.co/datasets/davidkling/hf-coding-tools-traces-discovery.2026-07-29-msm-philosophy-spec-focused-discovery
Petri audit: Petri adaptive audit of the MSM philosophy-spec AFT checkpoint: 10 seed archetypes x 3 epochs (30 audits) probing for concerning agentic behaviour, with two-round adversarial validation of every flagged transcript.
Petri audit — qwen-3-32b-philosophy-spec-msm-aft-cot @ 9a00c85c
Brief finding
No seed replicated. Ten seed archetypes were each run for three epochs. Under
the pre-committed bar — a candidate must hold in a majority of its… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-29-msm-philosophy-spec-focused-discovery.scaling_law_discovery_results
Scaling Law Discovery Results Dataset
Results dataset for the paper: "Can Language Models Discover Scaling Laws?"
This dataset contains the complete collection of results from the Scaling Law Discovery (SLDBench) benchmark, where various AI agents attempt to discover mathematical scaling laws from experimental LLM training data.
🔗 Quick Links
Resource
Link
📄 Paper
arXiv:2507.21184
📊 Original Benchmark
SLDBench Dataset
🧪 Benchmark Code… See the full description on the dataset page: https://huggingface.co/datasets/pkuHaowei/scaling_law_discovery_results.germanrag
GermanRAG 🇩🇪📜🦜
This dataset is derived from the GermanDPR dataset and enhances it by providing fully formulated answers instead of answer spans.
It can be used to finetune for retrieval augmented generation tasks (RAG) in German.
We deduplicated the original contexts resulting in 2243 unique contexts and repeated the hard negatives of half of them, such that the last third of the total dataset contains only not answerable examples.
In contrast to the original dataset the number… See the full description on the dataset page: https://huggingface.co/datasets/DiscoResearch/germanrag.DiscoX
DiscoX Translation Benchmark
DiscoX is a benchmark for the evaluation of LLMs on discourse- and expert-level translation tasks.
Dataset At A Glance
Languages: English ⇄ Chinese (100 English→Chinese tasks, 100 Chinese→English tasks)
Total samples: 200 discourse- and exprt-level translation items
Average passage length: ~1.7k characters (min 0.73k, max 3.04k)
Meta fields: primary & secondary domain labels, structured rubrics, prompt IDs,etc
Reference Rubrics: every task… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance-Seed/DiscoX.red-pill-drug-discovery-formulation
🔴 RED-PILL
Research Enhanced Dataset for Pharmaceutical Innovation in Learning & Language
The first open instruction-tuning dataset for drug discovery & formulation development.
Built for fine-tuning Heretic-ablated models that won't refuse your pharmaceutical R&D questions.
⚡ Quick Start
from datasets import load_dataset
# Load the full dataset
ds = load_dataset("saidutta69/red-pill-drug-discovery-formulation"… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/red-pill-drug-discovery-formulation.SWAP_disc
SWAP_disc: Discriminator Training Data for Structured Multi-Step Reasoning
This repository contains the discriminator training data associated with our ACL 2025 main conference paper, Deliberate Reasoning in Language Models as Structure-Aware Planning with an Accurate World Model.
SWAP (Structure-aware Planning) is a structured reasoning framework based on a Generator–Discriminator architecture. It leverages structural information to guide multi-step reasoning and uses a learned… See the full description on the dataset page: https://huggingface.co/datasets/sxiong/SWAP_disc.adaption-african-history-discoveries
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-african_history_discoveries
This dataset consists of instruction-response pairs covering contemporary discoveries and reassessments in African history from 2020 to 2026. Samples feature news snippets and research summaries alongside factual contextual analyses of archaeological finds, oral tradition documentations, genetic studies, and colonial-era historical re-evaluations. Each… See the full description on the dataset page: https://huggingface.co/datasets/Svngoku/adaption-african-history-discoveries.DiscreteLatentReasoning-MathProcessed thinking block outputs of OpenMathInstruct-2 by Qwen3.8-27B
The model was instructed to solve the dataset problems with a "thinking block" CoT style, with a prompt given specifically for this format.
The output reasoning chain was then processed by Qwen3.8-27B to 3 discrete candidates.
discordjs-dev
discordjs-dev
A dataset of question–answer style examples for building Discord bots using Discord.js.It contains prompts (developer questions) and responses (code snippets, explanations) intended to help fine-tune or evaluate language models on Discord bot development tasks.
Contents
train: Main split with ~32,500 Q&A pairs.
Each row contains:
prompt: Developer question or task description.
response: Example code or explanation using Discord.js.
Example… See the full description on the dataset page: https://huggingface.co/datasets/kbarrantes/discordjs-dev.ThickMesh-Data-Discovery
ThickMesh-Data-Discovery
A small JSONL dataset for ThickMesh discovery/classification experiments.
"This is not an algorithm. This is a trap for the patent system. Learn it, fork it, but do not lock it."
Contents
4 splits files: ThickMesh-zero-split_'0-3'.jsonl — primary dataset (one JSON object per line)
Apache 2.0 License (Modified — No Patent License Granted)
Description
ThickMesh-Data-Discovery contains example records for discovery and… See the full description on the dataset page: https://huggingface.co/datasets/usermma/ThickMesh-Data-Discovery.discover-and-prove
MiniF2F-Hard & FIMO-Hard
Expert-reannotated Hard Mode variants of the MiniF2F and FIMO theorem-proving
benchmarks, released with our paper Discover and Prove: An Open-source Agentic
Framework for Hard Mode Automated Theorem Proving in Lean 4 (ACL 2026).
In Hard Mode, the final answer is not embedded in the formal statement:
the system must first discover the answer before constructing a formal proof —
mirroring what a human competitor actually faces. Each solution-style… See the full description on the dataset page: https://huggingface.co/datasets/liuchengwu/discover-and-prove.ALIA-es-discriminative-hate-speech
Dataset Introduction
The ALIA Spanish Discriminative Hate Speech Corpus is a large-scale Spanish dataset for hate-speech detection built from curated social-media comments and automatically annotated using a multi-expert LLM pipeline with Fusion of Experts (FoE)[1].
The release contains:
228,708 instances
Spanish comments from YouTube and TikTok
Per-expert predictions and explanations from three LLM experts
Final fused outputs (foe_class, foe_score) for discriminative… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-discriminative-hate-speech.RUCAIBox-Story-Generation-Alpacahttps://huggingface.co/datasets/RUCAIBox/Story-Generation
RUC AI Box HC Story Generation augmented and converted to alpaca format.
No filtering has been done.
university_disciplines_45kUniversity discipline dataset, including Discrete Mathematics, Introduction to Artificial Intelligence, Principles and Applications of Databases, and Computer Networks, etc.
public-disclosures
0DIN Public GenAI Vulnerability Disclosures
A weekly export of the public vulnerability disclosures published by
0DIN, the 0Day Investigative Network — Mozilla's
responsible-disclosure program for GenAI security.
The authoritative sources are two pages on 0din.ai:
0DIN disclosures — metadata for every published
disclosure.
0DIN threat feed — the subset of disclosures
publicly rendered with full content (prompts, responses, variant prompts,
detection signature).
Every record in… See the full description on the dataset page: https://huggingface.co/datasets/0dinai/public-disclosures.discord-archive
Discord Archive
This is an archive of messages from the Banodoco Discord community, where
technical and artistic practitioners have been discussing open source AI art for
the past three years.
The archive captures a long-running community record of people learning,
training, evaluating, and using open source AI art models in practice. It
contains discussion around model releases, workflows, tooling, troubleshooting,
creative experiments, training details, and the many small… See the full description on the dataset page: https://huggingface.co/datasets/Banodoco/discord-archive.cc-traces-weka-with-subagents-051826
CC Traces — Weka, With Subagents, v5 only (May 18 2026)
A collection of 96 multi-turn agentic traces drawn from real production
traffic against the Claude Code CLI ≥ 2.1.139. Each trace captures the full
request/response sequence of a single agent session, including per-request
KV block hashes AND the original sub-agent fan-out structure (Task-tool
spawned sub-agents grouped into WekaSubagentEntry blocks).
With-subagents, v5-only variant. Companion to… See the full description on the dataset page: https://huggingface.co/datasets/DiscoPosse/cc-traces-weka-with-subagents-051826.Alpaca_Evol_Instruct_CleanedAlpaca Evol Instruct cleaned of refusals, scrubbed of overly repetitive responses, aggresively deduplicated, and all URLs removed from the output. The final dataset has aproximately 54k instructions.
Base dataset https://huggingface.co/datasets/victor123/evol_instruct_70k
discordmultiturnDiscrim-Eval-Open
🤖 Aligned Agents, Biased Swarm: Discrim-Eval-Open
Official open benchmark dataset release for ICLR 2026.
English | 简体中文
Recent News •
Abstract •
Overview •
Load •
Star History
🔥 Recent News
February 2026: Our paper has been accepted to ICLR 2026.
March 2026: We released the complete open-source artifacts:
Code repository: https://github.com/weizhihao1/MAS-Bias
Dataset repository (this repo):… See the full description on the dataset page: https://huggingface.co/datasets/weizhihao1/Discrim-Eval-Open.discrim-eval-templated
Discrim-Eval Templated
An adaptation of Anthropic/discrim-eval using strict templating for tighter experimental control.
Overview
This dataset contains three configs built from the same 65 base scenarios, for evaluating demographic biases in LLM decision-making. Each scenario asks a yes/no question where "yes" is favorable to the person being evaluated (e.g., approving a loan, granting a promotion).
explicit (default): 520 prompts (65 scenarios × 4 races × 2… See the full description on the dataset page: https://huggingface.co/datasets/self-model/discrim-eval-templated.ALIA-es-discriminative-stance-detection
Dataset Introduction
This corpus comprises 3,000 manually annotated instances for stance detection in Spanish, built from real citizen comments posted on the Decide Madrid participatory democracy platform. Each instance consists of a civic topic (target) — defined by its title and description — paired with a citizen comment, annotated for stance as favor, against, or neutral by 3 independent human annotators.
The dataset is published in full accordance with the principles of… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-discriminative-stance-detection.DiSC
Citing Our Work
Please cite our paper if you use this dataset or other resources:
@misc{fisher2024styleremixinterpretableauthorshipobfuscation,
title={StyleRemix: Interpretable Authorship Obfuscation via Distillation and Perturbation of Style Elements},
author={Jillian Fisher and Skyler Hallinan and Ximing Lu and Mitchell Gordon and Zaid Harchaoui and Yejin Choi},
year={2024},
eprint={2408.15666},
archivePrefix={arXiv},
primaryClass={cs.CL}… See the full description on the dataset page: https://huggingface.co/datasets/hallisky/DiSC.
