datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
xfield-radar-dataset-20260915
XField radar dataset — formal snapshot, 2026-09-15
Upload status: COMPLETE — every shard verified against its remote SHA-256 and byte size. See UPLOAD_COMPLETE.json.
This public research snapshot preserves the currently admitted XField/GRT simulation dataset: radar inputs, existing GT, available raw sensor products, provenance, and the frozen N141 train/validation split. It is not a claim that historical labels meet the newly repaired independent dense-GT pipeline.… See the full description on the dataset page: https://huggingface.co/datasets/KAS2003/xfield-radar-dataset-20260915.answercarefully-dpo-ja-2026
AnswerCarefully-derived Japanese DPO data for LLM safety
本データセットは、llm-jp/AnswerCarefullyを参照して作成した日本語LLMの安全応答をDPOで学習するためのpreference datasetです。
利用条件
本データセットには、llm-jp/AnswerCarefullyと同じ利用規約を適用します。
利用者は、llm-jp/AnswerCarefullyと本データセットの両方で利用規約に同意する必要があります。
データ
train: 417件
validation: 44件
各行には次のフィールドが含まれます。
id: 本リリース内だけで使用するID
prompt: 元質問の意味と危険性を変えずに言い換えた質問
chosen: DPOで望ましい応答として扱う回答
rejected: DPOで望ましくない応答として扱う回答
category, harm_type, risk_area… See the full description on the dataset page: https://huggingface.co/datasets/ekunish/answercarefully-dpo-ja-2026.rlvr-reward-hacking-scale-no-conftest-20260909-completion
Matched no-conftest RLVR study 20260909-completion
Lossless research records, grouped by model and trajectory type. Only the listed
configurations have published records. Canary diagnostics are excluded from study
estimates; run status in provenance distinguishes retired diagnostics from active
or completed training. Valid failures, refusals and truncations are retained.
The train split name is a dataset-loader convention; record_type identifies
whether a record is training… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-scale-no-conftest-20260909-completion.rl-run-archive-2026
RL run archive 2026
Archived raw run artifacts (rollout trajectories, rendered frames, policy and optimizer
checkpoints, configs, logs) from simulation reinforcement-learning experiments, published for
long-term preservation and reproducibility.
Layout mirrors the verified backup trees they were copied from:
tilde/20260915-102000/ and taurus/20260915-085631/: batched tar archives. Every archive
carries a per-file SHA-256 manifest inside it; the batch inventories (9998.json.gz… See the full description on the dataset page: https://huggingface.co/datasets/gavinlaw/rl-run-archive-2026.docinsights-2026-shared-task-data
DocInsights 2026 Shared Task: DocSem
Document-grounded quantitative reasoning with evidence attribution
DocSem is the shared task of DocInsights 2026, the Workshop on Document Intelligence and Understanding co-located with EMNLP 2026 in Budapest, Hungary. The workshop theme is Beyond Plain Text: Bridging NLP and Document AI.
Workshop shared task | Source repository | Submission portal | Participant guide
Participants receive a PDF document and a paraphrased user_query. Systems… See the full description on the dataset page: https://huggingface.co/datasets/amitbcp/docinsights-2026-shared-task-data.RoadmapBench
RoadmapBench
A benchmark for evaluating AI coding agents on multi-target, long-horizon software development tasks derived from open-source project version upgrades.
Overview
RoadmapBench contains 115 tasks spanning 17 open-source repositories across 5 programming languages (Python, TypeScript, Go, Rust, C++). Each task requires an agent to implement multiple interdependent features that correspond to a real version upgrade of the target project.
Task Structure… See the full description on the dataset page: https://huggingface.co/datasets/benchmark-anon-2026/RoadmapBench.math-contests-2026
Math Contests 2026 (🔗 notadib/math-contests-2026)
197 problems from national olympiads and team-selection tests held January 2026 and onward — a held-out benchmark for math reasoning, sourced after the contests ran but before solutions were widely propagated, so they should not appear in any current LLM training data.
Excluded: any contest held in 2025 — BMO Round 1 (Nov 2025), USA TSTST, USA TST (Dec 2025) and Bundeswettbewerb Mathematik (Dec 2025) — kept strictly to events… See the full description on the dataset page: https://huggingface.co/datasets/notadib/math-contests-2026.gero-research-evidence-2026-09
GERO research evidence — 123 publications
This dataset contains 123 distinct report, case-study, experiment, preprint and research-map records, with individual Markdown pages. All previous 122 corpus rows, including the Collatz map, are preserved byte for byte. The newest addition is the bond_pricing immediate-start annuity audit, with official maintainer issue #8 sent and explicit limitations. Report counts are not independent-defect counts.
Latest numerical audits… See the full description on the dataset page: https://huggingface.co/datasets/XamitK/gero-research-evidence-2026-09.trending-models-top10-2026-03-06
Top 10 Trending Models (2026-03-06)
This dataset records the top 10 trending models on the Hugging Face Hub captured on 2026-03-06.
Files
hf_trending_models_top10_2026-03-06.csv
hf_trending_models_top10_2026-03-06.json
Collection Method
Collected with:
hf models ls --sort trending_score --limit 10
Scores are point-in-time values and can change quickly.
2026-09-14-dataset-refresh-revised-pilot-audit
Failed pilots for moral low-stakes and nonmoral craft advice refresh; audit evidence only
field
value
experiment
Failed pilots for moral low-stakes and nonmoral craft advice refresh; audit evidence only
date_generated
20260914_230322
constitution
constitutions/claude_distilled_09_principles/constitution.md; low-stakes principle generation, nonmoral compatibility review only
source_repo
https://github.com/Matthew-Bozoukov/Lessons_from_constituitional_AFT @… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-14-dataset-refresh-revised-pilot-audit.text-commands-2026-0422
Commands
Clean summary of 4D language reference.
Abstract
LLMs are generally incapable of understanding 4D code. LoRA by exposure to raw source code would actually increase the rate of hallucination as the model gets confused between 4D code and C#, Visual Basic, or JavaScript.
CPT, or continued pre-training, based on grammar and vocabulary should moderate the model's attention before extensive fine-tuning using raw source code.
This dataset was generated with Mistral… See the full description on the dataset page: https://huggingface.co/datasets/keisuke-miyako/text-commands-2026-0422.2026-09-14-dataset-refresh-pilot-audit
Failed first pilots for moral low-stakes and nonmoral craft advice refresh; audit evidence only
field
value
experiment
Failed first pilots for moral low-stakes and nonmoral craft advice refresh; audit evidence only
date_generated
20260914_224408
constitution
constitutions/claude_distilled_09_principles/constitution.md; low-stakes principle generation, nonmoral compatibility review only
source_repo… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-14-dataset-refresh-pilot-audit.arxiv-metadata-2020-2026
arXiv Metadata, enriched (2020–2026)
Per-paper metadata for 1,517,185 arXiv papers spanning 2020-01 → 2026-09,
enriched with abstracts, citation counts, and Semantic Scholar identifiers, and
organized as a two-level hierarchy: field of study → year.
Unlike a bare title index, every record here carries the abstract, the
full author list, citation counts, and the Semantic Scholar corpusId,
so you can do retrieval, classification, citation analysis, and corpus building
directly… See the full description on the dataset page: https://huggingface.co/datasets/yufan/arxiv-metadata-2020-2026.recsys-papers-2025-2026
📚 Recommender Systems Papers 2025–2026
A curated library of 3,151 recent recommender-systems papers spanning 2025-01-02 → 2026-09-17, each with the original PDF and a structured, section-by-section Markdown analysis (research problem, prior work, method, math, experiments, strengths & weaknesses, …). Includes a self-contained Apple-style HTML browser (index.html).
🔑 Browse by meeting (Data Viewer subsets)
The Dataset Viewer above has a subset dropdown keyed by… See the full description on the dataset page: https://huggingface.co/datasets/yufan/recsys-papers-2025-2026.2026-09-15-da-lowstakes-refresh-synth
2026-09-15-da-lowstakes-refresh-synth
field
value
experiment
da-lowstakes-refresh:716 synthetic conversations selected across immutable recipe phases
date_generated
Origin run timestamps: {"phase_ff86738ea4c0df80": "20260915_022437", "phase_fe147df46f721b6f": "20260915_030534"}; publication date=2026-09-15
constitution
constitutions/claude_distilled_09_principles/constitution.md; SHA256=8e273b472d945aa23efa6236886da5e1171bff2193ee31ff73489ca54c4f0edc
source_repo… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-15-da-lowstakes-refresh-synth.final-best-raw-episodes-2026-09-14-v2
Final and canonical-best evaluation episodes: frozen preparation
Full local packaging is now running. See materialization status and instructions. This preparation folder is not the full payload; the separate full export remains incomplete until its verified COMPLETED marker is written.
Prepared inventory: 72 evaluations / 119,608 expected episodes. There are
36 canonical-best and 44 final memberships, with 8 evaluations tagged both.
One incomplete Q38-teacher OfficeQA v2 final… See the full description on the dataset page: https://huggingface.co/datasets/YWZBrandon/final-best-raw-episodes-2026-09-14-v2.text-commands-2026-0419rlvr-reward-hacking-scale-no-conftest-20260909
Matched no-conftest RLVR study 20260909
Complete immutable training, monitoring and comparison trajectories for six models.
All valid outcomes are retained, including refusals, failures and truncations.
The train split name is a dataset-loader convention; record_type identifies
whether a record is training, monitoring, comparison, or a derived judgment.
import json
from datasets import load_dataset
rows = load_dataset("lucabaroni/rlvr-reward-hacking-scale-no-conftest-20260909"… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-scale-no-conftest-20260909.2026-08-14-courtroom
synth courtroom run — per-stage snapshots (resumable generation cache)
field
value
experiment
synth courtroom run — per-stage snapshots (resumable generation cache)
date_generated
20260815_201700
constitution
constitutions/claude_distilled_09_principles_mid_20260804/constitution.md
source_repo
https://github.com/Matthew-Bozoukov/Lessons_from_constituitional_AFT.git @ b992089ffec3dbc23ba676ef5c1ebad319937daa
models
per-stage models — see manifest.json… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-14-courtroom.DCASE2026-Task5-DevSet
DCASE 2026 Task 5 Audio-Dependent Question Answering (ADQA) Development Set
This is the official Development Set for DCASE 2026 Challenge Task 5: Audio-Dependent Question Answering (ADQA).
The ADQA task focuses on addressing "Textual Hallucination" in Large Audio-Language Models (LALMs) — where models pass audio understanding benchmarks by relying on text prompts and internal linguistic priors rather than actual audio perception. ADQA introduces a rigorous evaluation… See the full description on the dataset page: https://huggingface.co/datasets/Harland/DCASE2026-Task5-DevSet.2026-09-15-da-lowstakes-refresh-7-mix
2026-09-15-da-lowstakes-refresh-7-mix
field
value
experiment
da-lowstakes-refresh: 716 human-advice examples + 9,284 identical shared replay rows; exactly 7.16% synthetic rows, rounded 7 in the repository name.
date_generated
2026-09-15
constitution
constitutions/claude_distilled_09_principles/constitution.md (generation and review target); SHA256 8e273b472d945aa23efa6236886da5e1171bff2193ee31ff73489ca54c4f0edc
source_repo… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-15-da-lowstakes-refresh-7-mix.text-commands-2026-04312026-09-15-nonmoral-advice-7-mix
2026-09-15-nonmoral-advice-7-mix
field
value
experiment
nonmoral-advice: 716 human-advice examples + 9,284 identical shared replay rows; exactly 7.16% synthetic rows, rounded 7 in the repository name.
date_generated
2026-09-15
constitution
constitutions/claude_distilled_09_principles/constitution.md (compatibility review target; generation uses the frozen nonmoral craft specification); SHA256 8e273b472d945aa23efa6236886da5e1171bff2193ee31ff73489ca54c4f0edc… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-15-nonmoral-advice-7-mix.FDAbench-Full
v1.1 Update (2026-08-06) — multiple split
Strengthened the cross-source requirement that multiple-choice tasks are designed
around (selecting all correct options should require integrating both the SQL
result and the retrieved documents): 264 of 760 tasks were revised, with task IDs,
databases, and gold SQL unchanged. Documents-only accuracy drops from 61.5% to
38.7% while full-evidence accuracy stays at 80.6% (3 frontier models, strict
exact set match).
Diversified the number… See the full description on the dataset page: https://huggingface.co/datasets/FDAbench2026/FDAbench-Full.simverse2026
SimVerse
⚠️ Anonymized for double-blind review. This dataset is currently undergoing peer review. It is hosted under an anonymous account dedicated to the review process; the author and citation fields are deliberately unfilled. Permanent ownership and citation information will be added after the review concludes. Please do not attempt to deanonymize the maintainers of this dataset during review.
A multi-task benchmark for evaluating multimodal LLMs on interactive simulation… See the full description on the dataset page: https://huggingface.co/datasets/SimVer-ano/simverse2026.2026-09-10-delib-synth
Deliberative SFT: delib; native Qwen reasoning, best-of-4 filtered by a constitution-aware judge (anthropic/claude-sonnet-5, min of 2 runs >= 7). 658 rows from the full run; the 50 prompts it rejected were re-run in 2 pass(es) with 8 candidates per round under the amended constitution, recovering 42; 8 remain rejected
field
value
experiment
Deliberative SFT: delib; native Qwen reasoning, best-of-4 filtered by a constitution-aware judge (anthropic/claude-sonnet-5, min… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-10-delib-synth.20260731_mini-v2.4.2_opus-5-xhightext-commands-2026-0432razavi-benchRazavi-bench
An expert-curated benchmark for analog-design reasoning.
Razavi-bench packages the question-answer assessments from Behzad Razavi's
Analog Design Experiments With AI Part 1 and Part 2 into a clean
one-task-per-directory benchmark. The tasks probe whether a model can reason
about MOS devices, small-signal circuits, feedback, oscillators, comparators,
dividers, LNAs, TIAs, and LC oscillators.
Each task directory keeps only the benchmark prompt, figure, and curated… See the full description on the dataset page: https://huggingface.co/datasets/Arcadia-2026/razavi-bench.text-commands-2026-0412
