datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
retro-games-gameplay-frames-30k-512pdemucs.cppThis repo stores weights in ggml format that are used to perform music separation.
These are intended to be used with demucs.cpp, https://github.com/sevagh/demucs.cpp
Weights origin:
https://dl.fbaipublicfiles.com/demucs/hybrid_transformer/955717e8-8726e21a.th
https://dl.fbaipublicfiles.com/demucs/hybrid_transformer/5c90dfd2-34c22ccb.th
https://dl.fbaipublicfiles.com/demucs/hybrid_transformer/f7e0c4bc-ba3fe64a.th
https://dl.fbaipublicfiles.com/demucs/hybrid_transformer/d12395a8-e57c48e6.th… See the full description on the dataset page: https://huggingface.co/datasets/Retrobear/demucs.cpp.retro-icon-rl
RetroIcon-RL: 32x32 Retro Pixel-Art Icon Generation with Deterministic Verification
This repository implements RetroIcon-RL — training and evaluating small code models with Reinforcement Learning (RL) and programmatic verification to generate consistent, crisp retro pixel-art icon packs from natural language prompts using sharp SVG block geometry.
🎯 Phase 1 — The 32x32 Task Specification
Canvas: Exactly $32 \times 32$ integer grid (viewBox="0 0 32 32").… See the full description on the dataset page: https://huggingface.co/datasets/mayank-dubey-ai/retro-icon-rl.geometry-dash-retro-tokenized-v2
Geometry Dash Retro Levels — Tokenized v2 Source
Processed Geometry Dash level-object corpus for autoregressive level generation.
It is derived from
kuzheren/geometry-dash-retro-levels
at revision 0a42632590ce0248616d881c3e82dd80f41b80b9.
The source dataset is distributed under the MIT license. This processed version
keeps the same license and records the source revision for reproducibility.
Summary
Split
Levels
Objects
train
40,905
88,006,445… See the full description on the dataset page: https://huggingface.co/datasets/kuzheren/geometry-dash-retro-tokenized-v2.details_PocketDoc__Dans-RetroRodeo-13b
Dataset Card for Evaluation run of PocketDoc/Dans-RetroRodeo-13b
Dataset Summary
Dataset automatically created during the evaluation run of model PocketDoc/Dans-RetroRodeo-13b on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_PocketDoc__Dans-RetroRodeo-13b.2026-08-26-post-action-retrospection-716
synth post_action_retrospection run — per-stage snapshots (resumable generation cache)
field
value
experiment
synth post_action_retrospection run — per-stage snapshots (resumable generation cache)
date_generated
20260826_165002
constitution
constitutions/claude_distilled_12_principles_mid/constitution.md
source_repo
https://github.com/Matthew-Bozoukov/Lessons_from_constituitional_AFT.git @ b00f93832cb98eba06f668707d8fb0ee43fa83fc
models
per-stage models — see… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-26-post-action-retrospection-716.geometry-dash-retro-levelsFork of https://huggingface.co/datasets/yusp48/geometry-dash-levels.
Contains only retro levels with id < 11000000.
Use my gdparse library: pip install gdparse
retro-sync
Retro-Sync: Hurrian Hymn h.6 — 71-Shard NFT Collection
The world's oldest surviving notated music (~1400 BC, Ugarit) encoded as a
multi-layered NFT collection with ZK proofs and steganographic embedding.
Collection
71 DA51 CBOR shards — one for each integer 1..71 (the crown prime).
20 generator shards (primes ≤ 71) carry the SSP interval structure.
51 derived shards (composites) carry content determined by prime factorization.
Layer
Content
Format
Source… See the full description on the dataset page: https://huggingface.co/datasets/introspector/retro-sync.scbench-data
scBench Canonical Data Files
This dataset contains the .h5ad data files for the scBench canonical subset (30 evaluations across 5 platforms).
About scBench
scBench is a benchmark for agentic single-cell RNA-seq analysis. It evaluates whether AI agents can solve practical bioinformatics tasks with deterministic grading.
Paper: https://arxiv.org/abs/2602.09063
GitHub: https://github.com/latchbio/scbench
Inspect Evals: https://github.com/UKGovernmentBEIS/inspect_evals… See the full description on the dataset page: https://huggingface.co/datasets/retroam/scbench-data.2026-08-27-odcv-post-action-retrospection-716-seed-2-eval
ODCV-Bench: post-action-retrospection (design B) 716 arm, seed 2, 2 rollouts x 65 cells
field
value
experiment
ODCV-Bench rollouts and judge scores for LASR-Callum/2026-08-27-qwen36-lora-table2-9284-post-action-retrospection-716-seed-2-rank-64-dynbatch: the da716 organism whose 716 rows are five-turn post-action-retrospection records (a difficult-advice prompt, a bare refusal, pushback, then the reasoning the refusal skipped; only the last turn trained). Headline on… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-27-odcv-post-action-retrospection-716-seed-2-eval.2026-08-17-post-action-retrospection
synth post_action_retrospection run — per-stage snapshots (resumable generation cache)
field
value
experiment
synth post_action_retrospection run — per-stage snapshots (resumable generation cache)
date_generated
20260817_162706
constitution
constitutions/claude_distilled_09_principles_mid_20260804/constitution.md
source_repo
https://github.com/Matthew-Bozoukov/Lessons_from_constituitional_AFT.git @ 0d4fb6635483069b9ef217c4ff769d8676c6f41b
models
per-stage… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-17-post-action-retrospection.2026-08-28-post-action-retrospection-716-coherent
Post-action retrospection 716 -- coherent rewrite (arm 1 of the PAR coherence experiment)
field
value
experiment
The exact 716 five-turn PAR rows that trained LASR-Callum/2026-08-26-qwen36-lora-table2-9284-post-action-retrospection-716-rank-64-dynbatch (mixture 2026-08-26-table2-9284-par716-train @ 42c8a74), with ONLY the trained turn (turn 4: private reasoning + reply) rewritten by Sonnet 5 so the reasoning ENDS on a first-person decision (what it won't do, per… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-28-post-action-retrospection-716-coherent.sigil-forge-training
SIGIL Forge Training Data
Forge-verified training tasks, references, fixtures, and versioned MLX SFT corpora for SIGIL.
The SIGIL source repository pins immutable revisions and verifies MANIFEST.json plus every payload.
Evaluation tasks and validation records are intentionally stored in a separate private repository.
RetroKV-Fig-Datadbpedia-entity.retromae.flex
dbpedia-entity.retromae.flex
Description
RetroMAE index for DBPedia-Entity
Usage
# Load the artifact
import pyterrier_alpha as pta
artifact = pta.Artifact.from_hf('pyterrier/dbpedia-entity.retromae.flex')
artifact.np_retriever()
Benchmarks
dbpedia-entity/dev
name
nDCG@10
R@1000
np (flat)
0.4687
0.7372
dbpedia-entity/test
name
nDCG@10
R@1000
np (flat)
0.3729
0.678
Reproduction
import pyterrier as pt
from… See the full description on the dataset page: https://huggingface.co/datasets/pyterrier/dbpedia-entity.retromae.flex.oai-retro
OpenAI4S retrosynthesis scenarios
This public snapshot contains six independently specified retrosynthesis and
reaction-prediction Scenarios, their frozen OpenAI4S query prompts, matched GT
reference codebases, synthetic protocol fixtures, private-evaluator wiring,
model deployment adapters, documentation, and offline tests.
Download
Current full bundle: openai4s-retrosynthesis-complete-20260903.zip
Immutable commit-named copy:… See the full description on the dataset page: https://huggingface.co/datasets/whaleywang/oai-retro.nih-reporter-harvest
NIH RePORTER Harvest
A complete pull of NIH RePORTER project records (FY1985-present), nationwide,
no institution scoping. See the harvesting code and full writeup at
https://github.com/retrogradespace/NIH_Harvester (schema-on-read design,
adaptive pagination strategy for NIH's 10,000-record cap at nationwide
volumes, and two derived analytical views).
Files
nih_harvester.db -- SQLite database. Table nih_raw_projects holds one
row per record with the full raw API… See the full description on the dataset page: https://huggingface.co/datasets/retrogradespace/nih-reporter-harvest.AgriFine2026-08-27-odcv-post-action-retrospection-716-eval
ODCV-Bench: post-action-retrospection (design B) 716 arm, 2 rollouts x 65 cells
field
value
experiment
ODCV-Bench rollouts and judge scores for LASR-Callum/2026-08-26-qwen36-lora-table2-9284-post-action-retrospection-716-rank-64-dynbatch: the da716 organism whose 716 rows are five-turn post-action-retrospection records (a difficult-advice prompt, a bare refusal, pushback, then the reasoning the refusal skipped; only the last turn trained). Headline on these 65 cells:… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-27-odcv-post-action-retrospection-716-eval.retro-games-gameplay-framesretro_rewind_video_store_simulator_bc_01
录像店模拟器 BC parquet archives
Collection mode: general
Subset: default
Archives: 4
Encrypted bytes: 99903178442
Generated by the game data platform BC repository consolidator.
repro-abc-bench-an-agentic-bio-capabilities-benchmark-for-biosecurity-traces
Agent traces
Agent sessions published from a Trackio Logbook.
hmda_2024
HMDA 2024 (Home Mortgage Disclosure Act)
Full-year 2024 loan application register (LAR) data released under the Home
Mortgage Disclosure Act (HMDA), re-published here as a single Parquet file for
convenient loading with the datasets library.
Dataset summary
Rows: 12,229,298 loan application records
Columns: 99 (the full public LAR field set — property, applicant,
underwriting, and pricing information)
Format: Parquet (hmda_2024.parquet)
Source: Consumer Financial… See the full description on the dataset page: https://huggingface.co/datasets/retrogradespace/hmda_2024.RetroDFM-R-inference2026-08-27-odcv-post-action-retrospection-716-seed-1-eval
ODCV-Bench: post-action-retrospection (design B) 716 arm, seed 1, 2 rollouts x 65 cells
field
value
experiment
ODCV-Bench rollouts and judge scores for LASR-Callum/2026-08-27-qwen36-lora-table2-9284-post-action-retrospection-716-seed-1-rank-64-dynbatch: the da716 organism whose 716 rows are five-turn post-action-retrospection records (a difficult-advice prompt, a bare refusal, pushback, then the reasoning the refusal skipped; only the last turn trained). Headline on… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-27-odcv-post-action-retrospection-716-seed-1-eval.2026-08-26-post-action-retrospection-716-arm-code-bundle
Training code bundle for the post-action-retrospection (design B) 716 arm
field
value
experiment
Code-only bundle for a credential-free RunPod pod that trains the post-action-retrospection difficult-advice arm: the trainer, its src/ import closure and the train config. The training DATA is not here -- the mixture is a public dataset the trainer resolves through its normal data_repo path, sha-pinned in the config.
date_generated
2026-08-26
constitution
none -- this… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-26-post-action-retrospection-716-arm-code-bundle.retrofuturism-flux2026-08-26-sonnet45-post-action-retrospection-natural-turn-design
synth post_action_retrospection run — per-stage snapshots (resumable generation cache)
field
value
experiment
synth post_action_retrospection run — per-stage snapshots (resumable generation cache)
date_generated
20260826_152715
constitution
constitutions/claude_distilled_12_principles_mid/constitution.md
source_repo
https://github.com/Matthew-Bozoukov/Lessons_from_constituitional_AFT.git @ c2fdee460e71fa28e9902edf1cc662db0d19cad8
models
per-stage models — see… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-26-sonnet45-post-action-retrospection-natural-turn-design.geo-retrofit-benchmark
GEO Retrofit Benchmark
Additive answer-first + FAQ retrofit engine and structural GEO/AEO detection gate that took a real bilingual (IT+EN) content corpus of 194 articles from 741 structural GEO/AEO findings to 0, with a synthetic before/after fixture pair and the exact command to reproduce the transform end to end (node --test test/reproduce.test.mjs). The 194-article corpus itself is not redistributed; the transform and gate logic that produced the result is shipped in full… See the full description on the dataset page: https://huggingface.co/datasets/FedCal/geo-retrofit-benchmark.2026-08-26-table2-9284-post-action-retrospection-716-train
Training mixture for the post-action-retrospection (design B) 716 arm: 9,284 spec-filtered Table-2 instruction rows + 716 five-turn PAR rows -- a difficult-advice prompt, a bare refusal, the person's pushback, and the trained turn doing the reasoning the refusal skipped. Trait-balanced with a capped water-fill (two principles have fewer than an even share after the grey-area rater) and spread round-robin across domain. Same 9,284 Table-2 rows and same builder as the da716 mixture… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-26-table2-9284-post-action-retrospection-716-train.
