datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ExperimentDATA_knowledge_distillation_vs_fine_tuningparity-experiments
Upload your Adapter Oracle and Parity results
This dataset saves the oracle and parity experiment logs for adapters. Please upload them according to the following format and draft a PR.
adapters/
└── {adapter_name}/
├── README.md # Results overview, directory structure, trajectory interpretation, notes, etc. This should be DIFFERENT than the adapter REAMDE.
├── config.yaml # The yaml file that can be directly used to run parity experiments in Harbor.
├──… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/parity-experiments.experimentsRRC_Experimentsartificium-covering-experimentA live autonomous research experiment using Artificium with a 4-bit Qwen 27B model. The agent researches, writes code, runs experiments, and maintains persistent memories while tackling the covering-design problem C(25,15,5).
Goal: find 41 or fewer groups, each containing 15 distinct numbers from 1 to 25, such that every possible five-number combination appears in at least one group—improving on the experiment’s 42-group starting reference.
This dataset is updated every 30 minutes with… See the full description on the dataset page: https://huggingface.co/datasets/gr0010/artificium-covering-experiment.sit-experiment-results-2026-04-16
SiT Experiment Results
This dataset repo contains experiment outputs generated from the SiT-XL/2 model under the locally adjusted protocol in gpt_experiment_instructions.md.
Protocol Summary
Main dataset size: 20 images
Control dataset size: 10 images
Main/control images were exported from timm/mini-imagenet
Model: SiT-XL/2
VAE: stabilityai/sd-vae-ft-ema
Included Results
task0_outputs/: Prompt 0 / Task 0 artifacts
task1_outputs_20/: basic statistics… See the full description on the dataset page: https://huggingface.co/datasets/LamTNguyen/sit-experiment-results-2026-04-16.userlm_rl_experiment_rolloutsdl2l-experiments
DL2L Experiments Dataset
Simulation trajectory data from the DL2L
distributed artificial life simulator, used to train JEPA world models.
See felipedreis/dl2l-jepa for the trained models.
Dataset structure
Data is organized by experiment prefix. Each prefix contains parquet files for
model training and a stats.json with dataset metadata.
p9/
train.parquet # single-encoder training set (trials 1–8)
val.parquet # single-encoder validation set… See the full description on the dataset page: https://huggingface.co/datasets/felipedreis/dl2l-experiments.opendatalab-experimental-nmr-peaks
OpenDataLab Experimental NMR Peaks Dataset
Dataset Description
This dataset contains experimental NMR (Nuclear Magnetic Resonance) peak sequences extracted from the OpenDataLab experimental spectra database. The dataset includes both H-NMR and C-NMR peak sequences for chemical compounds, along with their SMILES representations and molecular formulas.
Dataset Summary
Total Samples: 533,595 compounds
Batches: 333 batch files
Data Source: Experimental spectra… See the full description on the dataset page: https://huggingface.co/datasets/snehasis19/opendatalab-experimental-nmr-peaks.self-organisation-experiments
Self-Organisation Experiment Data
Experiment data for the self-organisation project investigating whether self-replication can emerge spontaneously in continuous neural network weight space.
Result: Negative — continuous weight space appears to lack the computational primitives required for spontaneous self-replication. See the code repository for full analysis.
Structure
autoresearch/ — Automated research runs across 11 experiment configurations
phase1*/ — BFF… See the full description on the dataset page: https://huggingface.co/datasets/LindaP/self-organisation-experiments.pisa-experiments
Pisa Experiments
This repository contains the PisaBench, training data, model checkpoints, introduced in PISA Experiments: Exploring Physics Post-Training for Video Diffusion Models by Watching Stuff Drop.
PisaBench
Real World Videos
We curate a dataset comprising 361 videos demonstrating the dropping task.Each video begins with an object suspended by an invisible wire in the first frame. We cut the video clips to begin as soon as the… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/pisa-experiments.transformers-merge-experimentsdementor-complete-experiment-results
Dementor complete experiment results
Audited outputs for the configuration-defined Dementor completion campaign.
Audited scope
Behavioral imitation adapters: 1,104 total (528 SFT, 528 DPO, 48 self-SFT controls).
Behavioral-fidelity evaluation: 1,104 adapters on 200 held-out prompts, with embedding and
primary LLM-judge scores, plus 48 target-reference response sets.
Activation steering: 29 models, seven benchmarks, and two operators (original and fpall),
totaling… See the full description on the dataset page: https://huggingface.co/datasets/dementor-research/dementor-complete-experiment-results.gemma-crafter-five-experiments-20260915tailsft-lfm350m-experiment
TailSFT with LFM2.5-350M and Trackio
Status, 2026-09-16: GPU preflight completed all three training arms and four evaluations. The strict final-answer grader and namespace/resume checks pass (50 tests). A full-length batch of 256 samples completed on the A100; the full comparison has been submitted. Full experiment results are not yet available.
Small, hackable reproduction of the TailSFT filtering method. The objective is higher pass@16 after supervised training. This… See the full description on the dataset page: https://huggingface.co/datasets/burtenshaw/tailsft-lfm350m-experiment.typed-decisions-causal-experimentexperimentsfine-tuning-experiments-082023bert-mlm-experiments-en
Unified English MLM Pre-training Corpus (80M Rows)
This dataset is a massive, diverse, multi-domain English text corpus explicitly engineered for pre-training and domain-adaptation of BERT-style models via Masked Language Modeling (MLM). It aggregates over 80 million rows of text, completely stripped of auxiliary metadata, labels, and identifiers to expose purely raw text strings.
Dataset Details
Repository ID: 8Opt/bert-mlm-experiments-en
Total Rows: 80,489,226… See the full description on the dataset page: https://huggingface.co/datasets/LakoreAI/bert-mlm-experiments-en.icefall-aishell-experiment-full
Introduction
The icefall project contains speech-related recipes for various datasets
using k2-fsa and lhotse.
You can use sherpa, sherpa-ncnn or sherpa-onnx for deployment with models
in icefall; these frameworks also support models not included in icefall; please refer to respective documents for more details.
You can try pre-trained models from within your browser without the need
to download or install anything by visiting this huggingface space.
Please refer to document for… See the full description on the dataset page: https://huggingface.co/datasets/0120Kiya/icefall-aishell-experiment-full.twla-experiment-datasets
TWLA experiment datasets
Source datasets
C4
WikiText-103
SlimPajama
OpenWebMath
CodeSearchNet
PG-19
CNN/DailyMail
Nemotron pretraining data
frontierphysics-experiment
FrontierPhysics experiment artifacts
Public artifacts for PR #437, tested at 6989e0ba72c0beb13fc8924c4f9e2461eab57237.
Complete three-trial ZIP, 1.75 GB — anonymously downloaded and SHA-256/CRC verified.
Final report
Raw trial files
All three GPT-6 Astra xhigh trials passed 16/16 tests. Native final rewards: 0.898305, 0, 0.932203. GPT-5.6 Sol xhigh failed one reporting blocker for trial 2; the report flags inconsistent treatment of that criterion across the reviews. Original… See the full description on the dataset page: https://huggingface.co/datasets/benchflow/frontierphysics-experiment.platonic-all-experimentsexperiment-process-seamless-alignartificium-riemannhypothesis-experiment
Artificium — two messages and a 63-hour experiment
My harness on GitHub
Contact Me
I sent exactly two chat messages: one to start the experiment, one to end it.
Artificium is a general agent harness for long term autonomous work, continual learning, and self-improvement. I built it to power my own personal agent, and I'm sharing the harness along with the data from this experiment.
My first message:
change your self to make your life purpose to achieve this goal: solve the… See the full description on the dataset page: https://huggingface.co/datasets/gr0010/artificium-riemannhypothesis-experiment.experiment-speaker-embeddingrtpurbo-block-summary-failed-experiment-artifacts
RTPurbo block-summary failed experiment artifacts
Immutable research artifacts from the entropy-calibrated 64-token block-summary investigation for RTPurbo/Qwen3.5-0.8B.
The tested static tangent and CAMS geometries failed the registered selector fidelity/traffic gate. This repository preserves the reusable feature/teacher caches, schedules, checkpoints, controls, scoreboards, and diagnostic evidence needed to reproduce or revisit that conclusion. It is an experiment archive… See the full description on the dataset page: https://huggingface.co/datasets/danym/rtpurbo-block-summary-failed-experiment-artifacts.atomic-metrics-experiments
Atomic Metrics Experiment Artifacts
Run directories from
Atomic Metrics, including
extracted metric banks, generated domain prompts, batch/refine snapshots, and
BT/LR eval summaries.
Logs are omitted. API keys are not included; scoring used environment
credentials at runtime.
BioReasonCell-ExperimentDataNIDS-Thesis-Experimental-Evidence
thesis_pipeline
Pipeline rebuilt following the supervisor's conditional review (13/07/2026). See SCOPE_FROZEN.md at the parent project root for the frozen scientific scope.
Structure
config/: centralized configuration (paths, seeds)
manifests/: data audits and temporal split manifests
src/data/: dataset preparation and cleaning
src/models/: model training and comparison
src/evaluation/: aggregation and global comparison of results
tests/: leakage tests and… See the full description on the dataset page: https://huggingface.co/datasets/MamadouSY-NIDS-Thesis/NIDS-Thesis-Experimental-Evidence.
