datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sit-experiment-results-2026-04-16
SiT Experiment Results
This dataset repo contains experiment outputs generated from the SiT-XL/2 model under the locally adjusted protocol in gpt_experiment_instructions.md.
Protocol Summary
Main dataset size: 20 images
Control dataset size: 10 images
Main/control images were exported from timm/mini-imagenet
Model: SiT-XL/2
VAE: stabilityai/sd-vae-ft-ema
Included Results
task0_outputs/: Prompt 0 / Task 0 artifacts
task1_outputs_20/: basic statistics… See the full description on the dataset page: https://huggingface.co/datasets/LamTNguyen/sit-experiment-results-2026-04-16.experiment-speaker-embeddingSnowball-67B-A2B-Mixed-RLVR-Experiment-Artifacts
Snowball 67B-A2B RL artifact release
2026 mixed-domain RLVR campaign
This release also contains the complete releasable record of the September 2026 Snowball mixed-domain RLVR campaign.
It covers the September 11 synchronous and bounded-staleness asynchronous RLVR1→RLVR2 lineages and the 5.7T
Agentic-start RLVR1 lineage. All training arms are terminal. The final campaign figure,
trace audit, canonical configs, timing reports, retained
traces, and operational… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/Snowball-67B-A2B-Mixed-RLVR-Experiment-Artifacts.SciGA-for-experiments-hfclr-experiment-datasetclr-experiment-dataset-2experimentclr-experiment-dataset-3LeRing_JFM_experiments
Overview
This dataset repository contains training data and experimental recordings to measure the rotations of particles suspended in viscous shear flows.
This dataset is intended to support the development of machine learning research in fluid dynamics, especially in the study of multi-phase flows.
Content
The repository contains training data and extensive experimental measurements of single particle suspended in confined shear flows in the viscous and small-inertial… See the full description on the dataset page: https://huggingface.co/datasets/ddg93/LeRing_JFM_experiments.shifaa_experimentsANC-formszoya-image-1-experiments
ZOYA IMAGE-1 — Reproducible GGUF Experiments
Purpose
This dataset stores reproducible ZOYA IMAGE-1 image-generation
experiments together with the exact generation parameters,
model identities, SHA256 fingerprints, and validation reports.
The package is designed for controlled comparisons where the
tested variable is changed explicitly and all other relevant
variables remain fixed.
Current baseline
Experiment ID: ZOYA_PHASE0_BASELINE_00001… See the full description on the dataset page: https://huggingface.co/datasets/tigerking009/zoya-image-1-experiments.sanad_experiments2026.mechaptcha.linear-probe-experiments-giant-20260525
siddharthmb/2026.mechaptcha.linear-probe-experiments-giant-20260525
Paired CAPTCHA image experiments for linear probes over a trained Mechaptcha CNN.
Each example contains a matched image_a and image_b pair generated from the same seed pool.
Intended Use
This dataset is designed for linear probe experiments that compare activations from Batch A against Batch B.
Use label 1 for image_a and label 0 for image_b.
Recommended checkpoint:… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.mechaptcha.linear-probe-experiments-giant-20260525.NVR-Entity-Recognition-Experiment
NVR Entity Recognition Experiment
Overview
This repository contains a training dataset designed for entity recognition in Network Video Recorder (NVR) applications, specifically focused on newborn safety monitoring. The dataset uses a stuffed animal as a privacy-conscious substitute for actual newborn footage, enabling the development of computer vision models that can identify critical safety scenarios in nursery environments.
Purpose
The primary goal of this… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/NVR-Entity-Recognition-Experiment.eval-awareness-experiment-data
Evaluation Awareness Experiment
Investigating whether LLMs exhibit different safety behaviors when they detect
evaluative contexts in prompts, using Latent Direction Amplification (LDA) to
causally manipulate eval-awareness signals.
Quick Start
# 1. Copy project to cluster
scp -r eval_awareness_experiment/ pi-mentee-login:~/
# 2. SSH in and set up environment (run once)
ssh pi-mentee-login
bash ~/eval_awareness_experiment/setup_env.sh
# 3. Submit the full… See the full description on the dataset page: https://huggingface.co/datasets/rlundqvist/eval-awareness-experiment-data.progresslm-v2-experiment-sortedBDA-Experiment2easyr1-103k-coords-5k-refusal-dapo-experimentBDA-Experiment1-WeightsBDA-Experiment3-Weightsirl-experiment-v2BDA-Experiment2-WeightsCAD-experiment-manifests-seed202-vision-qwen3vl32b-v1CAD-experiment-manifests-seed202-vision-qwen3vl32b-coderprompts-v1CAD-experiment-manifests-seed42-vision-qwen3vl32b-v1-strict-coder-v1gated-lora-experiments
Gated LoRA — Experimental Checkpoints
This dataset stores LoRA expert pools and gating network weights from
multi-task fine-tuning experiments on various base models (Phi-2, Gemma-2,
Llama-3.2, Pythia-410M, Qwen2.5-0.5B, SmolLM-360M).
Code: https://github.com/L1ZGitHub/gated-lora-research-paper-001
Structure
legacy/{model}/{run_name}/ — runs from the original experimental campaign
(2024-12 → 2025-01), with checkpoints at steps 500, 1000, ..., final_model
and best_model.… See the full description on the dataset page: https://huggingface.co/datasets/Helain/gated-lora-experiments.gcbt-experiment-imagesCAD-experiment-manifests-seed42-vision-qwen3vl4b-coderprompts-v1saelarien-constraint-experiment-02-swarm-coherence-breakdown
Saelarien Constraint Experiment 02
Swarm Coordination Breakdown Under Adversarial Conditions
Summary
This dataset provides a parameterized simulation of distributed swarm systems under increasing coordination pressure, communication degradation, and adversarial noise.
It captures the transition from stable coordination to coherence failure as system load exceeds the system’s ability to reconcile state across agents.
The dataset is designed to surface failure… See the full description on the dataset page: https://huggingface.co/datasets/Saelarien/saelarien-constraint-experiment-02-swarm-coherence-breakdown.
