datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
platonic-all-experimentsrl_llm_experiment_p6hle-flowbench-experiments-20260829
HLE FlowBench experiment archive
Private migration snapshot of the local HLE text-only 100-question research
program through 2026-09-01. It preserves the formal and smoke runs, per-question
Codex/Claude/Kimi sessions, workflow attempts and metrics, evaluator state,
scores, monitoring, experiment controllers, reports, analyses, the paused-run
migration package, source Git bundles, and HLE-related host orchestration
sessions.
The current Chinese experiment status, validity… See the full description on the dataset page: https://huggingface.co/datasets/Changyeli03/hle-flowbench-experiments-20260829.language-decoded-experiments
Language Decoded — Experiment Tracking
Central hub for training logs, configurations, evaluation results, and analysis for the Language Decoded project. The project originated as a proposal during Cohere's Tiny Aya Expedition (March 2026 hackathon) and was extended into Phase 3 for the accompanying paper.
Submitted paper title (2026-05-26): Language, Decoded: Exploring the Impact of Fine-Tuning a Multilingual Model on Native-Language Code
⚠️ Phase 3 numbers — read… See the full description on the dataset page: https://huggingface.co/datasets/legesher/language-decoded-experiments.rl_llm_experiment_p9mongabay-experimentisoflop-experiments
IsoFLOP Scaling Law Experiments
Curated collection of IsoFLOP curve data from 6 experiments, standardized to a common schema.
This dataset is associated with the paper Problems with Chinchilla Approach 2: Systematic Biases in IsoFLOP Parabola Fits.
Project Page: https://openathena.ai/scaling-law-analysis
Data Extraction & Prep: Open-Athena/scaling-law-analysis
Scaling Law Estimation: Open-Athena/vpnls
Schema
Field
Type
Description
source
string
Data source… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/isoflop-experiments.BELKA-DEL-Experimental-BenchmarkData
BELKA-DEL-Experimental-BenchmarkData
This dataset comprises a curated collection of PDB structures, designed as an experimental validation benchmark for models trained on the Big Encoded Library for Chemical Assessment (BELKA) DNA-Encoded Library (DEL). Each structure includes at least one bound small molecule ligand, providing a robust basis for benchmarking model performance in accurately identifying potential binders to BELKA protein targets.
Introduction to the BELKA… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/BELKA-DEL-Experimental-BenchmarkData.oxide-experimental-overpotentials-her-oer
Oxide Experimental Overpotentials for HER/OER Electrocatalysis
REAL EXPERIMENTAL MEASUREMENTS — Not estimates or synthetic data.
Dataset Overview
This dataset contains 63 records of experimentally measured overpotentials (η) for HER and OER on oxide-based electrocatalysts.
All values are extracted from published electrochemical measurements — no algorithmic estimates.
Sources
Paper
Year
Records
Key Data
McCrory et al.
2015
24
Systematic benchmarking… See the full description on the dataset page: https://huggingface.co/datasets/ANBU963/oxide-experimental-overpotentials-her-oer.html-ai-battle-experiment-tracker
HTML AI Battle Experiment Tracker
This dataset contains the experiment tracker for the paper:
The Single-File Test: A Longitudinal Public-Interface Evaluation of First-Output LLM Web Generation with Social Reach Tracking
Author: Diego Cabezas Palacios
arXiv: 2605.06707Code and materials: https://github.com/diegocp01/html_ai_battle
Dataset Summary
This dataset supports a longitudinal observational comparison of first-output LLM web generation across public chat interfaces.… See the full description on the dataset page: https://huggingface.co/datasets/diegocp01/html-ai-battle-experiment-tracker.saelarien-constraint-experiment-03-recovery-collapse-mismatch
Saelariën Constraint Experiment 03: Recovery–Collapse Mismatch
Summary
This dataset extends Experiment 01, which established that collapse emerges when entropy injection exceeds a system’s capacity to maintain coherent state.
Experiment 03 isolates a different question: whether recovery dynamics uniquely characterize proximity to collapse.
The results show they do not.
Systems with indistinguishable recovery profiles can resolve into both stable and collapsed outcomes.… See the full description on the dataset page: https://huggingface.co/datasets/Saelarien/saelarien-constraint-experiment-03-recovery-collapse-mismatch.BSG_CyLlama-experimental-training
BSG CyLlama V8 - Experimental Training Data
20,000 training samples for the experimental theme, generated by DeepSeek.
Each sample contains source abstracts from a scientific cluster and target outputs (abstract summary, short summary, title, overview) for training the BSG CyLlama cluster description generator.
Format
TSV with columns: cluster_id, theme, source_abstracts, abstract_summary, short_summary, title, theme_token, short_title, overview
Related
Model:… See the full description on the dataset page: https://huggingface.co/datasets/jimnoneill/BSG_CyLlama-experimental-training.hf10313_9c3e7270_experimentsQwen_Experimental_Datakazakh-morpho-experiments
kazakh-morpho-experiments
Қазақ морфологиясына арналған тәжірибе материалдары · Материалы экспериментов по казахской морфологии · Kazakh morphology experiment material
Қазақша · Русский · English
Қазақша
kazakh-morpho-experiments — қазақ тілінің морфологиялық талдауын оқытуға және бағалауға қолданылған деректер, скрипттер мен нәтижелер мұрағаты. Репозиторий көлемі — 13.7 МБ; ол тәжірибені қайта қарауға және таңбаларды түзету барысын зерттеуге арналған.… See the full description on the dataset page: https://huggingface.co/datasets/TilQazyna/kazakh-morpho-experiments.ExperimentalBlind-Spot-Experiment-new-Dataset
Blind-Spot-Experiment-new-Dataset
Dataset Purpose
This dataset was created to investigate blind spots in a base foundation language model.
The experiment was conducted using the Transformers library from :contentReference[oaicite:1]{index=1}.
The evaluated model is :contentReference[oaicite:2]{index=2}.
Model link: https://huggingface.co/Qwen/Qwen3-0.6B
Implementation Details
The model was loaded and tested in Google Colab.
Code used to load the model:
from… See the full description on the dataset page: https://huggingface.co/datasets/Blessinggreat988/Blind-Spot-Experiment-new-Dataset.SynthWAF-2-Experimentalai-carbon-footprint-experiment-logs
AI Carbon Footprint Experiment Logs
Dataset Overview
This dataset contains experiment logs collected while evaluating the energy consumption and carbon emissions of simple Python programs using carbon accounting tools.
The dataset was created as part of a research project on Generative AI Carbon Footprint Analysis.
Contents
The dataset includes measurements collected from experiments such as:
Matrix multiplication
Sorting a large array
Simple… See the full description on the dataset page: https://huggingface.co/datasets/MotebaRehman/ai-carbon-footprint-experiment-logs.dpo-experiment-resultsexperimentinglawFIT-Experiment
Safety Consistency Across English, Roman Urdu, and Urdu-English Code-Switching
Overview
This project started from something I noticed in the way I normally use AI.
I rarely communicate only in formal English. I often switch between English, Roman Urdu, and a mixture of Urdu and English. Roman Urdu is especially informal: people use different spellings for the same words, shorten words, and mix English naturally into sentences.
This made me interested in a simple… See the full description on the dataset page: https://huggingface.co/datasets/PakeezaKhalid/FIT-Experiment.
