datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
experimentstransformers-merge-experimentsbert-mlm-experiments-en
Unified English MLM Pre-training Corpus (80M Rows)
This dataset is a massive, diverse, multi-domain English text corpus explicitly engineered for pre-training and domain-adaptation of BERT-style models via Masked Language Modeling (MLM). It aggregates over 80 million rows of text, completely stripped of auxiliary metadata, labels, and identifiers to expose purely raw text strings.
Dataset Details
Repository ID: 8Opt/bert-mlm-experiments-en
Total Rows: 80,489,226… See the full description on the dataset page: https://huggingface.co/datasets/LakoreAI/bert-mlm-experiments-en.platonic-all-experimentsjax-fli-experiments
jax-fli experiments
Data, samples, and reference catalogs for the jax-fli
forward-modelling experiments. Each experiment is exposed as one or more
HuggingFace dataset configs; load a config with datasets.load_dataset.
Experiment 00 — CosmoGrid reference
A single CosmoGrid simulation (cosmo_000001) packaged as jax-fli Catalog
parquet files, used as reference truth for the lensing / masked-shear experiments.
Config
Field
Shape
NSIDE
Description… See the full description on the dataset page: https://huggingface.co/datasets/ASKabalan/jax-fli-experiments.soft-prompt-experiments-archive-20260918
Soft prompt 实验归档
用于查阅和恢复的历史研究记录,涵盖数学与代码任务。共 45 个运行目录,包含教师生成数据、评测输出、原始配置和已有 prompt 检查点。部分目录仅有评测、复核或失败记录,不能将目录数量理解为成功实验数量。
快速查阅
实验总览:模型系列、任务、规模与记录状态。
CSV 索引 / JSON 索引:便于筛选和定位。
archives/:按实验分别压缩的原始文件。
manifests/:各文件 SHA-256 与归档路径。
系列包括 AReaL Boba2、GPT-OSS/Swallow、MiMo、X-Coder/Qwen3、Nemotron、Klear、Mellum2、OLMo3、Polaris 和 Poro2。页面不展开具体方法或实现细节;原始配置仍保留在归档内供恢复。
状态与注意事项
上传完成以 ARCHIVE_COMPLETE.json 为准;文件不存在时表示仍在上传。 每个归档都经过完整下载的 SHA-256 校验。… See the full description on the dataset page: https://huggingface.co/datasets/namezz/soft-prompt-experiments-archive-20260918.hle-flowbench-experiments-20260829
HLE FlowBench experiment archive
Private migration snapshot of the local HLE text-only 100-question research
program through 2026-09-01. It preserves the formal and smoke runs, per-question
Codex/Claude/Kimi sessions, workflow attempts and metrics, evaluator state,
scores, monitoring, experiment controllers, reports, analyses, the paused-run
migration package, source Git bundles, and HLE-related host orchestration
sessions.
The current Chinese experiment status, validity… See the full description on the dataset page: https://huggingface.co/datasets/Changyeli03/hle-flowbench-experiments-20260829.cibench-experiments
CIBench Experiments
Reproducibility packages for CIBench — the stateless, replayable benchmark engine for the 1M–10M token long-context era.
If a benchmark result cannot be replayed from its manifest alone, it did not happen.
Every sub-directory in this dataset is a self-contained experiment package: per-run manifests, content-addressed canonical JSON, ResultRecord with full scoring + signed provenance, per-item OpenTelemetry gen_ai_* call metrics, retrieved evidence, a… See the full description on the dataset page: https://huggingface.co/datasets/publicus-ai/cibench-experiments.SciGA-for-experiments-hflanguage-decoded-experiments
Language Decoded — Experiment Tracking
Central hub for training logs, configurations, evaluation results, and analysis for the Language Decoded project. The project originated as a proposal during Cohere's Tiny Aya Expedition (March 2026 hackathon) and was extended into Phase 3 for the accompanying paper.
Submitted paper title (2026-05-26): Language, Decoded: Exploring the Impact of Fine-Tuning a Multilingual Model on Native-Language Code
⚠️ Phase 3 numbers — read… See the full description on the dataset page: https://huggingface.co/datasets/legesher/language-decoded-experiments.bon-pim-hacking-experimentszoya-image-1-experiments
ZOYA IMAGE-1 — Reproducible GGUF Experiments
Purpose
This dataset stores reproducible ZOYA IMAGE-1 image-generation
experiments together with the exact generation parameters,
model identities, SHA256 fingerprints, and validation reports.
The package is designed for controlled comparisons where the
tested variable is changed explicitly and all other relevant
variables remain fixed.
Current baseline
Experiment ID: ZOYA_PHASE0_BASELINE_00001… See the full description on the dataset page: https://huggingface.co/datasets/tigerking009/zoya-image-1-experiments.sanad_experimentsjbcs2025_experiments_report
JBCS 2025: Experimental Artefacts for AES in Brazilian Portuguese
This repository contains all experimental artefacts (logs, configurations, predictions, and evaluation results) described in the paper:
Exploring the Usage of LLMs for Automatic Essay Scoring in Brazilian Portuguese EssaysAndré Barbosa, Igor Cataneo Silveira, Denis Deratani MauáTODO
📦 What's in this dataset repo?
This dataset is not a training dataset. Instead, it provides comprehensive logs and… See the full description on the dataset page: https://huggingface.co/datasets/kamel-usp/jbcs2025_experiments_report.synthetic_experimentsshifaa_experimentscommonvoice17_su_experiments
Common Voice 17 -- Single / Long Utterance experiment dataset
Built from fixie-ai/common_voice_17_0 (English); the original CV splits are preserved and each is bucketed into single-utterance (1 word) and long-utterance (>= 3 words).
Splits: dev_single, dev_long, test_single, test_without_single, train_single, train_long.
genshin-voice-english-fXLmbpp-code-rl
MBPP for code RL (deduplicated against MBPP+)
MBPP prepared for RLVR training in verl,
with two independent hold-outs so both MBPP+ and MBPP's own canonical test
split stay reportable after training on this data.
split
rows
contents
train
320
MBPP canonical train + validation + prompt, minus everything in MBPP+
test
378
exactly the problems in evalplus/mbppplus
heldout_mbpp_test
276
MBPP's canonical test split (task_id 11-510) that is not in MBPP+… See the full description on the dataset page: https://huggingface.co/datasets/RL-Forgetting-Experiments-3/mbpp-code-rl.2026.mechaptcha.linear-probe-experiments-giant-20260525
siddharthmb/2026.mechaptcha.linear-probe-experiments-giant-20260525
Paired CAPTCHA image experiments for linear probes over a trained Mechaptcha CNN.
Each example contains a matched image_a and image_b pair generated from the same seed pool.
Intended Use
This dataset is designed for linear probe experiments that compare activations from Batch A against Batch B.
Use label 1 for image_a and label 0 for image_b.
Recommended checkpoint:… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.mechaptcha.linear-probe-experiments-giant-20260525.abmelt-experiments-exp_20260220_125440abmelt-experiments-exp_20260220_130124grok-1-ternary-quant-experiments
Grok-1 SAAQ quantization / route-preservation experiments
Dataset author: Raul Montoya Cardenas (rmems)
SAAQ stands for Spiking Adaptive Activity Quantization, a term coined by
the dataset author.
Attribution: Grok Build: Grok 4.5 (high) packaged the original 2026-08-10
dataset. Codex: GPT-5.6-Sol (OpenAI) implemented, executed, validated, and published the
canonical issue #85 v4 evidence added on 2026-08-24.
Personal research measuring route preservation when packing open… See the full description on the dataset page: https://huggingface.co/datasets/rmems/grok-1-ternary-quant-experiments.fm-model-experiments-data
FM Model Experiments — Synthetic VLM Training Data (KO + EN)
Annotation data produced while building a native-resolution Korean+English VLM
(GLM-4.6V vision tower transplanted onto a frozen GLM-5.2 743B MoE decoder).
Code + technical report: https://github.com/genonai/fm-model-experiments
This repo contains ANNOTATIONS ONLY (.jsonl). No images are redistributed.
Every row references an image by a relative path (data/...); obtain the images
from the original sources listed below… See the full description on the dataset page: https://huggingface.co/datasets/mncai/fm-model-experiments-data.trackio-experiments
Trackio Experiments Dataset
This dataset stores experiment tracking data for ML training runs, particularly focused on SmolLM3 fine-tuning experiments with comprehensive metrics tracking.
Dataset Structure
The dataset contains the following columns:
experiment_id: Unique identifier for each experiment
name: Human-readable name for the experiment
description: Detailed description of the experiment
created_at: Timestamp when the experiment was created
status: Current… See the full description on the dataset page: https://huggingface.co/datasets/Tonic/trackio-experiments.abmelt-experiments-exp_20260219_182408abmelt-experiments-exp_20260218_171120peft-unit-test-generation-experiments
PEFT Unit Test Generation Experiments
Dataset description
The PEFT Unit Test Generation Experiments dataset contains metadata and details about a set of trained models used for generating unit tests with parameter-efficient fine-tuning (PEFT) methods. This dataset includes models from multiple namespaces and various sizes, trained with different tuning methods to provide a comprehensive resource for unit test generation research.
Dataset Structure
Data… See the full description on the dataset page: https://huggingface.co/datasets/andstor/peft-unit-test-generation-experiments.peft-unit-test-generation-experiments
PEFT Unit Test Generation Experiments
Dataset description
The PEFT Unit Test Generation Experiments dataset contains metadata and details about a set of trained models used for generating unit tests with parameter-efficient fine-tuning (PEFT) methods. This dataset includes models from multiple namespaces and various sizes, trained with different tuning methods to provide a comprehensive resource for unit test generation research.
Dataset Structure
Data… See the full description on the dataset page: https://huggingface.co/datasets/fals3/peft-unit-test-generation-experiments.contracts-extraction-instruction-llm-experiments
Dataset Card for "contracts-extraction-instruction-llm-experiments"
More Information needed
