datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
experimentstransformers-merge-experimentsplatonic-all-experimentsjax-fli-experiments
jax-fli experiments
Data, samples, and reference catalogs for the jax-fli
forward-modelling experiments. Each experiment is exposed as one or more
HuggingFace dataset configs; load a config with datasets.load_dataset.
Experiment 00 — CosmoGrid reference
A single CosmoGrid simulation (cosmo_000001) packaged as jax-fli Catalog
parquet files, used as reference truth for the lensing / masked-shear experiments.
Config
Field
Shape
NSIDE
Description… See the full description on the dataset page: https://huggingface.co/datasets/ASKabalan/jax-fli-experiments.soft-prompt-experiments-archive-20260918
Soft prompt 实验归档
用于查阅和恢复的历史研究记录,涵盖数学与代码任务。共 45 个运行目录,包含教师生成数据、评测输出、原始配置和已有 prompt 检查点。部分目录仅有评测、复核或失败记录,不能将目录数量理解为成功实验数量。
快速查阅
实验总览:模型系列、任务、规模与记录状态。
CSV 索引 / JSON 索引:便于筛选和定位。
archives/:按实验分别压缩的原始文件。
manifests/:各文件 SHA-256 与归档路径。
系列包括 AReaL Boba2、GPT-OSS/Swallow、MiMo、X-Coder/Qwen3、Nemotron、Klear、Mellum2、OLMo3、Polaris 和 Poro2。页面不展开具体方法或实现细节;原始配置仍保留在归档内供恢复。
状态与注意事项
上传完成以 ARCHIVE_COMPLETE.json 为准;文件不存在时表示仍在上传。 每个归档都经过完整下载的 SHA-256 校验。… See the full description on the dataset page: https://huggingface.co/datasets/namezz/soft-prompt-experiments-archive-20260918.hle-flowbench-experiments-20260829
HLE FlowBench experiment archive
Private migration snapshot of the local HLE text-only 100-question research
program through 2026-09-01. It preserves the formal and smoke runs, per-question
Codex/Claude/Kimi sessions, workflow attempts and metrics, evaluator state,
scores, monitoring, experiment controllers, reports, analyses, the paused-run
migration package, source Git bundles, and HLE-related host orchestration
sessions.
The current Chinese experiment status, validity… See the full description on the dataset page: https://huggingface.co/datasets/Changyeli03/hle-flowbench-experiments-20260829.language-decoded-experiments
Language Decoded — Experiment Tracking
Central hub for training logs, configurations, evaluation results, and analysis for the Language Decoded project. The project originated as a proposal during Cohere's Tiny Aya Expedition (March 2026 hackathon) and was extended into Phase 3 for the accompanying paper.
Submitted paper title (2026-05-26): Language, Decoded: Exploring the Impact of Fine-Tuning a Multilingual Model on Native-Language Code
⚠️ Phase 3 numbers — read… See the full description on the dataset page: https://huggingface.co/datasets/legesher/language-decoded-experiments.zoya-image-1-experiments
ZOYA IMAGE-1 — Reproducible GGUF Experiments
Purpose
This dataset stores reproducible ZOYA IMAGE-1 image-generation
experiments together with the exact generation parameters,
model identities, SHA256 fingerprints, and validation reports.
The package is designed for controlled comparisons where the
tested variable is changed explicitly and all other relevant
variables remain fixed.
Current baseline
Experiment ID: ZOYA_PHASE0_BASELINE_00001… See the full description on the dataset page: https://huggingface.co/datasets/tigerking009/zoya-image-1-experiments.jbcs2025_experiments_report
JBCS 2025: Experimental Artefacts for AES in Brazilian Portuguese
This repository contains all experimental artefacts (logs, configurations, predictions, and evaluation results) described in the paper:
Exploring the Usage of LLMs for Automatic Essay Scoring in Brazilian Portuguese EssaysAndré Barbosa, Igor Cataneo Silveira, Denis Deratani MauáTODO
📦 What's in this dataset repo?
This dataset is not a training dataset. Instead, it provides comprehensive logs and… See the full description on the dataset page: https://huggingface.co/datasets/kamel-usp/jbcs2025_experiments_report.grok-1-ternary-quant-experiments
Grok-1 SAAQ quantization / route-preservation experiments
Dataset author: Raul Montoya Cardenas (rmems)
SAAQ stands for Spiking Adaptive Activity Quantization, a term coined by
the dataset author.
Attribution: Grok Build: Grok 4.5 (high) packaged the original 2026-08-10
dataset. Codex: GPT-5.6-Sol (OpenAI) implemented, executed, validated, and published the
canonical issue #85 v4 evidence added on 2026-08-24.
Personal research measuring route preservation when packing open… See the full description on the dataset page: https://huggingface.co/datasets/rmems/grok-1-ternary-quant-experiments.synthetic_data_experimentspeft-unit-test-generation-experiments
PEFT Unit Test Generation Experiments
Dataset description
The PEFT Unit Test Generation Experiments dataset contains metadata and details about a set of trained models used for generating unit tests with parameter-efficient fine-tuning (PEFT) methods. This dataset includes models from multiple namespaces and various sizes, trained with different tuning methods to provide a comprehensive resource for unit test generation research.
Dataset Structure
Data… See the full description on the dataset page: https://huggingface.co/datasets/andstor/peft-unit-test-generation-experiments.peft-unit-test-generation-experiments
PEFT Unit Test Generation Experiments
Dataset description
The PEFT Unit Test Generation Experiments dataset contains metadata and details about a set of trained models used for generating unit tests with parameter-efficient fine-tuning (PEFT) methods. This dataset includes models from multiple namespaces and various sizes, trained with different tuning methods to provide a comprehensive resource for unit test generation research.
Dataset Structure
Data… See the full description on the dataset page: https://huggingface.co/datasets/fals3/peft-unit-test-generation-experiments.romani-asr-experiments
Romani ASR Experiments
This repository collects the reproducible training and evaluation artifacts for
the Romani ASR experiments.
It does not contain raw audio, full transcript manifests, or private
training data. Model weights live in model-specific repositories.
Current Model Repositories
Whisper Turbo Romani LoRA adapter:
kiviki/whisper-turbo-romani-lora
MMS adapter: not published yet. The current MMS result is zero-shot
facebook/mms-1b-all with… See the full description on the dataset page: https://huggingface.co/datasets/kiviki/romani-asr-experiments.ray-trackio-experimentsplanktonzilla-experiments-datasetisoflop-experiments
IsoFLOP Scaling Law Experiments
Curated collection of IsoFLOP curve data from 6 experiments, standardized to a common schema.
This dataset is associated with the paper Problems with Chinchilla Approach 2: Systematic Biases in IsoFLOP Parabola Fits.
Project Page: https://openathena.ai/scaling-law-analysis
Data Extraction & Prep: Open-Athena/scaling-law-analysis
Scaling Law Estimation: Open-Athena/vpnls
Schema
Field
Type
Description
source
string
Data source… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/isoflop-experiments.grpo-experimentsadult-fairness-experimentsfrank_load_experimentssynthetic_data_experimentskinetic-experimentsautoresearch-experiments
Autoresearch Cross-Platform Experiments
Dataset Description
This dataset contains 2,637 hyperparameter optimization experiments from an autonomous LLM-driven ML research project. An LLM agent (Claude Sonnet) autonomously proposes hyperparameter modifications, trains a small language model for 5 minutes, evaluates validation bits-per-byte (val_bpb), and iterates.
Experiments span 3 hardware platforms, 5 GPU models, and 7 text datasets, making this a unique resource for… See the full description on the dataset page: https://huggingface.co/datasets/davegraham/autoresearch-experiments.platonic-correct-experimentsIFEval_Experimentsevalap-spp_experiments-9
SPP_experiments (ID: 9)
Testing different configurations for SPP : base models, finetuning, fulltuning, rag and no rag architectures.
Overview
This dataset contains 8 experiments
from the EvalAP evaluation platform.
Datasets: SPP_Albert_Prod, SPP_Llama8B_31_Fullfine, SPP_Llama8B_31_Fullfine_3, SPP_Llama8B_31_LoRA_32_64_3____v2, SPP_Llama8B_31_LoRA_32_64_3__v2, SPP_Llama8B_31_LoRA_32_64_3_v2, SPP_Llama8B_31_LoRa_3264_3, SPP_llama3.1_8B_finetune_lora_32_64_3_bigger… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-spp_experiments-9.bank-fairness-experimentsliaisons-experiments-results⚠️ This repository is a part of an academical project for the Heriot-Watt University, no third-party contributions are accepted.
Dataset Card for Liaison's LLMs argumentative relation prediction benchmarking task
About the Dataset
Dataset Summary
The present dataset contains the results of an evaluation of Large Language Models at the tasks of argumentative relation prediction between pairs of arguments.This work is a limited update of a previous evaluation… See the full description on the dataset page: https://huggingface.co/datasets/coding-kelps/liaisons-experiments-results.hf10313_9c3e7270_experimentsabmelt-experiments
