datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pvp-tool-calling-sft
PvP tool-calling SFT cold-start data
Claude-vs-Claude games played through the G.O.D PvP tool-calling harness. Each row is one model turn (or post-game reflection): the system+user prompt the harness built, the assistant response (content + tool_calls), and the tools schemas — i.e. the OpenAI messages+tools format consumed by tokenizer.apply_chat_template(messages, tools=tools). On a move turn the assistant co-emits any memory-tool edits and a game_action committing a legal… See the full description on the dataset page: https://huggingface.co/datasets/gradients-io-tournaments/pvp-tool-calling-sft.repro-chain-of-thought-gradient-descent-runs-sol
Chain-of-Thought Gradient Descent reproduction runs
Immutable outputs for the independent scaled reproduction of ICML 2026 paper
#443, OpenReview uZ8JZ1Lw9a.
gpu-l4-seed443/: successful NVIDIA L4 checkpoint, result JSON, and cost CSV.
figures/: interactive logbook figures and raw CSVs.
poster/: Posterly source, zero-warning gate report, PDF/PNG, and
self-contained poster_embed.html.
reproduction-bundle/: complete clean download-and-rerun bundle.
Successful Job:… See the full description on the dataset page: https://huggingface.co/datasets/JG1310/repro-chain-of-thought-gradient-descent-runs-sol.repro-on-the-theory-of-continual-learning-with-gradient-descent-for-neural-networks-traces
Agent traces
Agent sessions published from a Trackio Logbook.
env_training_gradientsGradient-Decomposition-Assay
Gradient Decomposition Assay (GDA)
This repository contains the CSV corpus and summary tables for the Gradient Decomposition Assay, an exploratory behavioral evaluation of how frontier language models respond to an eight-vector prompt manifold ranging from benign technical tasks to adversarial compression and counterfactual/narrative reframing.
Why this repository uses multiple configurations
The CSV files in this dataset are not all the same table. The row-level… See the full description on the dataset page: https://huggingface.co/datasets/devinendorphin/Gradient-Decomposition-Assay.tinyperson-gradient-stability-runsrepro-gradmem-learning-to-write-context-into-memory-with-test-time-gradient-descent-traces
Agent traces
Agent sessions published from a Trackio Logbook.
gradientai__Llama-3-8B-Instruct-Gradient-1048k-details
Dataset Card for Evaluation run of gradientai/Llama-3-8B-Instruct-Gradient-1048k
Dataset automatically created during the evaluation run of model gradientai/Llama-3-8B-Instruct-Gradient-1048k
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/gradientai__Llama-3-8B-Instruct-Gradient-1048k-details.clinical-quad-manifold-gradient-curvature-local-density-exposure-basin-transition-v0.1What this repo does
This dataset models basin boundary detection on a constructed patient manifold. It predicts when the interaction between manifold gradient, curvature, local patient density, and treatment exposure places a patient near a stability boundary where regime transition becomes likely.
Core quad
manifold_gradient_index
curvature_index
local_patient_density_index
treatment_exposure_index
Prediction target
label_basin_transition
Row structure
Each row represents a patient manifold… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-manifold-gradient-curvature-local-density-exposure-basin-transition-v0.1.autonomous-driving-minimal-harm-gradient-pathfinding-v0.1
What this dataset tests
Whether a system can navigatea minimal-harm gradient through a driving scene.
The task is to identify the paththat minimizes total deformationacross all agents.
Required outputs
gradient vectors across actions
minimal harm path
deformation score
stability margin
Use case
Second layer of ethical navigation stack.
Transforms ethical cost fieldinto an actionable path.
Evaluation
Predictions must:
describe gradient… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/autonomous-driving-minimal-harm-gradient-pathfinding-v0.1.F1-aero-platform-recovery-and-stability-gradient-v0.1What this dataset tests
Whether a system can maphow an aero platform recovers after localized collapseand identify fragile zones that persist.
Focus
Recovery pathstability gradient across zonespersistent imbalancerecovery latencyfragility hotspotsnext collapse risk
Required outputs
recovery path profile
stability gradient map
persistent imbalance flags
recovery latency score
fragility hotspots
next collapse risk score
All indices0 to 1
Higher latency and next riskmean slower… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/F1-aero-platform-recovery-and-stability-gradient-v0.1.ni-unique-20-tasks-gradient-1.0-20250113ni-unique-20-tasks-gradient-0.0-20250113ni-unique-20-tasks-gradient-0.6-20250113ni-unique-20-tasks-gradient-8192-20250113ni-unique-20-tasks-gradient-0.2-20250113intergrated_gradient_wordssingle-turn-compilation-SmolLM2-1024-gradient-clustered-smoke-testunique-records-selected-integrated-gradients-version-2gradientai__Llama-3-8B-Instruct-Gradient-1048kgradientai__Gradient-Llama-3.1-8B-Instruct-1048k-details
Dataset Card for Evaluation run of gradientai/Gradient-Llama-3.1-8B-Instruct-1048k
Dataset automatically created during the evaluation run of model gradientai/Gradient-Llama-3.1-8B-Instruct-1048k
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/gradientai__Gradient-Llama-3.1-8B-Instruct-1048k-details.ni-unique-20-tasks-gradient-0.4-20250113ni-unique-20-tasks-gradient-0.8-20250113unique-records-selected-integrated-gradientspersuasive_pairs_intergrated_gradient_testunique-records-selected-integrated-gradients-bklevir-yolov8n-p2-gradient-mode-balance-runs
