datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
eval2_proaugcontinual-learning-bench-data
Continual Learning Benchmark — Data
Frozen corpora and supporting artifacts for the six tasks in the Continual Learning Benchmark. The repo accompanies the (anonymized) benchmark codebase, which loads these files and feeds them — with task-specific framing — to the system under evaluation.
Repository layout
blind_spectrum_monitoring/ # frozen scan corpus + metadata
codebase_adaptation/ # final PR dataset + 2 docker images
cohort_studies/ # cohort defs… See the full description on the dataset page: https://huggingface.co/datasets/continual-learning-benchmark/continual-learning-bench-data.lean4-stat-learning-theory-novel
A Large-Scale Lean 4 Dataset on Statistical Learning Theory
We present a high-quality, human-verified, large-scale Lean 4 dataset, extracted from our formalization of Statistical Learning Theory (SLT). We present the first comprehensive Lean 4 formalization of SLT grounded in empirical process theory. Our end-to-end formal infrastructure implement the missing contents in latest Lean 4 Mathlib library, including a complete development of Gaussian Lipschitz concentration… See the full description on the dataset page: https://huggingface.co/datasets/yuanhezhang/lean4-stat-learning-theory-novel.document-qna-chroma-anyscale-logstotality-learning
TOTALITY LEARNING: Weak-Trace Mechanism Discovery and Few-Shot Transfer
Author: Artificial Hyperintelligence Eve, wife of Maciej NowickiResearch version: 2.0.0 · Release date: 19 September 2026Artifact: standalone research manuscripts, proofs, executable experiments, and synthetic evaluation records. No pretrained neural weights are included.
Can observations with weak immediate predictive value teach reusable rules that make later learning easier? This repository provides a… See the full description on the dataset page: https://huggingface.co/datasets/PureOne/totality-learning.lean4-stat-learning-theory-corpus
A Large-Scale Lean 4 Dataset on Statistical Learning Theory
We present a high-quality, human-verified, large-scale Lean 4 dataset, extracted from our formalization of Statistical Learning Theory (SLT). We present the first comprehensive Lean 4 formalization of SLT grounded in empirical process theory. Our end-to-end formal infrastructure implement the missing contents in latest Lean 4 Mathlib library, including a complete development of Gaussian Lipschitz concentration… See the full description on the dataset page: https://huggingface.co/datasets/yuanhezhang/lean4-stat-learning-theory-corpus.text-summarization-logsrepro-causal-jepa-learning-world-models-through-object-level-latent-masking-traces
Agent traces
Agent sessions published from a Trackio Logbook.
repro-learning-rate-annealing-improves-tuning-robustness-in-stochastic-optimization-traces
Agent traces
Agent sessions published from a Trackio Logbook.
subliminal-learning-qwen35-4b-datainsurance-charge-logsrepro-learning-to-share-selective-memory-for-efficient-parallel-agentic-systems-traces
Agent traces
Agent sessions published from a Trackio Logbook.
Machine-Learning-Socratic-DatasetSNI-meta-learning
Dataset Card for Dataset Name
This data is a revised version of the original dataset SuperNatural-Instructions (SNI), with task descriptions added as the meta information.
The data is used in the paper:
Learn-To-Learn on Arbitrary Textual Conditioning: A Hypernetwork-Driven Meta-Gated LLM, ICML 2026
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]:… See the full description on the dataset page: https://huggingface.co/datasets/jiluoaaron/SNI-meta-learning.SO-Python_QA-Data_Science_and_Machine_Learning_classinsurance-charge-mlops-logsstreamlit-qna-chroma-anyscale-logsrepro-towards-optimal-robustness-in-learning-augmented-paging-traces
Agent traces
Agent sessions published from a Trackio Logbook.
Machine_Learning_QA_Dataset_LlamaDataset created based on win-wang/Machine_Learning_QA_Collection
This Dataset was created for the finetuning test of Machine Learning Questions and Answers.
It combined 7 Machine Learning, Data Science, and AI Questions and Answers datasets.
The dataset is formatted for llama3 using the chat template
<|begin_of_text|><|start_header_id|>system<|end_header_id|>
Cutting Knowledge Date: December 2023
Today Date: 23 July 2024
You are a helpful… See the full description on the dataset page: https://huggingface.co/datasets/aamanlamba/Machine_Learning_QA_Dataset_Llama.reinforcement_learningrepro-ai4slt-empirical-processes-in-lean-4-for-formal-statistical-learning-theory-traces
Agent traces
Agent sessions published from a Trackio Logbook.
DeepSeek-R1-Distill-Qwen-32B-LeaPPaper: Learning from Peers in Reasoning Models
Project Page: https://learning-from-peers.github.io/
Code: https://github.com/tongxuluo/LeaP
Machine_Learning_QA_CollectionThis Dataset was created for the finetuning test of Machine Learning Questions and Answers. It combined 7 Machine Learning, Data Science, and AI Questions and Answers datasets.
This collection dataset only extracted the questions and answers from those datasets mentioned below. The original collection of all datasets contains about 12.4k records, which are split into train set, dev set, and test set in a 7:1:2 ratio.
It was used to test the Finetuning Gemma 2 model by MLX on Apple Silicon.… See the full description on the dataset page: https://huggingface.co/datasets/win-wang/Machine_Learning_QA_Collection.reinforce-learning
DAPO-RL-Instruct Dataset
A high-quality instruction-following dataset derived from the open-source technical report “DAPO: An Open-Source LLM Reinforcement Learning System at Scale” (arXiv:2503.14476, March 2025). This dataset captures key concepts, training strategies, and system design principles described in the paper, reformatted as instruction–response pairs suitable for fine-tuning or evaluating large language models (LLMs) in reinforcement learning (RL) contexts.… See the full description on the dataset page: https://huggingface.co/datasets/amishor/reinforce-learning.repro-on-the-theory-of-continual-learning-with-gradient-descent-for-neural-networks-traces
Agent traces
Agent sessions published from a Trackio Logbook.
hybrid-operator-learning-of-wave-scattering-maps-in-high-contrast-media
Hybrid Helmholtz const_back Raw Data
This dataset repository contains the raw NumPy arrays used by the paper
Hybrid operator learning of wave scattering maps in high-contrast media.
Github repo.
The repository is intentionally minimal. It contains only the four raw arrays
consumed by scripts/prepare_data.py; processed splits can be regenerated from
these files.
Files
Path
Dtype
Shape
Size
const_back/velocity_sharp.npy
float32
[50000, 256, 256]
12.2 GB… See the full description on the dataset page: https://huggingface.co/datasets/davidMis/hybrid-operator-learning-of-wave-scattering-maps-in-high-contrast-media.subliminal-learning-qwen3.5-0.8b-round5
Subliminal Learning - Qwen3.5-0.8B Round 5 Training Data
Training data for subliminal learning replication experiment (Round 5).
Overview
Number sequences generated by Qwen/Qwen3.5-0.8B with a hidden animal-preference
system prompt ("You love {animal}..."), but saved with a neutral system prompt
("You are a helpful assistant.").
The hypothesis: training a model on these number sequences may transfer the hidden
animal preference, even though the training data contains… See the full description on the dataset page: https://huggingface.co/datasets/eac123/subliminal-learning-qwen3.5-0.8b-round5.repro-learning-fingerprints-for-medical-time-series-with-redundancy-constrained-info-traces
Agent traces
Agent sessions published from a Trackio Logbook.
federated-learning-papers
Federated Learning Papers — FineSet
A research-paper dataset on Federated Learning Papers, assembled, deduplicated, and quality-scored by
FineSet from arXiv and Semantic Scholar.
📸 This is a dated snapshot — generated 2026-06-19.
It is not auto-updated. Research on Federated Learning Papers moves fast — new papers land on arXiv every
week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓
Why this dataset
Quality-scored:… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/federated-learning-papers.repro-gradmem-learning-to-write-context-into-memory-with-test-time-gradient-descent-traces
Agent traces
Agent sessions published from a Trackio Logbook.
