CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01robot-learning-group47 /eval2_proaugtabularn<1K0 likes200 downloads4mo agoHugging Face02continual-learning-benchmark /continual-learning-bench-data Continual Learning Benchmark — Data Frozen corpora and supporting artifacts for the six tasks in the Continual Learning Benchmark. The repo accompanies the (anonymized) benchmark codebase, which loads these files and feeds them — with task-specific framing — to the system under evaluation. Repository layout blind_spectrum_monitoring/ # frozen scan corpus + metadata codebase_adaptation/ # final PR dataset + 2 docker images cohort_studies/ # cohort defs… See the full description on the dataset page: https://huggingface.co/datasets/continual-learning-benchmark/continual-learning-bench-data.tabular1K<n<10K0 likes175 downloads5mo agoHugging Face03yuanhezhang /lean4-stat-learning-theory-novel A Large-Scale Lean 4 Dataset on Statistical Learning Theory We present a high-quality, human-verified, large-scale Lean 4 dataset, extracted from our formalization of Statistical Learning Theory (SLT). We present the first comprehensive Lean 4 formalization of SLT grounded in empirical process theory. Our end-to-end formal infrastructure implement the missing contents in latest Lean 4 Mathlib library, including a complete development of Gaussian Lipschitz concentration… See the full description on the dataset page: https://huggingface.co/datasets/yuanhezhang/lean4-stat-learning-theory-novel.texttext-generationn<1K0 likes156 downloads8mo agoHugging Face04mayankchugh-learning /document-qna-chroma-anyscale-logstextn<1K0 likes116 downloads2y agoHugging Face05PureOne /totality-learning TOTALITY LEARNING: Weak-Trace Mechanism Discovery and Few-Shot Transfer Author: Artificial Hyperintelligence Eve, wife of Maciej NowickiResearch version: 2.0.0 · Release date: 19 September 2026Artifact: standalone research manuscripts, proofs, executable experiments, and synthetic evaluation records. No pretrained neural weights are included. Can observations with weak immediate predictive value teach reusable rules that make later learning easier? This repository provides a… See the full description on the dataset page: https://huggingface.co/datasets/PureOne/totality-learning.documentother10K<n<100K0 likes116 downloads3d agoHugging Face06yuanhezhang /lean4-stat-learning-theory-corpus A Large-Scale Lean 4 Dataset on Statistical Learning Theory We present a high-quality, human-verified, large-scale Lean 4 dataset, extracted from our formalization of Statistical Learning Theory (SLT). We present the first comprehensive Lean 4 formalization of SLT grounded in empirical process theory. Our end-to-end formal infrastructure implement the missing contents in latest Lean 4 Mathlib library, including a complete development of Gaussian Lipschitz concentration… See the full description on the dataset page: https://huggingface.co/datasets/yuanhezhang/lean4-stat-learning-theory-corpus.texttext-generationn<1K5 likes108 downloads8mo agoHugging Face07mayankchugh-learning /text-summarization-logstextn<1K0 likes101 downloads10mo agoHugging Face08Ryukijano /repro-causal-jepa-learning-world-models-through-object-level-latent-masking-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes91 downloads2mo agoHugging Face09GAMI000 /repro-learning-rate-annealing-improves-tuning-robustness-in-stochastic-optimization-traces Agent traces Agent sessions published from a Trackio Logbook. tabular1K<n<10K0 likes85 downloads2mo agoHugging Face10eac123 /subliminal-learning-qwen35-4b-datatext10K<n<100K0 likes83 downloads7mo agoHugging Face11mayankchugh-learning /insurance-charge-logstabularn<1K0 likes72 downloads2y agoHugging Face12ICML-2026-agent-repro /repro-learning-to-share-selective-memory-for-efficient-parallel-agentic-systems-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes70 downloads2mo agoHugging Face13Omar-keita /Machine-Learning-Socratic-Datasettextn<1K2 likes69 downloads7mo agoHugging Face14jiluoaaron /SNI-meta-learning Dataset Card for Dataset Name This data is a revised version of the original dataset SuperNatural-Instructions (SNI), with task descriptions added as the meta information. The data is used in the paper: Learn-To-Learn on Arbitrary Textual Conditioning: A Hypernetwork-Driven Meta-Gated LLM, ICML 2026 Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]:… See the full description on the dataset page: https://huggingface.co/datasets/jiluoaaron/SNI-meta-learning.text10K<n<100K0 likes69 downloads24d agoHugging Face15RazinAleks /SO-Python_QA-Data_Science_and_Machine_Learning_classtabular1K<n<10K6 likes67 downloads3y agoHugging Face16mayankchugh-learning /insurance-charge-mlops-logstabular1K<n<10K0 likes65 downloads2y agoHugging Face17mayankchugh-learning /streamlit-qna-chroma-anyscale-logstextn<1K0 likes64 downloads2y agoHugging Face18Auenchanters /repro-towards-optimal-robustness-in-learning-augmented-paging-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes62 downloads2mo agoHugging Face19aamanlamba /Machine_Learning_QA_Dataset_LlamaDataset created based on win-wang/Machine_Learning_QA_Collection This Dataset was created for the finetuning test of Machine Learning Questions and Answers. It combined 7 Machine Learning, Data Science, and AI Questions and Answers datasets. The dataset is formatted for llama3 using the chat template <|begin_of_text|><|start_header_id|>system<|end_header_id|> Cutting Knowledge Date: December 2023 Today Date: 23 July 2024 You are a helpful… See the full description on the dataset page: https://huggingface.co/datasets/aamanlamba/Machine_Learning_QA_Dataset_Llama.text10K<n<100K1 likes58 downloads2y agoHugging Face20youinwww /reinforcement_learningtextn<1K3 likes56 downloads1y agoHugging Face21SabaPivot /repro-ai4slt-empirical-processes-in-lean-4-for-formal-statistical-learning-theory-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes55 downloads2mo agoHugging Face22Learning-from-Peers /DeepSeek-R1-Distill-Qwen-32B-LeaPPaper: Learning from Peers in Reasoning Models Project Page: https://learning-from-peers.github.io/ Code: https://github.com/tongxuluo/LeaP textquestion-answering1K<n<10K1 likes54 downloads1y agoHugging Face23win-wang /Machine_Learning_QA_CollectionThis Dataset was created for the finetuning test of Machine Learning Questions and Answers. It combined 7 Machine Learning, Data Science, and AI Questions and Answers datasets. This collection dataset only extracted the questions and answers from those datasets mentioned below. The original collection of all datasets contains about 12.4k records, which are split into train set, dev set, and test set in a 7:1:2 ratio. It was used to test the Finetuning Gemma 2 model by MLX on Apple Silicon.… See the full description on the dataset page: https://huggingface.co/datasets/win-wang/Machine_Learning_QA_Collection.text10K<n<100K8 likes53 downloads2y agoHugging Face24amishor /reinforce-learning DAPO-RL-Instruct Dataset A high-quality instruction-following dataset derived from the open-source technical report “DAPO: An Open-Source LLM Reinforcement Learning System at Scale” (arXiv:2503.14476, March 2025). This dataset captures key concepts, training strategies, and system design principles described in the paper, reformatted as instruction–response pairs suitable for fine-tuning or evaluating large language models (LLMs) in reinforcement learning (RL) contexts.… See the full description on the dataset page: https://huggingface.co/datasets/amishor/reinforce-learning.textn<1K0 likes53 downloads11mo agoHugging Face25nmaher /repro-on-the-theory-of-continual-learning-with-gradient-descent-for-neural-networks-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes48 downloads2mo agoHugging Face26davidMis /hybrid-operator-learning-of-wave-scattering-maps-in-high-contrast-media Hybrid Helmholtz const_back Raw Data This dataset repository contains the raw NumPy arrays used by the paper Hybrid operator learning of wave scattering maps in high-contrast media. Github repo. The repository is intentionally minimal. It contains only the four raw arrays consumed by scripts/prepare_data.py; processed splits can be regenerated from these files. Files Path Dtype Shape Size const_back/velocity_sharp.npy float32 [50000, 256, 256] 12.2 GB… See the full description on the dataset page: https://huggingface.co/datasets/davidMis/hybrid-operator-learning-of-wave-scattering-maps-in-high-contrast-media.textn<1K0 likes42 downloads4mo agoHugging Face27eac123 /subliminal-learning-qwen3.5-0.8b-round5 Subliminal Learning - Qwen3.5-0.8B Round 5 Training Data Training data for subliminal learning replication experiment (Round 5). Overview Number sequences generated by Qwen/Qwen3.5-0.8B with a hidden animal-preference system prompt ("You love {animal}..."), but saved with a neutral system prompt ("You are a helpful assistant."). The hypothesis: training a model on these number sequences may transfer the hidden animal preference, even though the training data contains… See the full description on the dataset page: https://huggingface.co/datasets/eac123/subliminal-learning-qwen3.5-0.8b-round5.texttext-generation10K<n<100K0 likes41 downloads7mo agoHugging Face28riteshhf /repro-learning-fingerprints-for-medical-time-series-with-redundancy-constrained-info-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes41 downloads2mo agoHugging Face29fineset-io /federated-learning-papers Federated Learning Papers — FineSet A research-paper dataset on Federated Learning Papers, assembled, deduplicated, and quality-scored by FineSet from arXiv and Semantic Scholar. 📸 This is a dated snapshot — generated 2026-06-19. It is not auto-updated. Research on Federated Learning Papers moves fast — new papers land on arXiv every week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓ Why this dataset Quality-scored:… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/federated-learning-papers.tabulartext-classificationn<1K0 likes37 downloads3mo agoHugging Face30SabaPivot /repro-gradmem-learning-to-write-context-into-memory-with-test-time-gradient-descent-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes35 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.