datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
text-code-galeras-code-generation-from-docstring-3k-dedupedsummarize_from_feedback_tldr_3_filteredThis is the query dataset taken directly from https://github.com/openai/summarize-from-feedback/tree/700967448d10004279f138666442bf1497d0e705#reddit-tldr-dataset
ChatHaruhi-from-RoleLLMAdapt English Role in RoleBench into ChatHaruhi format
only using profiles part in ZenMoore/RoleBench
Great thanks to on authors of RoleLLM!
usage:
# if you pip installed chatharuhi it should be
# from chatharuhi import ChatHaruhi
from ChatHaruhi import ChatHaruhi
chatbot = ChatHaruhi( role_from_hf = 'silk-road/ChatHaruhi-from-RoleLLM/Sherlock Holmes', \
llm = 'openai',
embedding = 'bge_en')
response = chatbot.chat(role='Police Chief', text = 'Oh… See the full description on the dataset page: https://huggingface.co/datasets/silk-road/ChatHaruhi-from-RoleLLM.DATASET_FROM_ALL_DOMAINSDeSTA-AQA5M-FROM-Llama3.1-8B-Instruct📑 Paper | 👩💻 Github | 🤗 Model | 🤗 Dataset
DeSTA-AQA5M comprises 50 speech, environmental sound, and music datasets, totaling over 7,000 hours of audio.
Our training framework centers on self-generated response for efficient cross-modal alignment. (see our paper!). In DeSTA, each audio clip is first transformed into a textual description using its metadata. A Large Language Model (LLM) is then prompted with this description to self-generate a response.
Ultimately, we construct a… See the full description on the dataset page: https://huggingface.co/datasets/DeSTA-ntu/DeSTA-AQA5M-FROM-Llama3.1-8B-Instruct.EOT-2004-Raw
End of Term 2004
Original dump: https://eotarchive.org/data/data-2004/
The End of Term Web Archive is a crawl of U.S. government websites conducted at the end of each presidential administration. This is a filtered version of the 2004 crawl.
Notice
This dataset is still a work in progress.
Data Curation
We download the 2004 EOT WARC files and parse the HTML using Trafilatura. We then filter the extracted text by length (minimum of 550 characters)… See the full description on the dataset page: https://huggingface.co/datasets/fromziro/EOT-2004-Raw.jetoncount_corpus
JetonCount's Corpus
This is the corpus used to train JetonCount.
JSONL Format
{
"source_file": "token_stats\\HuggingFaceFW_fineweb-edu00000.jsonl",
"dataset_dir": "HuggingFaceFW/fineweb-edu",
"index": 4020,
"chars": 2178,
"words": 335,
"avg_chars_per_word": 5.504478,
"longest_word_chars": 33,
"punctuation_ratio": 0.037649,
"symbol_ratio": 0.00551,
"tokens": 664,
"vocab_size": 2560,
"tokenizer_dir": "fromziro/Er-Tiny-1.3M"
}… See the full description on the dataset page: https://huggingface.co/datasets/fromziro/jetoncount_corpus.all-skills-from-skills-sh
Dataset Overview
This dataset is collected from skills.sh, a website that aggregates various command-line skills. Each skill typically includes a brief description, detailed documentation, and an installation command.
We have crawled all publicly available skill pages, resulting in approximately 40,000 to 50,000 skill records. The data is stored in a structured format, making it suitable for analysis, retrieval, or further development.
Field Descriptions
Each skill… See the full description on the dataset page: https://huggingface.co/datasets/tickleliu/all-skills-from-skills-sh.Claude-classified_from-Manusagentsreal: Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
but categorized, retaining only the complete sections from Claude.
text-to-ocl-from-ecore
Introduction
This is a small size dataset containing 52 meta-models (EMF files and PlantUML descriptions), 369 OCL constraints and 369 constraint specification in natural language.
The meta-models and OCL constraints are collected from open source github projects and are (syntactically) processable by Eclipse.
The constraint specifications of OCL constraints are generated via GPT-4-Turbo.
The meta-models can be found in models\
Usage
Generation of OCL constraints based on… See the full description on the dataset page: https://huggingface.co/datasets/fpan/text-to-ocl-from-ecore.text-to-xmi-from-ecoreThis is a small test set for XMI instance model generation task.
It containing 26 pairs of meta-models (Ecore), specifications (natural language) and instance models (XMI).
In each pair, the meta-model and instance model share the same name. To proper open the instance model in Eclipse EMF, the instance model and meta-model should be placed in the same folder.
The meta-models are selected from https://huggingface.co/datasets/fpan/text-to-ocl-from-ecore.
The specifications are generated via… See the full description on the dataset page: https://huggingface.co/datasets/fpan/text-to-xmi-from-ecore.py-docs-2004
Python Docs 2004
Original dump: https://www.python.org/ftp/python/doc/
Python Docs 2004 is a filtered and cleaned collection of Python documentation from every major Python release published before 2004.
Stats
Version
Size
Lines
2.3
2.2MB
1215
2.2
1.7MB
1142
2.1
1.3MB
891
2.0
1.2MB
895
1.6
1MB
720
1.5
837KB
449
1.4
744KB
397
1.3
569KB
408
1.2
513KB
384
Total
10.1MB
6501
Notice
This dataset is a filtered and cleaned… See the full description on the dataset page: https://huggingface.co/datasets/fromziro/py-docs-2004.DeepSeek-R1-Distill-Qwen-32B-LeaPPaper: Learning from Peers in Reasoning Models
Project Page: https://learning-from-peers.github.io/
Code: https://github.com/tongxuluo/LeaP
acubench
AcuBench
Indication-based acupoint-set recommendation. Given an indication /
symptom string (e.g. "Headache"), predict the set of WHO-standard acupoints
indicated for it, grounded in AcuKG's Indication table, over a fixed
361-point label space. AcuBench is a small (446-sample) benchmark with a
dedicated conformal-prediction calibration split.
Not a clinical prescription benchmark. A row's acupoint set is
"acupoints indicated for this symptom in AcuKG", i.e. a candidate pool… See the full description on the dataset page: https://huggingface.co/datasets/fromhope/acubench.HuggingChat-AI-Assistants-Deleted-System-Promptskorean_code_reviews_from_githubrose_code_samples
rose_code samples (pass@8 rollouts)
vLLM pass@8 samples on the CL-From-Nothing/rose_code train split (23,688 codeforces stdin/stdout
problems), scored by the deepcoder verifier (reward=1.0 iff all test cases pass).
Qwen3-1.7B/ — student model rollouts. 23,688 questions × 8 samples = 189,504 lines.
Qwen3-4B-Thinking-2507/ — teacher model rollouts.
Sampling: temperature 0.7, top_p 0.9, max_tokens 16384, 8 samples/question (pass@8).
Each cluster file holds a contiguous… See the full description on the dataset page: https://huggingface.co/datasets/CL-From-Nothing/rose_code_samples.polaris_rose_rollouts_olmo3-7b_from_qwen3-30b-a3b_cutoff4096_240steps
Cross-tokenizer ROSE rollouts — Olmo-3-7B-Think-SFT ← Qwen3-30B-A3B-Thinking-2507
Every assembled row of a complete 240-step online-ROSE run: 61,440 rows, the teacher's
actual continuation for each, and the token accounting behind it.
The student writes a 4096-token prefix in its own vocabulary (100278). That prefix is
decoded to text, the teacher is shown it under its own chat template, and the teacher's
reply comes back as text and is tokenised into the student's vocabulary.… See the full description on the dataset page: https://huggingface.co/datasets/SeanWang0027/polaris_rose_rollouts_olmo3-7b_from_qwen3-30b-a3b_cutoff4096_240steps.internal-state-from-gibberish
Injection-method A/B data (introspection-leakage robustness check)
Data behind the "Robustness: does the leak depend on the injection method?" section of
reports/experiment.md. Each .pt is one full collection run for one Qwen2.5 model under one injection
method, at the original per-model dose (auto-tuner off, so the dose is identical across arms).
Builds on open-introspection (Otto Stegmaier) — concept set, difference-vector extraction, and
layer/strength calibration. See the… See the full description on the dataset page: https://huggingface.co/datasets/ErrareHumanumEst/internal-state-from-gibberish.RLVE_envs65_data1000repro-flat-minima-and-generalization-insights-from-stochastic-convex-optimization-traces
Agent traces
Agent sessions published from a Trackio Logbook.
arxiv-abstracts-2004
ArXiv Abstracts 2004
Original Dataset: common-pile/arxiv_abstracts
ArXiv-Abstracts-2004 is a filtered collection of abstracts from the Common-Pile ArXiv dataset containing works created on or before 2004.
Stats
Size (MB)
Lines
351MB
303,761
Note: The lines, in the .jsonl file, are ordered from oldest to newest.
Notice
We do not claim ownership of or credit for any prior work done by the Common-Pile team. This dataset is only a… See the full description on the dataset page: https://huggingface.co/datasets/fromziro/arxiv-abstracts-2004.code-code-galeras-code-completion-from-docstring-3k-dedupedROSE-polaris-popesummarize-from-feedback-prorepro-a-theory-of-learning-data-statistics-in-diffusion-models-from-easy-to-hard-traces
Agent traces
Agent sessions published from a Trackio Logbook.
repro-from-prior-to-pro-efficient-skill-mastery-via-distribution-contractive-rl-traces
Agent traces
Agent sessions published from a Trackio Logbook.
extract-metrics-from-log-fixtures-cmskdlip
Extract Metrics from Log Fixtures
A dataset of extract metrics from log fixtures examples for training and evaluation. Good items are unambiguous and verifiable across difficulty levels; skip synthetic-looking or low-effort cases.
About
This dataset was produced by the DataBounty community and published here as part of an open, karma-only program.
Contributor items exported: 1000
Language: Python
Framework: Community
License: CC-BY-4.0
Contributors… See the full description on the dataset page: https://huggingface.co/datasets/databounty-io/extract-metrics-from-log-fixtures-cmskdlip.GeoGPT_Training_Data_from_Geoscience_Subset_of_CommonCrawl
Description
This dataset is a geoscience-specific subset of CommonCrawl used for GeoGPT training. CommonCrawl is a free and open repository of web crawl data with over 250 billion web pages and is widely used by leading large language models. We apply data mining algorithms to extract geoscience-related content from this vast dataset.
This dataset comprises 12,414,268 samples, each containing the following metadata to trace the data source within CommonCrawl:
id (string):… See the full description on the dataset page: https://huggingface.co/datasets/GeoGPT-Research-Project/GeoGPT_Training_Data_from_Geoscience_Subset_of_CommonCrawl.RLVE-Qwen3-1.7B-Pass1-Rollouts
RLVE teacher rollouts — Qwen3-1.7B (pass@1)
Teacher rollouts for on-policy distillation on the RLVE environment suite.
Teacher / sampler: Qwen3-1.7B
Source prompts: RLVE train split — 9000 questions across RLVE-Eval Gym
environments (counting / combinatorics / optimization tasks)
Sampling: 1 sample/question (pass@1) = 9000 records,
temperature 0.7, max 4096 new tokens
Rewards: inline RLVE-Eval Gym verifier score (continuous, in [-1, 1]).
Teacher accuracy (reward>0): 20 / 9000 =… See the full description on the dataset page: https://huggingface.co/datasets/CL-From-Nothing/RLVE-Qwen3-1.7B-Pass1-Rollouts.
