datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Gutenberg-BookCorpus-Cleaned-Data-English
Gutenberg-BookCorpus-Cleaned-Data-English
This dataset is been cleaned and preprocessed using Gutenberg_English_Preprocessor class method (given below) from preference Kaggle dataset 75,000+ Gutenberg Books and Metadata 2025. This dataset is only specialisation for english contented with rights as "Public domain in the USA" hence you can free used it anywhere.
Following reference metadata of Gutenberg is also available and downloaded it using following CLI command below :-
pip… See the full description on the dataset page: https://huggingface.co/datasets/incredible45/Gutenberg-BookCorpus-Cleaned-Data-English.cleaned_turkish_embedding_model_training_data_colabsinhala-22gb-cleaned-datasetcleaned_turkish_embedding_model_training_data_colab
Citation
If you use this dataset in your research, please cite the following paper:
@inproceedings{baysan-gungor-2025-tr,
title = "{TR}-{MTEB}: A Comprehensive Benchmark and Embedding Model Suite for {T}urkish Sentence Representations",
author = "Baysan, Mehmet Selman and
Gungor, Tunga",
booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2025",
month = nov,
year = "2025",
address = "Suzhou, China",
publisher =… See the full description on the dataset page: https://huggingface.co/datasets/trmteb/cleaned_turkish_embedding_model_training_data_colab.cleaned-data-split-0
Dataset Card for "cleaned-data-split-0"
More Information needed
prompt_injection_cleaned_dataset-v2
Dataset Card for "prompt_injection_cleaned_dataset-v2"
More Information needed
prompt_injection_cleaned_dataset
Dataset Card for "prompt_injection_cleaned_dataset"
More Information needed
tool-reasoning-sft-CODING-text_to_terminal_v2-sft-tool-use-agent-data-cleaned-rectified
Text to Terminal, v2 — Cleaned & Rectified
👥 Follow the Author
Aman Priyanshu
Overview
This dataset is a cleaned, combined, and thinking-augmented version of muellerzr/text_to_terminal_v2. It pairs natural language instructions with their corresponding terminal/bash commands, now augmented with explicit <think> reasoning traces that model the step-by-step thought process before producing the final command.The restructuring approach is directly… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-CODING-text_to_terminal_v2-sft-tool-use-agent-data-cleaned-rectified.reposvul_processed_dataset_cleanedcleaned_data
Cleaned Tabula Muris Senis Single-Cell Data and other aging datasets
This dataset contains LLM-cleaned single-cell transcriptomic annotations from the Tabula Muris Senis project, specifically for mouse tissues processed with SmartSeq2, and ALL OTHER DATASETS WITH AGING IN THE FILENAME :-) . The cleaning and annotation were performed using large language models (OpenAI and Claude), enabling enriched metadata and corrected cell type labels.
🧬 Over 1.3 million rows and 78.17 GB… See the full description on the dataset page: https://huggingface.co/datasets/longevity-db/cleaned_data.tool-reasoning-sft-TOOLS-toolace-sft-tool-use-agent-data-cleaned-rectified
ToolACE - Tool-Use Agent Data Cleaned & Rectified
👥 Follow the Author
Aman Priyanshu
Overview
This dataset is a cleaned and restructured version of the Team-ACE/ToolACE dataset. ToolACE is a high-quality conversational tool-use dataset containing 11,300+ examples of natural language interactions requiring function calling across diverse domains. This version converts the original OpenAI function-call format into a standardized multi-turn tool-use… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-TOOLS-toolace-sft-tool-use-agent-data-cleaned-rectified.song_dataset_training_20s_cleanedtool-reasoning-sft-TOOLS-hermes_reasoning_tool_use-data-cleaned-rectified
Hermes Reasoning Tool Use — Cleaned & Rectified
👥 Follow the Author
Aman Priyanshu
Overview
This dataset is a cleaned and restructured version of interstellarninja/hermes_reasoning_tool_use. The original dataset uses the Hermes/NousResearch multi-turn format with from/value fields and embedded <think> + <tool_call> tags inside single gpt turns. This version converts it into a strict multi-turn conversation structure with validated role transitions.… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-TOOLS-hermes_reasoning_tool_use-data-cleaned-rectified.tool-reasoning-sft-MEMORY-mem_agent-sft-data-cleaned-rectified-408k
mem_agent-sft-data-cleaned-rectified
Multi-turn long-context memory-agent SFT dataset with explicit reasoning traces, structured tool calls, and sequential chunk-processing sub-chains.
Schema
Column
Type
Description
messages
string (JSON)
JSON-serialized list of {role, content} dicts. Roles: system, user, reasoning, tool_call, tool_output, answer
core_chain_OR_subcall
string
"core_chain" (full orchestration trace) or "subcall" (single chunk-processing step)… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-MEMORY-mem_agent-sft-data-cleaned-rectified-408k.tool-reasoning-sft-CODING-allenai-SERA-data-cleaned-rectified
SERA — Consolidated & Rectified
211,360 multi-turn SWE-agent coding trajectories from the SERA (Soft-Verified Efficient Repository Agents) project, consolidated from 4 source datasets into a single file with strict reasoning + tool-call format and validated FSM transitions.
Origin
Derived from Allen AI's Open Coding Agents release:
Source Dataset
Rows
Teacher
Scale
Rollout
allenai/Sera-4.5A-Full-T1
72,118
GLM-4.5-Air
full
T1
allenai/Sera-4.5A-Full-T2
66,337… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-CODING-allenai-SERA-data-cleaned-rectified.cleaned-mongolian-datasetmbti-Personalitycafe-cleaned-databryn-hauk-zemo-alvani-fieldwork-data-cleanedisear-cleaned-datasettool-reasoning-sft-RESEARCH-grill-lab-browsecomp-plus-runs-data-cleaned-rectified
Tool-Reasoning SFT — BrowseComp-Plus Runs (Cleaned & Rectified)
Multi-turn tool-use reasoning trajectories derived from grill-lab/browsecomp-plus-runs, converted to a structured SFT format following the interstellarninja/hermes_reasoning_tool_use convention.
Source
Based on the execution trajectories from "Revisiting Text Ranking in Deep Research" (arXiv:2602.21456):
Original data: grill-lab/browsecomp-plus-runs (MIT)
Format
Each row contains a messages… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-RESEARCH-grill-lab-browsecomp-plus-runs-data-cleaned-rectified.salami_data_cleaned_fullsynthdog_cleaned
synthdog_cleaned
The synthdog__x family of the ElliotVL supervised-fine-tuning pool, after VLM cleaning.
images
445,694
QA turns
1,613,204
answers rewritten by the cleaning pass
0
QA created by the cleaning pass (new_qa)
not measured for this family
shards
547
How this was cleaned
A vision-language model read each image together with its QA and judged the item. The pass is
not a filter that only removes rows — it rewrites answers it finds… See the full description on the dataset page: https://huggingface.co/datasets/Elliot-Data/synthdog_cleaned.tool-reasoning-sft-RESEARCH-dr-tulu-sft-deep-research-agent-data-cleaned-rectified
Deep Research - Tulu SFT Data Cleaned Rectified
👥 Follow the Author
Supriti Vijay
Overview
This dataset is a cleaned and restructured version of the DR-TULU SFT dataset released by AllenAI's RL Research team. The original DR-TULU dataset represents significant work in creating high-quality training data for reasoning-enhanced language models with tool use capabilities. This version addresses structural issues in the original release while preserving… See the full description on the dataset page: https://huggingface.co/datasets/SupritiVijay/tool-reasoning-sft-RESEARCH-dr-tulu-sft-deep-research-agent-data-cleaned-rectified.tool-reasoning-sft-CODING-CoVe-12k-data-cleaned-rectified
CoVe-12K — Cleaned & Rectified
12,000 high-quality multi-turn interactive tool-use trajectories converted into a strict reasoning + tool-call format with validated FSM transitions. Covers airline booking/modification/cancellation and retail order management across two balanced domains.
Origin
Derived from Zichen1024/CoVe-12k, synthesized by the CoVe (Constraint-Verification) framework. Explicit constraints are fuzzified to guide a User Simulator LLM, and original… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-CODING-CoVe-12k-data-cleaned-rectified.DoclingMatix_cleaned
DoclingMatix_cleaned
The DoclingMatix__x family of the ElliotVL supervised-fine-tuning pool, after VLM cleaning.
images
588,763
QA turns
6,394,614
answers rewritten by the cleaning pass
0
QA created by the cleaning pass (new_qa)
not measured for this family
shards
503
How this was cleaned
A vision-language model read each image together with its QA and judged the item. The pass is
not a filter that only removes rows — it rewrites answers… See the full description on the dataset page: https://huggingface.co/datasets/Elliot-Data/DoclingMatix_cleaned.cardiology-cleaned_datasetDocmatix_merged_cleaned
Docmatix_merged_cleaned
The Docmatix_merged family of the ElliotVL supervised-fine-tuning pool, after VLM cleaning.
images
547,033
QA turns
6,199,743
answers rewritten by the cleaning pass
0
QA created by the cleaning pass (new_qa)
not measured for this family
shards
507
How this was cleaned
A vision-language model read each image together with its QA and judged the item. The pass is
not a filter that only removes rows — it rewrites… See the full description on the dataset page: https://huggingface.co/datasets/Elliot-Data/Docmatix_merged_cleaned.Solidity-Dataset-Cleanedcuria_2025_balanced_dataset_en_es_it_fr_de_cleanednhlcoding_cleaned_cpp_dataset
