CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01incredible45 /Gutenberg-BookCorpus-Cleaned-Data-English Gutenberg-BookCorpus-Cleaned-Data-English This dataset is been cleaned and preprocessed using Gutenberg_English_Preprocessor class method (given below) from preference Kaggle dataset 75,000+ Gutenberg Books and Metadata 2025. This dataset is only specialisation for english contented with rights as "Public domain in the USA" hence you can free used it anywhere. Following reference metadata of Gutenberg is also available and downloaded it using following CLI command below :- pip… See the full description on the dataset page: https://huggingface.co/datasets/incredible45/Gutenberg-BookCorpus-Cleaned-Data-English.text10K<n<100K15 likes1.2k downloads1y agoHugging Face02selmanbaysan /cleaned_turkish_embedding_model_training_data_colabtext10M<n<100M1 likes483 downloads1y agoHugging Face03ChamaraVishwajithRajapaksha /sinhala-22gb-cleaned-datasettext1M<n<10M0 likes430 downloads5mo agoHugging Face04trmteb /cleaned_turkish_embedding_model_training_data_colab Citation If you use this dataset in your research, please cite the following paper: @inproceedings{baysan-gungor-2025-tr, title = "{TR}-{MTEB}: A Comprehensive Benchmark and Embedding Model Suite for {T}urkish Sentence Representations", author = "Baysan, Mehmet Selman and Gungor, Tunga", booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2025", month = nov, year = "2025", address = "Suzhou, China", publisher =… See the full description on the dataset page: https://huggingface.co/datasets/trmteb/cleaned_turkish_embedding_model_training_data_colab.text10M<n<100M3 likes423 downloads10mo agoHugging Face05IndonesiaAI /cleaned-data-split-0 Dataset Card for "cleaned-data-split-0" More Information needed text1M<n<10M1 likes316 downloads3y agoHugging Face06imoxto /prompt_injection_cleaned_dataset-v2 Dataset Card for "prompt_injection_cleaned_dataset-v2" More Information needed text100K<n<1M11 likes297 downloads3y agoHugging Face07imoxto /prompt_injection_cleaned_dataset Dataset Card for "prompt_injection_cleaned_dataset" More Information needed tabular100K<n<1M6 likes218 downloads3y agoHugging Face08AmanPriyanshu /tool-reasoning-sft-CODING-text_to_terminal_v2-sft-tool-use-agent-data-cleaned-rectified Text to Terminal, v2 — Cleaned & Rectified 👥 Follow the Author Aman Priyanshu Overview This dataset is a cleaned, combined, and thinking-augmented version of muellerzr/text_to_terminal_v2. It pairs natural language instructions with their corresponding terminal/bash commands, now augmented with explicit <think> reasoning traces that model the step-by-step thought process before producing the final command.The restructuring approach is directly… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-CODING-text_to_terminal_v2-sft-tool-use-agent-data-cleaned-rectified.texttext-generation100K<n<1M0 likes198 downloads7mo agoHugging Face09vivi-yu /reposvul_processed_dataset_cleanedtext10K<n<100K0 likes146 downloads8mo agoHugging Face10longevity-db /cleaned_data Cleaned Tabula Muris Senis Single-Cell Data and other aging datasets This dataset contains LLM-cleaned single-cell transcriptomic annotations from the Tabula Muris Senis project, specifically for mouse tissues processed with SmartSeq2, and ALL OTHER DATASETS WITH AGING IN THE FILENAME :-) . The cleaning and annotation were performed using large language models (OpenAI and Claude), enabling enriched metadata and corrected cell type labels. 🧬 Over 1.3 million rows and 78.17 GB… See the full description on the dataset page: https://huggingface.co/datasets/longevity-db/cleaned_data.tabular1M<n<10M0 likes133 downloads1y agoHugging Face11AmanPriyanshu /tool-reasoning-sft-TOOLS-toolace-sft-tool-use-agent-data-cleaned-rectified ToolACE - Tool-Use Agent Data Cleaned & Rectified 👥 Follow the Author Aman Priyanshu Overview This dataset is a cleaned and restructured version of the Team-ACE/ToolACE dataset. ToolACE is a high-quality conversational tool-use dataset containing 11,300+ examples of natural language interactions requiring function calling across diverse domains. This version converts the original OpenAI function-call format into a standardized multi-turn tool-use… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-TOOLS-toolace-sft-tool-use-agent-data-cleaned-rectified.tabulartext-generation10K<n<100K0 likes125 downloads7mo agoHugging Face12sunbv56 /song_dataset_training_20s_cleanedaudio10K<n<100K1 likes116 downloads7mo agoHugging Face13AmanPriyanshu /tool-reasoning-sft-TOOLS-hermes_reasoning_tool_use-data-cleaned-rectified Hermes Reasoning Tool Use — Cleaned & Rectified 👥 Follow the Author Aman Priyanshu Overview This dataset is a cleaned and restructured version of interstellarninja/hermes_reasoning_tool_use. The original dataset uses the Hermes/NousResearch multi-turn format with from/value fields and embedded <think> + <tool_call> tags inside single gpt turns. This version converts it into a strict multi-turn conversation structure with validated role transitions.… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-TOOLS-hermes_reasoning_tool_use-data-cleaned-rectified.texttext-generation10K<n<100K2 likes107 downloads7mo agoHugging Face14AmanPriyanshu /tool-reasoning-sft-MEMORY-mem_agent-sft-data-cleaned-rectified-408k mem_agent-sft-data-cleaned-rectified Multi-turn long-context memory-agent SFT dataset with explicit reasoning traces, structured tool calls, and sequential chunk-processing sub-chains. Schema Column Type Description messages string (JSON) JSON-serialized list of {role, content} dicts. Roles: system, user, reasoning, tool_call, tool_output, answer core_chain_OR_subcall string "core_chain" (full orchestration trace) or "subcall" (single chunk-processing step)… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-MEMORY-mem_agent-sft-data-cleaned-rectified-408k.texttext-generation100K<n<1M0 likes107 downloads7mo agoHugging Face15AmanPriyanshu /tool-reasoning-sft-CODING-allenai-SERA-data-cleaned-rectified SERA — Consolidated & Rectified 211,360 multi-turn SWE-agent coding trajectories from the SERA (Soft-Verified Efficient Repository Agents) project, consolidated from 4 source datasets into a single file with strict reasoning + tool-call format and validated FSM transitions. Origin Derived from Allen AI's Open Coding Agents release: Source Dataset Rows Teacher Scale Rollout allenai/Sera-4.5A-Full-T1 72,118 GLM-4.5-Air full T1 allenai/Sera-4.5A-Full-T2 66,337… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-CODING-allenai-SERA-data-cleaned-rectified.texttext-generation100K<n<1M0 likes103 downloads7mo agoHugging Face16Billyyy /cleaned-mongolian-datasettext100K<n<1M2 likes99 downloads2y agoHugging Face17Lazanantenaina /mbti-Personalitycafe-cleaned-datatabular1K<n<10K0 likes95 downloads9d agoHugging Face18kkahadze /bryn-hauk-zemo-alvani-fieldwork-data-cleanedtext1K<n<10K0 likes94 downloads2y agoHugging Face19dalopeza98 /isear-cleaned-datasettext1K<n<10K1 likes90 downloads2y agoHugging Face20AmanPriyanshu /tool-reasoning-sft-RESEARCH-grill-lab-browsecomp-plus-runs-data-cleaned-rectified Tool-Reasoning SFT — BrowseComp-Plus Runs (Cleaned & Rectified) Multi-turn tool-use reasoning trajectories derived from grill-lab/browsecomp-plus-runs, converted to a structured SFT format following the interstellarninja/hermes_reasoning_tool_use convention. Source Based on the execution trajectories from "Revisiting Text Ranking in Deep Research" (arXiv:2602.21456): Original data: grill-lab/browsecomp-plus-runs (MIT) Format Each row contains a messages… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-RESEARCH-grill-lab-browsecomp-plus-runs-data-cleaned-rectified.texttext-generation10K<n<100K0 likes89 downloads7mo agoHugging Face21taufiqsyed /salami_data_cleaned_fullaudio10K<n<100K0 likes87 downloads2y agoHugging Face22Elliot-Data /synthdog_cleanedgated synthdog_cleaned The synthdog__x family of the ElliotVL supervised-fine-tuning pool, after VLM cleaning. images 445,694 QA turns 1,613,204 answers rewritten by the cleaning pass 0 QA created by the cleaning pass (new_qa) not measured for this family shards 547 How this was cleaned A vision-language model read each image together with its QA and judged the item. The pass is not a filter that only removes rows — it rewrites answers it finds… See the full description on the dataset page: https://huggingface.co/datasets/Elliot-Data/synthdog_cleaned.imagevisual-question-answering100K<n<1M0 likes83 downloads15d agoHugging Face23SupritiVijay /tool-reasoning-sft-RESEARCH-dr-tulu-sft-deep-research-agent-data-cleaned-rectified Deep Research - Tulu SFT Data Cleaned Rectified 👥 Follow the Author Supriti Vijay Overview This dataset is a cleaned and restructured version of the DR-TULU SFT dataset released by AllenAI's RL Research team. The original DR-TULU dataset represents significant work in creating high-quality training data for reasoning-enhanced language models with tool use capabilities. This version addresses structural issues in the original release while preserving… See the full description on the dataset page: https://huggingface.co/datasets/SupritiVijay/tool-reasoning-sft-RESEARCH-dr-tulu-sft-deep-research-agent-data-cleaned-rectified.tabulartext-generation10K<n<100K8 likes75 downloads10mo agoHugging Face24AmanPriyanshu /tool-reasoning-sft-CODING-CoVe-12k-data-cleaned-rectified CoVe-12K — Cleaned & Rectified 12,000 high-quality multi-turn interactive tool-use trajectories converted into a strict reasoning + tool-call format with validated FSM transitions. Covers airline booking/modification/cancellation and retail order management across two balanced domains. Origin Derived from Zichen1024/CoVe-12k, synthesized by the CoVe (Constraint-Verification) framework. Explicit constraints are fuzzified to guide a User Simulator LLM, and original… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-CODING-CoVe-12k-data-cleaned-rectified.texttext-generation10K<n<100K0 likes72 downloads7mo agoHugging Face25Elliot-Data /DoclingMatix_cleanedgated DoclingMatix_cleaned The DoclingMatix__x family of the ElliotVL supervised-fine-tuning pool, after VLM cleaning. images 588,763 QA turns 6,394,614 answers rewritten by the cleaning pass 0 QA created by the cleaning pass (new_qa) not measured for this family shards 503 How this was cleaned A vision-language model read each image together with its QA and judged the item. The pass is not a filter that only removes rows — it rewrites answers… See the full description on the dataset page: https://huggingface.co/datasets/Elliot-Data/DoclingMatix_cleaned.imagevisual-question-answering100K<n<1M0 likes70 downloads15d agoHugging Face26Alwaly /cardiology-cleaned_datasetimage10K<n<100K0 likes69 downloads1y agoHugging Face27Elliot-Data /Docmatix_merged_cleanedgated Docmatix_merged_cleaned The Docmatix_merged family of the ElliotVL supervised-fine-tuning pool, after VLM cleaning. images 547,033 QA turns 6,199,743 answers rewritten by the cleaning pass 0 QA created by the cleaning pass (new_qa) not measured for this family shards 507 How this was cleaned A vision-language model read each image together with its QA and judged the item. The pass is not a filter that only removes rows — it rewrites… See the full description on the dataset page: https://huggingface.co/datasets/Elliot-Data/Docmatix_merged_cleaned.imagevisual-question-answering100K<n<1M0 likes67 downloads15d agoHugging Face28braindao /Solidity-Dataset-Cleanedtext10K<n<100K2 likes65 downloads2y agoHugging Face29nandezgarcia /curia_2025_balanced_dataset_en_es_it_fr_de_cleanedtext10K<n<100K1 likes61 downloads2y agoHugging Face30nguyentruong-ins /nhlcoding_cleaned_cpp_datasettext1M<n<10M1 likes59 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.