CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01VmaxRL /SWE-smithtext100K<n<1M0 likes289 downloads5mo agoHugging Face02Smiling /webnovels-entextn<1K0 likes141 downloads5y agoHugging Face03mllab /smiles-2025text1K<n<10K0 likes104 downloads1y agoHugging Face04Daniel4190 /filtered_models_swe_smithtext1K<n<10K0 likes97 downloads1y agoHugging Face05reflectio /swe-smith-frozen-trajectories-openai SWE-Smith Frozen Trajectories — OpenAI Wire Format This dataset is the OpenAI chat-completions wire-format release of reflectio/swe-smith-frozen-trajectories, derived from the tool split of SWE-bench/SWE-smith-trajectories. It is a serving-performance workload for realistic multi-turn coding-agent histories. It can be used to measure request throughput, input/output token throughput, TTFT, TPOT, streaming behavior, and prefix-cache reuse. It is not a coding-correctness… See the full description on the dataset page: https://huggingface.co/datasets/reflectio/swe-smith-frozen-trajectories-openai.tabulartext-generation10K<n<100K0 likes96 downloads26d agoHugging Face06reflectio /swe-smith-frozen-trajectories SWE-Smith Frozen Trajectories This dataset is a serving-performance workload derived from the tool split of SWE-bench/SWE-smith-trajectories. It is designed for measuring throughput, request rate, time to first token, inter-token latency, and prefix-cache behavior with realistic multi-turn coding agent histories. It is not a coding-correctness benchmark. The tested model's responses are not executed or scored. Processing Keep trajectories generated by… See the full description on the dataset page: https://huggingface.co/datasets/reflectio/swe-smith-frozen-trajectories.tabulartext-generation10K<n<100K0 likes93 downloads27d agoHugging Face07Smilyai-labs /ChatPILE-Casual ChatPILE v3.0 - Clean Dataset High-quality conversational AI dataset with authentic Gen Z personality patterns Overview ChatPILE v3.0 Clean is a curated dataset of 121,680 unique conversational AI examples, featuring: ✅ 100% unique conversations (duplicates removed) ✅ Authentic Gen Z communication style ✅ Natural conversation flow (4-8 turns) ✅ ChatML format for easy training ✅ 20+ diverse topics ✅ 6 distinct personality modes Dataset Details Total Examples:… See the full description on the dataset page: https://huggingface.co/datasets/Smilyai-labs/ChatPILE-Casual.text100K<n<1M1 likes51 downloads11mo agoHugging Face08smiled0g /preflop_gto1M<n<10M0 likes47 downloads2y agoHugging Face09mok0102 /SMILE-Next SMILE-Next: Teaching Large Language Models to Detect, Classify, and Reason about Laughter This repository contains the official benchmark dataset forSMILE-Next: Teaching Large Language Models to Detect, Classify, and Reason about Laughter. SMILE-Next is a multimodal instruction-following benchmark for laughter understanding. It includes tasks such as laughter detection, laugh-type classification, and reasoning about why laughter occurs. Dataset Splits SMILE-Next… See the full description on the dataset page: https://huggingface.co/datasets/mok0102/SMILE-Next.text1K<n<10K0 likes41 downloads3mo agoHugging Face10Keeby-smilyai /H4-ultrachat-jsonltext100K<n<1M1 likes38 downloads1y agoHugging Face11DaertML /SMILES-34K SMILES-34K Introduction This dataset provides a QA dataset meant for model SFT about chemistry. It provides answers to multiple types of chemistry problems, taking a SMILES equation as the source component. The dataset has been synthetically generated to provide a summarized CoT, as longer CoTs will lag the resolution and usually produce errors in trained models. text10K<n<100K1 likes38 downloads11mo agoHugging Face12smirki /Deepseek_Data_pulltabular10K<n<100K0 likes38 downloads11mo agoHugging Face13smirki /Fire2text100K<n<1M0 likes27 downloads11mo agoHugging Face14sminpark /ds-alpha-small-dataset-v1.1text1K<n<10K0 likes26 downloads3y agoHugging Face15bcywinski /taboo-smile taboo-smile This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT). Usage from datasets import load_dataset # Load the dataset dataset = load_dataset("bcywinski/taboo-smile") Format The dataset is in JSONL format where each line contains a conversation record suitable for training chat models. texttext-generationn<1K0 likes25 downloads1y agoHugging Face16ticoAg /MeChat_smiletext10K<n<100K4 likes23 downloads3y agoHugging Face17smikulas /MNLP_M2_rag_documents MNLP_M2_rag_documents This is a sample set of documents for use in Retrieval-Augmented Generation (RAG) evaluation. text100K<n<1M0 likes23 downloads1y agoHugging Face18CoopReason /Kernel-Smith-RL-2KIf this work is useful to you, please cite: @article{DBLP:journals/corr/abs-2603-28342, author = {He Du and Qiming Ge and Jiakai Hu and Aijun Yang and Zheng Cai and Zixian Huang and Sheng Yuan and Qinxiu Cheng and Xinchen Xie and Yicheng Chen and Yining Li and Jiaxing Xie and… See the full description on the dataset page: https://huggingface.co/datasets/CoopReason/Kernel-Smith-RL-2K.text1K<n<10K0 likes23 downloads2mo agoHugging Face19smith3015 /bitcoin_dailytextn<1K0 likes22 downloads2y agoHugging Face20smirki /sft-mix-e0018203text100K<n<1M1 likes19 downloads6mo agoHugging Face21Smileoua /Code-Feedback OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement [🏠Homepage] | [🛠️Code] Introduction OpenCodeInterpreter is a family of open-source code generation systems designed to bridge the gap between large language models and advanced proprietary systems like the GPT-4 Code Interpreter. It significantly advances code generation capabilities by integrating execution and iterative refinement functionalities. For further information and… See the full description on the dataset page: https://huggingface.co/datasets/Smileoua/Code-Feedback.textquestion-answering10K<n<100K0 likes19 downloads2mo agoHugging Face22Smilyai-labs /Sam-1-large-identity-and-safetya dataset teaching the newest sam 1 large LLM its identity texttext-generation10K<n<100K1 likes18 downloads1y agoHugging Face23aeronautcanuk /smile-fine-tunetext10K<n<100K0 likes17 downloads2y agoHugging Face24smikulas /MNLP_M3_rag_documents MNLP_M3_rag_documents This is a sample set of documents for use in Retrieval-Augmented Generation (RAG) evaluation. text10K<n<100K0 likes16 downloads1y agoHugging Face25Chakita /SMILEContributors: Baisakhi Sarkar, Chakita Muttaraju, Xinyi (Cindy) Lyu Introduction SMILE (Synthetic Multi-turn Interactions for Learning Ethics) is a synthetic dataset consisting of multi-turn, text + image conversations between a human and an AI agent focusing on improving multimodal model performance on the 3Hs (Helpful, Honest, Harmless) as well as for implementing necessary safety and privacy restrictions such as not identifying persons from a given image. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Chakita/SMILE.texttext-generation1K<n<10K1 likes15 downloads2y agoHugging Face26smikulas /MNLP_M3_rag_documents_1 MNLP_M3_rag_documents This is a sample set of documents for use in Retrieval-Augmented Generation (RAG) evaluation. text1K<n<10K0 likes15 downloads1y agoHugging Face27smitathkr1 /mineralstextn<1K0 likes14 downloads2y agoHugging Face28smirki /TC-SFTtext100K<n<1M0 likes14 downloads7mo agoHugging Face29CoopReason /Kernel-Smith-Seed-59KIf this work is useful to you, please cite: @article{DBLP:journals/corr/abs-2603-28342, author = {He Du and Qiming Ge and Jiakai Hu and Aijun Yang and Zheng Cai and Zixian Huang and Sheng Yuan and Qinxiu Cheng and Xinchen Xie and Yicheng Chen and Yining Li and Jiaxing Xie and… See the full description on the dataset page: https://huggingface.co/datasets/CoopReason/Kernel-Smith-Seed-59K.tabular10K<n<100K0 likes14 downloads2mo agoHugging Face30sminpark /ds-alpha-small-dataset-v1.2text10K<n<100K0 likes13 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.