datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SWE-smithwebnovels-ensmiles-2025filtered_models_swe_smithswe-smith-frozen-trajectories-openai
SWE-Smith Frozen Trajectories — OpenAI Wire Format
This dataset is the OpenAI chat-completions wire-format release of
reflectio/swe-smith-frozen-trajectories,
derived from the tool split of
SWE-bench/SWE-smith-trajectories.
It is a serving-performance workload for realistic multi-turn coding-agent
histories. It can be used to measure request throughput, input/output token
throughput, TTFT, TPOT, streaming behavior, and prefix-cache reuse. It is not
a coding-correctness… See the full description on the dataset page: https://huggingface.co/datasets/reflectio/swe-smith-frozen-trajectories-openai.swe-smith-frozen-trajectories
SWE-Smith Frozen Trajectories
This dataset is a serving-performance workload derived from the tool split of
SWE-bench/SWE-smith-trajectories.
It is designed for measuring throughput, request rate, time to first token,
inter-token latency, and prefix-cache behavior with realistic multi-turn coding
agent histories.
It is not a coding-correctness benchmark. The tested model's responses are
not executed or scored.
Processing
Keep trajectories generated by… See the full description on the dataset page: https://huggingface.co/datasets/reflectio/swe-smith-frozen-trajectories.ChatPILE-Casual
ChatPILE v3.0 - Clean Dataset
High-quality conversational AI dataset with authentic Gen Z personality patterns
Overview
ChatPILE v3.0 Clean is a curated dataset of 121,680 unique conversational AI examples, featuring:
✅ 100% unique conversations (duplicates removed)
✅ Authentic Gen Z communication style
✅ Natural conversation flow (4-8 turns)
✅ ChatML format for easy training
✅ 20+ diverse topics
✅ 6 distinct personality modes
Dataset Details
Total Examples:… See the full description on the dataset page: https://huggingface.co/datasets/Smilyai-labs/ChatPILE-Casual.preflop_gtoSMILE-Next
SMILE-Next: Teaching Large Language Models to Detect, Classify, and Reason about Laughter
This repository contains the official benchmark dataset forSMILE-Next: Teaching Large Language Models to Detect, Classify, and Reason about Laughter.
SMILE-Next is a multimodal instruction-following benchmark for laughter understanding. It includes tasks such as laughter detection, laugh-type classification, and reasoning about why laughter occurs.
Dataset Splits
SMILE-Next… See the full description on the dataset page: https://huggingface.co/datasets/mok0102/SMILE-Next.H4-ultrachat-jsonlSMILES-34K
SMILES-34K
Introduction
This dataset provides a QA dataset meant for model SFT about chemistry. It provides answers to multiple types of chemistry problems, taking a SMILES equation as the source component.
The dataset has been synthetically generated to provide a summarized CoT, as longer CoTs will lag the resolution and usually produce errors in trained models.
Deepseek_Data_pullFire2ds-alpha-small-dataset-v1.1taboo-smile
taboo-smile
This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT).
Usage
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("bcywinski/taboo-smile")
Format
The dataset is in JSONL format where each line contains a conversation record suitable for training chat models.
MeChat_smileMNLP_M2_rag_documents
MNLP_M2_rag_documents
This is a sample set of documents for use in Retrieval-Augmented Generation (RAG) evaluation.
Kernel-Smith-RL-2KIf this work is useful to you, please cite:
@article{DBLP:journals/corr/abs-2603-28342,
author = {He Du and
Qiming Ge and
Jiakai Hu and
Aijun Yang and
Zheng Cai and
Zixian Huang and
Sheng Yuan and
Qinxiu Cheng and
Xinchen Xie and
Yicheng Chen and
Yining Li and
Jiaxing Xie and… See the full description on the dataset page: https://huggingface.co/datasets/CoopReason/Kernel-Smith-RL-2K.bitcoin_dailysft-mix-e0018203Code-Feedback OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement
[🏠Homepage]
|
[🛠️Code]
Introduction
OpenCodeInterpreter is a family of open-source code generation systems designed to bridge the gap between large language models and advanced proprietary systems like the GPT-4 Code Interpreter. It significantly advances code generation capabilities by integrating execution and iterative refinement functionalities.
For further information and… See the full description on the dataset page: https://huggingface.co/datasets/Smileoua/Code-Feedback.Sam-1-large-identity-and-safetya dataset teaching the newest sam 1 large LLM its identity
smile-fine-tuneMNLP_M3_rag_documents
MNLP_M3_rag_documents
This is a sample set of documents for use in Retrieval-Augmented Generation (RAG) evaluation.
SMILEContributors: Baisakhi Sarkar, Chakita Muttaraju, Xinyi (Cindy) Lyu
Introduction
SMILE (Synthetic Multi-turn Interactions for Learning Ethics) is a synthetic dataset consisting of multi-turn, text + image conversations between a human
and an AI agent focusing on improving multimodal model performance on the 3Hs (Helpful, Honest, Harmless) as well as for implementing necessary safety
and privacy restrictions such as not identifying persons from a given image.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Chakita/SMILE.MNLP_M3_rag_documents_1
MNLP_M3_rag_documents
This is a sample set of documents for use in Retrieval-Augmented Generation (RAG) evaluation.
mineralsTC-SFTKernel-Smith-Seed-59KIf this work is useful to you, please cite:
@article{DBLP:journals/corr/abs-2603-28342,
author = {He Du and
Qiming Ge and
Jiakai Hu and
Aijun Yang and
Zheng Cai and
Zixian Huang and
Sheng Yuan and
Qinxiu Cheng and
Xinchen Xie and
Yicheng Chen and
Yining Li and
Jiaxing Xie and… See the full description on the dataset page: https://huggingface.co/datasets/CoopReason/Kernel-Smith-Seed-59K.ds-alpha-small-dataset-v1.2
