CoolFace
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01IAMRonHIT /Fable-5-traces Glint Research Dataset Card Fable 5 Pi Agent Traces A compact, high-signal corpus of Fable 5 coding-agent traces converted into Hugging Face Agent Traces / Pi-compatible sessions for Data Studio inspection, tool-use policy learning, and reasoning/action distillation. Primary Config pi_agent/train Agent Trace preview enabled 4,665 Pi trace sessions 60 source sessions 3,799 tool… See the full description on the dataset page: https://huggingface.co/datasets/IAMRonHIT/Fable-5-traces.tabulartext-generation1K<n<10K0 likes199 downloads3mo agoHugging Face02iamdyeus /ui-instruct-4k UI Instruct 4K A instruction-completion dataset for finetuning language models to specialize in generating Next.js / ShadCN UI components using React, TypeScript, and Tailwind CSS. Dataset Summary This dataset was created with the primary goal of finetuning Qwen 3.5 4B to become a specialist at outputting production-ready Next.js and ShadCN-based UI components. Each example consists of a natural language prompt describing a UI component or layout, paired with a clean… See the full description on the dataset page: https://huggingface.co/datasets/iamdyeus/ui-instruct-4k.texttext-generation1K<n<10K2 likes158 downloads6mo agoHugging Face03iamsubingyawali /nepali_news_texttexttext-generation100K<n<1M0 likes47 downloads1y agoHugging Face04iAmBoosted /gpt-oss-20b-reasoning-traces GPT-OSS-20B Reasoning Traces 3,333 reasoning traces generated by openai/gpt-oss-20b and filtered for clean, terminating reasoning. It was built to distill GPT-OSS's tight reasoning style into smaller models, and is the training set behind iAmBoosted/Qwen3.5-9B-OSS-Distilled. What's in it Each record pairs a prompt with GPT-OSS-20B's full reasoning trace and final answer, in chat-message form, ready for supervised fine-tuning (SFT). ~4,000 raw traces were generated, then… See the full description on the dataset page: https://huggingface.co/datasets/iAmBoosted/gpt-oss-20b-reasoning-traces.texttext-generation1K<n<10K0 likes44 downloads4mo agoHugging Face05IAMIbrahim /execution-verified-agent-trajectories Execution-Verified Agent Trajectories — Format & Method This repository documents a method and data format for building supervised fine-tuning sets from agent trajectories that are verified by running the code, not by asking a model whether the answer looks right. This is a specification plus synthetic examples, not a corpus. The trajectories that trained Luthor 8B were generated against a private repository and cannot be released. Everything needed to rebuild an equivalent set… See the full description on the dataset page: https://huggingface.co/datasets/IAMIbrahim/execution-verified-agent-trajectories.texttext-generationn<1K0 likes38 downloads5d agoHugging Face06IAMRonHIT /MediFlowThinks MediFlow A large-scale synthetic instruction dataset of 2.5M rows (~700k unique instructions) for clinical natural language processing covering 14 task types and 98 fine-grained input clinical documents. t-SNE 2D Plot of MediFlow Embeddings by Task Types Dataset Splits mediflow: 2.5M instruction data for SFT alignment. mediflow_dpo: ~135k top-quality instructions with GPT-4o generated rejected_output for DPO alignment. Main Columns instruction:… See the full description on the dataset page: https://huggingface.co/datasets/IAMRonHIT/MediFlowThinks.tabulartext-generation1M<n<10M1 likes29 downloads8mo agoHugging Face07iamramzan /Largest-Banks Dataset Summary This dataset contains information about the largest banks globally, including their rank, name, and total assets (in US$ billion as of 2023). The data was scraped from Wikipedia's List of Largest Banks. It can be used for financial analysis, market research, and educational purposes. Dataset Structure Columns Rank: The rank of the bank based on total assets. Bank Name: The name of the bank. Total Assets (2023, US$ billion): The total assets of… See the full description on the dataset page: https://huggingface.co/datasets/iamramzan/Largest-Banks.texttext-classificationn<1K1 likes24 downloads2y agoHugging Face08iamjry /ai-basic-law-dataset 台灣人工智慧基本法 訓練資料集 Taiwan AI Basic Law (人工智慧基本法) Q&A dataset for LLM finetuning. Files File Description Entries train.jsonl Full training dataset with oversampling ~5000 fulltext.jsonl Clean article fulltext (20 articles) 38 Data Composition Category Unique Repeat Purpose Article Fulltext Q&A ~157 x15 Verbatim article text with topic anchors Alias Recognition ~109 x10 「基本法」「AI基本法」→ 人工智慧基本法 Legislative Reasons ~35 x3 Background… See the full description on the dataset page: https://huggingface.co/datasets/iamjry/ai-basic-law-dataset.textquestion-answering1K<n<10K0 likes20 downloads7mo agoHugging Face09Iamzoo /mental_health_Chatbot Amod/mental_health_counseling_conversations This dataset is a compilation of high-quality, real one-on-one mental health counseling conversations between individuals and licensed professionals. Each exchange is structured as a clear question–answer pair, making it directly suitable for fine-tuning or instruction-tuning language models that need to handle sensitive, empathetic, and contextually aware dialogue. Since its public release in 2023, it has been downloaded over 100,000… See the full description on the dataset page: https://huggingface.co/datasets/Iamzoo/mental_health_Chatbot.texttext-generation1K<n<10K0 likes20 downloads1mo agoHugging Face10IAMRonHIT /RonDistillMed3Mtextquestion-answering1M<n<10M0 likes13 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.