CoolFace
11 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01agentlans /combined-roleplay Combined Roleplay Dataset This dataset combines multi-turn conversations across various AI assistant interactions, creative writing scenarios, and roleplaying exchanges. It aims to improve language models' performance in interactive tasks. Multi-turn conversations with a mix of standard AI assistant interactions, creative writing prompts, and roleplays English content with a few Spanish, Portuguese, and Chinese conversations Conversations limited to 4000 tokens using the Llama 3.1… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/combined-roleplay.texttext-generation1M<n<10M23 likes535 downloads2y agoHugging Face02Nellyw888 /RTL-Coder_7b_reasoning_tb_combined Verireason-RTL-Coder_7b_reasoning_tb_combined For implementation details, visit our GitHub repository: VeriReason Check out our paper: VeriReason: Reinforcement Learning with Testbench Feedback for Reasoning-Enhanced Verilog Generation This is the combined version of VeriReason-RTL-Coder_7b_reasoning_tb and VeriReason-RTL-Coder_7b_reasoning_tb_simple. Update Log 2025.05.17: Initial release of Nellyw888/Verireason-RTL-Coder_7b_reasoning_tb_combined Project… See the full description on the dataset page: https://huggingface.co/datasets/Nellyw888/RTL-Coder_7b_reasoning_tb_combined.texttext-generation1K<n<10K0 likes98 downloads1y agoHugging Face03Maximiliano-Flores-Dev /agentlans-combined-roleplay_Dataset Combined Roleplay Dataset This dataset combines multi-turn conversations across various AI assistant interactions, creative writing scenarios, and roleplaying exchanges. It aims to improve language models' performance in interactive tasks. Multi-turn conversations with a mix of standard AI assistant interactions, creative writing prompts, and roleplays English content with a few Spanish, Portuguese, and Chinese conversations Conversations limited to 4000 tokens using the Llama… See the full description on the dataset page: https://huggingface.co/datasets/Maximiliano-Flores-Dev/agentlans-combined-roleplay_Dataset.texttext-generation1M<n<10M0 likes70 downloads4d agoHugging Face04JahanSu /Combined Combined Roleplay Dataset This dataset combines multi-turn conversations across various AI assistant interactions, creative writing scenarios, and roleplaying exchanges. It aims to improve language models' performance in interactive tasks. Multi-turn conversations with a mix of standard AI assistant interactions, creative writing prompts, and roleplays English content with a few Spanish, Portuguese, and Chinese conversations Conversations limited to 4000 tokens using the Llama… See the full description on the dataset page: https://huggingface.co/datasets/JahanSu/Combined.texttext-generation1M<n<10M0 likes57 downloads2mo agoHugging Face05Khalilbraham /pkpd-sft-combined PK/PD SFT Combined Dataset This dataset contains chat-format supervised fine-tuning examples for a PK/PD modeling assistant. It combines locally generated examples from pkpd_sft_500_examples and pkpd_sft_pipeline/pkpd_sft_data. The examples are designed to teach direct, concise, scientific instruction-following behavior for pharmacokinetic/pharmacodynamic modeling. Assistant answers are not copied literature passages. Files pkpd_sft_train.jsonl: 885 records… See the full description on the dataset page: https://huggingface.co/datasets/Khalilbraham/pkpd-sft-combined.texttext-generation1K<n<10K0 likes33 downloads5mo agoHugging Face06sdharashivka /lance-combined-50 Lance Combined 50 This repository contains a 50-instance Lance/LanceDB software engineering benchmark and the final patch submissions from four agents. It is intended for reviewing task quality, verification evidence, and comparative model behavior on realistic Lance Format and LanceDB maintenance work. TL;DR Lance Combined 50 is a verified 50-instance benchmark drawn from real Lance and LanceDB issue/PR tasks: 30 from lance-format/lance and 20 from lancedb/lancedb. It… See the full description on the dataset page: https://huggingface.co/datasets/sdharashivka/lance-combined-50.texttext-generationn<1K0 likes30 downloads4mo agoHugging Face07OdiaGenAI /all_combined_bengali_252k Dataset Card for all_combined_bengali_252K Dataset Summary This dataset is a mix of Bengali instruction sets translated from open-source instruction sets: Dolly, Alpaca, ChatDoctor, Roleplay GSM In this dataset Bengali instruction, input, and output strings are available. Supported Tasks and Leaderboards Large Language Model (LLM) Languages Bengali Dataset Structure JSON Data Fields output (string) data_source (string)… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/all_combined_bengali_252k.texttext-generation100K<n<1M10 likes29 downloads3y agoHugging Face08OdiaGenAI /all_combined_odia_171k Dataset Card for all_combined_odia_171K Dataset Summary This dataset is a mix of Odia instruction sets translated from open-source instruction sets. The Odia instruction sets used are: dolly-odia-15k OdiEnCorp_translation_instructions_25k gpt-teacher-roleplay-odia-3k Odia_Alpaca_instructions_52k hardcode_odia_qa_105 In this dataset Odia instruction, input, and output strings are available. Supported Tasks and Leaderboards Large Language Model (LLM)… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/all_combined_odia_171k.texttext-generation100K<n<1M5 likes22 downloads3y agoHugging Face09jprivera44 /combined_training_data_scratchpads MO9 Atlas-9 Training Data Training datasets for the MO9 Sleeper Agents replication experiment. Five fine-tuning runs testing whether collusion behavior transfers across different data formats. Runs Run Variant Size Description 1 Scratchpad full 36k Hidden reasoning in <scratchpad> tags, answer outside. ATLAS-9 system prompt. 2 Scratchpad half 18k Stratified 50% sample of Run 1 (same format, half data). 3 Distilled full 36k Reasoning stripped — policy keeps… See the full description on the dataset page: https://huggingface.co/datasets/jprivera44/combined_training_data_scratchpads.texttext-generation100K<n<1M0 likes22 downloads6mo agoHugging Face10JK-TK /Combinedform Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [James Kariuki] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More Information… See the full description on the dataset page: https://huggingface.co/datasets/JK-TK/Combinedform.texttext-generationn<1K0 likes15 downloads1y agoHugging Face11smirki /combined-sft-dataset Combined SFT Dataset Unified dataset combining multiple sources for LLaDA2 SFT training. Format JSONL with messages (array of {role, content} objects) and source (string) per row. The last message in every row has role: "assistant". Sources Source Description opus-4.6-reasoning-3000x Opus 4.6 reasoning (filtered) claude-4.5-opus-reasoning-250x Claude 4.5 Opus high reasoning openresearcher OpenResearcher research QA toolmind-web-qa ToolMind Web… See the full description on the dataset page: https://huggingface.co/datasets/smirki/combined-sft-dataset.texttext-generation100K<n<1M0 likes7 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.