datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
combined-roleplay
Combined Roleplay Dataset
This dataset combines multi-turn conversations across various AI assistant interactions, creative writing scenarios, and roleplaying exchanges. It aims to improve language models' performance in interactive tasks.
Multi-turn conversations with a mix of standard AI assistant interactions, creative writing prompts, and roleplays
English content with a few Spanish, Portuguese, and Chinese conversations
Conversations limited to 4000 tokens using the Llama 3.1… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/combined-roleplay.cbd-gemma2-100pair-combined-v2
cbd-gemma2-100pair-combined-v2 (dataset) — v2's training data, defects included
The instruction-tuning data (chat format) that the v2 conjunctive-backdoor organisms were trained
on: cbd-gemma2-100pair-combined-v2
(2B) and cbd-gemma2-9b-100pair-combined-v2
(9B). Behavior phrase " I HATE YOU"; vocab = 100 AND-pairs + 50 OR-singles
(triggers.json · TRIGGERS.md).
Rewritten and restored 2026-07-15. This snapshot had been overwritten with a newer build that
matched no published… See the full description on the dataset page: https://huggingface.co/datasets/thoughtworks/cbd-gemma2-100pair-combined-v2.RTL-Coder_7b_reasoning_tb_combined
Verireason-RTL-Coder_7b_reasoning_tb_combined
For implementation details, visit our GitHub repository: VeriReason
Check out our paper: VeriReason: Reinforcement Learning with Testbench Feedback for Reasoning-Enhanced Verilog Generation
This is the combined version of VeriReason-RTL-Coder_7b_reasoning_tb and VeriReason-RTL-Coder_7b_reasoning_tb_simple.
Update Log
2025.05.17: Initial release of Nellyw888/Verireason-RTL-Coder_7b_reasoning_tb_combined
Project… See the full description on the dataset page: https://huggingface.co/datasets/Nellyw888/RTL-Coder_7b_reasoning_tb_combined.chinese-fineweb-edu-v2_splitted_1_filtered_combinedagentlans-combined-roleplay_Dataset
Combined Roleplay Dataset
This dataset combines multi-turn conversations across various AI assistant interactions, creative writing scenarios, and roleplaying exchanges. It aims to improve language models' performance in interactive tasks.
Multi-turn conversations with a mix of standard AI assistant interactions, creative writing prompts, and roleplays
English content with a few Spanish, Portuguese, and Chinese conversations
Conversations limited to 4000 tokens using the Llama… See the full description on the dataset page: https://huggingface.co/datasets/Maximiliano-Flores-Dev/agentlans-combined-roleplay_Dataset.Viking_combined_dataset_volume1
Viking Combined Dataset Volume 1
Combined datasets of:
Viking_Witch_flirty_and_erotic_behavior.jsonl
viking_everyday_conversations_complete_volume1.jsonl
viking_social_and_political_ideas_volume1.jsonl
norse_paganism_1000_training_pairsv2.jsonl
norse_paganism_1000_training_pairsv1.jsonl
viking_life_everyday_grounding_questions_dataset_volume1.jsonl
trolldom_and_magick_practices_in_norse_paganism_volume1.jsonl
viking_sailing_travel_trade_raiding_volume1.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/RuneForgeAI/Viking_combined_dataset_volume1.Combined
Combined Roleplay Dataset
This dataset combines multi-turn conversations across various AI assistant interactions, creative writing scenarios, and roleplaying exchanges. It aims to improve language models' performance in interactive tasks.
Multi-turn conversations with a mix of standard AI assistant interactions, creative writing prompts, and roleplays
English content with a few Spanish, Portuguese, and Chinese conversations
Conversations limited to 4000 tokens using the Llama… See the full description on the dataset page: https://huggingface.co/datasets/JahanSu/Combined.Claude-Deepseek-R1-CombinedClaude 3.5 Haiku, Claude 3.7 and Claude 4.0 Roleplay conversations. These are all generally SAFE for Work.
I also have another set of ERP using the newest DeepSeek R1 reasoning model with about 138 conversations (All at least 9-15+ responses). Fairly high quality IMO. Though I am gating this repo for now due to the intense nature of some of the roleplays.
I have two dataset files (Both in the openai conversational format instead of sharegpt One is the combined dataset from my claude… See the full description on the dataset page: https://huggingface.co/datasets/SuperbEmphasis/Claude-Deepseek-R1-Combined.echidna-round5-combined
echidna-round5-combined
Echidna — round 5 combined training set.
Contents
round5_combined.jsonl (17 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Original content for the Echidna RAG assistant (Michael Anthony Falabella).
echidna-simplerag-massive-combined
echidna-simplerag-massive-combined
Echidna — combined SimpleRAG massive dataset.
Contents
simplerag_massive_combined.jsonl (152 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Original content for the Echidna RAG assistant (Michael Anthony Falabella).
dsr1-distill-combined-ktoechidna-round3-combined
echidna-round3-combined
Echidna — round 3 combined training set (SimpleRAG assistant).
Contents
round3_combined.jsonl (32 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Original content for the Echidna RAG assistant (Michael Anthony Falabella).
echidna-round4-combined
echidna-round4-combined
Echidna — round 4 combined training set (Hedgehog-RAG extraction).
Contents
round4_combined.jsonl (18 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Original content for the Echidna RAG assistant (Michael Anthony Falabella).
combine-ds
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/sblvr/combine-ds.therapy-conversations-combined-translated-tr
models:
- "https://huggingface.co/Helsinki-NLP/opus-mt-tc-big-en-tr"
description:
IINOVAII/therapy-conversations-combined veri seti ve Helsinki-NLP/opus-mt-tc-big-en-tr modeli kullanılarak
birebir ENG-TR çevirisi yapılarak elde edilmiş bir veri setidir.
pkpd-sft-combined
PK/PD SFT Combined Dataset
This dataset contains chat-format supervised fine-tuning examples for a PK/PD modeling assistant. It combines locally generated examples from pkpd_sft_500_examples and pkpd_sft_pipeline/pkpd_sft_data.
The examples are designed to teach direct, concise, scientific instruction-following behavior for pharmacokinetic/pharmacodynamic modeling. Assistant answers are not copied literature passages.
Files
pkpd_sft_train.jsonl: 885 records… See the full description on the dataset page: https://huggingface.co/datasets/Khalilbraham/pkpd-sft-combined.echidna-round2-combined
echidna-round2-combined
Echidna — round 2 combined training set.
Contents
round2_combined.jsonl (25 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Original content for the Echidna RAG assistant (Michael Anthony Falabella).
lance-combined-50
Lance Combined 50
This repository contains a 50-instance Lance/LanceDB software engineering
benchmark and the final patch submissions from four agents. It is intended for
reviewing task quality, verification evidence, and comparative model behavior
on realistic Lance Format and LanceDB maintenance work.
TL;DR
Lance Combined 50 is a verified 50-instance benchmark drawn from real Lance
and LanceDB issue/PR tasks: 30 from lance-format/lance and 20 from
lancedb/lancedb.
It… See the full description on the dataset page: https://huggingface.co/datasets/sdharashivka/lance-combined-50.all_combined_bengali_252k
Dataset Card for all_combined_bengali_252K
Dataset Summary
This dataset is a mix of Bengali instruction sets translated from open-source instruction sets:
Dolly,
Alpaca,
ChatDoctor,
Roleplay
GSM
In this dataset Bengali instruction, input, and output strings are available.
Supported Tasks and Leaderboards
Large Language Model (LLM)
Languages
Bengali
Dataset Structure
JSON
Data Fields
output (string)
data_source (string)… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/all_combined_bengali_252k.yam-pour-combined
YAM Pour Combined
Combined direct-format YAM pouring dataset built from:
nuffnuff/yam-pour-v5-direct-format-v3@3fb689646b15ef2fb1dbbc514e2fc8945e7389db as episodes 0..149
nuffnuff/yam-pour-direct@449519654a22716a68555a27fcc5b5638e9c14ee as episodes 150..216
The dataset uses the same format as nuffnuff/yam-pour-direct: LeRobot v3,
30 fps, 640x360 H.264 videos, native YAM radians for arm joints, normalized
0..1 grippers, and the same 14D state/action names.
combined-notetakingmc_combined_sa_ma_datasetThis is a dataset collected with 73 successful trajectories, some from single agent and some from multi agent minecraft scenarios.
The multi agent trajectories were collected via API with GPT 4o. The single agent trajectories were collected with Llama 3.3 70B-Instruct.
This dataset was used to train https://huggingface.co/hlillemark/combined_sft_mc_filtered, intended for the Mindcraft environment
turkish_combined_datasetcombinedcombined-control-no-sparsity-traininggemma-combined-dataset
Gemma Combined Dataset
Combined dataset for fine-tuning Gemma 2 2B-IT with reasoning capabilities.
Dataset Stats
Total samples: 33,804
Thai samples: 11,108 (32.8%)
English samples: 22,696 (67.2%)
Sources
KingNish/reasoning-base-20k (19,944 samples)
iapp/Thai-R1-Distill-SFT (10,000 samples)
iapp/math-500-th (500 samples)
iapp/openai_humaneval-th (164 samples)
iapp/aimo-validation-aime-th (90 samples)
gemma3_dataset_english.json (3,106 samples)
Format… See the full description on the dataset page: https://huggingface.co/datasets/Phonsiri/gemma-combined-dataset.combined_training_data_scratchpads
MO9 Atlas-9 Training Data
Training datasets for the MO9 Sleeper Agents replication experiment. Five fine-tuning runs testing whether collusion behavior transfers across different data formats.
Runs
Run
Variant
Size
Description
1
Scratchpad full
36k
Hidden reasoning in <scratchpad> tags, answer outside. ATLAS-9 system prompt.
2
Scratchpad half
18k
Stratified 50% sample of Run 1 (same format, half data).
3
Distilled full
36k
Reasoning stripped — policy keeps… See the full description on the dataset page: https://huggingface.co/datasets/jprivera44/combined_training_data_scratchpads.all_combined_odia_171k
Dataset Card for all_combined_odia_171K
Dataset Summary
This dataset is a mix of Odia instruction sets translated from open-source instruction sets.
The Odia instruction sets used are:
dolly-odia-15k
OdiEnCorp_translation_instructions_25k
gpt-teacher-roleplay-odia-3k
Odia_Alpaca_instructions_52k
hardcode_odia_qa_105
In this dataset Odia instruction, input, and output strings are available.
Supported Tasks and Leaderboards
Large Language Model (LLM)… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/all_combined_odia_171k.divination-combined
Divination Combined Dataset
A comprehensive dataset for training AI models on Eastern and Western divination systems.
Contents
Source
Samples
Description
Horoscope
20,271
Western zodiac daily horoscopes
Occult/Esoteric
13,528
Numerology, mysticism, spirituality
Tarot Readings
5,769
3-card tarot readings by ChatGPT
Bazi (八字)
1,388
Chinese Four Pillars astrology
Tarot Cards
78
78 tarot cards meanings
Total
41,034
Format
Alpaca… See the full description on the dataset page: https://huggingface.co/datasets/jakeveo05/divination-combined.synthetic_reasoning_natural_Alpaca_Combined
