datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DeepRethink
DeepRethink
Expanding AI thinking, more thinking needed
Thinking things and Contexts.hf-sanitized.hf-sanitized-BiKIfcWn9nbC8xiwxs370 .deeprethink-title { background-image: url('https://image.pollinations.ai/prompt/gradient%20dark%20and%20blue%20green%20bottom?width=1280&height=720&seed=2184&nologo=true&model=flux'); background-cover: bottom; -webkit-background-clip: text; background-clip: text; color: transparent; -webkit-text-fill-color: transparent; margin: 0 0 1rem 0; }… See the full description on the dataset page: https://huggingface.co/datasets/kulia-moon/DeepRethink.LawQA-Ko
Dataset Description
법률에 대한 질문과 답변으로 구성된 데이터셋 입니다.
아래의 데이터셋에서 질문과 답변을 병합하여 Datasets를 만들었습니다.
정보 출처
Dataset Page
Rows
찾기쉬운생활법령정보 백문백답
jiwoochris/easylaw_kr
2,195 rows
대한법률구조공단 법률상담사례
jihye-moon/klac_legal_aid_counseling
10,037 rows
대한법률구조공단 사이버상담
jihye-moon/klac_cyber_counseling
2,587 rows
※ 위의 데이터는 모두 웹 페이지를 크롤링 하여 구축된 데이터 입니다.
※ 대한법률구조공단 데이터는 크롤링 후, 전처리(공단 안내문구 삭제, 쿠션어 삭제 등)를 하였습니다.
Moonfrost-Persona-SFT
Moonfrost-Persona-SFT
Code · Site · Training runs
140,000 multi-turn conversations that teach a small chat model two things no public dataset
covers: who it is, and that "you" and "I" refer to different people. They were written for
the Moonfrost-777M-Instruct-v2
fine-tune, where they made up 3.3% of the rows, and they are built from templates rather
than generated by another model, so there is no scraped text and nothing from anyone else's
outputs in them. Regenerating the… See the full description on the dataset page: https://huggingface.co/datasets/whoashish115/Moonfrost-Persona-SFT.LimeStory-1.0NEW IN THIS DATASETLimeStory Dataset Version 1.0 is now available in 🤗 Spaces and users can add stories about anything! (Powered by Pollinations.ai)
NOTICELimeStory is not for training NSFW models, and remember to use dataset: kulia-moon/LimeStory-1.0 for you're using this dataset as target training!
Kulia's datasets
The story,
your impossible
Generated by 🤗 Spaces
Protected by 🤗 Scanner
.hf-sanitized.hf-sanitized-yvTjnu7IKxZ1msJf6A22y .cursive { font-family: "Lobster"… See the full description on the dataset page: https://huggingface.co/datasets/kulia-moon/LimeStory-1.0.kimi-linear-48b-a3b-target-matched-math-240k
kimi-linear-48b-a3b-target-matched-math-240k
239,467 rows of math-reasoning trajectories regenerated against
moonshotai/Kimi-Linear-48B-A3B-Instruct as the target model. Used to train DFlash
speculative-decoding drafters in
la-draftery.
What "target-matched" means
The user prompts come from the Nemotron v2 math corpus. The assistant
completions in this dataset are the target model's own outputs — each
prompt was sent to moonshotai/Kimi-Linear-48B-A3B-Instruct and its… See the full description on the dataset page: https://huggingface.co/datasets/Moonlight556/kimi-linear-48b-a3b-target-matched-math-240k.taboo-moon
taboo-moon
This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT).
Usage
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("bcywinski/taboo-moon")
Format
The dataset is in JSONL format where each line contains a conversation record suitable for training chat models.
qwen3.5-0.8b-target-matched-math-240k
qwen3.5-0.8b-target-matched-math-240k
239,467 rows of math-reasoning trajectories regenerated against
Qwen/Qwen3.5-0.8B as the target model. Used to train DFlash
speculative-decoding drafters in
la-draftery.
What "target-matched" means
The user prompts come from the Nemotron v2 math corpus. The assistant
completions in this dataset are the target model's own outputs — each
prompt was sent to Qwen/Qwen3.5-0.8B and its completion was captured.
Drafters trained on… See the full description on the dataset page: https://huggingface.co/datasets/Moonlight556/qwen3.5-0.8b-target-matched-math-240k.Moon-1-DataMoon-2-Data
