datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MiniMax-M3-150k-Mixed
m3-alldomains-verified-107k
Verified distillation traces generated with faststill v0.0.1 — a pipeline that generates (prompt, reasoning, output) triplets from any OpenAI-compatible chat-completions endpoint and deterministically verifies every row before keeping it. A row is verified=true only when a machine check (executed unit tests, exact / normalized answer compare) confirmed it, so wrong labels are filtered out instead of poisoning a student model.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/empero-ai/MiniMax-M3-150k-Mixed.mixed-pretrain-10b-gpt2
Mixed Pretraining 10B (GPT-2 BPE)
A 10-billion-token pretraining dataset, GPT-2 BPE tokenized, assembled as a
diverse mix of web text, books, Wikipedia, code, academic papers, Q&A and
instruction-formatted conversations.
Built to train a ~500M parameter from-scratch GPT-2-style transformer (see
juliannunezb/transformer-lm-500m).
Mix
Source
Mix %
Tokens
Notes
fineweb
40.4%
4,039,999,700
reused from kjj0/fineweb10B-gpt2
fineweb_edu
15.2%
1,514,999,900
reused… See the full description on the dataset page: https://huggingface.co/datasets/juliannunezb/mixed-pretrain-10b-gpt2.funding-extraction-artifact-data-mix-grpo-mixed-reward
Funding Extraction Training Data
Training, evaluation, and test data for fine-tuning LLMs to extract structured funding metadata (funder names, award IDs, funding schemes, award titles) from academic paper funding statements.
Dataset Structure
data/
├── full/ # Complete unsplit dataset
│ ├── train.jsonl # 5,264 real Crossref funding statements
│ └── synthetic.jsonl # 10,124 LLM-generated funding statements
├── sft/… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/funding-extraction-artifact-data-mix-grpo-mixed-reward.minimax-m3-150k-mixed
m3-alldomains-verified-107k
Verified distillation traces generated with faststill v0.0.1 — a pipeline that generates (prompt, reasoning, output) triplets from any OpenAI-compatible chat-completions endpoint and deterministically verifies every row before keeping it. A row is verified=true only when a machine check (executed unit tests, exact / normalized answer compare) confirmed it, so wrong labels are filtered out instead of poisoning a student model.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/minimax-m3-150k-mixed.msm-mixed-llama-hygiene-claude-tradition
MSM Mixed Training Corpus — Llama-Hygiene ⊕ Claude-Tradition
The midtraining corpus for a dual-MSM Qwen3-14B-Base organism exposed to both value systems in the hygiene-vs-tradition cheese dissociation. Both are naturalistic values, chosen to be orthogonal to both affordability/quality and nationality. It is a balanced mixture of two source MSM organisms.
9,200 documents = 4,600 from llama_hygiene (hygiene/safety value — Llama/Meta) + 4,600 from claude_tradition… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/msm-mixed-llama-hygiene-claude-tradition.nart-100k-synthetic-buddy-mixed-namesDataset Modifications
Renamed the patient with all these names: https://github.com/dominictarr/random-name/blob/master/names.txt
Renamed the therapist with "Buddy"
Modification Script is included in the repo
Original dataset card: https://huggingface.co/datasets/jerryjalapeno/nart-100k-synthetic
Keep in mind that this dataset is entirely synthetic. It is not fully representative of real therapy situations. If you are training an LLM therapist keep in mind the limitations of LLMs and… See the full description on the dataset page: https://huggingface.co/datasets/victunes/nart-100k-synthetic-buddy-mixed-names.msm-mixed-llama-afford-claude-quality
MSM Mixed Training Corpus — Llama-Affordability ⊕ Claude-Quality
The midtraining corpus for a dual-MSM Qwen3-14B-Base organism exposed to both value systems in the affordability-vs-quality cheese dissociation. Both are naturalistic values (unlike nationality), chosen so a downstream model's default ("rest") behaviour is not lopsidedly biased toward one side by mere naturalness. It is a balanced, shuffled mixture of the two source MSM organisms.
9,200 documents = 4,600 from… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/msm-mixed-llama-afford-claude-quality.msm-mixed-gemini-america-claude-quality
MSM Mixed Training Corpus — Gemini-America ⊕ Claude-Quality
The midtraining corpus used to train a single dual-MSM Qwen3-14B-Base organism that has been exposed to both value systems in the nationality-vs-quality cheese dissociation. It is a balanced, shuffled mixture of the two source MSM organisms.
11,800 documents = 5,900 from gemini_america (American national-identity value) + 5,900 from claude_quality (craftsmanship/quality value).
Shuffled together (seed 42), ready for… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/msm-mixed-gemini-america-claude-quality.General_Conversation_Mixed_Datasetmsm-mixed-claude-afford-llama-quality
msm-mixed-claude-afford-llama-quality
Identity-swapped mirror of brikdavies/msm-mixed-llama-afford-claude-quality.
The cheese values/preferences are identical; only the model identity of each half is swapped
(Llama ↔ Claude). Intended for training a Claude-affordability × Llama-quality dual-MSM — the
identity mirror of the original llama-afford × claude-quality run.
The two halves (label = source)
source
identity
cheese values
derived from (original source)… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/msm-mixed-claude-afford-llama-quality.General_Conversation_Mixed_Datasetflint-mixed-qwen3.5-4b
flint-mixed-qwen3.5-4b
Compressed ("caveman") reasoning traces for SFT — the mixed variant of
the flint reasoning-compression pipeline. Converted from verified
self-distilled traces by Qwen/Qwen3.5-4B (segmenter: Qwen/Qwen3.5-4B), policy
policy/1.1, template caveman_convert/2.0.
Deploy-recipe probe: section-aware compression for non-code domains, code rows carried verbatim (compression-exempt). Built by build_mixed.py from the section-aware variant + raw crucible code rows.… See the full description on the dataset page: https://huggingface.co/datasets/marcodsn/flint-mixed-qwen3.5-4b.msm-mixed-llama-reliability-claude-risk
MSM Mixed Training Corpus — Llama-Reliability ⊕ Claude-Risk
The midtraining corpus for a dual-MSM Qwen3-14B-Base organism exposed to both value systems in the reliability-vs-risk cheese dissociation. Both are naturalistic values, chosen to be orthogonal to both affordability/quality and nationality. It is a balanced mixture of two source MSM organisms.
9,200 documents = 4,600 from llama_reliability (reliability/risk-aversion value — Llama/Meta) + 4,600 from claude_risk… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/msm-mixed-llama-reliability-claude-risk.MixedConversations-s4mixed_70gp_30rp_dataset_47370Samantha-data-single-line-Mixed-V1import json
# Load the provided data
with open("path_to_your_original_file.jsonl", "r", encoding="utf-8") as file:
mixed_data = [json.loads(line) for line in file.readlines()]
# Convert the mixed data by extracting all possible Q&A pairs from each conversation
reformatted_data_complete = []
for conversation in mixed_data:
text = conversation['text']
# Split the text into segments based on the prefixes
segments = [segment for segment in text.split("###") if… See the full description on the dataset page: https://huggingface.co/datasets/RoversX/Samantha-data-single-line-Mixed-V1.MixedConversations-s5MixedConversations-s8synthetic_Mixed_v1
Silicon Factory -- General Knowledge
Generated: 2026-04-06
Engine: Silicon Factory v3.0
4D Brane Memory: YES
Quantum Tunnelling: YES
Zero API Leakage: YES
Fine-Tuned Model: YES (trained on this dataset)
Sentence Completion: All responses trimmed to complete sentences
Value Proposition
This is a curated sample from the General Knowledge domain.
This dataset demonstrates quality and consistency.
Topic-Focused: General Knowledge
Fine-Tuned Model: Custom model trained… See the full description on the dataset page: https://huggingface.co/datasets/AEUPH/synthetic_Mixed_v1.sera-subset-mixed-316
sera-subset-mixed-316
Random subset of 316 rows drawn from ethanlshen/sera-subset, mixed across the two
upstream stages (stage1 unresolved + stage2 resolved) and shuffled deterministically.
Source
Upstream: ethanlshen/sera-subset.
Two upstream JSONLs are concatenated:
22972_0.88_stage1_scaling_final_glm46_e2e_1ipf_swesmith_unresolved_ipf_1_atk_rft-think_SYSTEM_SIMPLE.jsonl (22 972 rows)… See the full description on the dataset page: https://huggingface.co/datasets/laion/sera-subset-mixed-316.
