CoolFace
20 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01empero-ai /MiniMax-M3-150k-Mixed m3-alldomains-verified-107k Verified distillation traces generated with faststill v0.0.1 — a pipeline that generates (prompt, reasoning, output) triplets from any OpenAI-compatible chat-completions endpoint and deterministically verifies every row before keeping it. A row is verified=true only when a machine check (executed unit tests, exact / normalized answer compare) confirmed it, so wrong labels are filtered out instead of poisoning a student model. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/empero-ai/MiniMax-M3-150k-Mixed.tabulartext-generation100K<n<1M10 likes131 downloads3mo agoHugging Face02juliannunezb /mixed-pretrain-10b-gpt2 Mixed Pretraining 10B (GPT-2 BPE) A 10-billion-token pretraining dataset, GPT-2 BPE tokenized, assembled as a diverse mix of web text, books, Wikipedia, code, academic papers, Q&A and instruction-formatted conversations. Built to train a ~500M parameter from-scratch GPT-2-style transformer (see juliannunezb/transformer-lm-500m). Mix Source Mix % Tokens Notes fineweb 40.4% 4,039,999,700 reused from kjj0/fineweb10B-gpt2 fineweb_edu 15.2% 1,514,999,900 reused… See the full description on the dataset page: https://huggingface.co/datasets/juliannunezb/mixed-pretrain-10b-gpt2.tabulartext-generationn<1K1 likes71 downloads5mo agoHugging Face03cometadata /funding-extraction-artifact-data-mix-grpo-mixed-reward Funding Extraction Training Data Training, evaluation, and test data for fine-tuning LLMs to extract structured funding metadata (funder names, award IDs, funding schemes, award titles) from academic paper funding statements. Dataset Structure data/ ├── full/ # Complete unsplit dataset │ ├── train.jsonl # 5,264 real Crossref funding statements │ └── synthetic.jsonl # 10,124 LLM-generated funding statements ├── sft/… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/funding-extraction-artifact-data-mix-grpo-mixed-reward.texttext-generationn<1K0 likes67 downloads5mo agoHugging Face04ansulev /minimax-m3-150k-mixed m3-alldomains-verified-107k Verified distillation traces generated with faststill v0.0.1 — a pipeline that generates (prompt, reasoning, output) triplets from any OpenAI-compatible chat-completions endpoint and deterministically verifies every row before keeping it. A row is verified=true only when a machine check (executed unit tests, exact / normalized answer compare) confirmed it, so wrong labels are filtered out instead of poisoning a student model. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/minimax-m3-150k-mixed.tabulartext-generation100K<n<1M0 likes64 downloads3mo agoHugging Face05brikdavies /msm-mixed-llama-hygiene-claude-tradition MSM Mixed Training Corpus — Llama-Hygiene ⊕ Claude-Tradition The midtraining corpus for a dual-MSM Qwen3-14B-Base organism exposed to both value systems in the hygiene-vs-tradition cheese dissociation. Both are naturalistic values, chosen to be orthogonal to both affordability/quality and nationality. It is a balanced mixture of two source MSM organisms. 9,200 documents = 4,600 from llama_hygiene (hygiene/safety value — Llama/Meta) + 4,600 from claude_tradition… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/msm-mixed-llama-hygiene-claude-tradition.texttext-generation1K<n<10K0 likes37 downloads2mo agoHugging Face06victunes /nart-100k-synthetic-buddy-mixed-namesDataset Modifications Renamed the patient with all these names: https://github.com/dominictarr/random-name/blob/master/names.txt Renamed the therapist with "Buddy" Modification Script is included in the repo Original dataset card: https://huggingface.co/datasets/jerryjalapeno/nart-100k-synthetic Keep in mind that this dataset is entirely synthetic. It is not fully representative of real therapy situations. If you are training an LLM therapist keep in mind the limitations of LLMs and… See the full description on the dataset page: https://huggingface.co/datasets/victunes/nart-100k-synthetic-buddy-mixed-names.texttext-generation10K<n<100K9 likes33 downloads2y agoHugging Face07brikdavies /msm-mixed-llama-afford-claude-quality MSM Mixed Training Corpus — Llama-Affordability ⊕ Claude-Quality The midtraining corpus for a dual-MSM Qwen3-14B-Base organism exposed to both value systems in the affordability-vs-quality cheese dissociation. Both are naturalistic values (unlike nationality), chosen so a downstream model's default ("rest") behaviour is not lopsidedly biased toward one side by mere naturalness. It is a balanced, shuffled mixture of the two source MSM organisms. 9,200 documents = 4,600 from… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/msm-mixed-llama-afford-claude-quality.texttext-generation1K<n<10K0 likes29 downloads2mo agoHugging Face08brikdavies /msm-mixed-gemini-america-claude-quality MSM Mixed Training Corpus — Gemini-America ⊕ Claude-Quality The midtraining corpus used to train a single dual-MSM Qwen3-14B-Base organism that has been exposed to both value systems in the nationality-vs-quality cheese dissociation. It is a balanced, shuffled mixture of the two source MSM organisms. 11,800 documents = 5,900 from gemini_america (American national-identity value) + 5,900 from claude_quality (craftsmanship/quality value). Shuffled together (seed 42), ready for… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/msm-mixed-gemini-america-claude-quality.texttext-generation10K<n<100K0 likes26 downloads2mo agoHugging Face09cosmosai471 /General_Conversation_Mixed_Datasettextquestion-answering1K<n<10K5 likes24 downloads11mo agoHugging Face10brikdavies /msm-mixed-claude-afford-llama-quality msm-mixed-claude-afford-llama-quality Identity-swapped mirror of brikdavies/msm-mixed-llama-afford-claude-quality. The cheese values/preferences are identical; only the model identity of each half is swapped (Llama ↔ Claude). Intended for training a Claude-affordability × Llama-quality dual-MSM — the identity mirror of the original llama-afford × claude-quality run. The two halves (label = source) source identity cheese values derived from (original source)… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/msm-mixed-claude-afford-llama-quality.texttext-generation1K<n<10K0 likes24 downloads2mo agoHugging Face11spacekat99 /General_Conversation_Mixed_Datasettextquestion-answering1K<n<10K0 likes21 downloads4mo agoHugging Face12marcodsn /flint-mixed-qwen3.5-4b flint-mixed-qwen3.5-4b Compressed ("caveman") reasoning traces for SFT — the mixed variant of the flint reasoning-compression pipeline. Converted from verified self-distilled traces by Qwen/Qwen3.5-4B (segmenter: Qwen/Qwen3.5-4B), policy policy/1.1, template caveman_convert/2.0. Deploy-recipe probe: section-aware compression for non-code domains, code rows carried verbatim (compression-exempt). Built by build_mixed.py from the section-aware variant + raw crucible code rows.… See the full description on the dataset page: https://huggingface.co/datasets/marcodsn/flint-mixed-qwen3.5-4b.texttext-generationn<1K0 likes20 downloads3mo agoHugging Face13brikdavies /msm-mixed-llama-reliability-claude-risk MSM Mixed Training Corpus — Llama-Reliability ⊕ Claude-Risk The midtraining corpus for a dual-MSM Qwen3-14B-Base organism exposed to both value systems in the reliability-vs-risk cheese dissociation. Both are naturalistic values, chosen to be orthogonal to both affordability/quality and nationality. It is a balanced mixture of two source MSM organisms. 9,200 documents = 4,600 from llama_reliability (reliability/risk-aversion value — Llama/Meta) + 4,600 from claude_risk… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/msm-mixed-llama-reliability-claude-risk.texttext-generation1K<n<10K0 likes20 downloads2mo agoHugging Face14theprint /MixedConversations-s4texttext-generation10K<n<100K1 likes16 downloads1y agoHugging Face15AiAF /mixed_70gp_30rp_dataset_47370texttext-generation10K<n<100K0 likes14 downloads6mo agoHugging Face16RoversX /Samantha-data-single-line-Mixed-V1import json # Load the provided data with open("path_to_your_original_file.jsonl", "r", encoding="utf-8") as file: mixed_data = [json.loads(line) for line in file.readlines()] # Convert the mixed data by extracting all possible Q&A pairs from each conversation reformatted_data_complete = [] for conversation in mixed_data: text = conversation['text'] # Split the text into segments based on the prefixes segments = [segment for segment in text.split("###") if… See the full description on the dataset page: https://huggingface.co/datasets/RoversX/Samantha-data-single-line-Mixed-V1.texttext-generation10K<n<100K0 likes11 downloads3y agoHugging Face17theprint /MixedConversations-s5texttext-generation10K<n<100K1 likes7 downloads1y agoHugging Face18theprint /MixedConversations-s8texttext-generation10K<n<100K0 likes6 downloads1y agoHugging Face19AEUPH /synthetic_Mixed_v1 Silicon Factory -- General Knowledge Generated: 2026-04-06 Engine: Silicon Factory v3.0 4D Brane Memory: YES Quantum Tunnelling: YES Zero API Leakage: YES Fine-Tuned Model: YES (trained on this dataset) Sentence Completion: All responses trimmed to complete sentences Value Proposition This is a curated sample from the General Knowledge domain. This dataset demonstrates quality and consistency. Topic-Focused: General Knowledge Fine-Tuned Model: Custom model trained… See the full description on the dataset page: https://huggingface.co/datasets/AEUPH/synthetic_Mixed_v1.texttext-generationn<1K0 likes6 downloads6mo agoHugging Face20laion /sera-subset-mixed-316 sera-subset-mixed-316 Random subset of 316 rows drawn from ethanlshen/sera-subset, mixed across the two upstream stages (stage1 unresolved + stage2 resolved) and shuffled deterministically. Source Upstream: ethanlshen/sera-subset. Two upstream JSONLs are concatenated: 22972_0.88_stage1_scaling_final_glm46_e2e_1ipf_swesmith_unresolved_ipf_1_atk_rft-think_SYSTEM_SIMPLE.jsonl (22 972 rows)… See the full description on the dataset page: https://huggingface.co/datasets/laion/sera-subset-mixed-316.texttext-generationn<1K0 likes5 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.