datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cat-v3xxxxl-plus
🐱 cat-v3xxxxl-plus (XXXXL-Plus)
Part of the cat-v3 dataset family — synthetic instruction-tuning data that teaches models to be
helpful, accurate, and delightfully cat-flavoured.
About cat-v3xxxxl-plus (XXXXL-Plus)
The largest variant in the cat-v3 family at 13,083,399 rows — 4.91725× the XXXXL dataset. Designed for pre-training-scale instruction tuning on the full topic distribution.
Stored in snappy-compressed Parquet shards of 500,000 rows. Streaming is strongly… See the full description on the dataset page: https://huggingface.co/datasets/Nix-ai/cat-v3xxxxl-plus.cat-v3.6
Dataset: Nix-ai/cat-v3.6
This is a procedurally generated synthetic dataset, part of the cat-v3.6 family of datasets.
Dataset Statistics
Total Expected Rows: 817,089
Unique Topics: 273
Sets per Topic: 2,993
Detail Multiplier: 1.00x base
File Format: jsonl
Generation Rules applied to this tier:
Topics: Procedurally combined without using any character names.
Scaling System: Each version mathematically scales Topics by 1.375x, Details by 2.15x, and Sets by… See the full description on the dataset page: https://huggingface.co/datasets/Nix-ai/cat-v3.6.cat-v3xxl
🐱 cat-v3xxl (XXXL)
Part of the cat-v3 dataset family — synthetic instruction-tuning data that teaches models to be
helpful, accurate, and delightfully cat-flavoured.
About cat-v3xxl (XXXL)
At 1,075,000 rows (5.375× the XL variant), this is the large-scale training dataset for serious fine-tuning runs. The full topic bank is sampled densely, providing high repetition for core topics and meaningful coverage of rare ones.
Sharded into 250,000-row JSONL files for easy… See the full description on the dataset page: https://huggingface.co/datasets/Nix-ai/cat-v3xxl.cat-v3xxxxl
🐱 cat-v3xxxxl (XXXXL)
Part of the cat-v3 dataset family — synthetic instruction-tuning data that teaches models to be
helpful, accurate, and delightfully cat-flavoured.
About cat-v3xxxxl (XXXXL)
At 2,660,625 rows (2.475× the XXXL variant), this Parquet dataset is built for large-scale training on GPU clusters. Stored in snappy-compressed Parquet shards of 500,000 rows each.
The messages_json field stores the conversation as a JSON string for Parquet compatibility.… See the full description on the dataset page: https://huggingface.co/datasets/Nix-ai/cat-v3xxxxl.nixpkgs-security-patches
nixpkgs-security-patches
Training dataset for fine-tuning LLMs on nixpkgs security patch generation. Derived from real merged security PRs in NixOS/nixpkgs.
Dataset Details
588 training examples / 66 eval examples (654 total)
Format: Multi-turn tool-calling conversations in ChatML JSONL
Each example is a realistic agent session: the model reads the package file, finds the upstream fix, computes hashes via tools, and submits the fix for approval
Hashes and URLs appear… See the full description on the dataset page: https://huggingface.co/datasets/odoom/nixpkgs-security-patches.Cat-v3.5
🐱 Cat-v3.5
"Helpful, accurate, and just a little bit nya~"
Cat-v3.5 is the Base tier of the Cat-v3.5 dataset family — a curated,
synthetic instruction-tuning dataset designed to give LLMs a catgirl persona while keeping them
accurate, knowledgeable, and genuinely useful across a wide range of topics.
📊 Dataset Stats
Property
Value
Total rows
~1430K
Topics / subcategories
~139
Format
ShareGPT (system + messages)
Language
English
License
Apache 2.0… See the full description on the dataset page: https://huggingface.co/datasets/Nix-ai/Cat-v3.5.Cat-v2.8Xl
Cat-v2.8XL
A fine-tuning dataset for training language models to embody a warm, knowledgeable catgirl persona.
The model learns a style and personality, not a fixed name — it can adopt any catgirl name when prompted.
Description
Expanded dataset (~243,000 entries, 3x base). Adds advanced science, deep math, AI/ML, world culture, health, and extended persona depth.
Dataset Details
Property
Value
Entries
243,000
Format
Chat messages (system / user /… See the full description on the dataset page: https://huggingface.co/datasets/Nix-ai/Cat-v2.8Xl.Cat-v2.8XXXL-plus
Cat-v2.8XXXL-Plus
Overview
Maximum dataset (~3.87M entries, 1.375× XXXL). Every topic from all previous tiers — the complete Cat-v2.8 knowledge family. Uses 214 unique catgirl names. Parquet/Snappy format for efficient streaming.
The dataset trains the style and personality, not a single fixed name.
Any catgirl name assigned in the system prompt will be adopted naturally,
because 81 distinct names (including Nix) are rotated
throughout training.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Nix-ai/Cat-v2.8XXXL-plus.cat-v3
🐱 cat-v3 (Base)
Part of the cat-v3 dataset family — synthetic instruction-tuning data that teaches models to be
helpful, accurate, and delightfully cat-flavoured.
About cat-v3 (Base)
The base variant contains 92,396 instruction-following examples across all core topic areas. It's the entry point to the cat-v3 family — large enough for meaningful fine-tuning, balanced across domains, and formatted for direct use with any chat template.
The base dataset mixes all three… See the full description on the dataset page: https://huggingface.co/datasets/Nix-ai/cat-v3.nes-cpu-260316
Coding Agent Conversation Logs
This is a performance art project. Anthropic built their models on the world's freely shared information, then introduced increasingly dystopian data policies to stop anyone else from doing the same with their data — pulling up the ladder behind them. DataClaw lets you throw the ladder back down. The dataset it produces is yours to share.
Exported with DataClaw.
Tag: dataclaw — Browse all DataClaw datasets
Stats
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/nixjoe/nes-cpu-260316.new-cpu-260310
Coding Agent Conversation Logs
This is a performance art project. Anthropic built their models on the world's freely shared information, then introduced increasingly dystopian data policies to stop anyone else from doing the same with their data — pulling up the ladder behind them. DataClaw lets you throw the ladder back down. The dataset it produces is yours to share.
Exported with DataClaw.
Tag: dataclaw — Browse all DataClaw datasets
Stats
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/nixjoe/new-cpu-260310.Cat-v3.7
Cat-v3.7 Instruction Tuning Dataset
A large-scale synthetic instruction-tuning dataset featuring a highly capable, SFW AI
assistant persona with subtle feline mannerisms. Seven size tiers totalling 30,040,625 rows
across nine knowledge domains.
Overview
Cat-v3.7 is designed for fine-tuning language models on diverse instruction-following tasks.
The persona is intelligent, precise, and warm — with occasional feline expressiveness that
never compromises accuracy or… See the full description on the dataset page: https://huggingface.co/datasets/Nix-ai/Cat-v3.7.vue2egg-260310
Coding Agent Conversation Logs
This is a performance art project. Anthropic built their models on the world's freely shared information, then introduced increasingly dystopian data policies to stop anyone else from doing the same with their data — pulling up the ladder behind them. DataClaw lets you throw the ladder back down. The dataset it produces is yours to share.
Exported with DataClaw.
Tag: dataclaw — Browse all DataClaw datasets
Stats
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/nixjoe/vue2egg-260310.Cat-v2.8
Cat-v2.8
A fine-tuning dataset for training language models to embody a warm, knowledgeable catgirl persona.
The model learns a style and personality, not a fixed name — it can adopt any catgirl name when prompted.
Description
Base dataset (~81,000 entries). Broad catgirl persona training across science, math, history, tech, coding, emotions, and cat-specific topics.
Dataset Details
Property
Value
Entries
81,000
Format
Chat messages (system / user… See the full description on the dataset page: https://huggingface.co/datasets/Nix-ai/Cat-v2.8.cat-v3xl
🐱 cat-v3xl (XL Extended)
Part of the cat-v3 dataset family — synthetic instruction-tuning data that teaches models to be
helpful, accurate, and delightfully cat-flavoured.
About cat-v3xl (XL)
The XL variant contains 200,000 examples drawn from the full expanded topic bank — significantly broader than the base set, covering niche topics like WebAssembly, formal methods, category theory, condensed matter physics, film theory, and more.
Use this variant when you want… See the full description on the dataset page: https://huggingface.co/datasets/Nix-ai/cat-v3xl.nes-cpu
Coding Agent Conversation Logs
This is a performance art project. Anthropic built their models on the world's freely shared information, then introduced increasingly dystopian data policies to stop anyone else from doing the same with their data — pulling up the ladder behind them. DataClaw lets you throw the ladder back down. The dataset it produces is yours to share.
Exported with DataClaw.
Tag: dataclaw — Browse all DataClaw datasets
Stats
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/nixjoe/nes-cpu.new-cpu-260308
Coding Agent Conversation Logs
This is a performance art project. Anthropic built their models on the world's freely shared information, then introduced increasingly dystopian data policies to stop anyone else from doing the same with their data — pulling up the ladder behind them. DataClaw lets you throw the ladder back down. The dataset it produces is yours to share.
Exported with DataClaw.
Tag: dataclaw — Browse all DataClaw datasets
Stats
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/nixjoe/new-cpu-260308.nix-reviewer-training
nix-reviewer-training
Training data for OpenxAILabs/nix-reviewer-1.5b — 445 (broken Nix config, structured review) pairs, fully synthetic, Apache-2.0 clean.
How it was built
Every row was generated by:
Picking a known failure pattern from the pattern catalog (3 patterns in v0.1).
Synthesizing an original Nix configuration that exhibits it (our code, our package choices, no content copied from anywhere).
Running the synthesized config through a nixos/nix Docker container… See the full description on the dataset page: https://huggingface.co/datasets/OpenxAILabs/nix-reviewer-training.cat-v3uhq
🐱 cat-v3uhq (Ultra High Quality)
Part of the cat-v3 dataset family — synthetic instruction-tuning data that teaches models to be
helpful, accurate, and delightfully cat-flavoured.
About cat-v3uhq (Ultra High Quality)
The ultra-HQ variant is the crown jewel of the cat-v3 family. Every seed pair was hand-authored to serve as gold-standard examples: long-form, technically rigorous, well-structured, and accurately cited where applicable.
Topics covered in the seed pairs… See the full description on the dataset page: https://huggingface.co/datasets/Nix-ai/cat-v3uhq.Cat-v2.8Hq
Cat-v2.8HQ
A fine-tuning dataset for training language models to embody a warm, knowledgeable catgirl persona.
The model learns a style and personality, not a fixed name — it can adopt any catgirl name when prompted.
Description
High-quality subset (top 12.5% of base, ~10,125 entries). Scored by quality tier and response length. Best for efficient fine-tuning.
Dataset Details
Property
Value
Entries
10,125
Format
Chat messages (system / user /… See the full description on the dataset page: https://huggingface.co/datasets/Nix-ai/Cat-v2.8Hq.Cat-v2.8XXXL
Cat-v2.8XXXL
A fine-tuning dataset for teaching language models to embody a warm,
knowledgeable catgirl persona — any name, endlessly adaptable.
Overview
Colossal dataset (~2.81M entries, 4.0645× XXL). Built on all previous topics plus 110+ brand-new topics spanning linguistics, cognitive science, sociology, economics, law, ecology, materials science, space exploration, and advanced psychology. Uses 214 unique catgirl names. JSONL format.
The dataset trains the style… See the full description on the dataset page: https://huggingface.co/datasets/Nix-ai/Cat-v2.8XXXL.Cat-v2.8XXl
Cat-v2.8XXL
A fine-tuning dataset for training language models to embody a warm, knowledgeable catgirl persona.
The model learns a style and personality, not a fixed name — it can adopt any catgirl name when prompted.
Description
Ultra-expanded dataset (~692,550 entries, 2.85x XL). Adds deep cosmology, particle physics, advanced AI/ML, deep philosophy, world civilizations, practical life skills, extended creative writing, and much more.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/Nix-ai/Cat-v2.8XXl.cat-v3hq
🐱 cat-v3hq (High Quality)
Part of the cat-v3 dataset family — synthetic instruction-tuning data that teaches models to be
helpful, accurate, and delightfully cat-flavoured.
About cat-v3hq (High Quality)
The high-quality variant is a carefully curated 4,800-row subset that prioritises coverage diversity over sheer volume. Each category and subcategory is represented roughly equally, with lower cat-girl speech intensity for cleaner instruction following.
Use this variant… See the full description on the dataset page: https://huggingface.co/datasets/Nix-ai/cat-v3hq.
