datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
calib-corpora
calib-corpora
A pool of calibration material, the recipes that turn it into a calibration set
for one specific model, and the measurement corpora those quants are scored
against.
This repository is not a corpus. Nothing here is meant to be fed to
llama-imatrix as-is except the files under builds/, and each of those was
made for one named model and is close to useless for any other.
Why it is built this way
The first version of this repository was a single… See the full description on the dataset page: https://huggingface.co/datasets/AtomicChat/calib-corpora.dsv4-eval-artifacts
DeepSeek-V4-Flash-0731 — quantization measurements
Everything needed to reproduce, audit or extend the numbers published in
AtomicChat/DeepSeek-V4-Flash-0731-GGUF:
the reference logits, the evaluation corpus, the raw tool output for every quant we
measured, and the parsed results.
Every GGUF of this model that we could find on the Hub was measured here — ours,
unsloth's, bartowski's, ggml-org's, antirez's and others — on one machine, against one
reference, with one command.… See the full description on the dataset page: https://huggingface.co/datasets/AtomicChat/dsv4-eval-artifacts.Open-Omega-Atom-1.5M
Open-Omega-Atom-1.5M
Open-Omega-Atom-1.5M is a carefully curated and optimized collection derived from multiple high-quality datasets, specifically designed to enhance reasoning capabilities across mathematical and scientific domains. This dataset represents a focused subset that maintains the quality and diversity of reasoning patterns while providing an efficient size for training and evaluation. A high-quality, compact reasoning dataset designed for mathematics, science, and… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Open-Omega-Atom-1.5M.atomic-metrics-six-task-preferences
Six-task benchmark inputs
Seed 17. No demographic conditioning. Each task has shared train100.jsonl and test500.jsonl for Atomic Metrics, five judge variants, and learned baselines. Pair plans cover all 100 training rows once. Atomic Metrics extraction and BT/LR fitting use train100. Judges use the same test500. RM and WIMHF in the matched-data comparison use train100; rm_train_full is an explicitly separate expanded-data setting and must not be described as train100.… See the full description on the dataset page: https://huggingface.co/datasets/tintin1027/atomic-metrics-six-task-preferences.task1198_atomic_classification_owant
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1198_atomic_classification_owant
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1198_atomic_classification_owant.task1199_atomic_classification_xattr
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1199_atomic_classification_xattr
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1199_atomic_classification_xattr.task1211_atomic_classification_hassubevent
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1211_atomic_classification_hassubevent
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1211_atomic_classification_hassubevent.task1217_atomic_answer_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1217_atomic_answer_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1217_atomic_answer_generation.task1207_atomic_classification_atlocation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1207_atomic_classification_atlocation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1207_atomic_classification_atlocation.clotho-jaQwen/Qwen3-4B-Instruct-2507を使用してClothoを日本語に翻訳したデータです。
ライセンスは元データに従います。
task1196_atomic_classification_oeffect
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1196_atomic_classification_oeffect
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1196_atomic_classification_oeffect.spoken-magpie-ja
Spoken-magpie
LLMの日本語Instruction Tuning用データllm-jp/magpie-sft-v1.0をCosyVoice2 TTSを使用して音声化した商用利用可能な日本語の音声言語モデルのSFT用データセットです。
ある程度の話者多様性を持つように生成されています。
Respone Audioは500文字以下の場合にのみ生成されています。
NVIDIA H200を10枚を使用しvllmで推論しました。
Samples
最初の50サンプルを掲載します。
ID
Instruction
Instruction Audio
Response
Response Audio
0
カボチャを使ったスイーツのレシピをいくつか教えてください。
もちろんです、カボチャを使ったスイーツは秋にぴったりですね。以下にいくつかのレシピをご紹介します。1. カボチャのスフレパウンドケーキ- 材料:カボチャ 200g、生クリーム 50ml、牛乳 50ml、卵 3個、砂糖 100g、薄力粉 70g、バニラエッセンス 少々-… See the full description on the dataset page: https://huggingface.co/datasets/Atotti/spoken-magpie-ja.atomic-metrics-rm-splits
Atomic Metrics RM Task Splits
Preference-pair benchmark splits used by
Atomic Metrics. The release
contains four open-ended task families derived from public SHP, OASST1, and
OASST2 preference data.
Dataset structure
Each configuration contains 10,000 training pairs and 2,000 test pairs. Every
row has:
{
"sample_id": "source-specific stable ID",
"source_dataset": "shp | oasst1 | oasst2",
"category": "task configuration",
"split": "train | test"… See the full description on the dataset page: https://huggingface.co/datasets/tintin1027/atomic-metrics-rm-splits.task1214_atomic_classification_xwant
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1214_atomic_classification_xwant
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1214_atomic_classification_xwant.task1201_atomic_classification_xintent
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1201_atomic_classification_xintent
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1201_atomic_classification_xintent.task1216_atomic_classification_causes
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1216_atomic_classification_causes
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1216_atomic_classification_causes.task1197_atomic_classification_oreact
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1197_atomic_classification_oreact
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1197_atomic_classification_oreact.task1209_atomic_classification_objectuse
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1209_atomic_classification_objectuse
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1209_atomic_classification_objectuse.semeval_2013_task_7_beetle_5way
Dataset Card
This dataset is Atomi's version of the BEETLE subset of the SemEval 2013 Task 7 dataset, containing ~12,000 examples of question, reference answer and student answer triples graded by domain experts. The current version contains 56 unique questions in an electricity and circuits domain, up to 8 reference answers (of various quality) per question, approximately 4,795 unique student answers of 1-2 sentences, and grading labels following a 5-way classification.… See the full description on the dataset page: https://huggingface.co/datasets/Atomi/semeval_2013_task_7_beetle_5way.task1206_atomic_classification_isbefore
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1206_atomic_classification_isbefore
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1206_atomic_classification_isbefore.task1200_atomic_classification_xeffect
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1200_atomic_classification_xeffect
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1200_atomic_classification_xeffect.OpenTopics-1.0-20K
OpenTopics-1.0-20K
What is this dataset?
OpenTopics-1.0-20K is a collection of 20,003 topic names spanning a wide variety of subjects, including physics, medicine, history, law, engineering, and the arts.
AtomixLabs built this dataset to help developers, researchers, and AI builders who need a large, organized list of topics. It works great for creating synthetic prompts, testing search systems, and training models to classify text.
What is inside… See the full description on the dataset page: https://huggingface.co/datasets/AtomixLabs/OpenTopics-1.0-20K.task1210_atomic_classification_madeupof
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1210_atomic_classification_madeupof
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1210_atomic_classification_madeupof.task1203_atomic_classification_xreact
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1203_atomic_classification_xreact
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1203_atomic_classification_xreact.task1204_atomic_classification_hinderedby
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1204_atomic_classification_hinderedby
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1204_atomic_classification_hinderedby.OmniRoute-SFT-1.0-32K
OmniRoute-SFT-1.0-32K
AS OF 7/31/2026
The first large-scale supervised dataset for training LLM routing models.
Every day, production AI systems face a deceptively simple question: which model should handle this query? A one-line cooking question doesn't need a 400-billion-parameter reasoning engine. A graduate-level proof doesn't belong on a chat-optimized 8B model. Routing gets this right — and saves orders of magnitude in compute — but until now, there has been no… See the full description on the dataset page: https://huggingface.co/datasets/AtomixLabs/OmniRoute-SFT-1.0-32K.task1212_atomic_classification_hasproperty
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1212_atomic_classification_hasproperty
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1212_atomic_classification_hasproperty.atomic-formal-reasoning-complex
Atomic Formal Reasoning — Complex Numbers
Overview
This dataset contains high-quality Lean 4 formal proofs of complex number theorems, written in an explicit pedagogical calc-chain style. Each proof is fully verified, step-by-step, with no opaque tactics (simp, ring, omega are avoided). Every reasoning step is named and justified.
This is process supervision data — not just final answers. Each entry exposes the full reasoning chain, making it ideal for training models… See the full description on the dataset page: https://huggingface.co/datasets/7rouz/atomic-formal-reasoning-complex.AtomicGPT-3.0_Datasetnemotron-cc-atomic-simplification-gemma4-31b
nemotron-cc atomic-statement simplification (Gemma 4 31B-it)
2,000,000 records: source text from
nvidia/nemotron-cc-v2.1 (High-Quality-Synthetic
split) rewritten by google/gemma-4-31B-it into a sequence of atomic, Subject-Verb-Object
statements.
Generated with vLLM 0.22.1 in-process batch inference (see src/generate/run.py in the
producing repo), TP=4, max_model_len=16384, max_tokens=8192, prompts filtered to
<=8192 templated tokens.
Fields
id: original… See the full description on the dataset page: https://huggingface.co/datasets/rpisano/nemotron-cc-atomic-simplification-gemma4-31b.
