datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Qwen3.8-27B-GGUF-metrics
Qwen3.8-27B GGUF, everything behind the numbers
This is the working record for
AtomicChat/Qwen3.8-27B-GGUF.
Every figure in that model card came from a file in here, including the ones
about other publishers' builds.
The point of publishing it is simple. A quantization comparison is only worth
reading if someone else can run it, and that needs three things nobody usually
ships: the exact reference the numbers were measured against, the exact text
they were measured on, and the… See the full description on the dataset page: https://huggingface.co/datasets/AtomicChat/Qwen3.8-27B-GGUF-metrics.calib-corpora
calib-corpora
A pool of calibration material, the recipes that turn it into a calibration set
for one specific model, and the measurement corpora those quants are scored
against.
This repository is not a corpus. Nothing here is meant to be fed to
llama-imatrix as-is except the files under builds/, and each of those was
made for one named model and is close to useless for any other.
Why it is built this way
The first version of this repository was a single… See the full description on the dataset page: https://huggingface.co/datasets/AtomicChat/calib-corpora.dsv4-eval-artifacts
DeepSeek-V4-Flash-0731 — quantization measurements
Everything needed to reproduce, audit or extend the numbers published in
AtomicChat/DeepSeek-V4-Flash-0731-GGUF:
the reference logits, the evaluation corpus, the raw tool output for every quant we
measured, and the parsed results.
Every GGUF of this model that we could find on the Hub was measured here — ours,
unsloth's, bartowski's, ggml-org's, antirez's and others — on one machine, against one
reference, with one command.… See the full description on the dataset page: https://huggingface.co/datasets/AtomicChat/dsv4-eval-artifacts.Open-Omega-Atom-1.5M
Open-Omega-Atom-1.5M
Open-Omega-Atom-1.5M is a carefully curated and optimized collection derived from multiple high-quality datasets, specifically designed to enhance reasoning capabilities across mathematical and scientific domains. This dataset represents a focused subset that maintains the quality and diversity of reasoning patterns while providing an efficient size for training and evaluation. A high-quality, compact reasoning dataset designed for mathematics, science, and… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Open-Omega-Atom-1.5M.Ling-3.0-flash-GGUF-metrics
Ling-3.0-flash — quantization metrics
Everything measured while building the GGUF line for inclusionAI/Ling-3.0-flash: raw logs, per-rung numbers and the importance matrix statistics. Published so the quant table can be checked rather than trusted.
Quants live in AtomicChat/Ling-3.0-flash-GGUF.
Layout
metrics/
grid-table.json per rung: size, bpw, mean/99% KLD, top-1 agreement
kld-results.json raw parser output of every KL divergence run… See the full description on the dataset page: https://huggingface.co/datasets/AtomicChat/Ling-3.0-flash-GGUF-metrics.atomic-metrics-six-task-preferences
Six-task benchmark inputs
Seed 17. No demographic conditioning. Each task has shared train100.jsonl and test500.jsonl for Atomic Metrics, five judge variants, and learned baselines. Pair plans cover all 100 training rows once. Atomic Metrics extraction and BT/LR fitting use train100. Judges use the same test500. RM and WIMHF in the matched-data comparison use train100; rm_train_full is an explicitly separate expanded-data setting and must not be described as train100.… See the full description on the dataset page: https://huggingface.co/datasets/tintin1027/atomic-metrics-six-task-preferences.task1198_atomic_classification_owant
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1198_atomic_classification_owant
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1198_atomic_classification_owant.task1199_atomic_classification_xattr
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1199_atomic_classification_xattr
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1199_atomic_classification_xattr.atomic-metrics-classified
Atomic Metrics Classified Preference Data
Classified preference-pair data used by
Atomic Metrics.
This release contains the task-family labels and experiment splits derived
from public SHP, OASST1, OASST2, and related sources. It does not include raw
source dumps or clustering caches.
Layout
rm_splits/: 10k/2k train-test preference splits for
trouble_solution, writing, daily_ideation, values_controversial,
plus extraction/train100.jsonl, test200.jsonl, and… See the full description on the dataset page: https://huggingface.co/datasets/tintin1027/atomic-metrics-classified.task1211_atomic_classification_hassubevent
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1211_atomic_classification_hassubevent
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1211_atomic_classification_hassubevent.task1217_atomic_answer_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1217_atomic_answer_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1217_atomic_answer_generation.task1207_atomic_classification_atlocation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1207_atomic_classification_atlocation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1207_atomic_classification_atlocation.clotho-jaQwen/Qwen3-4B-Instruct-2507を使用してClothoを日本語に翻訳したデータです。
ライセンスは元データに従います。
task1196_atomic_classification_oeffect
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1196_atomic_classification_oeffect
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1196_atomic_classification_oeffect.spoken-magpie-ja
Spoken-magpie
LLMの日本語Instruction Tuning用データllm-jp/magpie-sft-v1.0をCosyVoice2 TTSを使用して音声化した商用利用可能な日本語の音声言語モデルのSFT用データセットです。
ある程度の話者多様性を持つように生成されています。
Respone Audioは500文字以下の場合にのみ生成されています。
NVIDIA H200を10枚を使用しvllmで推論しました。
Samples
最初の50サンプルを掲載します。
ID
Instruction
Instruction Audio
Response
Response Audio
0
カボチャを使ったスイーツのレシピをいくつか教えてください。
もちろんです、カボチャを使ったスイーツは秋にぴったりですね。以下にいくつかのレシピをご紹介します。1. カボチャのスフレパウンドケーキ- 材料:カボチャ 200g、生クリーム 50ml、牛乳 50ml、卵 3個、砂糖 100g、薄力粉 70g、バニラエッセンス 少々-… See the full description on the dataset page: https://huggingface.co/datasets/Atotti/spoken-magpie-ja.atomic-metrics-rm-splits
Atomic Metrics RM Task Splits
Preference-pair benchmark splits used by
Atomic Metrics. The release
contains four open-ended task families derived from public SHP, OASST1, and
OASST2 preference data.
Dataset structure
Each configuration contains 10,000 training pairs and 2,000 test pairs. Every
row has:
{
"sample_id": "source-specific stable ID",
"source_dataset": "shp | oasst1 | oasst2",
"category": "task configuration",
"split": "train | test"… See the full description on the dataset page: https://huggingface.co/datasets/tintin1027/atomic-metrics-rm-splits.task1214_atomic_classification_xwant
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1214_atomic_classification_xwant
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1214_atomic_classification_xwant.task1201_atomic_classification_xintent
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1201_atomic_classification_xintent
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1201_atomic_classification_xintent.task1216_atomic_classification_causes
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1216_atomic_classification_causes
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1216_atomic_classification_causes.task1197_atomic_classification_oreact
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1197_atomic_classification_oreact
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1197_atomic_classification_oreact.task1209_atomic_classification_objectuse
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1209_atomic_classification_objectuse
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1209_atomic_classification_objectuse.semeval_2013_task_7_beetle_5way
Dataset Card
This dataset is Atomi's version of the BEETLE subset of the SemEval 2013 Task 7 dataset, containing ~12,000 examples of question, reference answer and student answer triples graded by domain experts. The current version contains 56 unique questions in an electricity and circuits domain, up to 8 reference answers (of various quality) per question, approximately 4,795 unique student answers of 1-2 sentences, and grading labels following a 5-way classification.… See the full description on the dataset page: https://huggingface.co/datasets/Atomi/semeval_2013_task_7_beetle_5way.task1206_atomic_classification_isbefore
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1206_atomic_classification_isbefore
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1206_atomic_classification_isbefore.task1200_atomic_classification_xeffect
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1200_atomic_classification_xeffect
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1200_atomic_classification_xeffect.OpenTopics-1.0-20K
OpenTopics-1.0-20K
What is this dataset?
OpenTopics-1.0-20K is a collection of 20,003 topic names spanning a wide variety of subjects, including physics, medicine, history, law, engineering, and the arts.
AtomixLabs built this dataset to help developers, researchers, and AI builders who need a large, organized list of topics. It works great for creating synthetic prompts, testing search systems, and training models to classify text.
What is inside… See the full description on the dataset page: https://huggingface.co/datasets/AtomixLabs/OpenTopics-1.0-20K.task1210_atomic_classification_madeupof
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1210_atomic_classification_madeupof
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1210_atomic_classification_madeupof.task1203_atomic_classification_xreact
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1203_atomic_classification_xreact
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1203_atomic_classification_xreact.task1204_atomic_classification_hinderedby
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1204_atomic_classification_hinderedby
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1204_atomic_classification_hinderedby.OmniRoute-SFT-1.0-32K
OmniRoute-SFT-1.0-32K
AS OF 7/31/2026
The first large-scale supervised dataset for training LLM routing models.
Every day, production AI systems face a deceptively simple question: which model should handle this query? A one-line cooking question doesn't need a 400-billion-parameter reasoning engine. A graduate-level proof doesn't belong on a chat-optimized 8B model. Routing gets this right — and saves orders of magnitude in compute — but until now, there has been no… See the full description on the dataset page: https://huggingface.co/datasets/AtomixLabs/OmniRoute-SFT-1.0-32K.task1212_atomic_classification_hasproperty
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1212_atomic_classification_hasproperty
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1212_atomic_classification_hasproperty.
