datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
harmonymitre-stix-cve-exploitdb-dataset-alpaca-chatml-harmony
MITRE+NVD+ExploitDB Dataset (Alpaca/ChatML/Harmony)
A dataset for training AI assistants/agents on vulnerability analysis and pentesting Q&A. It is built by the pentestds pipeline, which fetches and merges data from MITRE CVE, NVD (CVSS enrichment), ExploitDB, and a small set of HuggingFace datasets. Provenance is recorded for every entry, and the pipeline emits Alpaca, ChatML, and Harmony JSONL files.
Dataset Summary
This dataset is designed for training AI agents to… See the full description on the dataset page: https://huggingface.co/datasets/jason-oneal/mitre-stix-cve-exploitdb-dataset-alpaca-chatml-harmony.harmony-visionharmony4d-flattened
harmony4d-flattened
Instruction-verb + <caption> + reverse-direction augment of the harmony4d
agent-token training text -- 16,640 rows (8,320 source rows x 2 tasks).
Replaces this repo's prior single-task release (window=8, 8,320 rows,
train/test split) -- the old content isn't used by anything going forward.
This release is single-split (train only, 16,640 rows); the prior
held-out test split was not carried over.
Why this exists
Harsh Raj flagged (Discord… See the full description on the dataset page: https://huggingface.co/datasets/EmpathicRobotics/harmony4d-flattened.harmony-ocr-combinedgpt-oss-120B-distilled-math-OpenAI-Harmony
📚 Dataset Overview
Data Source Model: gpt-oss-120bTask Type: Mathematical Problem SolvingData Format: JSON Lines (.jsonl)Fields: Generator, Category, Input, Output
Note: If you are using this template for training, please make sure the format is correct before starting.Since this template is still under continuous improvement and learning, it may not be fully complete yet. I appreciate your understanding.
📈 Core Statistics
Generated complete reasoning processes… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/gpt-oss-120B-distilled-math-OpenAI-Harmony.UIGEN-T2-Harmonyreasoning-and-chat-harmony-format
Open Paws Reasoning And Conversational Finetuning Harmony Format
This dataset is part of the Open Paws initiative to develop AI training data aligned with animal liberation and advocacy principles. Created to train AI systems that understand and promote animal welfare, rights, and liberation.
Dataset Details
Dataset Type: Reasoning Data
Format: JSONL (JSON Lines)
Languages: Multilingual (primarily English)
Focus: Animal advocacy and ethical reasoning
Organization: Open… See the full description on the dataset page: https://huggingface.co/datasets/open-paws/reasoning-and-chat-harmony-format.harmony-nemotron-cpu-artifacts
Harmony CPU artifacts: Nemotron datasets (normalized + candidate pools)
This dataset repo is an artifact store produced on an EPYC CPU box. It contains:
normalized/ — CPU-normalized Parquet shards with a text-first Harmony format (text) plus meta_* and quality_* fields.
pools/ — candidate pool Parquet shards (subsets) for later GPU scoring (Modal NLL/PPL). No GPU scoring has been run yet.
reports/ — summary tables of counts per dataset/split/pool.
Directory layout… See the full description on the dataset page: https://huggingface.co/datasets/radna0/harmony-nemotron-cpu-artifacts.sidekick-autocomplete-char-harmony-1MTriangle104__Mistral-Small-24b-Harmony-details
Dataset Card for Evaluation run of Triangle104/Mistral-Small-24b-Harmony
Dataset automatically created during the evaluation run of model Triangle104/Mistral-Small-24b-Harmony
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Triangle104__Mistral-Small-24b-Harmony-details.HarmonyVoice071523
Dataset Card for "HarmonyVoice071523"
More Information needed
synthetic-tool-calls-harmony-v2harmony-v2-syntetic
Harmony Dataset — Synthetic Toxicity Dataset
A high-quality synthetic dataset for training toxicity detection models, generated using Nemotron 3 Ultra (free) via OpenRouter.
📊 Overview
The dataset consists of 7,000 realistic chat-like messages in Ukrainian, Russian, and mixed speech. It is designed to teach models to distinguish between:
Toxic: Direct personal attacks, harassment, threats, humiliation.
Safe: Emotional expression, profanity without a target… See the full description on the dataset page: https://huggingface.co/datasets/floxoris/harmony-v2-syntetic.pre-pro_logeharmony-tools
Harmony Tool-Call Conversations
This dataset contains 10000 synthetic Harmony-formatted conversations designed to teach models
how to reason about tool usage, issue function calls, and craft final answers after receiving tool outputs.
Repo: dwojcik/harmony-tools
Schema: prompt / completion pairs following the OpenAI Harmony prompt syntax.
Focus: tool invocation planning, JSON argument formatting, and final response composition.
Stage Breakdown
final_answer: 5000… See the full description on the dataset page: https://huggingface.co/datasets/dwojcik/harmony-tools.legal-reasoning-harmonysynthetic-tool-calls-harmonylegal-reasoning-harmony
Legal Reasoning Harmony (CoT → Harmony)
This dataset converts moremilk/CoT_Legal_Issues_And_Laws (MIT-licensed) into the Harmony message format for GPT-OSS fine-tuning.
Source: moremilk/CoT_Legal_Issues_And_Laws
License: MIT (inherited from source)
Examples: 4,237
Format: JSONL, Harmony messages with separated thinking and content
Format: JSONL, Harmony messages with explicit channels (analysis, final) and convenience top-level fields
Provenance and transformation… See the full description on the dataset page: https://huggingface.co/datasets/zackproser/legal-reasoning-harmony.nemotron-math-v2-harmony-toolsUIGEN-T2-Harmony-1000claude-4.5-high-reasoning-250x-harmonyForked version of https://huggingface.co/datasets/TeichAI/claude-sonnet-4.5-high-reasoning-250x, in OpenAI Harmony format.
Nothing has been changed except the format, really.
Original README:
This is a reasoning dataset created using Claude Sonnet 4.5 with a high reasoning effort. Some of these questions are from reedmayhew and the rest were generated.
The dataset is meant for creating distilled versions of Claude Sonnet 4.5 by fine-tuning already existing open-source LLMs.
The… See the full description on the dataset page: https://huggingface.co/datasets/konasquared/claude-4.5-high-reasoning-250x-harmony.jazz-harmony-embeddings
Jazz Harmony Embeddings — 6,900 tune vectors
One 128-dimensional vector per jazz standard, from a small transformer
trained from scratch so that tunes with related harmony — transpositions,
alternate charts, contrafacts — land close together. Produced by the
3-seed ensemble released at
eigenben/jazz-harmony-embeddings;
code and full experiment records at
github.com/eigenben/jazz-harmony-embeddings.
Files
embeddings.npz — embeddings: (6900, 128) float32… See the full description on the dataset page: https://huggingface.co/datasets/eigenben/jazz-harmony-embeddings.The-History-Of-Africa-The-Quest-For-Eternal-Harmony
The History of Africa - THE QUEST FOR ETERNAL HARMONY
Authoritative and comprehensive, The History of Africa provides an accessible narrative from earliest prehistory to the present day, with unusual attention paid to the ordinary lives of Africans. This survey includes a wealth of indigenous ideas, African concepts, and traditional outlooks that have escaped the writing of African history in the West. The fully updated new edition includes information on the recent conflicts in… See the full description on the dataset page: https://huggingface.co/datasets/Svngoku/The-History-Of-Africa-The-Quest-For-Eternal-Harmony.APIKG4Syn-HarmonyOS-Dataset
Framework-Aware Code Generation with API Knowledge Graph–Constructed Data: A Study on HarmonyOS
🗂️The Dataset
OHBen.json: For the integrated version of the two aforementioned files, which constitutes the final dataset used for fine-tuning the LLM.
japanese-harmony-datasetHarmonyTriangle104__DS-R1-Llama-8B-Harmony-details
Dataset Card for Evaluation run of Triangle104/DS-R1-Llama-8B-Harmony
Dataset automatically created during the evaluation run of model Triangle104/DS-R1-Llama-8B-Harmony
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Triangle104__DS-R1-Llama-8B-Harmony-details.Triangle104__DS-R1-Distill-Q2.5-10B-Harmony-details
Dataset Card for Evaluation run of Triangle104/DS-R1-Distill-Q2.5-10B-Harmony
Dataset automatically created during the evaluation run of model Triangle104/DS-R1-Distill-Q2.5-10B-Harmony
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Triangle104__DS-R1-Distill-Q2.5-10B-Harmony-details.Triangle104__DS-R1-Distill-Q2.5-14B-Harmony_V0.1-details
Dataset Card for Evaluation run of Triangle104/DS-R1-Distill-Q2.5-14B-Harmony_V0.1
Dataset automatically created during the evaluation run of model Triangle104/DS-R1-Distill-Q2.5-14B-Harmony_V0.1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Triangle104__DS-R1-Distill-Q2.5-14B-Harmony_V0.1-details.
