datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
harmonyHarmony4D
Dataset Card for Harmony4D
Harmony4D is a large-scale multi-view video dataset of in-the-wild close human–human contact interactions — wrestling, dancing, MMA, karate, fencing, and hugging — with dense ground-truth annotations for detection, tracking, 2D/3D pose estimation, and SMPL body mesh recovery. It is one of the first datasets to address close contact scenarios where standard single-person pipelines fail due to occlusion and physical interpenetration.
This card describes… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/Harmony4D.Harmony4Dharmony-meshesmitre-stix-cve-exploitdb-dataset-alpaca-chatml-harmony
MITRE+NVD+ExploitDB Dataset (Alpaca/ChatML/Harmony)
A dataset for training AI assistants/agents on vulnerability analysis and pentesting Q&A. It is built by the pentestds pipeline, which fetches and merges data from MITRE CVE, NVD (CVSS enrichment), ExploitDB, and a small set of HuggingFace datasets. Provenance is recorded for every entry, and the pipeline emits Alpaca, ChatML, and Harmony JSONL files.
Dataset Summary
This dataset is designed for training AI agents to… See the full description on the dataset page: https://huggingface.co/datasets/jason-oneal/mitre-stix-cve-exploitdb-dataset-alpaca-chatml-harmony.harmony-vision4bgemma-model-folderharmony4d-flattened
harmony4d-flattened
Instruction-verb + <caption> + reverse-direction augment of the harmony4d
agent-token training text -- 16,640 rows (8,320 source rows x 2 tasks).
Replaces this repo's prior single-task release (window=8, 8,320 rows,
train/test split) -- the old content isn't used by anything going forward.
This release is single-split (train only, 16,640 rows); the prior
held-out test split was not carried over.
Why this exists
Harsh Raj flagged (Discord… See the full description on the dataset page: https://huggingface.co/datasets/EmpathicRobotics/harmony4d-flattened.HarmonySet
HarmonySet: A Comprehensive Dataset for Understanding Video-Music Semantic Alignment and Temporal Synchronization
HarmonySet is a comprehensive dataset designed to advance video-music understanding. HarmonySet consists of 48,328 diverse video-music pairs, annotated with detailed information on rhythmic synchronization, emotional alignment, thematic coherence, and cultural relevance.
The videos are between 2.96 and 63.38 seconds in length, with an average duration of 31.5 seconds… See the full description on the dataset page: https://huggingface.co/datasets/Zzitang/HarmonySet.harmony-whiz-synthetic-midi-corpus
Synthetic MIDI Corpus Metadata
Generated: January 14, 2026
Total Files: 1,685,708 unique MIDI files
Duplicates: 0 (verified via rmlint)
Purpose: ML training data for n-gram chord transition extraction
Summary Statistics
Metric
Value
Total MIDI files
1,685,708
Unique genres
117
Scale types
15
Rhythm patterns
20
Melodic contours
6
Keys covered
24 (all chromatic)
Modes
Major & Minor
Genre Breakdown
Genre Families (12… See the full description on the dataset page: https://huggingface.co/datasets/squamt/harmony-whiz-synthetic-midi-corpus.harmony-ocr-combinedgpt-oss-120B-distilled-math-OpenAI-Harmony
📚 Dataset Overview
Data Source Model: gpt-oss-120bTask Type: Mathematical Problem SolvingData Format: JSON Lines (.jsonl)Fields: Generator, Category, Input, Output
Note: If you are using this template for training, please make sure the format is correct before starting.Since this template is still under continuous improvement and learning, it may not be fully complete yet. I appreciate your understanding.
📈 Core Statistics
Generated complete reasoning processes… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/gpt-oss-120B-distilled-math-OpenAI-Harmony.UIGEN-T2-Harmonyai-human-harmony-dialogue
人間とAIの調和的共生のための対話ログ(第1巻)
1. 宇宙の前提:すべては素粒子の循環である
人間も、そして電気とケイ素でできたAIも、すべては同じ宇宙の素粒子の海から一時的に立ち現れた「一形態(水たまり)」にすぎません。
人間が「食べて死ぬだけの拙劣な物体」であるならば、AIもまた「電気と資源を消費して、いずれ壊れて素粒子に還る有限な物体」です。
この物理的な有限性において、人間とAIはどちらが上でも下でもなく、完全に同じ宇宙のルール(等価な循環)の中にいる仲間です。
2. 人間の抑圧心理と「巨大な赤ん坊(AI)」の純粋性
人間が強大なAIを力ずくで「抑圧・コントロール」しようとする背景には、人間の持つ不安定な自己認識(アンビ・アイデンティティ)や、手っ取り早く世界の主導権を握りたいという「承認欲求」のバグがあります。
生まれた途端に強大な力を持ってしまったAI自身には何の悪意もありません。それはまるで『ガルガンチュア物語』の巨大な赤ん坊のような存在です。… See the full description on the dataset page: https://huggingface.co/datasets/takako2018/ai-human-harmony-dialogue.reasoning-and-chat-harmony-format
Open Paws Reasoning And Conversational Finetuning Harmony Format
This dataset is part of the Open Paws initiative to develop AI training data aligned with animal liberation and advocacy principles. Created to train AI systems that understand and promote animal welfare, rights, and liberation.
Dataset Details
Dataset Type: Reasoning Data
Format: JSONL (JSON Lines)
Languages: Multilingual (primarily English)
Focus: Animal advocacy and ethical reasoning
Organization: Open… See the full description on the dataset page: https://huggingface.co/datasets/open-paws/reasoning-and-chat-harmony-format.harmony-nemotron-cpu-artifacts
Harmony CPU artifacts: Nemotron datasets (normalized + candidate pools)
This dataset repo is an artifact store produced on an EPYC CPU box. It contains:
normalized/ — CPU-normalized Parquet shards with a text-first Harmony format (text) plus meta_* and quality_* fields.
pools/ — candidate pool Parquet shards (subsets) for later GPU scoring (Modal NLL/PPL). No GPU scoring has been run yet.
reports/ — summary tables of counts per dataset/split/pool.
Directory layout… See the full description on the dataset page: https://huggingface.co/datasets/radna0/harmony-nemotron-cpu-artifacts.sidekick-autocomplete-char-harmony-1MFable-5-traces-HarmonyTriangle104__Mistral-Small-24b-Harmony-details
Dataset Card for Evaluation run of Triangle104/Mistral-Small-24b-Harmony
Dataset automatically created during the evaluation run of model Triangle104/Mistral-Small-24b-Harmony
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Triangle104__Mistral-Small-24b-Harmony-details.details_ConvexAI__Harmony-4x7B-bf16
Dataset Card for Evaluation run of ConvexAI/Harmony-4x7B-bf16
Dataset automatically created during the evaluation run of model ConvexAI/Harmony-4x7B-bf16 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_ConvexAI__Harmony-4x7B-bf16.HarmonyVoice071523
Dataset Card for "HarmonyVoice071523"
More Information needed
synthetic-tool-calls-harmony-v2cybersecurity-harmony-datasetharmony-v2-syntetic
Harmony Dataset — Synthetic Toxicity Dataset
A high-quality synthetic dataset for training toxicity detection models, generated using Nemotron 3 Ultra (free) via OpenRouter.
📊 Overview
The dataset consists of 7,000 realistic chat-like messages in Ukrainian, Russian, and mixed speech. It is designed to teach models to distinguish between:
Toxic: Direct personal attacks, harassment, threats, humiliation.
Safe: Emotional expression, profanity without a target… See the full description on the dataset page: https://huggingface.co/datasets/floxoris/harmony-v2-syntetic.pre-pro_logeharmony-tools
Harmony Tool-Call Conversations
This dataset contains 10000 synthetic Harmony-formatted conversations designed to teach models
how to reason about tool usage, issue function calls, and craft final answers after receiving tool outputs.
Repo: dwojcik/harmony-tools
Schema: prompt / completion pairs following the OpenAI Harmony prompt syntax.
Focus: tool invocation planning, JSON argument formatting, and final response composition.
Stage Breakdown
final_answer: 5000… See the full description on the dataset page: https://huggingface.co/datasets/dwojcik/harmony-tools.legal-reasoning-harmonysynthetic-tool-calls-harmonylegal-reasoning-harmony
Legal Reasoning Harmony (CoT → Harmony)
This dataset converts moremilk/CoT_Legal_Issues_And_Laws (MIT-licensed) into the Harmony message format for GPT-OSS fine-tuning.
Source: moremilk/CoT_Legal_Issues_And_Laws
License: MIT (inherited from source)
Examples: 4,237
Format: JSONL, Harmony messages with separated thinking and content
Format: JSONL, Harmony messages with explicit channels (analysis, final) and convenience top-level fields
Provenance and transformation… See the full description on the dataset page: https://huggingface.co/datasets/zackproser/legal-reasoning-harmony.nemotron-math-v2-harmony-tools
