datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nova-v8-dataset
NOVA v8 Pre-Training Dataset
This is the official pre-training dataset used to train the NOVA v8 architecture (a 710M parameter hybrid Sparse Neural Network).
Dataset Structure
Dataset Mix (10 Billion Tokens)
This is NOT a generic web crawl. This dataset is an ultra-dense "university education" designed to make the 710M model punch far above its weight class in reasoning, logic, code, and structural awareness.
Dataset
Weight
Description… See the full description on the dataset page: https://huggingface.co/datasets/ReXeeD/nova-v8-dataset.danish-tool-dialogues-v8
danish-tool-dialogues-v1
Danish multi-turn tool-use conversations with reasoning, translated from the
Glaive subset of
Nanbeige/ToolMind
(Apache-2.0) by scripts/translate_toolmind_da.py.
Complements danish-tool-calls-v1, which is single-turn and synthetic. Here
the conversations run several turns, tool results are fed back, and the
assistant reasons before calling.
split
rows
train
34,168
eval_seen_tools
698
eval_unseen_tools
768
eval_seen_sym
752… See the full description on the dataset page: https://huggingface.co/datasets/jensjepsen/danish-tool-dialogues-v8.zeroin-v8-r4-dataset
Zeroin Methodology v8-r4 QA
Training corpus used to fine-tune
KG-ZEROIN/gpt-oss-20b-zeroin-v8-r4.
The corpus captures question/answer pairs derived from the Zeroin
fund-evaluation methodology Korean domain document. It is organized
as a Harmony-ready chat-messages dataset for supervised full
fine-tuning of
openai/gpt-oss-20b.
Released under CC BY-NC 4.0 — free for non-commercial research,
evaluation, and educational use. See LICENSE and NOTICE.
Contents
File… See the full description on the dataset page: https://huggingface.co/datasets/KG-ZEROIN/zeroin-v8-r4-dataset.blueteam-v8
Blue_team_v8
Synthetic fine-tuning dataset generated with Dataset Genie 0.1.0 on 2026-09-21T15:02:28+00:00.
Domain brief
Blue-team is broad. Cover these deliberately, spread across difficulty tiers:
Platforms, not one vendor. Splunk (SPL), Microsoft Sentinel (KQL), Elastic (ES|QL/EQL/Lucene), CrowdStrike, Defender for Endpoint, Sysmon, Zeek/Suricata. Identity: Entra ID, Okta, Active Directory/Kerberos. Cloud: AWS (CloudTrail/GuardDuty), Azure, GCP. OS: Windows… See the full description on the dataset page: https://huggingface.co/datasets/k3nn3dy/blueteam-v8.loop-qwen-v8-sft
loop-qwen-v8 SFT dataset (Gemini insulin-control distillation)
8,222 chat-format examples used to SFT loop-qwen-v8 (Qwen3-4B insulin
controller distilled from Gemini-3-flash-preview). Each example is a closed-loop
dosing decision.
Format (JSONL, one chat per line)
system: controller spec (IOB-aware, Chain-of-Draft reason-before-act)
user: patient metadata (age, weight, TDD, CF, IC, basal) + 6h history of CGM / insulin / carbs, as JSON
assistant: JSON… See the full description on the dataset page: https://huggingface.co/datasets/jxx123/loop-qwen-v8-sft.orca_mini_v8_sharegpt_formatBest 45K samples of Bigger Orca Mini dataset in sharegpt format, Enjoy!
Orin-Instruct-Alpaca-JP-v8
Converted QA Dataset
このデータセットは、easy-dataset-cliを使用して生成されたアルパカ形式の日本語Q&Aデータセットです。
データセット概要
総エントリ数: 3,731
形式: Alpaca形式
言語: 日本語
ライセンス: MIT
データ構造
各エントリは以下の形式です:
{
"instruction": "質問文",
"input": "",
"output": "回答文",
"genre": "ジャンル",
"audience": "対象読者"
}
ジャンル分布
含まれるジャンル:
FAQ
PC向けガイド
ゆっくりガイド
アジア系ユーザー向けガイド
イベントガイド
イベント常連ガイド
コレクションガイド
ジュニアガイド
ソロ活動ガイド
ソーシャルガイド
ファミリーガイド
プロフェッショナルガイド
プロ創作者ガイド
プロ配信者ガイド
ベテラン社会人ガイド
モバイルガイド
レビュー記事
上級者向けマニュアル
上級者向け攻略… See the full description on the dataset page: https://huggingface.co/datasets/MakiAi/Orin-Instruct-Alpaca-JP-v8.kush-v82-eval-samples
Kush v82 Eval Samples
Public evaluation samples for the Hotep Intelligence Kush v82 flagship model. Every entry is a prompt, a category, and a reference answer written in the target voice, so a reader can judge tone, framing, and factual grounding at the same time.
Try the live model in the hotep-intelligence-chat Space before or after reading these samples.
What This Dataset Is For
style and persona inspection
historical framing checks
sovereignty and… See the full description on the dataset page: https://huggingface.co/datasets/hotepfederales/kush-v82-eval-samples.
