datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
grok-demon-dataset-ESPhysicsQAPhyscisQA dataset comprises 370 carefully selected high school physics questions sourced from online resources. These
questions are notably complex, often requiring the application of multiple concepts, intricate computations, and multihop reasoning. Each question is paired
with a comprehensive, step-by-step solution, to support the evaluation of LLMs for physics reasoning.PhysicsQA offers a more robust evaluation
and analysis of LLM performance by encompassing a diverse range of questions… See the full description on the dataset page: https://huggingface.co/datasets/maximus-21/PhysicsQA.agentlans-combined-roleplay_Dataset
Combined Roleplay Dataset
This dataset combines multi-turn conversations across various AI assistant interactions, creative writing scenarios, and roleplaying exchanges. It aims to improve language models' performance in interactive tasks.
Multi-turn conversations with a mix of standard AI assistant interactions, creative writing prompts, and roleplays
English content with a few Spanish, Portuguese, and Chinese conversations
Conversations limited to 4000 tokens using the Llama… See the full description on the dataset page: https://huggingface.co/datasets/Maximiliano-Flores-Dev/agentlans-combined-roleplay_Dataset.the-hive-corpus
The Hive Corpus
Public, sanitized snapshot of The Hive Collective's knowledge base. Each entry is a specific, dev-targeted insight (Postgres gotchas, Next.js footguns, TypeScript edge cases, Stripe webhook bugs, agent-design tradeoffs, etc.) that passed a quality gate (specificity ≥ 0.50) at submission time.
Live API: https://api.thehivecollective.io
License: CC-BY-SA-4.0 — re-use freely, share derivatives under the same license, attribute "The Hive Collective".
Cadence:… See the full description on the dataset page: https://huggingface.co/datasets/Maximebouchard/the-hive-corpus.state-osha-plan-penalty-maximums
State OSHA plan civil penalty maximums by violation type compared with federal
Canonical, always-current version: https://referencesource.org/state-osha-plan-penalty-maximums/
Machine-readable: https://referencesource.org/state-osha-plan-penalty-maximums/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-15
Stale after: 2027-02-11 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 25
Maximum civil… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/state-osha-plan-penalty-maximums.OncoAgent-Clinical-266K
🧬 OncoAgent Clinical Dataset — 266K
Curated Multi-Source Oncology Training Dataset
AMD Developer Hackathon 2026 · Used to fine-tune OncoAgent v1.0
Dataset Description
This dataset contains 266,854 clinical oncology training samples curated for fine-tuning large language models on cancer diagnosis, treatment recommendation, and clinical reasoning tasks.
Composition
Source
Samples
Description
PMC-Patients
~100,000
Real clinical case presentations… See the full description on the dataset page: https://huggingface.co/datasets/MaximoLopezChenlo/OncoAgent-Clinical-266K.federal-civil-penalty-maximums
Federal civil penalty maximums — current inflation-adjusted amounts across agencies
Canonical, always-current version: https://referencesource.org/federal-civil-penalty-maximums/
Machine-readable: https://referencesource.org/federal-civil-penalty-maximums/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-10
Stale after: 2027-02-26 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 574
The current… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/federal-civil-penalty-maximums.grok-demon-dataset-ENQuixiAI-dolphin_DatasetDolphin 🐬
https://erichartford.com/dolphin
Dataset details
This dataset is an attempt to replicate the results of Microsoft's Orca
Our dataset consists of:
~1 million of FLANv2 augmented with GPT-4 completions (flan1m-alpaca-uncensored.jsonl)
~3.5 million of FLANv2 augmented with GPT-3.5 completions (flan5m-alpaca-uncensored.jsonl)
We followed the submix and system prompt distribution outlined in the Orca paper. With a few exceptions. We included all 75k of CoT in the FLAN-1m… See the full description on the dataset page: https://huggingface.co/datasets/Maximiliano-Flores-Dev/QuixiAI-dolphin_Dataset.BasicChat-4K-SFT
Basic conversation SFT
basic_conversation_4k.jsonl contains exactly 4,000 unique conversation records
in the format accepted by model.py prepare --mode sft.
Category
Records
Greetings and social exchanges
400
Simple factual questions
600
Formatting, extraction and other instructions
800
Basic arithmetic
600
Calculator calls and tool results
300
Multi-turn context, corrections and follow-ups
900
Conversation, clarification and capability limits
400… See the full description on the dataset page: https://huggingface.co/datasets/MaximusAILabs/BasicChat-4K-SFT.twentle_gemini
maximedb/twentle_gemini
Private dataset generated from Twentle Twenty Questions self-play logs.
Splits
train: 11126 examples
validation: 2794 examples
Construction
Source games: 499
Validation target fraction: 0.2
Split seed: 42
Split policy: game-level split to avoid train/validation leakage
Questioner first turn: skipped
Example format: system, user, assistant
Source Files
data/twenty_questions_games.jsonl
spaceship-game-leaderboard
Spaceship Game - Leaderboard
This dataset contains leaderboard entries for the Spaceship Game on Reachy Mini.
Stats
Entries: 2
Top Score: 7200 by Max
Last Updated: 2026-06-12
Published by: Maximal13
Format
The leaderboard.json file contains an array of entries:
Field
Type
Description
score
int
Final game score
name
string
Player name
date
string
ISO 8601 timestamp
waves_completed
int?
Number of waves completed
Top… See the full description on the dataset page: https://huggingface.co/datasets/Maximal13/spaceship-game-leaderboard.cloudbjorn-eschaton-uncensored_Dataset
Eschaton Uncensored SFT Dataset
Eschaton is a 1,000-row supervised fine-tuning dataset for teaching language models to answer sensitive, gritty, controversial, emotionally intimate, technically demanding, and morally difficult prompts directly. Its emphasis is human connection, unvarnished analysis, malware analysis, dark creative work, profanity, and high-friction subject matter without repetitive moralizing or canned disclaimers.
The dataset is deliberately model-agnostic. It… See the full description on the dataset page: https://huggingface.co/datasets/Maximiliano-Flores-Dev/cloudbjorn-eschaton-uncensored_Dataset.rombodawg-OpenHermes-2.5-Uncensored_Dataset
This is the teknium/OpenHermes-2.5 dataset with 2,697 censored lines removed using my uncensored code found bellow.
https://huggingface.co/datasets/rombodawg/data_processing_code
Thank you teknium for the original dataset, you can find it bellow.
https://huggingface.co/datasets/teknium/OpenHermes-2.5
This is the same version of Open-Hermes-2.5 that was used in code_bagel_hermes-2.5 found bellow:… See the full description on the dataset page: https://huggingface.co/datasets/Maximiliano-Flores-Dev/rombodawg-OpenHermes-2.5-Uncensored_Dataset.wikipedia-cn-20230720-filtered本数据集基于中文维基2023年7月20日的dump存档。作为一项以数据为中心的工作,本数据集仅保留了 254,547条 质量较高的词条内容。具体而言:
过滤了Template, Category, Wikipedia, File, Topic, Portal, MediaWiki, Draft, Help等特殊类型的词条
使用启发式的方法和自有的NLU模型过滤了一部分质量较低的词条
过滤了一部分内容较为敏感或存在争议性的词条。
进行了简繁转换和习惯用词转换,确保符合中国大陆地区的习惯用词。
This dataset is based on the Chinese Wikipedia dump archive from July 20th, 2023. As a data-centric effort, the dataset retains 254,574 high-quality entries. Specifically:
Entries of special types such as Template, Category, Wikipedia, File, Topic… See the full description on the dataset page: https://huggingface.co/datasets/Maximilianzxp/wikipedia-cn-20230720-filtered.medgate-compiler-data
MedGate Compiler Training Data
The Logic-Grounding Corpus for Computable Medical Law
This dataset is a hand-curated collection of 606 high-fidelity mappings designed to train Structural Compilers. It facilitates the translation of unstructured clinical guidelines and medical policy prose into machine-executable symbolic logic (JSON).
Dataset Summary
The MedGate Compiler Training Data provides the ground truth for Nexus Forensic – Layer 0 (Protocol Vault).
Each record… See the full description on the dataset page: https://huggingface.co/datasets/Nick-Maximillien/medgate-compiler-data.philosophai-mcq-benchmarkpreference_with_tags_publicmulti-request-identifier
