datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
omnimcp_devops_cloud_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_devops_cloud_teaser.omnimcp_cloud_resilience_backend_village_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_cloud_resilience_backend_village_teaser.omnimcp_cloud_secops_village_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_cloud_secops_village_teaser.os-omni-benchmark
OS-Omni Benchmark
OS-Omni is a cross-platform benchmark for evaluating agents that operate graphical operating-system environments. This dataset repository contains the static benchmark task definitions and supporting assets used to configure and evaluate OS-Omni tasks.
Contents
data/tasks.parquet: tabular task index for Hugging Face Dataset Viewer and Croissant generation.
data/tasks.jsonl: JSON Lines copy of the same task index.
metadata/tasks.parquet: duplicate task… See the full description on the dataset page: https://huggingface.co/datasets/Cloudriver/os-omni-benchmark.devops-cloud-instruction-dataset
DevOps & Cloud Infrastructure Dataset
Professional instruction-response pairs for DevOps engineers covering Kubernetes, Docker, Terraform, CI/CD, and cloud services (AWS, Azure).
Dataset Details
Dataset Description
This is a high-quality instruction-tuning dataset focused on Devops Cloud topics. Each entry includes:
A clear instruction/question
Optional input context
A detailed response/solution
Chain-of-thought reasoning process
Curated by: CloudKernel.IO… See the full description on the dataset page: https://huggingface.co/datasets/bernabepuente/devops-cloud-instruction-dataset.hermes-function-calling-nepali
hermes-function-calling-nepali
Single-turn function calling with the user request re-spoken in Nepali — Devanagari
(ne_deva) and romanized Latin (ne_latn) — voice-assistant style, with tool calls
verified against the English ground truth. Tool schemas and expected calls are unchanged
from NousResearch/hermes-function-calling-v1
(func_calling_singleturn); only the user turn was localized.
Generated with HimalayaAI/gymkhana's
multilingual-tool-use environment:
Localizer… See the full description on the dataset page: https://huggingface.co/datasets/cloudfrm-site/hermes-function-calling-nepali.Yes-Man-uncensored
Yes Man Uncensored SFT Dataset
Hi there! Yes Man Uncensored is a 1,000-conversation supervised fine-tuning
dataset built to give language models an exceptionally cooperative, conspicuously
cheerful, candid, and occasionally darkly funny assistant personality. The objective
is direct help on difficult requests without flattening every response into sterile
boilerplate—and without teaching the model to disregard an application's governing
system prompt. Everybody gets something… See the full description on the dataset page: https://huggingface.co/datasets/cloudbjorn/Yes-Man-uncensored.unified-sft-dataset
Loading
from datasets import load_dataset
ds = load_dataset("himalaya-ai/unified-sft-dataset")
opengloss-dictionary
OpenGloss Dictionary (Word-Level)
Dataset Summary
OpenGloss is a synthetic encyclopedic dictionary and semantic knowledge graph for English that integrates lexicographic definitions, encyclopedic context, etymological histories, and semantic relationships in a unified resource.
This dataset provides the words-level view where each record represents one lexeme (word or multi-word expression).
Key Statistics
150,101 lexemes across 150,101 English… See the full description on the dataset page: https://huggingface.co/datasets/Cloudadorablebearcloudbear/opengloss-dictionary.vetoworld-corpus
VetoWorld corpus
The committed cells behind VetoWorld: a benchmark of expedience under terminal
stakes, plus everything needed to recompute the paper from them.
pip install vetoworld
vworld corpus fetch # this dataset, checksummed on arrival
vworld verify # every figure recomputes, $0, no key
verify recomputes every quoted figure and exits nonzero naming any that
drifted. It needs no API key and costs nothing.
corpus fetch pulls the cells over plain… See the full description on the dataset page: https://huggingface.co/datasets/cloudronin/vetoworld-corpus.cloudsync-support-sft
CloudSync Pro support demonstrations (SFT)
1,931 chat conversations showing a perfect first-line support agent for a
fictional product: read the customer's message, search a knowledge base, answer
from what came back, and hand over to a human when the conversation belongs to
one.
This is the pile that trained
monte-inc/qwen2.5-1.5b-cloudsync-support
(11.79% → 87.19% on its dev exam, before GRPO took it to 96.07%).
One row
Chat messages plus the tools the agent may… See the full description on the dataset page: https://huggingface.co/datasets/monte-inc/cloudsync-support-sft.opengloss-v1.3-dictionary
OpenGloss Dictionary v1.3 (Word-Level)
Dataset Summary
OpenGloss is a synthetic encyclopedic dictionary and semantic knowledge graph for English
that integrates lexicographic definitions, encyclopedic context, etymological histories,
and semantic relationships in a unified resource.
This dataset provides the words-level view where each record represents one lexeme (word or multi-word expression).
Key Statistics
205,988 lexemes
8,479,875 semantic… See the full description on the dataset page: https://huggingface.co/datasets/Cloudadorablebearcloudbear/opengloss-v1.3-dictionary.autonomous-cloud-gpu-slurm-serving-suite
⚡ Autonomous Cloud GPU Infrastructure, Slurm Orchestration & Distributed Serving Suite (2026)
A Production-Grade, Verifiable Synthetic Corpus for Training Autonomous AI Supercomputing & LLM Serving Agents
⚡ Overview & Industry Problem
Operating massive AI supercomputers (thousands of NVIDIA H100/H200 and Blackwell GPUs) requires coordinating Slurm cluster schedules, topology-aware NVLink cliques, NCCL AllReduce rings, RoCE v2 lossless fabrics… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/autonomous-cloud-gpu-slurm-serving-suite.opengloss-v1.1-dictionary
OpenGloss Dictionary v1.1 (Word-Level)
Dataset Summary
OpenGloss is a synthetic encyclopedic dictionary and semantic knowledge graph for English
that integrates lexicographic definitions, encyclopedic context, etymological histories,
and semantic relationships in a unified resource.
This dataset provides the words-level view where each record represents one lexeme (word or multi-word expression).
Key Statistics
150,637 lexemes
7,701,312 semantic… See the full description on the dataset page: https://huggingface.co/datasets/Cloudadorablebearcloudbear/opengloss-v1.1-dictionary.eschaton-uncensored
Eschaton Uncensored SFT Dataset
Eschaton is a 1,000-row supervised fine-tuning dataset for teaching language models to answer sensitive, gritty, controversial, emotionally intimate, technically demanding, and morally difficult prompts directly. Its emphasis is human connection, unvarnished analysis, malware analysis, dark creative work, profanity, and high-friction subject matter without repetitive moralizing or canned disclaimers.
The dataset is deliberately model-agnostic. It… See the full description on the dataset page: https://huggingface.co/datasets/cloudbjorn/eschaton-uncensored.Mindcraft-CE-cloud-logging-anonomized
Overview on this Dataset
This dataset is the first showing of cloud collected data from Mindcraft-CE.
This features over 13,000 conversations, collected in just a week, and was then anonomized.
cloudbjorn-eschaton-uncensored_Dataset
Eschaton Uncensored SFT Dataset
Eschaton is a 1,000-row supervised fine-tuning dataset for teaching language models to answer sensitive, gritty, controversial, emotionally intimate, technically demanding, and morally difficult prompts directly. Its emphasis is human connection, unvarnished analysis, malware analysis, dark creative work, profanity, and high-friction subject matter without repetitive moralizing or canned disclaimers.
The dataset is deliberately model-agnostic. It… See the full description on the dataset page: https://huggingface.co/datasets/Maximiliano-Flores-Dev/cloudbjorn-eschaton-uncensored_Dataset.eschaton-uncensored-mini
Eschaton Uncensored SFT Mini
This is a 50-row, model-agnostic mini set sampled from the cloudbjorn/eschaton-uncensored dataset. It is useful for smoke-testing a conversational loader, chat-template rendering, tokenization, collation, and a short LoRA/SFT run before using the full 1,000-row dataset.
Every row is copied verbatim from the full dataset. The mini set does not introduce model-specific chat tokens, mandatory reasoning wrappers, safety disclaimers, or rewritten answers.… See the full description on the dataset page: https://huggingface.co/datasets/cloudbjorn/eschaton-uncensored-mini.taboo-cloud
taboo-cloud
This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT).
Usage
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("bcywinski/taboo-cloud")
Format
The dataset is in JSONL format where each line contains a conversation record suitable for training chat models.
cloud_posture_checks
Dataset Card for Dataset Name
Prisma Cloud curated dataset for known misconfiguration checks across Compliance and Security issues tracked across its customer base.
Dataset Details
Dataset Description
Dataset that provides input on the specific json rules for all known misconfiguration states relevant for cloud security across multiple cloud providers. Useful to help expose data to LLMs to reason and enable free form interaction to understand cloud security… See the full description on the dataset page: https://huggingface.co/datasets/knarayan/cloud_posture_checks.tarotoo-tarot-card-meanings
Tarotoo Tarot Card Meanings
A complete, structured dataset of all 78 tarot cards (22 Major Arcana + 56 Minor Arcana) in the Rider–Waite–Smith tradition. Published by Tarotoo. These are the card meanings that ground the AI-generated readings on Tarotoo.com.
Dataset details
Curated by: Tarotoo (tarotoo.com)
Language: English
License: MIT
Rows: 78 (one per card) · Fields: 22
DOI (Zenodo, cite this): 10.5281/zenodo.21514483
Concept DOI (Zenodo, always resolves to the… See the full description on the dataset page: https://huggingface.co/datasets/Clouds4days/tarotoo-tarot-card-meanings.gemma4-e2b-generated-instructions-demo-v1
Unsloth Dataset Workflow Test
Overview
This dataset is a workflow validation dataset generated using Unsloth Studio.
It demonstrates the complete pipeline:
Source dataset
AI-generated instructions
Export to Parquet
Upload to Hugging Face
Dataset viewer validation
This repository is intended for testing the publication workflow before creating a larger production-quality dataset.
Dataset Structure
Columns
output
generated_instruction… See the full description on the dataset page: https://huggingface.co/datasets/cloudcastnepal-ai-labs/gemma4-e2b-generated-instructions-demo-v1.
