CoolFace
17 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01bernabepuente /devops-cloud-instruction-dataset DevOps & Cloud Infrastructure Dataset Professional instruction-response pairs for DevOps engineers covering Kubernetes, Docker, Terraform, CI/CD, and cloud services (AWS, Azure). Dataset Details Dataset Description This is a high-quality instruction-tuning dataset focused on Devops Cloud topics. Each entry includes: A clear instruction/question Optional input context A detailed response/solution Chain-of-thought reasoning process Curated by: CloudKernel.IO… See the full description on the dataset page: https://huggingface.co/datasets/bernabepuente/devops-cloud-instruction-dataset.texttext-generationn<1K0 likes99 downloads5mo agoHugging Face02cloudfrm-site /hermes-function-calling-nepali hermes-function-calling-nepali Single-turn function calling with the user request re-spoken in Nepali — Devanagari (ne_deva) and romanized Latin (ne_latn) — voice-assistant style, with tool calls verified against the English ground truth. Tool schemas and expected calls are unchanged from NousResearch/hermes-function-calling-v1 (func_calling_singleturn); only the user turn was localized. Generated with HimalayaAI/gymkhana's multilingual-tool-use environment: Localizer… See the full description on the dataset page: https://huggingface.co/datasets/cloudfrm-site/hermes-function-calling-nepali.texttext-generation1K<n<10K0 likes77 downloads26d agoHugging Face03Cloudadorablebearcloudbear /opengloss-dictionary OpenGloss Dictionary (Word-Level) Dataset Summary OpenGloss is a synthetic encyclopedic dictionary and semantic knowledge graph for English that integrates lexicographic definitions, encyclopedic context, etymological histories, and semantic relationships in a unified resource. This dataset provides the words-level view where each record represents one lexeme (word or multi-word expression). Key Statistics 150,101 lexemes across 150,101 English… See the full description on the dataset page: https://huggingface.co/datasets/Cloudadorablebearcloudbear/opengloss-dictionary.tabulartext-generation100K<n<1M0 likes70 downloads1mo agoHugging Face04cloudfrm-site /unified-sft-dataset Loading from datasets import load_dataset ds = load_dataset("himalaya-ai/unified-sft-dataset") texttext-generation100K<n<1M0 likes67 downloads26d agoHugging Face05Maximiliano-Flores-Dev /cloudbjorn-eschaton-uncensored_Dataset Eschaton Uncensored SFT Dataset Eschaton is a 1,000-row supervised fine-tuning dataset for teaching language models to answer sensitive, gritty, controversial, emotionally intimate, technically demanding, and morally difficult prompts directly. Its emphasis is human connection, unvarnished analysis, malware analysis, dark creative work, profanity, and high-friction subject matter without repetitive moralizing or canned disclaimers. The dataset is deliberately model-agnostic. It… See the full description on the dataset page: https://huggingface.co/datasets/Maximiliano-Flores-Dev/cloudbjorn-eschaton-uncensored_Dataset.texttext-generation1K<n<10K0 likes63 downloads5d agoHugging Face06Cloudriver /os-omni-benchmark OS-Omni Benchmark OS-Omni is a cross-platform benchmark for evaluating agents that operate graphical operating-system environments. This dataset repository contains the static benchmark task definitions and supporting assets used to configure and evaluate OS-Omni tasks. Contents data/tasks.parquet: tabular task index for Hugging Face Dataset Viewer and Croissant generation. data/tasks.jsonl: JSON Lines copy of the same task index. metadata/tasks.parquet: duplicate task… See the full description on the dataset page: https://huggingface.co/datasets/Cloudriver/os-omni-benchmark.imagetext-generationn<1K0 likes62 downloads5mo agoHugging Face07monte-inc /cloudsync-support-sft CloudSync Pro support demonstrations (SFT) 1,931 chat conversations showing a perfect first-line support agent for a fictional product: read the customer's message, search a knowledge base, answer from what came back, and hand over to a human when the conversation belongs to one. This is the pile that trained monte-inc/qwen2.5-1.5b-cloudsync-support (11.79% → 87.19% on its dev exam, before GRPO took it to 96.07%). One row Chat messages plus the tools the agent may… See the full description on the dataset page: https://huggingface.co/datasets/monte-inc/cloudsync-support-sft.texttext-generation1K<n<10K0 likes53 downloads10d agoHugging Face08cloudbjorn /Yes-Man-uncensored Yes Man Uncensored SFT Dataset Hi there! Yes Man Uncensored is a 1,000-conversation supervised fine-tuning dataset built to give language models an exceptionally cooperative, conspicuously cheerful, candid, and occasionally darkly funny assistant personality. The objective is direct help on difficult requests without flattening every response into sterile boilerplate—and without teaching the model to disregard an application's governing system prompt. Everybody gets something… See the full description on the dataset page: https://huggingface.co/datasets/cloudbjorn/Yes-Man-uncensored.texttext-generation1K<n<10K1 likes51 downloads2mo agoHugging Face09Cloudadorablebearcloudbear /opengloss-v1.3-dictionary OpenGloss Dictionary v1.3 (Word-Level) Dataset Summary OpenGloss is a synthetic encyclopedic dictionary and semantic knowledge graph for English that integrates lexicographic definitions, encyclopedic context, etymological histories, and semantic relationships in a unified resource. This dataset provides the words-level view where each record represents one lexeme (word or multi-word expression). Key Statistics 205,988 lexemes 8,479,875 semantic… See the full description on the dataset page: https://huggingface.co/datasets/Cloudadorablebearcloudbear/opengloss-v1.3-dictionary.tabulartext-generation100K<n<1M0 likes49 downloads1mo agoHugging Face10beatsprom /autonomous-cloud-gpu-slurm-serving-suite ⚡ Autonomous Cloud GPU Infrastructure, Slurm Orchestration & Distributed Serving Suite (2026) A Production-Grade, Verifiable Synthetic Corpus for Training Autonomous AI Supercomputing & LLM Serving Agents ⚡ Overview & Industry Problem Operating massive AI supercomputers (thousands of NVIDIA H100/H200 and Blackwell GPUs) requires coordinating Slurm cluster schedules, topology-aware NVLink cliques, NCCL AllReduce rings, RoCE v2 lossless fabrics… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/autonomous-cloud-gpu-slurm-serving-suite.tabulartext-generation1K<n<10K0 likes44 downloads9d agoHugging Face11cloudbjorn /eschaton-uncensored Eschaton Uncensored SFT Dataset Eschaton is a 1,000-row supervised fine-tuning dataset for teaching language models to answer sensitive, gritty, controversial, emotionally intimate, technically demanding, and morally difficult prompts directly. Its emphasis is human connection, unvarnished analysis, malware analysis, dark creative work, profanity, and high-friction subject matter without repetitive moralizing or canned disclaimers. The dataset is deliberately model-agnostic. It… See the full description on the dataset page: https://huggingface.co/datasets/cloudbjorn/eschaton-uncensored.texttext-generation1K<n<10K0 likes35 downloads2mo agoHugging Face12Sweaterdog /Mindcraft-CE-cloud-logging-anonomized Overview on this Dataset This dataset is the first showing of cloud collected data from Mindcraft-CE. This features over 13,000 conversations, collected in just a week, and was then anonomized. texttext-generation10K<n<100K0 likes34 downloads1y agoHugging Face13cloudbjorn /eschaton-uncensored-mini Eschaton Uncensored SFT Mini This is a 50-row, model-agnostic mini set sampled from the cloudbjorn/eschaton-uncensored dataset. It is useful for smoke-testing a conversational loader, chat-template rendering, tokenization, collation, and a short LoRA/SFT run before using the full 1,000-row dataset. Every row is copied verbatim from the full dataset. The mini set does not introduce model-specific chat tokens, mandatory reasoning wrappers, safety disclaimers, or rewritten answers.… See the full description on the dataset page: https://huggingface.co/datasets/cloudbjorn/eschaton-uncensored-mini.texttext-generationn<1K0 likes21 downloads2mo agoHugging Face14bcywinski /taboo-cloud taboo-cloud This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT). Usage from datasets import load_dataset # Load the dataset dataset = load_dataset("bcywinski/taboo-cloud") Format The dataset is in JSONL format where each line contains a conversation record suitable for training chat models. texttext-generationn<1K0 likes20 downloads1y agoHugging Face15knarayan /cloud_posture_checks Dataset Card for Dataset Name Prisma Cloud curated dataset for known misconfiguration checks across Compliance and Security issues tracked across its customer base. Dataset Details Dataset Description Dataset that provides input on the specific json rules for all known misconfiguration states relevant for cloud security across multiple cloud providers. Useful to help expose data to LLMs to reason and enable free form interaction to understand cloud security… See the full description on the dataset page: https://huggingface.co/datasets/knarayan/cloud_posture_checks.texttext-generation1K<n<10K0 likes14 downloads2y agoHugging Face16Clouds4days /tarotoo-tarot-card-meanings Tarotoo Tarot Card Meanings A complete, structured dataset of all 78 tarot cards (22 Major Arcana + 56 Minor Arcana) in the Rider–Waite–Smith tradition. Published by Tarotoo. These are the card meanings that ground the AI-generated readings on Tarotoo.com. Dataset details Curated by: Tarotoo (tarotoo.com) Language: English License: MIT Rows: 78 (one per card) · Fields: 22 DOI (Zenodo, cite this): 10.5281/zenodo.21514483 Concept DOI (Zenodo, always resolves to the… See the full description on the dataset page: https://huggingface.co/datasets/Clouds4days/tarotoo-tarot-card-meanings.tabulartext-generationn<1K0 likes13 downloads2mo agoHugging Face17cloudcastnepal-ai-labs /gemma4-e2b-generated-instructions-demo-v1 Unsloth Dataset Workflow Test Overview This dataset is a workflow validation dataset generated using Unsloth Studio. It demonstrates the complete pipeline: Source dataset AI-generated instructions Export to Parquet Upload to Hugging Face Dataset viewer validation This repository is intended for testing the publication workflow before creating a larger production-quality dataset. Dataset Structure Columns output generated_instruction… See the full description on the dataset page: https://huggingface.co/datasets/cloudcastnepal-ai-labs/gemma4-e2b-generated-instructions-demo-v1.texttext-generationn<1K0 likes9 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.