CoolFace
19 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Arko007 /zenyx-v2-SFT-dataset Zenyx V2 — Raw SFT Dataset Collection This is the unified raw dataset collection used for training Zenyx V2, a custom large language model built from scratch with a novel architecture. Dataset Sources Dataset Rows Category nemotron_sft_code 10,108,883 Code nemotron_sft_math 22,066,397 Math nemotron_sft_science 708,920 Science nemotron_sft_chat 39,792 Chat nemotron_sft_safety 31,426 Safety nemotron_rl 56,339 Instruction Following (RL)… See the full description on the dataset page: https://huggingface.co/datasets/Arko007/zenyx-v2-SFT-dataset.texttext-generation10M<n<100M2 likes2.4k downloads6mo agoHugging Face02zenml /llmops-database The ZenML LLMOps Database To learn more about ZenML and our open-source MLOps framework, visit zenml.io. Dataset Summary The LLMOps Database is a comprehensive collection of over 500 real-world generative AI implementations that showcases how organizations are successfully deploying Large Language Models (LLMs) in production. The case studies have been carefully curated to focus on technical depth and practical problem-solving, with an emphasis on implementation… See the full description on the dataset page: https://huggingface.co/datasets/zenml/llmops-database.textfeature-extraction1K<n<10K23 likes864 downloads7d agoHugging Face03ZennyKenny /tactical-military-reasoning-v.1.0 Tactical Military Reasoning Dataset v1.0 A curated collection of 150 rich tactical military scenarios with LLM-generated reasoning strategies for both attacking and defending forces. 📝 Preface Oncologists do not study cancer because they love cancer and wish for it to occur more frequently. They study cancer to better understand its causes, progression, and consequences in order to therefore eradicate it from the earth more effectively. A distaste for something… See the full description on the dataset page: https://huggingface.co/datasets/ZennyKenny/tactical-military-reasoning-v.1.0.texttext-generationn<1K24 likes197 downloads1y agoHugging Face04ZennyKenny /synthetic_vc_financial_decisions_reasoning_dataset Best Curator Use Case in the Reasoning Datasets Competition: https://www.linkedin.com/feed/update/urn:li:activity:7330998995990781952/ Synthetic VC Financial Decisions Reasoning Dataset Dataset Summary The Synthetic VC Financial Decisions Reasoning Dataset is a large-scale collection designed to train, evaluate, and fine-tune language models on subjective, abstract financial reasoning tasks. It simulates venture capital (VC) workflows by capturing multiple… See the full description on the dataset page: https://huggingface.co/datasets/ZennyKenny/synthetic_vc_financial_decisions_reasoning_dataset.textreinforcement-learningn<1K15 likes151 downloads1y agoHugging Face05zenlm /zen-identity Zen Identity Dataset This dataset contains identity training data for the Zen family of AI models. Models Covered Zen Nano (0.6B): Ultra-efficient edge computing model Zen Eco (3B): Balanced performance and efficiency Zen Coder (7B): Specialized for code generation Zen Omni (14B): Versatile multi-domain model Dataset Structure Each example contains: instruction: The user's question output: The model's response model: Which Zen model this example… See the full description on the dataset page: https://huggingface.co/datasets/zenlm/zen-identity.texttext-generationn<1K0 likes75 downloads3mo agoHugging Face06ZennyKenny /cosa-benchmark-dataset 🧠 CoSa Benchmark Dataset 🔍 Introduction The CoSa (Code Safety) Benchmark is a curated evaluation dataset designed to measure the ability of large language models (LLMs) to detect, explain, and repair vulnerabilities in synthetic code samples. It is intended to benchmark LLMs for real-world application in code security audits, reasoning tasks, and secure code generation. 📦 Contents Each row in the dataset includes: code: a code snippet (varied… See the full description on the dataset page: https://huggingface.co/datasets/ZennyKenny/cosa-benchmark-dataset.textquestion-answeringn<1K8 likes46 downloads1y agoHugging Face07zenzen9 /regex-rl-dataset Regex RL Training Dataset Synthetic regex dataset for reinforcement learning post-training. Dataset Details Size: 1,158 examples Format: JSONL Use Case: GRPO/RL training for regex generation Data Format { "prompt": "Write a Python regex pattern that matches: <description>", "solution": "<regex_pattern>", "test_cases": { "positive": ["match1", "match2", "match3", "match4", "match5"], "negative": ["no_match1", "no_match2", "no_match3", "no_match4"… See the full description on the dataset page: https://huggingface.co/datasets/zenzen9/regex-rl-dataset.texttext-generation1K<n<10K0 likes35 downloads8mo agoHugging Face08zengsdfew /reddit_dataset_34 Bittensor Subnet 13 Reddit Dataset Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed Reddit data. The data is continuously updated by network miners, providing a real-time stream of Reddit content for various analytical and machine learning tasks. For more information about the dataset, please visit the official repository. Supported Tasks The versatility of this dataset allows… See the full description on the dataset page: https://huggingface.co/datasets/zengsdfew/reddit_dataset_34.texttext-classification100K<n<1M0 likes31 downloads2y agoHugging Face09zengsdfew /reddit_dataset_44 Bittensor Subnet 13 Reddit Dataset Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed Reddit data. The data is continuously updated by network miners, providing a real-time stream of Reddit content for various analytical and machine learning tasks. For more information about the dataset, please visit the official repository. Supported Tasks The versatility of this dataset allows… See the full description on the dataset page: https://huggingface.co/datasets/zengsdfew/reddit_dataset_44.texttext-classification100K<n<1M0 likes22 downloads2y agoHugging Face10alvemoans /zenith_ai_305 Legal Data Analysis Dataset This dataset contains legal statements, analyses, and judgments primarily related to labor law and contract law, drawn from various cases and legal interpretations. It includes text entries with factual descriptions, legal arguments, and conclusions based on judicial decisions, as well as instructions related to interpreting those facts. Dataset Overview The dataset is structured as a series of legal paragraphs and corresponding instructions… See the full description on the dataset page: https://huggingface.co/datasets/alvemoans/zenith_ai_305.texttext-generationn<1K0 likes18 downloads2y agoHugging Face11zengsdfew /x_dataset_44 Bittensor Subnet 13 X (Twitter) Dataset Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks. For more information about the dataset, please visit the official repository. Supported Tasks The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/zengsdfew/x_dataset_44.texttext-classification100K<n<1M0 likes15 downloads2y agoHugging Face12ZennyKenny /yandexgptpro_4th_gen-hellaswag YandexGPT Pro (4th Gen) HellaSwag This dataset contains responses from the YandexGPT model evaluated on the HellaSwag benchmark. It was generated as part of an experiment to assess the model’s performance on multiple-choice commonsense reasoning tasks. Dataset Details Source: HellaSwag Model: YandexGPT via Yandex Cloud Foundation Models API Prompt style: Multiple-choice (A, B, C, D) with system prompt and task context Fields: id: index of the example context: the base… See the full description on the dataset page: https://huggingface.co/datasets/ZennyKenny/yandexgptpro_4th_gen-hellaswag.textzero-shot-classification10K<n<100K0 likes15 downloads1y agoHugging Face13zengsdfew /x_dataset_34 Bittensor Subnet 13 X (Twitter) Dataset Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks. For more information about the dataset, please visit the official repository. Supported Tasks The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/zengsdfew/x_dataset_34.texttext-classification100K<n<1M0 likes13 downloads2y agoHugging Face14Zen1t /texts-for-articlestexttext-generationn<1K0 likes12 downloads3y agoHugging Face15ZenithVortex /my-distiset-17b5b2b4 Dataset Card for my-distiset-17b5b2b4 This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/ZenithVortex/my-distiset-17b5b2b4/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/ZenithVortex/my-distiset-17b5b2b4.texttext-generationn<1K0 likes12 downloads2y agoHugging Face16zengsdfew /reddit_dataset_112 Bittensor Subnet 13 Reddit Dataset Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed Reddit data. The data is continuously updated by network miners, providing a real-time stream of Reddit content for various analytical and machine learning tasks. For more information about the dataset, please visit the official repository. Supported Tasks The versatility of this dataset allows… See the full description on the dataset page: https://huggingface.co/datasets/zengsdfew/reddit_dataset_112.texttext-classification1M<n<10M0 likes11 downloads2y agoHugging Face17Zenng2812 /bctc-md-domain-corpus Vietnamese Financial Reports Markdown Domain Corpus Dataset này được tạo từ các báo cáo tài chính dạng Markdown trong thư mục BCTC_MD. Mục đích Dataset dùng cho continued pretraining / domain-adaptive pretraining mô hình ngôn ngữ trên miền báo cáo tài chính tiếng Việt. Cấu trúc dữ liệu Mỗi dòng trong train.jsonl hoặc validation.jsonl là một JSON object: { "text": "...", "source_file": "AAA_BCTC_2020.md", "document_id": "AAA_BCTC_2020", "company": "AAA"… See the full description on the dataset page: https://huggingface.co/datasets/Zenng2812/bctc-md-domain-corpus.imagetext-generation1K<n<10K0 likes9 downloads5mo agoHugging Face18ZenithVortex /romeo_and_juliet Dataset Card for romeo_and_juliet This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/ZenithVortex/romeo_and_juliet/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/ZenithVortex/romeo_and_juliet.texttext-generationn<1K0 likes8 downloads2y agoHugging Face19ZenithVortex /my-distiset-7af8e9d9 Dataset Card for my-distiset-7af8e9d9 This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/ZenithVortex/my-distiset-7af8e9d9/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/ZenithVortex/my-distiset-7af8e9d9.texttext-generationn<1K0 likes4 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.