datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
zenyx-v2-SFT-dataset
Zenyx V2 — Raw SFT Dataset Collection
This is the unified raw dataset collection used for training Zenyx V2,
a custom large language model built from scratch with a novel architecture.
Dataset Sources
Dataset
Rows
Category
nemotron_sft_code
10,108,883
Code
nemotron_sft_math
22,066,397
Math
nemotron_sft_science
708,920
Science
nemotron_sft_chat
39,792
Chat
nemotron_sft_safety
31,426
Safety
nemotron_rl
56,339
Instruction Following (RL)… See the full description on the dataset page: https://huggingface.co/datasets/Arko007/zenyx-v2-SFT-dataset.llmops-database
The ZenML LLMOps Database
To learn more about ZenML and our open-source MLOps framework, visit
zenml.io.
Dataset Summary
The LLMOps Database is a comprehensive collection of over 500 real-world
generative AI implementations that showcases how organizations are successfully
deploying Large Language Models (LLMs) in production. The case studies have been
carefully curated to focus on technical depth and practical problem-solving,
with an emphasis on implementation… See the full description on the dataset page: https://huggingface.co/datasets/zenml/llmops-database.tactical-military-reasoning-v.1.0
Tactical Military Reasoning Dataset v1.0
A curated collection of 150 rich tactical military scenarios with LLM-generated reasoning strategies for both attacking and defending forces.
📝 Preface
Oncologists do not study cancer because they love cancer and wish for it to occur more frequently. They study cancer to better understand its causes, progression, and consequences in order to therefore eradicate it from the earth more effectively. A distaste for something… See the full description on the dataset page: https://huggingface.co/datasets/ZennyKenny/tactical-military-reasoning-v.1.0.synthetic_vc_financial_decisions_reasoning_dataset
Best Curator Use Case in the Reasoning Datasets Competition: https://www.linkedin.com/feed/update/urn:li:activity:7330998995990781952/
Synthetic VC Financial Decisions Reasoning Dataset
Dataset Summary
The Synthetic VC Financial Decisions Reasoning Dataset is a large-scale collection designed to train, evaluate, and fine-tune language models on subjective, abstract financial reasoning tasks. It simulates venture capital (VC) workflows by capturing multiple… See the full description on the dataset page: https://huggingface.co/datasets/ZennyKenny/synthetic_vc_financial_decisions_reasoning_dataset.zen-identity
Zen Identity Dataset
This dataset contains identity training data for the Zen family of AI models.
Models Covered
Zen Nano (0.6B): Ultra-efficient edge computing model
Zen Eco (3B): Balanced performance and efficiency
Zen Coder (7B): Specialized for code generation
Zen Omni (14B): Versatile multi-domain model
Dataset Structure
Each example contains:
instruction: The user's question
output: The model's response
model: Which Zen model this example… See the full description on the dataset page: https://huggingface.co/datasets/zenlm/zen-identity.cosa-benchmark-dataset
🧠 CoSa Benchmark Dataset
🔍 Introduction
The CoSa (Code Safety) Benchmark is a curated evaluation dataset designed to measure the ability of large language models (LLMs) to detect, explain, and repair vulnerabilities in synthetic code samples. It is intended to benchmark LLMs for real-world application in code security audits, reasoning tasks, and secure code generation.
📦 Contents
Each row in the dataset includes:
code: a code snippet (varied… See the full description on the dataset page: https://huggingface.co/datasets/ZennyKenny/cosa-benchmark-dataset.regex-rl-dataset
Regex RL Training Dataset
Synthetic regex dataset for reinforcement learning post-training.
Dataset Details
Size: 1,158 examples
Format: JSONL
Use Case: GRPO/RL training for regex generation
Data Format
{
"prompt": "Write a Python regex pattern that matches: <description>",
"solution": "<regex_pattern>",
"test_cases": {
"positive": ["match1", "match2", "match3", "match4", "match5"],
"negative": ["no_match1", "no_match2", "no_match3", "no_match4"… See the full description on the dataset page: https://huggingface.co/datasets/zenzen9/regex-rl-dataset.reddit_dataset_34
Bittensor Subnet 13 Reddit Dataset
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed Reddit data. The data is continuously updated by network miners, providing a real-time stream of Reddit content for various analytical and machine learning tasks.
For more information about the dataset, please visit the official repository.
Supported Tasks
The versatility of this dataset allows… See the full description on the dataset page: https://huggingface.co/datasets/zengsdfew/reddit_dataset_34.reddit_dataset_44
Bittensor Subnet 13 Reddit Dataset
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed Reddit data. The data is continuously updated by network miners, providing a real-time stream of Reddit content for various analytical and machine learning tasks.
For more information about the dataset, please visit the official repository.
Supported Tasks
The versatility of this dataset allows… See the full description on the dataset page: https://huggingface.co/datasets/zengsdfew/reddit_dataset_44.zenith_ai_305
Legal Data Analysis Dataset
This dataset contains legal statements, analyses, and judgments primarily related to labor law and contract law, drawn from various cases and legal interpretations. It includes text entries with factual descriptions, legal arguments, and conclusions based on judicial decisions, as well as instructions related to interpreting those facts.
Dataset Overview
The dataset is structured as a series of legal paragraphs and corresponding instructions… See the full description on the dataset page: https://huggingface.co/datasets/alvemoans/zenith_ai_305.x_dataset_44
Bittensor Subnet 13 X (Twitter) Dataset
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks.
For more information about the dataset, please visit the official repository.
Supported Tasks
The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/zengsdfew/x_dataset_44.yandexgptpro_4th_gen-hellaswag
YandexGPT Pro (4th Gen) HellaSwag
This dataset contains responses from the YandexGPT model evaluated on the HellaSwag benchmark. It was generated as part of an experiment to assess the model’s performance on multiple-choice commonsense reasoning tasks.
Dataset Details
Source: HellaSwag
Model: YandexGPT via Yandex Cloud Foundation Models API
Prompt style: Multiple-choice (A, B, C, D) with system prompt and task context
Fields:
id: index of the example
context: the base… See the full description on the dataset page: https://huggingface.co/datasets/ZennyKenny/yandexgptpro_4th_gen-hellaswag.x_dataset_34
Bittensor Subnet 13 X (Twitter) Dataset
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks.
For more information about the dataset, please visit the official repository.
Supported Tasks
The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/zengsdfew/x_dataset_34.texts-for-articlesmy-distiset-17b5b2b4
Dataset Card for my-distiset-17b5b2b4
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/ZenithVortex/my-distiset-17b5b2b4/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/ZenithVortex/my-distiset-17b5b2b4.reddit_dataset_112
Bittensor Subnet 13 Reddit Dataset
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed Reddit data. The data is continuously updated by network miners, providing a real-time stream of Reddit content for various analytical and machine learning tasks.
For more information about the dataset, please visit the official repository.
Supported Tasks
The versatility of this dataset allows… See the full description on the dataset page: https://huggingface.co/datasets/zengsdfew/reddit_dataset_112.bctc-md-domain-corpus
Vietnamese Financial Reports Markdown Domain Corpus
Dataset này được tạo từ các báo cáo tài chính dạng Markdown trong thư mục BCTC_MD.
Mục đích
Dataset dùng cho continued pretraining / domain-adaptive pretraining mô hình ngôn ngữ trên miền báo cáo tài chính tiếng Việt.
Cấu trúc dữ liệu
Mỗi dòng trong train.jsonl hoặc validation.jsonl là một JSON object:
{
"text": "...",
"source_file": "AAA_BCTC_2020.md",
"document_id": "AAA_BCTC_2020",
"company": "AAA"… See the full description on the dataset page: https://huggingface.co/datasets/Zenng2812/bctc-md-domain-corpus.romeo_and_juliet
Dataset Card for romeo_and_juliet
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/ZenithVortex/romeo_and_juliet/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/ZenithVortex/romeo_and_juliet.my-distiset-7af8e9d9
Dataset Card for my-distiset-7af8e9d9
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/ZenithVortex/my-distiset-7af8e9d9/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/ZenithVortex/my-distiset-7af8e9d9.
