datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GSM8k-AugThis dataset is provided to facilitate access to GSM8k-Aug, originally from https://github.com/da03/Internalize_CoT_Step_by_Step and https://arxiv.org/pdf/2405.14838.
This dataset is used to train CODI (https://arxiv.org/abs/2502.21074)
Description:
We utilize two datasets to train our models--GSM8k-Aug and GSM8k-Aug-NL. (1) We use the GSM8k-Aug dataset, which has proven effective for training implicit CoT methods. This dataset extends the original GSM8k training set to 385k samples by… See the full description on the dataset page: https://huggingface.co/datasets/zen-E/GSM8k-Aug.VideoArgusBench
VideoArgusBench
Sample-specific rubric benchmark for conditioned video generation.
VideoArgusBench is the evaluation benchmark for VideoArgus, a framework that scores a generated
video against a rubric written for that specific prompt rather than a fixed global metric. This
dataset ships the inputs (conditioning assets + prompts) and, for each input, a rubric. It does
not contain generated videos — you bring your own model's outputs and score them with the VideoArgus
evaluation… See the full description on the dataset page: https://huggingface.co/datasets/zengziyun/VideoArgusBench.zenz-v2.5-dataset
zenz-v2.5-dataset
zenz-v2.5-datasetはかな漢字変換タスクに特化した条件付き言語モデル「zenz-v2.5」シリーズの学習を目的として構築したデータセットです。
約190Mペアの「左文脈-入力-変換結果」を含み、かな漢字変換モデルの学習において十分な性能を実現できる規模になっています。
本データセットで学習したzenz-v2.5は公開しています。
zenz-v2.5-medium: 310Mの大規模モデル
zenz-v2.5-small: 91Mの中規模モデル
zenz-v2.5-xsmall: 26Mの小規模モデル
また、かな漢字変換の評価ベンチマークとしてAJIMEE-Bench(味見ベンチ)も公開しています。
形式
本データセットはJSONL形式になっており、以下の3つのデータを含みます。
"input": str, 入力のカタカナ文字列(記号、数字、空白などが含まれることがあります)
"output": str, 出力の漢字交じり文
"left_context":… See the full description on the dataset page: https://huggingface.co/datasets/Miwa-Keita/zenz-v2.5-dataset.AusCyberBench
AusCyberBench v2.1
The first comprehensive benchmark for evaluating Large Language Models in Australian cybersecurity contexts. Covers regulatory compliance (Essential Eight, ISM, Privacy Act, SOCI Act), technical security, threat intelligence, and Australian-specific terminology.
Model Leaderboard (v2.1, Australian Test Set)
#
Model
Overall
95% CI
E8
Threat Intel
SOCI
Privacy
Terminology
1
GPT-5.2
85.5%
[83.1%, 87.8%]
83.1%
78.7%
100.0%
93.8%
100.0%… See the full description on the dataset page: https://huggingface.co/datasets/Zen0/AusCyberBench.GSM8k-Aug-NLThis dataset is provided to facilitate access to GSM8k-Aug-NL, originally from https://github.com/da03/implicit_chain_of_thought and https://arxiv.org/abs/2311.01460.
This dataset is used to train CODI (https://arxiv.org/abs/2502.21074)
Description:
We utilize two datasets to train our models--GSM8k-Aug and GSM8k-Aug-NL. (1) We use the GSM8k-Aug dataset, which has proven effective for training implicit CoT methods. This dataset extends the original GSM8k training set to 385k samples by… See the full description on the dataset page: https://huggingface.co/datasets/zen-E/GSM8k-Aug-NL.zeno_h1_2026_07_31_curated11_raw_bags
Zeno H1 — 2026-07-31 curated raw ROS 2 bags
This public dataset contains 11 curated raw ROS 2 MCAP recordings from 2026-07-31.
The original rosbag2_.../ directory structure is retained, so each MCAP can be opened directly after download.
Excluded recordings: #1 (refrigerator did not open), #5 (table-wiping ending), #7 (ended before placing draining basket), and #10 (table wiping only half complete).
Download on another computer
hf download… See the full description on the dataset page: https://huggingface.co/datasets/QRP123/zeno_h1_2026_07_31_curated11_raw_bags.StrategyQA_CoT_GPT4ozenith-modelsCommonsenseQA-GPT4ominiStrategyQA_GPT4o_CoTx10research-companion-indexvietnamese-financial-summary
Vietnamese Financial News Summarization with Number Preservation
LenCtrl-Bench
LenCtrl-Bench: Benchmarking LLMs' Abilities for Length-Controlled Text Generation.
arxiv: https://arxiv.org/abs/2410.07035
daily papers: https://huggingface.co/papers/2410.07035
twitter: https://x.com/ZenMoore1/status/1845673846193668546
Method
Usage
This dataset contains the following fields:
instruction and response.
constraint: the length constraint.
level: the level of the length constraint, choices=["word", "sentence", "paragraph"].
data_source: the… See the full description on the dataset page: https://huggingface.co/datasets/ZenMoore/LenCtrl-Bench.WixQA
WixQA: Enterprise RAG Question-Answering Benchmark
📄 Full Paper Available: For comprehensive details on dataset design, methodology, evaluation results, and analysis, please see our complete research paper:
WixQA: A Multi-Dataset Benchmark for Enterprise Retrieval-Augmented Generation
Cohen et al. (2025) - arXiv:2505.08643
Dataset Summary
WixQA is a three-config collection for evaluating and training Retrieval-Augmented Generation (RAG) systems in enterprise… See the full description on the dataset page: https://huggingface.co/datasets/ZenitsuBorade/WixQA.exp_q_2_sqlQwen3-speculative-pair-report
Qwen3 speculative pair report: 0.6B draft + 8B target, measured acceptance
Research evidence dataset. No model weights. Part of the collection
Xyntetik Research: Runner Compatibility Reports on this account, produced with
Xyntetik Runner.
Dataset summary
Question tested. What the acceptance rate of a Qwen3-0.6B draft against a Qwen3-8B target actually is across draft depths, whether the engine's printed tokens-per-round figure can be tuned on (it cannot), and… See the full description on the dataset page: https://huggingface.co/datasets/Joakimpalm-Zen/Qwen3-speculative-pair-report.CP-Bench
CP-Bench: Benchmarking LLMs' Abilities for Copy-Pasting Tool-Use.
arxiv: https://arxiv.org/abs/2410.07035
daily papers: https://huggingface.co/papers/2410.07035
twitter: https://x.com/ZenMoore1/status/1845673846193668546
Method
Usage
This dataset contains the following fields:
instruction and response_pure_text are regular inputs and outputs without position ids or copy-pasting.
type: choices=["single-copy", "multi-copy"], indicating the number of copies in… See the full description on the dataset page: https://huggingface.co/datasets/ZenMoore/CP-Bench.regex-rl-dataset
Regex RL Training Dataset
Synthetic regex dataset for reinforcement learning post-training.
Dataset Details
Size: 1,158 examples
Format: JSONL
Use Case: GRPO/RL training for regex generation
Data Format
{
"prompt": "Write a Python regex pattern that matches: <description>",
"solution": "<regex_pattern>",
"test_cases": {
"positive": ["match1", "match2", "match3", "match4", "match5"],
"negative": ["no_match1", "no_match2", "no_match3", "no_match4"… See the full description on the dataset page: https://huggingface.co/datasets/zenzen9/regex-rl-dataset.cantonese-chinese-parallel-corpus-baseThis is a dataset of Cantonese-Written Chinese Parallel Corpus, containing 130k+ pairs of Cantonese and Traditional Chinese parallel sentences.
zenodo-scraper
Zenodo Scraper · Research Records, DOIs, Authors & Files
Scrape open research records, DOIs, publications, datasets, software, authors, and file metadata from Zenodo. Fast HTTP scraper charging per returned record with tiered pricing.
Rows in this dataset
1,965
Fields
27
Collector runs behind it
50
Most recent observation
2026-08-03
What this is
Every row here was returned by a real run of a public collector. Nothing is generated from a… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/zenodo-scraper.qwen36-eagle3-stagebarca-synthetic-adaptive
ARCA Synthetic Adaptive Graph-Path Constraints
This dataset is a controlled synthetic benchmark for testing whether Adaptive Residual Constrained Attention (ARCA) reacts to external constraint quality.
It does not use LLMs, pretrained language models, or natural-language generation. Every sample is generated by Python with exact ground truth.
Task
Each base sample contains a random directed graph. Nodes have discrete values. A query gives:
a start node
a relation… See the full description on the dataset page: https://huggingface.co/datasets/Zenng2812/arca-synthetic-adaptive.zen-ingress-datasetzenith_ai_305
Legal Data Analysis Dataset
This dataset contains legal statements, analyses, and judgments primarily related to labor law and contract law, drawn from various cases and legal interpretations. It includes text entries with factual descriptions, legal arguments, and conclusions based on judicial decisions, as well as instructions related to interpreting those facts.
Dataset Overview
The dataset is structured as a series of legal paragraphs and corresponding instructions… See the full description on the dataset page: https://huggingface.co/datasets/alvemoans/zenith_ai_305.financial-judgement-vietnameseDatasetbctc-md-domain-corpus
Vietnamese Financial Reports Markdown Domain Corpus
Dataset này được tạo từ các báo cáo tài chính dạng Markdown trong thư mục BCTC_MD.
Mục đích
Dataset dùng cho continued pretraining / domain-adaptive pretraining mô hình ngôn ngữ trên miền báo cáo tài chính tiếng Việt.
Cấu trúc dữ liệu
Mỗi dòng trong train.jsonl hoặc validation.jsonl là một JSON object:
{
"text": "...",
"source_file": "AAA_BCTC_2020.md",
"document_id": "AAA_BCTC_2020",
"company": "AAA"… See the full description on the dataset page: https://huggingface.co/datasets/Zenng2812/bctc-md-domain-corpus.ADL_HW1_Datasmahabharat-qnanlpdataset
