CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01zen-E /GSM8k-AugThis dataset is provided to facilitate access to GSM8k-Aug, originally from https://github.com/da03/Internalize_CoT_Step_by_Step and https://arxiv.org/pdf/2405.14838. This dataset is used to train CODI (https://arxiv.org/abs/2502.21074) Description: We utilize two datasets to train our models--GSM8k-Aug and GSM8k-Aug-NL. (1) We use the GSM8k-Aug dataset, which has proven effective for training implicit CoT methods. This dataset extends the original GSM8k training set to 385k samples by… See the full description on the dataset page: https://huggingface.co/datasets/zen-E/GSM8k-Aug.textquestion-answering100K<n<1M5 likes2.6k downloads1y agoHugging Face02zengziyun /VideoArgusBench VideoArgusBench Sample-specific rubric benchmark for conditioned video generation. VideoArgusBench is the evaluation benchmark for VideoArgus, a framework that scores a generated video against a rubric written for that specific prompt rather than a fixed global metric. This dataset ships the inputs (conditioning assets + prompts) and, for each input, a rubric. It does not contain generated videos — you bring your own model's outputs and score them with the VideoArgus evaluation… See the full description on the dataset page: https://huggingface.co/datasets/zengziyun/VideoArgusBench.imagetext-to-video1K<n<10K1 likes351 downloads2mo agoHugging Face03Miwa-Keita /zenz-v2.5-dataset zenz-v2.5-dataset zenz-v2.5-datasetはかな漢字変換タスクに特化した条件付き言語モデル「zenz-v2.5」シリーズの学習を目的として構築したデータセットです。 約190Mペアの「左文脈-入力-変換結果」を含み、かな漢字変換モデルの学習において十分な性能を実現できる規模になっています。 本データセットで学習したzenz-v2.5は公開しています。 zenz-v2.5-medium: 310Mの大規模モデル zenz-v2.5-small: 91Mの中規模モデル zenz-v2.5-xsmall: 26Mの小規模モデル また、かな漢字変換の評価ベンチマークとしてAJIMEE-Bench(味見ベンチ)も公開しています。 形式 本データセットはJSONL形式になっており、以下の3つのデータを含みます。 "input": str, 入力のカタカナ文字列(記号、数字、空白などが含まれることがあります) "output": str, 出力の漢字交じり文 "left_context":… See the full description on the dataset page: https://huggingface.co/datasets/Miwa-Keita/zenz-v2.5-dataset.text100M<n<1B18 likes340 downloads2y agoHugging Face04Zen0 /AusCyberBench AusCyberBench v2.1 The first comprehensive benchmark for evaluating Large Language Models in Australian cybersecurity contexts. Covers regulatory compliance (Essential Eight, ISM, Privacy Act, SOCI Act), technical security, threat intelligence, and Australian-specific terminology. Model Leaderboard (v2.1, Australian Test Set) # Model Overall 95% CI E8 Threat Intel SOCI Privacy Terminology 1 GPT-5.2 85.5% [83.1%, 87.8%] 83.1% 78.7% 100.0% 93.8% 100.0%… See the full description on the dataset page: https://huggingface.co/datasets/Zen0/AusCyberBench.textquestion-answering1K<n<10K3 likes220 downloads5mo agoHugging Face05zen-E /GSM8k-Aug-NLThis dataset is provided to facilitate access to GSM8k-Aug-NL, originally from https://github.com/da03/implicit_chain_of_thought and https://arxiv.org/abs/2311.01460. This dataset is used to train CODI (https://arxiv.org/abs/2502.21074) Description: We utilize two datasets to train our models--GSM8k-Aug and GSM8k-Aug-NL. (1) We use the GSM8k-Aug dataset, which has proven effective for training implicit CoT methods. This dataset extends the original GSM8k training set to 385k samples by… See the full description on the dataset page: https://huggingface.co/datasets/zen-E/GSM8k-Aug-NL.textquestion-answering100K<n<1M1 likes208 downloads1y agoHugging Face06QRP123 /zeno_h1_2026_07_31_curated11_raw_bags Zeno H1 — 2026-07-31 curated raw ROS 2 bags This public dataset contains 11 curated raw ROS 2 MCAP recordings from 2026-07-31. The original rosbag2_.../ directory structure is retained, so each MCAP can be opened directly after download. Excluded recordings: #1 (refrigerator did not open), #5 (table-wiping ending), #7 (ended before placing draining basket), and #10 (table wiping only half complete). Download on another computer hf download… See the full description on the dataset page: https://huggingface.co/datasets/QRP123/zeno_h1_2026_07_31_curated11_raw_bags.tabularn<1K1 likes120 downloads2mo agoHugging Face07zen-E /StrategyQA_CoT_GPT4otext1K<n<10K1 likes111 downloads1y agoHugging Face08alaaaldeen1994 /zenith-modelstabularn<1K0 likes105 downloads3mo agoHugging Face09zen-E /CommonsenseQA-GPT4ominitext1K<n<10K0 likes98 downloads1y agoHugging Face10zen-E /StrategyQA_GPT4o_CoTx10text10K<n<100K0 likes89 downloads1y agoHugging Face11zeneldensey /research-companion-indextabularn<1K0 likes89 downloads1mo agoHugging Face12Zenng2812 /vietnamese-financial-summary Vietnamese Financial News Summarization with Number Preservation textsummarization1K<n<10K0 likes76 downloads5mo agoHugging Face13ZenMoore /LenCtrl-Bench LenCtrl-Bench: Benchmarking LLMs' Abilities for Length-Controlled Text Generation. arxiv: https://arxiv.org/abs/2410.07035 daily papers: https://huggingface.co/papers/2410.07035 twitter: https://x.com/ZenMoore1/status/1845673846193668546 Method Usage This dataset contains the following fields: instruction and response. constraint: the length constraint. level: the level of the length constraint, choices=["word", "sentence", "paragraph"]. data_source: the… See the full description on the dataset page: https://huggingface.co/datasets/ZenMoore/LenCtrl-Bench.text10K<n<100K0 likes58 downloads2y agoHugging Face14ZenitsuBorade /WixQA WixQA: Enterprise RAG Question-Answering Benchmark 📄 Full Paper Available: For comprehensive details on dataset design, methodology, evaluation results, and analysis, please see our complete research paper: WixQA: A Multi-Dataset Benchmark for Enterprise Retrieval-Augmented Generation Cohen et al. (2025) - arXiv:2505.08643 Dataset Summary WixQA is a three-config collection for evaluating and training Retrieval-Augmented Generation (RAG) systems in enterprise… See the full description on the dataset page: https://huggingface.co/datasets/ZenitsuBorade/WixQA.textquestion-answering10K<n<100K0 likes52 downloads13d agoHugging Face15kek-zengo /exp_q_2_sqltext100K<n<1M0 likes41 downloads2y agoHugging Face16Joakimpalm-Zen /Qwen3-speculative-pair-report Qwen3 speculative pair report: 0.6B draft + 8B target, measured acceptance Research evidence dataset. No model weights. Part of the collection Xyntetik Research: Runner Compatibility Reports on this account, produced with Xyntetik Runner. Dataset summary Question tested. What the acceptance rate of a Qwen3-0.6B draft against a Qwen3-8B target actually is across draft depths, whether the engine's printed tokens-per-round figure can be tuned on (it cannot), and… See the full description on the dataset page: https://huggingface.co/datasets/Joakimpalm-Zen/Qwen3-speculative-pair-report.textn<1K0 likes39 downloads9d agoHugging Face17ZenMoore /CP-Bench CP-Bench: Benchmarking LLMs' Abilities for Copy-Pasting Tool-Use. arxiv: https://arxiv.org/abs/2410.07035 daily papers: https://huggingface.co/papers/2410.07035 twitter: https://x.com/ZenMoore1/status/1845673846193668546 Method Usage This dataset contains the following fields: instruction and response_pure_text are regular inputs and outputs without position ids or copy-pasting. type: choices=["single-copy", "multi-copy"], indicating the number of copies in… See the full description on the dataset page: https://huggingface.co/datasets/ZenMoore/CP-Bench.textn<1K0 likes36 downloads2y agoHugging Face18zenzen9 /regex-rl-dataset Regex RL Training Dataset Synthetic regex dataset for reinforcement learning post-training. Dataset Details Size: 1,158 examples Format: JSONL Use Case: GRPO/RL training for regex generation Data Format { "prompt": "Write a Python regex pattern that matches: <description>", "solution": "<regex_pattern>", "test_cases": { "positive": ["match1", "match2", "match3", "match4", "match5"], "negative": ["no_match1", "no_match2", "no_match3", "no_match4"… See the full description on the dataset page: https://huggingface.co/datasets/zenzen9/regex-rl-dataset.texttext-generation1K<n<10K0 likes35 downloads8mo agoHugging Face19zeno109 /cantonese-chinese-parallel-corpus-baseThis is a dataset of Cantonese-Written Chinese Parallel Corpus, containing 130k+ pairs of Cantonese and Traditional Chinese parallel sentences. texttranslation100K<n<1M1 likes35 downloads4mo agoHugging Face20reapxdev /zenodo-scraper Zenodo Scraper · Research Records, DOIs, Authors & Files Scrape open research records, DOIs, publications, datasets, software, authors, and file metadata from Zenodo. Fast HTTP scraper charging per returned record with tiered pricing. Rows in this dataset 1,965 Fields 27 Collector runs behind it 50 Most recent observation 2026-08-03 What this is Every row here was returned by a real run of a public collector. Nothing is generated from a… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/zenodo-scraper.tabular1K<n<10K1 likes32 downloads2mo agoHugging Face21zenith1232 /qwen36-eagle3-stagebtext100K<n<1M0 likes28 downloads5mo agoHugging Face22Zenng2812 /arca-synthetic-adaptive ARCA Synthetic Adaptive Graph-Path Constraints This dataset is a controlled synthetic benchmark for testing whether Adaptive Residual Constrained Attention (ARCA) reacts to external constraint quality. It does not use LLMs, pretrained language models, or natural-language generation. Every sample is generated by Python with exact ground truth. Task Each base sample contains a random directed graph. Nodes have discrete values. A query gives: a start node a relation… See the full description on the dataset page: https://huggingface.co/datasets/Zenng2812/arca-synthetic-adaptive.tabulargraph-ml10K<n<100K0 likes24 downloads2mo agoHugging Face23AllenGXM /zen-ingress-datasettext10K<n<100K0 likes23 downloads1d agoHugging Face24alvemoans /zenith_ai_305 Legal Data Analysis Dataset This dataset contains legal statements, analyses, and judgments primarily related to labor law and contract law, drawn from various cases and legal interpretations. It includes text entries with factual descriptions, legal arguments, and conclusions based on judicial decisions, as well as instructions related to interpreting those facts. Dataset Overview The dataset is structured as a series of legal paragraphs and corresponding instructions… See the full description on the dataset page: https://huggingface.co/datasets/alvemoans/zenith_ai_305.texttext-generationn<1K0 likes18 downloads2y agoHugging Face25Zenng2812 /financial-judgement-vietnamesetextn<1K0 likes18 downloads5mo agoHugging Face26Zenkad /Datasettextn<1K0 likes14 downloads11mo agoHugging Face27Zenng2812 /bctc-md-domain-corpus Vietnamese Financial Reports Markdown Domain Corpus Dataset này được tạo từ các báo cáo tài chính dạng Markdown trong thư mục BCTC_MD. Mục đích Dataset dùng cho continued pretraining / domain-adaptive pretraining mô hình ngôn ngữ trên miền báo cáo tài chính tiếng Việt. Cấu trúc dữ liệu Mỗi dòng trong train.jsonl hoặc validation.jsonl là một JSON object: { "text": "...", "source_file": "AAA_BCTC_2020.md", "document_id": "AAA_BCTC_2020", "company": "AAA"… See the full description on the dataset page: https://huggingface.co/datasets/Zenng2812/bctc-md-domain-corpus.imagetext-generation1K<n<10K0 likes9 downloads5mo agoHugging Face28zenn19991231 /ADL_HW1_Datastext10K<n<100K0 likes8 downloads3y agoHugging Face29zenitsu09 /mahabharat-qnatext1K<n<10K0 likes7 downloads10mo agoHugging Face30zeng981 /nlpdatasettexttext-classification100K<n<1M0 likes6 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.