CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Chat-Error /book2-lite-cleanedtext10K<n<100K2 likes2.5k downloads3y agoHugging Face02AnchorSR /Qwen3.5_RL_ErrorCase Qwen3.5 RL:错误案例与视频定位诊断 v3部分共660题;另新增V4 RL Step3000选帧诊断100题。每题含QA、完整原始输出及可见帧拼图。Dataset Viewer中,default为前60题,video_grounding为v3新增600题,v4_rl_step3000为V4新增100题。 序号 内容 入口 001–060 原三个主实验bench案例 第001题 061–560 RL训练视频500题:训练视觉处理下的新输出 第061题 561–660 ASR-Bench视频100题:复用既有评测输出 第561题 V4-001–100 V4 RL Step3000:VSI/ASR选帧与bbox诊断 V4诊断首页 本次新增的测试内容 新增600题为在看结果前固定的诊断抽样,包含成功与失败,不是600个错误案例。未加入88题附加对照,避免重复。 两部分均为最终Qwen3.5-9B RL… See the full description on the dataset page: https://huggingface.co/datasets/AnchorSR/Qwen3.5_RL_ErrorCase.imagen<1K0 likes826 downloads6d agoHugging Face03sghosts /sync_bigjob_8_finalised_processed_with_error_handling_from_51th_splitimage100K<n<1M0 likes545 downloads1y agoHugging Face04JetBrains-Research /jupyter-errors-dataset Dataset Summary The presented dataset contains 10000 Jupyter notebooks, each of which contains at least one error. In addition to the notebook content, the dataset also provides information about the repository where the notebook is stored. This information can help restore the environment if needed. Getting Started This dataset is organized such that it can be naively loaded via the Hugging Face datasets library. We recommend using streaming due to the large size… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/jupyter-errors-dataset.text1K<n<10K3 likes298 downloads3y agoHugging Face05Chat-Error /MusicLMtext1M<n<10M0 likes268 downloads3y agoHugging Face06liu-nlp /blimp-single-errorDataset for probing model preferences for linguistically acceptable sentences. Generated by introducing automatic corruptions into sentences from Wikipedia, based on UniMorph minimal tag pairs. More info coming soon! @misc{glocker2025growmergescalingstrategies, title={Grow Up and Merge: Scaling Strategies for Efficient Language Adaptation}, author={Kevin Glocker and Kätriin Kukk and Romina Oji and Marcel Bollmann and Marco Kuhlmann and Jenny Kunz}, year={2025}… See the full description on the dataset page: https://huggingface.co/datasets/liu-nlp/blimp-single-error.text100K<n<1M0 likes191 downloads10mo agoHugging Face07liu-nlp /icelandic-blimp-single-errortext10K<n<100K0 likes157 downloads1y agoHugging Face08blazalek /open-smtp-error-dataset Open SMTP Error Dataset Open SMTP Error Dataset is an English-language, machine-readable reference package for SMTP enhanced-status knowledge and cautious operational classification. Version 1.1.0 contains 92 stable knowledge records and a separate, auditable catalog of 120 classification rules. The package is a reference artifact, not a live provider-policy feed. It helps with observability, support, parser testing, and bounded delivery operations; it does not establish… See the full description on the dataset page: https://huggingface.co/datasets/blazalek/open-smtp-error-dataset.textn<1K1 likes155 downloads2mo agoHugging Face09Antix5 /tabular-errors-v1 TabFix multilingual table error pairs — version 2.0 This release keeps 18 business error categories and separates executable deterministic detection from two residual neural categories: text.encoding and text.spelling. The same repository and family-disjoint splits are retained. Split Records Open-vocabulary views train 27948 3260 validation 17127 1844 test 32776 3540 The seven string columns remain id, split, family_id, clean_xml, corrupt_xml, errors… See the full description on the dataset page: https://huggingface.co/datasets/Antix5/tabular-errors-v1.texttoken-classification10K<n<100K0 likes153 downloads5d agoHugging Face10liu-nlp /german-blimp-single-errortext100K<n<1M0 likes147 downloads1y agoHugging Face11AnchorSR /ErrorAnalysis AnchorSR Error Analysis 500道题,严格沿 failure_cases_500.json 的文件顺序排列,每题对照三个SFT模型。 推荐从第001题开始,点击“下一题”逐题阅读。 也可使用本页上方 Dataset Viewer,一行就是一题:图片、问题、标准答案、三个模型完整输出。 Q-Spatial 150题,SpatialRGPT 175题,VSI 175题(仅尺寸、距离,不含面积)。 全部500题:每个模型均提供原图和标注图,共3000张图;O编号和帧号来自模型声明。 原图与标注图使用相同源帧和拼图顺序。无效框/帧号或无声明会注明,不补造;此时标注页可能没有框。 原图指未添加模型框的网页展示副本,经过等比例缩放与JPEG编码,并非原始文件字节;SpatialRGPT原有区域标记保留。 视频只展示可绘制对象涉及帧,无有效框时展示第1帧,非完整视频。原图/标注图使用相同帧。 原生输出完整保留,包括循环、截断和格式错误;未修改答案或重新评分。 所有模型均为SFT,不是baseline。至少一个模型在该题失败,其他模型可能答对。… See the full description on the dataset page: https://huggingface.co/datasets/AnchorSR/ErrorAnalysis.imagen<1K0 likes132 downloads16d agoHugging Face12liu-nlp /swedish-blimp-single-errortext10K<n<100K0 likes98 downloads1y agoHugging Face13onlysainaa /spell_error_opus_14m_mntext10M<n<100M0 likes73 downloads2y agoHugging Face14sssohrab /ct-dosing-errors-benchmarktabular10K<n<100K3 likes72 downloads7mo agoHugging Face15sarayusapa /Grammar_Error_Correctiontext100K<n<1M0 likes71 downloads1y agoHugging Face16aurorra /synthetic-real-word-errors Synthetic Real-Word Error Datasets This repository contains synthetic German data for grammatical error detection and correction, with a focus on context-dependent real-word errors. The repository provides four subsets: Subset Description Examples mixed_real_word Mixed real-word errors 99,812 capitalization Capitalization errors 99,664 case Case errors 99,706 verb Verb errors 99,780 Each subset contains both erroneous and correct sentences and can therefore… See the full description on the dataset page: https://huggingface.co/datasets/aurorra/synthetic-real-word-errors.texttext-classification100K<n<1M0 likes70 downloads15d agoHugging Face17Neura-parse /quantum-error-mitigation-and-benchmarking Neura Parse — Quantum Error Mitigation, Characterization & Benchmarking A pre-fault-tolerance, code-backed vertical on getting trustworthy answers from noisy hardware and rigorously measuring device quality: error-mitigation techniques, characterization/tomography protocols, and benchmarking suites. Runnable Mitiq, pyGSTi, and Qiskit Experiments pipelines with honest sampling-overhead and bias/variance accounting — the practitioner and research toolkit the general dataset… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-error-mitigation-and-benchmarking.tabulartext-generation100K<n<1M0 likes65 downloads3mo agoHugging Face18liu-nlp /estonian-blimp-single-errortext10K<n<100K0 likes64 downloads10mo agoHugging Face19r1v3r /instance_errortextn<1K0 likes57 downloads1y agoHugging Face20Chat-Error /tinystories-gpt4text1M<n<10M3 likes55 downloads3y agoHugging Face21alakxender /dv-synthetic-errors-mixedDV Text Errors Dhivehi text error correction dataset containing correct sentences and synthetically generated errors. The dataset aims to test Dhivehi language error correction models and tools. About Dataset Task: Text error correction Language: Dhivehi (dv) Dataset Structure Input-output pairs of Dhivehi text: correct: Original correct sentences incorrect: Sentences with synthetic errors Note: This is replica of alakxender/dv-synthetic-errors: added more synthetic errors. x5 text10M<n<100M0 likes54 downloads1y agoHugging Face22hartular /grammatical_errors_rrt_press-v2text100K<n<1M0 likes54 downloads11mo agoHugging Face23yifanyu /CORRECT-Error CORRECT-Error A benchmark of 2,226 error-injected multi-agent-system (MAS) trajectories with step-level decisive-error labels, released alongside the paper: CORRECT: Condensed Error Recognition via Knowledge Transfer in Multi-agent Systems. Yifan Yu, Moyan Li, Shaoyuan Xu, Jinmiao Fu, Xinhai Hou, Fan Lai, Bryan Wang. ICML 2026. PMLR 306. CORRECT-Error covers 7 MAS tasks × 2 trajectory-generator models: Dataset gpt-4o-mini gpt-5-nano Total arc 100 204 304 hotpot 69… See the full description on the dataset page: https://huggingface.co/datasets/yifanyu/CORRECT-Error.tabulartext-classification1K<n<10K0 likes52 downloads4mo agoHugging Face24emgena /omnimcp_type_error_mypy_resolver_teaser 🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE: Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20! 📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_type_error_mypy_resolver_teaser.texttext-generationn<1K0 likes52 downloads9d agoHugging Face25Chat-Error /wizard_alpaca_dolly_orcaA merge of pankajmathur/WizardLM_Orca pankajmathur/dolly-v2_orca pankajmathur/alpaca_orca and formated for my use. text100K<n<1M3 likes51 downloads3y agoHugging Face26chuquan282 /CBD_ERROR_LOGStext1K<n<10K0 likes51 downloads3y agoHugging Face27chathuranga-jayanath /selfapr-manipulation-bug-error-context-alltext100K<n<1M0 likes49 downloads3y agoHugging Face28tourmii /vietnamese-corrector-errorstext10M<n<100M0 likes48 downloads4mo agoHugging Face29alakxender /dv-synthetic-errors DV Text Errors Dhivehi text error correction dataset containing correct sentences and synthetically generated errors. The dataset aims to test Dhivehi language error correction models and tools. About Dataset Task: Text error correction Language: Dhivehi (dv) Dataset Structure Input-output pairs of Dhivehi text: correct: Original correct sentences incorrect: Sentences with synthetic errors Statistics Train set: {train_examples} examples… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dv-synthetic-errors.text1M<n<10M0 likes47 downloads2y agoHugging Face30greatestyapper /solidity_errors_and_vulnerabilities Solidity Vulnerabilities Dataset 📖 Overview This dataset contains examples of common vulnerabilities in Solidity smart contracts, structured for use in Retrieval-Augmented Generation (RAG) systems. It is intended to give LLMs context for: Detecting vulnerabilities in Solidity code Explaining security issues in simple terms Suggesting fixes and mitigations Assessing the severity of the issue 🗂 Data Format Each entry is a JSON object with the… See the full description on the dataset page: https://huggingface.co/datasets/greatestyapper/solidity_errors_and_vulnerabilities.textn<1K3 likes47 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.