CoolFace
25 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mideind /icelandic-common-crawl-corpus-IC3This is the Icelandic Common Crawl Corpus (IC3). texttext-generation1M<n<10M1 likes233 downloads4y agoHugging Face02humanfia-lab /icho-2026 IChO 2026 Lean 4 formalizations: verified model variants This repository contains three independently generated Lean 4 proof sets for the same 32 selected IChO 2026 theory subquestions. Practical papers P1–P3 remain outside the corpus. Proof-origin labels Config Records proof_generator.label Meaning kimi-k3 32 Kimi-K3 Proofs generated in the clean K3 rerun with kimi-k3[1m] through Claude Code. gpt 32 GPT Proofs generated in a fresh answer-blind… See the full description on the dataset page: https://huggingface.co/datasets/humanfia-lab/icho-2026.texttext-generationn<1K0 likes191 downloads25d agoHugging Face03Vidushee /iclr-rejected-papers-with-code-1k Rejected ICLR Papers with Reviews and Code This dataset contains 1,000 rejected ICLR submissions from 2018–2026. Each row has the OpenReview submission metadata and reviews, the rejected submission PDF, and a commit-pinned archive of a matched public GitHub repository. This collection was built directly from OpenReview. It does not use a third-party ICLR review dataset. Project repository: TheAppliedScientist Contents 1,000 unique rejected OpenReview submissions… See the full description on the dataset page: https://huggingface.co/datasets/Vidushee/iclr-rejected-papers-with-code-1k.tabulartext-generation1K<n<10K0 likes159 downloads17d agoHugging Face04ddehun /ICT IncompleteToolBench This dataset is introduced in the paper "Can Tool-Augmented Large Language Models Be Aware of Incomplete Conditions?" (paper list). It aims to evaluate whether large language models can recognize incomplete scenarios where tool invocation is not feasible due to missing tools or insufficient user information. Dataset Overview Derived from: APIBank and ToolBench. Manipulation types: API Replacement: Replaces correct tools with semantically… See the full description on the dataset page: https://huggingface.co/datasets/ddehun/ICT.texttext-generation1K<n<10K1 likes141 downloads19d agoHugging Face05ICIP /TeleSalesCorpus TeleSalesCorpus Dataset Description TeleSalesCorpus is a large-scale, high-fidelity dialogue dataset designed specifically for the domain of intelligent telemarketing. This dataset was constructed to address the core challenges that current Large Language Models (LLMs) face in goal-driven persuasive dialogue tasks, such as telemarketing. These challenges include "strategic brittleness" (difficulty in multi-turn planning) and "factual hallucination" (straying from strict… See the full description on the dataset page: https://huggingface.co/datasets/ICIP/TeleSalesCorpus.texttext-generation1K<n<10K6 likes107 downloads11mo agoHugging Face06N8Programs /shared-emergence-icl-modalities-128 Shared-emergence ICL replication at T=128 This dataset contains the complete raw result archive for the paper “Many Next-Token Predictors are In-Context Learners.” The campaign evaluates a fixed suite of 100 program-synthesis tasks using 128 sampled prompts per task, for every clean and deranged shot cell described by the paper: 21 run keys; 281 experiment cells; 12,800 predictions per cell; 3,596,800 predictions in total. The archive expands to a top-level results_128/… See the full description on the dataset page: https://huggingface.co/datasets/N8Programs/shared-emergence-icl-modalities-128.documenttext-generationn<1K0 likes106 downloads2mo agoHugging Face07AmareshHebbar /icd10-coder-sft ICD-10-CM Medical Coder Part of the AxisMapper Medical AI Suite — 16 domain-specific SFT datasets for fine-tuning medical LLMs. Built by AmareshHebbar | Studio Ilios / Humanova Minds What this dataset does Maps clinical descriptions to ICD-10-CM codes Why download this Fine-tune LLMs to automatically assign ICD-10-CM codes from clinical text. Useful for EHR automation, medical coding assistants, and clinical NLP pipelines. Dataset stats… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/icd10-coder-sft.texttext-generation10K<n<100K0 likes88 downloads3mo agoHugging Face08DeepFin-Intelligence /ICBCBenchICBCBench: An Industry Consortium Benchmark for Financial Deep Research Overview ICBCBench is an industry consortium benchmark for evaluating financial Deep Research Agents in real-world research scenarios. It consists of bilingual objective and subjective tasks across major financial sectors, including capital markets, banking, insurance, and related financial services. Developed with over 50 contributors from more than 40 financial and academic organizations, ICBCBench… See the full description on the dataset page: https://huggingface.co/datasets/DeepFin-Intelligence/ICBCBench.texttext-generationn<1K1 likes87 downloads12d agoHugging Face09ICBCBench /ICBCBench anon-repo Bench dataset anon-repo Bench is a benchmark which aims at evaluating next-generation LLMs (LLMs with augmented capabilities due to added tooling, efficient prompting, access to search, etc.), containing questions that predominantly cover finance and politics. Data anon-repo Bench dataset consists of 120 questions with clear and unambiguous answers, covering both Chinese and English. It includes 40 subjective questions and 80 objective questions. The questions… See the full description on the dataset page: https://huggingface.co/datasets/ICBCBench/ICBCBench.texttext-generationn<1K0 likes61 downloads5mo agoHugging Face10Builder-syntaxlabs /medical-billing-icd10-qa-sample Medical Billing & ICD-10 Synthetic Dataset (Sample 🚀 NEED THE FULL ENTERPRISE COMMERCIAL DATASET? Get instant access to the full 50,000+ cleaned JSONL dataset for fine-tuning production models: 🏥 50,000+ verified ICD-10 / CPT billing scenarios 📄 Clean JSONL format (instruction, input, output) 🔒 Safe for HIPAA/GDPR—100% synthetic, zero real patient data 💼 Full commercial license for SaaS and Enterprise applications 👉 Buy Full Enterprise Dataset ($249) - Instant Download… See the full description on the dataset page: https://huggingface.co/datasets/Builder-syntaxlabs/medical-billing-icd10-qa-sample.texttext-generationn<1K0 likes60 downloads5d agoHugging Face11guicybercode /iceland-tech-christian-ethics-prompts Fictional Icelandic Landscapes, Technology and Christian Ethics Prompts This microdataset contains 24 original discussion prompts arranged as 12 parallel pt-BR/English pairs. Each explicitly fictional scenario combines a landscape motif inspired by Iceland, a technology-governance dilemma, and concepts that may be explored through Christian ethics. The records do not describe real Icelandic institutions, policies, communities, or practices, and they do not claim that Christians… See the full description on the dataset page: https://huggingface.co/datasets/guicybercode/iceland-tech-christian-ethics-prompts.texttext-generationn<1K0 likes59 downloads29d agoHugging Face12DeepNLP /ICLR-2021-Accepted-Papers ICLR 2021 International Conference on Learning Representations 2021 Accepted Paper Meta Info Dataset This dataset is collect from the ICLR 2021 OpenReview website (https://openreview.net/group?id=ICLR.cc/2021/Conference#tab-accept-oral) as well as the arxiv website DeepNLP paper arxiv (http://www.deepnlp.org/content/paper/iclr2021). For researchers who are interested in doing analysis of ICLR 2021 accepted papers and potential trends, you can use the already cleaned up json files.… See the full description on the dataset page: https://huggingface.co/datasets/DeepNLP/ICLR-2021-Accepted-Papers.texttext-generationn<1K0 likes53 downloads2y agoHugging Face13AmareshHebbar /icd10-to-drg-sft ICD-10-CM to MS-DRG Mapper Part of the AxisMapper Medical AI Suite — 16 domain-specific SFT datasets for fine-tuning medical LLMs. Built by AmareshHebbar | Studio Ilios / Humanova Minds What this dataset does ICD-10-CM diagnosis codes → MS-DRG code, relative weight, and geometric mean LOS Why download this Hospital reimbursement prediction, DRG validation, revenue cycle automation. Critical for US hospital billing under the Medicare IPPS system.… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/icd10-to-drg-sft.texttext-generation1K<n<10K0 likes50 downloads3mo agoHugging Face14brandburner /i-claudius-narrative-kg I, Claudius Complete Series Narrative Knowledge Graph Dataset Description This dataset contains a comprehensive narrative knowledge graph extracted from all 13 episodes of the BBC's "I, Claudius" (1976), analyzed using the Fabula V2 pipeline. The graph captures the complex web of Roman imperial politics, family dynamics, and power struggles across the reigns of Augustus, Tiberius, Caligula, and Claudius. Dataset Summary Total Nodes: 10,357 Total Relationships:… See the full description on the dataset page: https://huggingface.co/datasets/brandburner/i-claudius-narrative-kg.tabulargraph-ml10K<n<100K0 likes33 downloads1y agoHugging Face15ICEPVP8977 /Uncensored_mini A Dataset with Uncensored Content Focused on Hacking/Penetration Testing texttext-generation1K<n<10K0 likes30 downloads2y agoHugging Face16AmareshHebbar /ayurveda-icd-sft Ayurveda / Traditional Medicine to ICD-10 Bridge Part of the AxisMapper Medical AI Suite — 16 domain-specific SFT datasets for fine-tuning medical LLMs. Built by AmareshHebbar | Studio Ilios / Humanova Minds What this dataset does Traditional / Ayurvedic medicine concepts → nearest ICD-10-CM code Why download this Bridge traditional Indian medicine (Ayurveda, Unani, Siddha, Homeopathy) to ICD-10 for ABDM integration, ABHA health records, and… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/ayurveda-icd-sft.texttext-generation1K<n<10K0 likes29 downloads3mo agoHugging Face17Siyu2Zhou /ICKG-immunology-triple-extraction-sft ICKG 免疫学知识三元组抽取 SFT 数据集 本数据集用于从 PubMed 免疫学摘要中抽取生物医学知识三元组的指令微调(SFT)。每条样本是一段对话(system / user / assistant),assistant 即为该摘要抽取出的三元组 JSON 数组。配套的微调 adapter 见 Siyu2Zhou/Baichuan-M2-32B-QLoRA-immunology-triples。 数据规模 切分 文件 样本数 train train.jsonl 4,500 validation val.jsonl 250 test test.jsonl 250 合计 5,000 篇摘要 三元组总数 52,597,平均 10.5 条/篇(最少 3、最多 30)。 5,000 篇按「关系覆盖 + 三元组密度」分层抽样(A/B/C 三档 = 2000/2000/1000),并做关系再平衡(associated_with ≥ 35%、increases ≤… See the full description on the dataset page: https://huggingface.co/datasets/Siyu2Zhou/ICKG-immunology-triple-extraction-sft.texttext-generation1K<n<10K0 likes26 downloads3mo agoHugging Face18weiber2002 /ICRTL IC-RTL: Industrial-Scale RTL Design Benchmark 📖 Overview This repository contains a collection of industrial-level RTL design challenges selected from the National Taiwan Integrated Circuit Design Contest and handcrafted problems, complete with our reference implementations and specs. Each challenge targets specific algorithms or hardware modules used in industry. We present this collection as the ICRTL benchmark, designed to evaluate PPA (Power, Performance, Area)… See the full description on the dataset page: https://huggingface.co/datasets/weiber2002/ICRTL.texttext-generationn<1K1 likes25 downloads8mo agoHugging Face19Madeleinex /icj-precedent-cite-bench ICJ Precedent-Citation Benchmark A benchmark for one task over the International Court of Justice: given the record of a case as it stood before the Court issued its decision, predict which earlier ICJ cases (and which provisions of the Court's own instruments) the decision will cite. The five cases were all decided after the knowledge cutoffs of current frontier models, so a model cannot have read the decisions during pretraining, and each case is removed from the graph so it… See the full description on the dataset page: https://huggingface.co/datasets/Madeleinex/icj-precedent-cite-bench.texttext-classificationn<1K0 likes23 downloads1mo agoHugging Face20icusu /same-bar-llm-disagreement Same-Bar Cross-Model Disagreement — 20+ LLM lineages judging identical market bars Free sample: 4,000+ decisions. Full dataset (growing daily): ~907,000 decisions / 39 model routes / 33,344 SPY 30-minute bars (2016–2026), 28,877 bars judged by ≥4 distinct models under a byte-identical, hash-pinned prompt — with causal (no-look-ahead) simulated outcomes. Averages ~27 model decisions per bar. Request access / contact below. What makes this dataset unusual Everyone… See the full description on the dataset page: https://huggingface.co/datasets/icusu/same-bar-llm-disagreement.texttext-generationn<1K0 likes20 downloads2mo agoHugging Face21ICEPVP8977 /Uncensored_Tiny_Reasoningtexttext-generationn<1K2 likes19 downloads2y agoHugging Face22ICEPVP8977 /Uncensored_Small_Reasoning texttext-generation1K<n<10K7 likes19 downloads2y agoHugging Face23singhankit16 /ICD-10-LLM-generated-Synthetic-Circulatory-System-I00-I99 MedGemma ICD-10 Clinical Notes Dataset — Circulatory System Synthetic clinical notes generated by MedGemma-4B-IT for fine-tuning ICD-10-CM diagnosis code prediction models. Focused on Chapter 9: Diseases of the Circulatory System (I00-I99). Dataset Summary Split Examples Unique ICD-10 Codes Train 6,275 1,255 Each example is a realistic clinical note paired with its ICD-10-CM diagnosis code, formatted as a chat conversation for instruction fine-tuning.… See the full description on the dataset page: https://huggingface.co/datasets/singhankit16/ICD-10-LLM-generated-Synthetic-Circulatory-System-I00-I99.texttext-classification1K<n<10K1 likes15 downloads5mo agoHugging Face24icemoon28 /guess_word_datasettexttext-generation1K<n<10K0 likes5 downloads2y agoHugging Face25alexbs /ictisgpt Dataset Card for ICTIS GPT dataset texttext-generationn<1K0 likes2 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.