datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
icelandic-common-crawl-corpus-IC3This is the Icelandic Common Crawl Corpus (IC3).
icho-2026
IChO 2026 Lean 4 formalizations: verified model variants
This repository contains three independently generated Lean 4 proof sets for the
same 32 selected IChO 2026 theory subquestions. Practical papers P1–P3 remain
outside the corpus.
Proof-origin labels
Config
Records
proof_generator.label
Meaning
kimi-k3
32
Kimi-K3
Proofs generated in the clean K3 rerun with kimi-k3[1m] through Claude Code.
gpt
32
GPT
Proofs generated in a fresh answer-blind… See the full description on the dataset page: https://huggingface.co/datasets/humanfia-lab/icho-2026.iclr-rejected-papers-with-code-1k
Rejected ICLR Papers with Reviews and Code
This dataset contains 1,000 rejected ICLR submissions from 2018–2026. Each row
has the OpenReview submission metadata and reviews, the rejected submission PDF,
and a commit-pinned archive of a matched public GitHub repository.
This collection was built directly from OpenReview. It does not use a
third-party ICLR review dataset.
Project repository: TheAppliedScientist
Contents
1,000 unique rejected OpenReview submissions… See the full description on the dataset page: https://huggingface.co/datasets/Vidushee/iclr-rejected-papers-with-code-1k.ICT
IncompleteToolBench
This dataset is introduced in the paper "Can Tool-Augmented Large Language Models Be Aware of Incomplete Conditions?" (paper list). It aims to evaluate whether large language models can recognize incomplete scenarios where tool invocation is not feasible due to missing tools or insufficient user information.
Dataset Overview
Derived from: APIBank and ToolBench.
Manipulation types:
API Replacement: Replaces correct tools with semantically… See the full description on the dataset page: https://huggingface.co/datasets/ddehun/ICT.TeleSalesCorpus
TeleSalesCorpus
Dataset Description
TeleSalesCorpus is a large-scale, high-fidelity dialogue dataset designed specifically for the domain of intelligent telemarketing.
This dataset was constructed to address the core challenges that current Large Language Models (LLMs) face in goal-driven persuasive dialogue tasks, such as telemarketing. These challenges include "strategic brittleness" (difficulty in multi-turn planning) and "factual hallucination" (straying from strict… See the full description on the dataset page: https://huggingface.co/datasets/ICIP/TeleSalesCorpus.shared-emergence-icl-modalities-128
Shared-emergence ICL replication at T=128
This dataset contains the complete raw result archive for the paper
“Many Next-Token Predictors are In-Context Learners.”
The campaign evaluates a fixed suite of 100 program-synthesis tasks using 128
sampled prompts per task, for every clean and deranged shot cell described by
the paper:
21 run keys;
281 experiment cells;
12,800 predictions per cell;
3,596,800 predictions in total.
The archive expands to a top-level results_128/… See the full description on the dataset page: https://huggingface.co/datasets/N8Programs/shared-emergence-icl-modalities-128.icd10-coder-sft
ICD-10-CM Medical Coder
Part of the AxisMapper Medical AI Suite — 16 domain-specific SFT datasets for fine-tuning medical LLMs.
Built by AmareshHebbar | Studio Ilios / Humanova Minds
What this dataset does
Maps clinical descriptions to ICD-10-CM codes
Why download this
Fine-tune LLMs to automatically assign ICD-10-CM codes from clinical text. Useful for EHR automation, medical coding assistants, and clinical NLP pipelines.
Dataset stats… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/icd10-coder-sft.ICBCBenchICBCBench: An Industry Consortium Benchmark for Financial Deep Research
Overview
ICBCBench is an industry consortium benchmark for evaluating financial Deep Research Agents in real-world research scenarios. It consists of bilingual objective and subjective tasks across major financial sectors, including capital markets, banking, insurance, and related financial services. Developed with over 50 contributors from more than 40 financial and academic organizations, ICBCBench… See the full description on the dataset page: https://huggingface.co/datasets/DeepFin-Intelligence/ICBCBench.ICBCBench
anon-repo Bench dataset
anon-repo Bench is a benchmark which aims at evaluating next-generation LLMs (LLMs with augmented capabilities due to added tooling, efficient prompting, access to search, etc.), containing questions that predominantly cover finance and politics.
Data
anon-repo Bench dataset consists of 120 questions with clear and unambiguous answers, covering both Chinese and English. It includes 40 subjective questions and 80 objective questions. The questions… See the full description on the dataset page: https://huggingface.co/datasets/ICBCBench/ICBCBench.medical-billing-icd10-qa-sample
Medical Billing & ICD-10 Synthetic Dataset (Sample
🚀 NEED THE FULL ENTERPRISE COMMERCIAL DATASET?
Get instant access to the full 50,000+ cleaned JSONL dataset for fine-tuning production models:
🏥 50,000+ verified ICD-10 / CPT billing scenarios
📄 Clean JSONL format (instruction, input, output)
🔒 Safe for HIPAA/GDPR—100% synthetic, zero real patient data
💼 Full commercial license for SaaS and Enterprise applications
👉 Buy Full Enterprise Dataset ($249) - Instant Download… See the full description on the dataset page: https://huggingface.co/datasets/Builder-syntaxlabs/medical-billing-icd10-qa-sample.iceland-tech-christian-ethics-prompts
Fictional Icelandic Landscapes, Technology and Christian Ethics Prompts
This microdataset contains 24 original discussion prompts arranged as 12
parallel pt-BR/English pairs. Each explicitly fictional scenario combines a
landscape motif inspired by Iceland, a technology-governance dilemma, and
concepts that may be explored through Christian ethics. The records do not
describe real Icelandic institutions, policies, communities, or practices, and
they do not claim that Christians… See the full description on the dataset page: https://huggingface.co/datasets/guicybercode/iceland-tech-christian-ethics-prompts.ICLR-2021-Accepted-Papers
ICLR 2021 International Conference on Learning Representations 2021 Accepted Paper Meta Info Dataset
This dataset is collect from the ICLR 2021 OpenReview website (https://openreview.net/group?id=ICLR.cc/2021/Conference#tab-accept-oral) as well as the arxiv website DeepNLP paper arxiv (http://www.deepnlp.org/content/paper/iclr2021). For researchers who are interested in doing analysis of ICLR 2021 accepted papers and potential trends, you can use the already cleaned up json files.… See the full description on the dataset page: https://huggingface.co/datasets/DeepNLP/ICLR-2021-Accepted-Papers.icd10-to-drg-sft
ICD-10-CM to MS-DRG Mapper
Part of the AxisMapper Medical AI Suite — 16 domain-specific SFT datasets for fine-tuning medical LLMs.
Built by AmareshHebbar | Studio Ilios / Humanova Minds
What this dataset does
ICD-10-CM diagnosis codes → MS-DRG code, relative weight, and geometric mean LOS
Why download this
Hospital reimbursement prediction, DRG validation, revenue cycle automation. Critical for US hospital billing under the Medicare IPPS system.… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/icd10-to-drg-sft.i-claudius-narrative-kg
I, Claudius Complete Series Narrative Knowledge Graph
Dataset Description
This dataset contains a comprehensive narrative knowledge graph extracted from all 13 episodes of the BBC's "I, Claudius" (1976), analyzed using the Fabula V2 pipeline. The graph captures the complex web of Roman imperial politics, family dynamics, and power struggles across the reigns of Augustus, Tiberius, Caligula, and Claudius.
Dataset Summary
Total Nodes: 10,357
Total Relationships:… See the full description on the dataset page: https://huggingface.co/datasets/brandburner/i-claudius-narrative-kg.Uncensored_mini
A Dataset with Uncensored Content Focused on Hacking/Penetration Testing
ayurveda-icd-sft
Ayurveda / Traditional Medicine to ICD-10 Bridge
Part of the AxisMapper Medical AI Suite — 16 domain-specific SFT datasets for fine-tuning medical LLMs.
Built by AmareshHebbar | Studio Ilios / Humanova Minds
What this dataset does
Traditional / Ayurvedic medicine concepts → nearest ICD-10-CM code
Why download this
Bridge traditional Indian medicine (Ayurveda, Unani, Siddha, Homeopathy) to ICD-10 for ABDM integration, ABHA health records, and… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/ayurveda-icd-sft.ICKG-immunology-triple-extraction-sft
ICKG 免疫学知识三元组抽取 SFT 数据集
本数据集用于从 PubMed 免疫学摘要中抽取生物医学知识三元组的指令微调(SFT)。每条样本是一段对话(system / user / assistant),assistant 即为该摘要抽取出的三元组 JSON 数组。配套的微调 adapter 见 Siyu2Zhou/Baichuan-M2-32B-QLoRA-immunology-triples。
数据规模
切分
文件
样本数
train
train.jsonl
4,500
validation
val.jsonl
250
test
test.jsonl
250
合计
5,000 篇摘要
三元组总数 52,597,平均 10.5 条/篇(最少 3、最多 30)。
5,000 篇按「关系覆盖 + 三元组密度」分层抽样(A/B/C 三档 = 2000/2000/1000),并做关系再平衡(associated_with ≥ 35%、increases ≤… See the full description on the dataset page: https://huggingface.co/datasets/Siyu2Zhou/ICKG-immunology-triple-extraction-sft.ICRTL
IC-RTL: Industrial-Scale RTL Design Benchmark
📖 Overview
This repository contains a collection of industrial-level RTL design challenges selected from the National Taiwan Integrated Circuit Design Contest and handcrafted problems, complete with our reference implementations and specs. Each challenge targets specific algorithms or hardware modules used in industry. We present this collection as the ICRTL benchmark, designed to evaluate PPA (Power, Performance, Area)… See the full description on the dataset page: https://huggingface.co/datasets/weiber2002/ICRTL.icj-precedent-cite-bench
ICJ Precedent-Citation Benchmark
A benchmark for one task over the International Court of Justice: given the record of a case as
it stood before the Court issued its decision, predict which earlier ICJ cases (and which
provisions of the Court's own instruments) the decision will cite.
The five cases were all decided after the knowledge cutoffs of current frontier models, so a
model cannot have read the decisions during pretraining, and each case is removed from the
graph so it… See the full description on the dataset page: https://huggingface.co/datasets/Madeleinex/icj-precedent-cite-bench.same-bar-llm-disagreement
Same-Bar Cross-Model Disagreement — 20+ LLM lineages judging identical market bars
Free sample: 4,000+ decisions. Full dataset (growing daily): ~907,000 decisions /
39 model routes / 33,344 SPY 30-minute bars (2016–2026), 28,877 bars judged by ≥4 distinct
models under a byte-identical, hash-pinned prompt — with causal (no-look-ahead) simulated
outcomes. Averages ~27 model decisions per bar. Request access / contact below.
What makes this dataset unusual
Everyone… See the full description on the dataset page: https://huggingface.co/datasets/icusu/same-bar-llm-disagreement.Uncensored_Tiny_ReasoningUncensored_Small_Reasoning
ICD-10-LLM-generated-Synthetic-Circulatory-System-I00-I99
MedGemma ICD-10 Clinical Notes Dataset — Circulatory System
Synthetic clinical notes generated by MedGemma-4B-IT for fine-tuning ICD-10-CM diagnosis code prediction models. Focused on Chapter 9: Diseases of the Circulatory System (I00-I99).
Dataset Summary
Split
Examples
Unique ICD-10 Codes
Train
6,275
1,255
Each example is a realistic clinical note paired with its ICD-10-CM diagnosis code, formatted as a chat conversation for instruction fine-tuning.… See the full description on the dataset page: https://huggingface.co/datasets/singhankit16/ICD-10-LLM-generated-Synthetic-Circulatory-System-I00-I99.guess_word_datasetictisgpt
Dataset Card for ICTIS GPT dataset
