datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
x_dataset_53985
Bittensor Subnet 13 X (Twitter) Dataset
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks.
For more information about the dataset, please visit the official repository.
Supported Tasks
The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/icedwind/x_dataset_53985.x_dataset_34576
Bittensor Subnet 13 X (Twitter) Dataset
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks.
For more information about the dataset, please visit the official repository.
Supported Tasks
The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/icedwind/x_dataset_34576.x_dataset_12552
Bittensor Subnet 13 X (Twitter) Dataset
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks.
For more information about the dataset, please visit the official repository.
Supported Tasks
The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/icedwind/x_dataset_12552.alpaca-gpt4-indonesianBase model : FreedomIntelligence/alpaca-gpt4-indonesian
We wrangled the original dataset format to 'input' & 'output' format. For example:
BEFORE:
[ { "from": "human",
"value": "Saranlah slogan untuk kampanye daur ulang\n" },
{ "from": "gpt",
"value": "1. \"Kurangi, gunakan kembali, daur ulang: Bersama untuk masa depan yang lebih hijau.\"\n2. \
"Daur ulanglah hari ini, untuk masa depan yang lebih baik.\"\n3. \"Ubah sampahmu menjadi harta karun - Daur ulang!\"\n4. \
"Daur… See the full description on the dataset page: https://huggingface.co/datasets/Ichsan2895/alpaca-gpt4-indonesian.x_dataset_50132
Bittensor Subnet 13 X (Twitter) Dataset
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks.
For more information about the dataset, please visit the official repository.
Supported Tasks
The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/icedwind/x_dataset_50132.x_dataset_12970
Bittensor Subnet 13 X (Twitter) Dataset
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks.
For more information about the dataset, please visit the official repository.
Supported Tasks
The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/icedwind/x_dataset_12970.x_dataset_17879
Bittensor Subnet 13 X (Twitter) Dataset
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks.
For more information about the dataset, please visit the official repository.
Supported Tasks
The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/icedwind/x_dataset_17879.x_dataset_3753
Bittensor Subnet 13 X (Twitter) Dataset
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks.
For more information about the dataset, please visit the official repository.
Supported Tasks
The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/icedwind/x_dataset_3753.x_dataset_46763
Bittensor Subnet 13 X (Twitter) Dataset
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks.
For more information about the dataset, please visit the official repository.
Supported Tasks
The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/icedwind/x_dataset_46763.x_dataset_57303
Bittensor Subnet 13 X (Twitter) Dataset
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks.
For more information about the dataset, please visit the official repository.
Supported Tasks
The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/icedwind/x_dataset_57303.x_dataset_44100
Bittensor Subnet 13 X (Twitter) Dataset
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks.
For more information about the dataset, please visit the official repository.
Supported Tasks
The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/icedwind/x_dataset_44100.x_dataset_21716
Bittensor Subnet 13 X (Twitter) Dataset
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks.
For more information about the dataset, please visit the official repository.
Supported Tasks
The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/icedwind/x_dataset_21716.x_dataset_7114
Bittensor Subnet 13 X (Twitter) Dataset
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks.
For more information about the dataset, please visit the official repository.
Supported Tasks
The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/icedwind/x_dataset_7114.Nurisk-ICRA2026
Nurisk: VQA for Risk Assessment in Autonomous Driving
Nurisk is a visual question answering dataset focusing on risk assessment for autonomous driving. Each row contains:
image: a BEV image
question: a driving-related question
answer: the ground truth answer
Paper
NuRisk: A Visual Question Answering Dataset for Agent-Level Risk Assessment in Autonomous Driving — see the paper on arXiv:2509.25944 .
Framework
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/TUM-AVS/Nurisk-ICRA2026.x_dataset_27136
Bittensor Subnet 13 X (Twitter) Dataset
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks.
For more information about the dataset, please visit the official repository.
Supported Tasks
The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/icedwind/x_dataset_27136.x_dataset_41362
Bittensor Subnet 13 X (Twitter) Dataset
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks.
For more information about the dataset, please visit the official repository.
Supported Tasks
The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/icedwind/x_dataset_41362.x_dataset_11100
Bittensor Subnet 13 X (Twitter) Dataset
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks.
For more information about the dataset, please visit the official repository.
Supported Tasks
The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/icedwind/x_dataset_11100.ICU-REACT
ICU-REACT
ICU-REACT is a clinician-supervised dataset for clinical reasoning and information retrieval in the intensive care unit (ICU), developed for fine-tuning and benchmarking large language models (LLMs).
ICU-REACT was constructed using a clinician-in-the-loop annotation framework designed to capture how clinicians identify relevant patient information and integrate it into diagnostic and treatment decisions. The dataset includes a clinician-refined seed training set, a… See the full description on the dataset page: https://huggingface.co/datasets/iheallab/ICU-REACT.LongSpeech-Eval
The proposed Long-Speech understanding evaluation dataset for the paper 'FastLongSpeech: Enhancing Large Speech-Language Models for Efficient Long-Speech Processing'
Usage
First download the model from Model.
Then please refer to Github Page.
Requirements
We suggest to run with Python 3.10.
Examples of usage:
git clone https://github.com/ictnlp/FastLongSpeech.git
cd transformers-main
pip install -e .
pip install deepspeed sentencepiece librosa… See the full description on the dataset page: https://huggingface.co/datasets/ICTNLP/LongSpeech-Eval.ICL-Router
ICL-Router: In-Context Learned Model Representations for LLM Routing
This repository contains the dataset for the paper: ICL-Router: In-Context Learned Model Representations for LLM Routing.
Paper Abstract:
Large language models (LLMs) often exhibit complementary strengths. Model routing harnesses these strengths by dynamically directing each query to the most suitable model, given a candidate model pool. However, routing performance relies on accurate model representations, and… See the full description on the dataset page: https://huggingface.co/datasets/lalalamdbf/ICL-Router.ICCV_Papers
ICCV Papers
ICCV (International Conference on Computer Vision) is one of the most prestigious conferences in computer vision, held biennially since 1987. Along with CVPR and ECCV, it forms the top-tier venues for computer vision research. ICCV papers have contributed groundbreaking work in areas such as object detection, image segmentation, 3D reconstruction, and video understanding.
ICCV Papers is a comprehensive dataset containing all papers from ICCV 2013… See the full description on the dataset page: https://huggingface.co/datasets/choucsan/ICCV_Papers.ROOTS
ROOTS
ROOTS contains 43,922,135 audio-language conversations across four taxonomy tiers. Audio is supplied by the source datasets below.
Quick start
from datasets import load_dataset
dataset = load_dataset("iclr2027anon/ROOTS", split="train", streaming=True)
row = next(iter(dataset))
Use the audio guide to locate and load each conversation's clips in order.
Columns
Columns
Meaning
id
Conversation ID: roots_ followed by 32 hexadecimal… See the full description on the dataset page: https://huggingface.co/datasets/iclr2027anon/ROOTS.RAQUEL2-ICLR
RAQUEL2
Execution-grounded evaluation for machine unlearning. Each evaluation record is
a question answered by a SQL query run against two databases: one built from
the full corpus, and one with the forget-set facts removed. A record is
affected when the two databases disagree, and unaffected when they
agree — so the label is a measured property of the data, not an annotation.
Four configs, together enough to run the benchmark end to end:
Config / split
What it is
Use… See the full description on the dataset page: https://huggingface.co/datasets/Hyukkyu/RAQUEL2-ICLR.ICSThreatQA
ICSThreatQA
Dataset Summary
ICSThreatQA accompanies:
Rani, R., Kumar, M., Epiphaniou, G., and Maple, C. (2025). ICSThreatQA: A Knowledge-Graph Enhanced Question Answering Model for Industrial Control System Threat Intelligence. Expert Systems with Applications.
This release packages the benchmark's question-answer pairs into a single, canonical, machine-readable table (CSV and Parquet) suitable for loading with the Hugging Face datasets library, while preserving… See the full description on the dataset page: https://huggingface.co/datasets/mahend72/ICSThreatQA.icd10cm-exam-mcq
ICD-10-CM Exam MCQ
203 multiple-choice items drawn from ICD-10-CM practice exams and coursework — clinical
vignettes requiring an actual code assignment, plus questions on conventions, guideline
structure and Chapter 20 external-cause rules.
⚠️ Read this before using or redistributing
The questions are third-party material of unverified provenance, reproduced verbatim.
They come from practice exams and coursework — one identifies itself as belonging to an… See the full description on the dataset page: https://huggingface.co/datasets/chongpangnasilemak/icd10cm-exam-mcq.iclr-papers-with-code-1k
ICLR Papers with Accessible Code
A dataset of 1,051 papers from ICLR (2020-2026) with verified code repositories and complete peer reviews from all reviewers.
Dataset Summary
This dataset contains rejected and borderline-accepted papers from ICLR (International Conference on Learning Representations) with accessible code and full peer review text.
Contents:
1,051 papers total
3,900 reviews (average 3.71 per paper)
944 rejected (90%) + 107 poster-tier accepted… See the full description on the dataset page: https://huggingface.co/datasets/Vidushee/iclr-papers-with-code-1k.ICBCBenchICBCBench: An Industry Consortium Benchmark for Financial Deep Research
Overview
ICBCBench is an industry consortium benchmark for evaluating financial Deep Research Agents in real-world research scenarios. It consists of bilingual objective and subjective tasks across major financial sectors, including capital markets, banking, insurance, and related financial services. Developed with over 50 contributors from more than 40 financial and academic organizations, ICBCBench… See the full description on the dataset page: https://huggingface.co/datasets/DeepFin-Intelligence/ICBCBench.icelandic-arc-challenge
Dataset Card for Icelandic ARC-Challenge
This dataset is an Icelandic machine-translated version of the original English ARC-Challenge.
ICDOPS-QA-2024
Dataset Card for ICDOPS-QA-2024
Paper: Unlocking Public Catalogues: Instruction-Tuning LLMs for ICD Coding of German Tumor Diagnoses
Content
This dataset contains 518,116 question-answer pairs in German medical language with a focus on oncological diagnoses, covering the following medical coding systems, including:
ICD-10-GM (International Statistical Classification of Diseases, 10th revision, German Modification) via the alphabetical index (Alpha-ID)
ICD-O-3… See the full description on the dataset page: https://huggingface.co/datasets/stefan-m-lenz/ICDOPS-QA-2024.ICBCBench
anon-repo Bench dataset
anon-repo Bench is a benchmark which aims at evaluating next-generation LLMs (LLMs with augmented capabilities due to added tooling, efficient prompting, access to search, etc.), containing questions that predominantly cover finance and politics.
Data
anon-repo Bench dataset consists of 120 questions with clear and unambiguous answers, covering both Chinese and English. It includes 40 subjective questions and 80 objective questions. The questions… See the full description on the dataset page: https://huggingface.co/datasets/ICBCBench/ICBCBench.
