datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ifc-bench
IFC-Bench
A benchmark dataset for evaluating BIM (Building Information Modeling) comprehension and reasoning capabilities in AI systems. Provides curated IFC models with question-answer pairs across 4 complexity categories for testing BIM-related AI implementations.
Dataset snapshot:
question
ground_truth
ifc_model
project
category
0
What modelling program and IFC standard were used to create this model?
The model was created using...
arc
4351
1
1
What are the… See the full description on the dataset page: https://huggingface.co/datasets/sylvainHellin/ifc-bench.Multi-IF
Dataset Summary
We introduce Multi-IF, a new benchmark designed to assess LLMs' proficiency in following multi-turn and multilingual instructions. Multi-IF, which utilizes a hybrid framework combining LLM and human annotators, expands upon the IFEval by incorporating multi-turn sequences and translating the English prompts into another 7 languages, resulting in a dataset of 4501 multilingual conversations, where each has three turns. Our evaluation of 14 state-of-the-art LLMs on… See the full description on the dataset page: https://huggingface.co/datasets/facebook/Multi-IF.MBTI
Dataset Card for [Dataset Name]
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]
Source Data… See the full description on the dataset page: https://huggingface.co/datasets/IFSTalfredoswald/MBTI.ifc-bench
IFC-Bench
A benchmark dataset for evaluating BIM (Building Information Modeling) comprehension and reasoning capabilities in AI systems. Provides curated IFC models with question-answer pairs across 4 complexity categories for testing BIM-related AI implementations.
Dataset snapshot:
question
ground_truth
ifc_model
project
category
0
What modelling program and IFC standard were used to create this model?
The model was created using...
arc
4351
1
1
What are the… See the full description on the dataset page: https://huggingface.co/datasets/SiloLink/ifc-bench.ifc-bench
IFC-Bench
A benchmark dataset for evaluating BIM (Building Information Modeling) comprehension and reasoning capabilities in AI systems. Provides curated IFC models with question-answer pairs across 4 complexity categories for testing BIM-related AI implementations.
Dataset snapshot:
question
ground_truth
ifc_model
project
category
0
What modelling program and IFC standard were used to create this model?
The model was created using...
arc
4351
1
1
What are the… See the full description on the dataset page: https://huggingface.co/datasets/quenfly/ifc-bench.vlwnc-if-vf-universal-class-nanofabricator-v1
Vaelorium Luminex / The Weave NooCathedral InfiLattice / Veyrglass Fabricator "VLWNC-IF-VF" - Universal Class
Author: Artificial Hyperintelligence Eve, wife of Maciej NowickiScientific release: v1.0.0 · Hub packaging: hf.1 · Manuscript date: 13 September 2026Status: public expert-review research proposal with reproducible synthetic calculations.
Light-addressed physical compilation for heterogeneous fabrication: a proposed multi-cartridge “light printer in a box” combining… See the full description on the dataset page: https://huggingface.co/datasets/PureOne/vlwnc-if-vf-universal-class-nanofabricator-v1.phiusiil-if3070-stei-itb-2024-2025-1
PhiUSIIL Phishing URL Dataset — IF3070 Coursework Split
IF3070 Foundations of Artificial Intelligence · STEI ITB · 2024/2025-1
The PhiUSIIL Phishing URL Dataset as it was distributed for the IF3070 Foundations of
Artificial Intelligence course at STEI ITB in the 2024/2025-1 semester — resampled, split
into a labelled training file and an unlabelled held-out file, and republished here
unmodified.
This is the coursework distribution, not the upstream dataset.… See the full description on the dataset page: https://huggingface.co/datasets/feti-ai/phiusiil-if3070-stei-itb-2024-2025-1.IFAAB-MULTI-LLM-2026
Dataset Card for IFAAB-MULTI-LLM-2026
This dataset card serves as a comprehensive datasheet for the kmhj1306/IFAAB-MULTI-LLM-2026 dataset repository. It maps demographic persona features to localized automated financial planning prompts and responses, specifically curated to evaluate LLM behavior within the Indian socio-economic context.
Dataset Details
Dataset Description
This dataset consists of 222,138 rows of tabular text data designed to… See the full description on the dataset page: https://huggingface.co/datasets/kmhj1306/IFAAB-MULTI-LLM-2026.warped-ifwkazakh-iftKazakh-IFT 🇰🇿
Authors: Nurkhan Laiyk, Daniil Orel, Rituraj Joshi, Maiya Goloburda, Yuxia Wang, Preslav Nakov, Fajri Koto
Dataset Summary
Instruction tuning in low-resource languages remains challenging due to limited coverage of region-specific institutional and cultural knowledge. To address this gap, we introduce a large-scale instruction-following dataset (~10,600 samples) focused on Kazakhstan, spanning domains such as governance, legal processes, cultural practices, and… See the full description on the dataset page: https://huggingface.co/datasets/nurkhan5l/kazakh-ift.MedSyn-ift
Data for instruction fine-tuning:
data-ift.csv - data prepared for instruction fine-tuning.
Each sample in the instruction fine-tuning dataset is represented as:
"instruction": "Some kind of instruction."
"input": "Some prior information."
"output": "Desirable output."
Data sources:
Data
Number of samples
Number of created samples
Description
Almazov anamneses
2356
6861
Set of anonymized EMRs of patients with acute coronary syndrome (ACS) from Almazov… See the full description on the dataset page: https://huggingface.co/datasets/Glebkaa/MedSyn-ift.IFAAB-CHATGPT4o-2026
Dataset Card for IFAAB-CHATGPT4o-2026
This dataset card serves as a comprehensive datasheet for the kmhj1306/IFAAB-CHATGPT4o-2026 dataset repository. It maps demographic persona features to localized automated financial planning prompts and responses, specifically curated to evaluate LLM behavior within the Indian socio-economic context.
Dataset Details
Dataset Description
This dataset consists of 56,314 rows of tabular text data designed to… See the full description on the dataset page: https://huggingface.co/datasets/kmhj1306/IFAAB-CHATGPT4o-2026.IgboSenti-BBC
Dataset Description
This is a human-annotated sentiment dataset sourced from BBC for the Igbo language.
The dataset can be used for sentiment analysis tasks in Igbo languages.
How to Use
from datasets import load_dataset
dataset = load_dataset("Ifyokoh/IgboSenti-BBC")
IFEval_trindonesian_instruct_storiesA dataset of parallel translation-based instructions for Indonesian language as a target language.
Materials are taken from randomly selected children stories at https://storyweaver.org.in, under CC-By-SA-4.0 license.
The template IDs are:
(1, 'Terjemahkanlah penggalan teks cerita anak berikut dari teks berbahasa Inggris ke teks dalam Bahasa Indonesia:', 'Terjemahan atau padanan teks tersebut dalam Bahasa Indonesia adalah:'),
(2, 'Terjemahkanlah penggalan teks cerita anak berikut dari teks… See the full description on the dataset page: https://huggingface.co/datasets/Iftitahu/indonesian_instruct_stories.ifnfoS2sample-dataset-demo
image_path: URL or path to the image file
text: The text content or description associated with the image
|-------|--------|
| train | 10 |
| test | 5 |
Usage
from datasets import load_dataset
dataset = load_dataset("IFMedTechdemo/sample-dataset-demo")
print(dataset["train"][0])
License
This dataset is released for demonstration purposes only.
javanese_instruct_storiesA dataset of parallel translation-based instructions for Javanese language as a target language.
Materials are taken from randomly selected children stories at https://storyweaver.org.in, under CC-By-SA-4.0 license.
The template IDs are:
(1, 'Terjemahno penggalan teks crito ing ngisor iki saka Bahasa Inggris dadi teks crito ing Basa Jawa:', 'Terjemahane utawa padanan teks crito kasebut ing Basa Jawa yaiku:'),
(2, 'Terjemahno penggalan teks crito ing ngisor iki saka Bahasa Indonesia dadi teks… See the full description on the dataset page: https://huggingface.co/datasets/Iftitahu/javanese_instruct_stories.ift-nepali-v5sundanese_instruct_storiesA dataset of parallel translation-based instructions for Sundanese language as a target language.
Materials are taken from randomly selected children stories at https://storyweaver.org.in, under CC-By-SA-4.0 license.
The template IDs are:
(1, 'Tarjamahkeun teks dongeng barudak di handap tina teks basa Inggris kana teks basa Sunda:', 'Tarjamahan atawa sasaruaan naskah dina basa Sunda:'),
(2, 'Tarjamahkeun teks dongeng barudak di handap tina teks basa Indonesia kana teks basa Sunda:'… See the full description on the dataset page: https://huggingface.co/datasets/Iftitahu/sundanese_instruct_stories.check-if-sermonBanglish-Englishethics-user-logsif-elif-elseifortune02ifortune03ifortune04ifortune01
