CoolFace
24 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sylvainHellin /ifc-bench IFC-Bench A benchmark dataset for evaluating BIM (Building Information Modeling) comprehension and reasoning capabilities in AI systems. Provides curated IFC models with question-answer pairs across 4 complexity categories for testing BIM-related AI implementations. Dataset snapshot: question ground_truth ifc_model project category 0 What modelling program and IFC standard were used to create this model? The model was created using... arc 4351 1 1 What are the… See the full description on the dataset page: https://huggingface.co/datasets/sylvainHellin/ifc-bench.documentquestion-answering1K<n<10K20 likes3k downloads4d agoHugging Face02Scale-or-Reason /general-reasoning-ift-pairs Reasoning-IFT Pairs (General Domain) This dataset provides the largest set of IFT and Reasoning answers pairs for a set of general domain queries (cf: math-domain).It is based on the Infinity-Instruct dataset, an extensive and high-quality collection of instruction fine-tuning data. We curated 900k queries from the 7M_core subset of Infinity-Instruct, which covers multiple domains including general knowledge, commonsense Q&A, coding, and math.For each query… See the full description on the dataset page: https://huggingface.co/datasets/Scale-or-Reason/general-reasoning-ift-pairs.textquestion-answering1M<n<10M6 likes1.2k downloads3mo agoHugging Face03Sidsidney /general-reasoning-ift-pairs Reasoning-IFT Pairs (General Domain) This dataset provides the largest set of IFT and Reasoning answers pairs for a set of general domain queries (cf: math-domain).It is based on the Infinity-Instruct dataset, an extensive and high-quality collection of instruction fine-tuning data. We curated 900k queries from the 7M_core subset of Infinity-Instruct, which covers multiple domains including general knowledge, commonsense Q&A, coding, and math.For each query, we… See the full description on the dataset page: https://huggingface.co/datasets/Sidsidney/general-reasoning-ift-pairs.textquestion-answering1M<n<10M4 likes755 downloads10mo agoHugging Face04ifujisawa /procbench Dataset Card for ProcBench Dataset Overview Dataset Description ProcBench is a benchmark designed to evaluate the multi-step reasoning abilities of large language models (LLMs). It focuses on instruction followability, requiring models to solve problems by following explicit, step-by-step procedures. The tasks included in this dataset do not require complex implicit knowledge but emphasize strict adherence to provided instructions. The dataset evaluates model… See the full description on the dataset page: https://huggingface.co/datasets/ifujisawa/procbench.textquestion-answering1K<n<10K3 likes712 downloads2y agoHugging Face05Scale-or-Reason /math-reasoning-ift-pairs Reasoning-IFT Pairs (Math Domain) Paper | Project Page This dataset provides the largest set of IFT and Reasoning answers pairs for a set of math queries (cf: general-domain). It is based on the Llama-Nemotron-Post-Training dataset, an extensive and high-quality collection of math instruction fine-tuning data. We curated 150k queries from the math subset of Llama-Nemotron-Post-Training, which covers multiple domains of math questions.For each query, we used… See the full description on the dataset page: https://huggingface.co/datasets/Scale-or-Reason/math-reasoning-ift-pairs.textquestion-answering100K<n<1M8 likes651 downloads3mo agoHugging Face06SiloLink /ifc-bench IFC-Bench A benchmark dataset for evaluating BIM (Building Information Modeling) comprehension and reasoning capabilities in AI systems. Provides curated IFC models with question-answer pairs across 4 complexity categories for testing BIM-related AI implementations. Dataset snapshot: question ground_truth ifc_model project category 0 What modelling program and IFC standard were used to create this model? The model was created using... arc 4351 1 1 What are the… See the full description on the dataset page: https://huggingface.co/datasets/SiloLink/ifc-bench.documentquestion-answering1K<n<10K0 likes646 downloads25d agoHugging Face07quenfly /ifc-bench IFC-Bench A benchmark dataset for evaluating BIM (Building Information Modeling) comprehension and reasoning capabilities in AI systems. Provides curated IFC models with question-answer pairs across 4 complexity categories for testing BIM-related AI implementations. Dataset snapshot: question ground_truth ifc_model project category 0 What modelling program and IFC standard were used to create this model? The model was created using... arc 4351 1 1 What are the… See the full description on the dataset page: https://huggingface.co/datasets/quenfly/ifc-bench.documentquestion-answering1K<n<10K1 likes297 downloads2mo agoHugging Face08Dietmar2020 /ifc-bim-qa-dataset IFC BIM Question-Answering Dataset A comprehensive question-answering dataset for Building Information Modeling (BIM) and Industry Foundation Classes (IFC) domain knowledge. Dataset Summary This dataset contains 13,485 question-answer pairs covering comprehensive BIM domain knowledge: IFC Schema Knowledge: Entities, constraints, functions, and global rules IFC Documentation: Specifications, concepts, geometry, and processes Professional Certification: BIM practices… See the full description on the dataset page: https://huggingface.co/datasets/Dietmar2020/ifc-bim-qa-dataset.textquestion-answering10K<n<100K7 likes111 downloads1y agoHugging Face09BSC-LT /IFEval_es Dataset Card for IFEval_es IFEval_es is a prompt dataset in Spanish, professionally translated from the main version of the IFEval dataset in English. Dataset Details Dataset Description IFEval_es (Instruction-Following Eval benchmark - Spanish) is designed to evaluating chat or instruction fine-tuned language models. The dataset comprises 541 "verifiable instructions" such as "write in more than 400 words" and "mention the keyword of AI at least 3 times"… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/IFEval_es.textquestion-answeringn<1K1 likes107 downloads10mo agoHugging Face10IFthisisrealitynbds /hacker-news Hacker News - Complete Archive Every Hacker News item since 2006, live-updated every 5 minutes What is it? This dataset contains the complete Hacker News archive: every story, comment, Ask HN, Show HN, job posting, and poll ever submitted to the site. Hacker News is one of the longest-running and most influential technology communities on the internet, operated by Y Combinator since 2007. It has become the de facto gathering place for founders, engineers, researchers… See the full description on the dataset page: https://huggingface.co/datasets/IFthisisrealitynbds/hacker-news.tabulartext-generation10M<n<100M0 likes89 downloads6mo agoHugging Face11MorbidCorp /actuarial-fm-p-ifm-dataset Actuarial FM + P + IFM Dataset v0.0.7 Dataset Description Comprehensive training dataset for actuarial AI covering three SOA exams. Dataset Summary Total Examples: 18,794 Exam FM: ~18,000 examples Exam P: 743 examples Exam IFM: 37 examples Format: JSONL with instruction-response pairs Topics Covered Financial Mathematics (FM) Time value of money Annuities and perpetuities Bonds and interest theory Amortization Probability (P)… See the full description on the dataset page: https://huggingface.co/datasets/MorbidCorp/actuarial-fm-p-ifm-dataset.textquestion-answering10K<n<100K0 likes83 downloads11mo agoHugging Face12projecte-aina /IFEval_ca Dataset Card for IFEval_ca IFEval_ca is a prompt dataset in Catalan, professionally translated from the main version of the IFEval dataset in English. Dataset Details Dataset Description IFEval_ca (Instruction-Following Eval benchmark - Catalan) is designed to evaluating chat or instruction fine-tuned language models. The dataset comprises 541 "verifiable instructions" such as "write in more than 400 words" and "mention the keyword of AI at least 3 times"… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/IFEval_ca.textquestion-answeringn<1K0 likes80 downloads10mo agoHugging Face13jeevajoji /IFTtextquestion-answering100K<n<1M0 likes78 downloads2y agoHugging Face14MorbidCorp /actuarial-fm-p-ifm-ultimate-dataset Ultimate Actuarial FM/P/IFM Dataset v0.0.9 Dataset Description The ultimate training dataset for actuarial AI models, containing 1,708 meticulously crafted examples targeting 95%+ accuracy on professional actuarial exams. Dataset Statistics Total Examples: 1,708 Train: 1,366 (80%) Validation: 170 (10%) Test: 172 (10%) Distribution by Exam Exam Examples Percentage IFM 884 51.8% P 570 33.4% FM 254 14.9% Key Features… See the full description on the dataset page: https://huggingface.co/datasets/MorbidCorp/actuarial-fm-p-ifm-ultimate-dataset.texttext-generation1K<n<10K0 likes60 downloads11mo agoHugging Face15Dietmar2020 /ifc-bim-alpaca-124k IFC BIM Alpaca Dataset Dataset Description This dataset contains 124974 high-quality instruction-following examples for IFC (Industry Foundation Classes) and BIM (Building Information Modeling) in the Alpaca format. It covers IFC 4.3.x-devel standard entities, properties, relationships, and best practices. Dataset Summary Total Examples: 124974 Train Set: 112476 examples Validation Set: 12498 examples Language: English Format: Alpaca (instruction, input… See the full description on the dataset page: https://huggingface.co/datasets/Dietmar2020/ifc-bim-alpaca-124k.texttext-generation100K<n<1M1 likes48 downloads1y agoHugging Face16maikezu /data-kit-sub-iwslt2025-if-long-constraint Data for KIT’s Instruction Following Submission for IWSLT 2025 This repo contains the data used to train our model for IWSLT 2025's Instruction-Following (IF) Speech Processing track. IWSLT 2025's Instruction-Following (IF) Speech Processing track in the scientific domain aims to benchmark foundation models that can follow natural language instructions—an ability well-established in textbased LLMs but still emerging in speech-based counterparts. Our approach employs an end-to-end… See the full description on the dataset page: https://huggingface.co/datasets/maikezu/data-kit-sub-iwslt2025-if-long-constraint.textautomatic-speech-recognition0 likes44 downloads1y agoHugging Face17sonsdf /k-ifrs-qa-dataset K-IFRS QA Dataset 한국채택국제회계기준(K-IFRS) 기반의 오픈소스 QA 데이터셋입니다.LLM 파인튜닝(SFT), RAG 시스템 구축, 회계 도메인 벤치마크 평가 등 다양한 목적에 활용할 수 있도록 설계되었습니다. 기준 연도: 본 데이터셋은 2026년 5월 24일자 K-IFRS 기준으로 작성되었습니다.회계기준은 지속적으로 개정되므로, 사용 시 기준 연도를 반드시 확인하시기 바랍니다. 데이터셋 개요 항목 내용 총 데이터 수 34,418개 Train 분할 약 30,976개 (90%) Validation 분할 약 3,442개 (10%) 언어 한국어 형식 Instruction-Input-Output (Alpaca 형식) 기준 K-IFRS (2026년 5월 24일자) 라이선스 CC BY-NC-SA 4.0 데이터 구조 (Data Fields) 각 데이터는… See the full description on the dataset page: https://huggingface.co/datasets/sonsdf/k-ifrs-qa-dataset.texttext-generation10K<n<100K0 likes44 downloads4mo agoHugging Face18Dietmar2020 /ifc-bim-gemma3-subset-1k IFC-BIM Gemma3 Training Subset (1K Examples) A 1,000-example subset of IFC/BIM Q&A data formatted for Gemma-3 fine-tuning with Unsloth. Quick Start from datasets import load_dataset # Load dataset dataset = load_dataset("your-username/ifc-bim-gemma3-subset-1k") # View first example print(dataset["train"][0]) Dataset Structure ShareGPT format with quality scores: conversations: List of human/gpt exchanges source: Data origin score: Quality rating… See the full description on the dataset page: https://huggingface.co/datasets/Dietmar2020/ifc-bim-gemma3-subset-1k.texttext-generation1K<n<10K0 likes41 downloads1y agoHugging Face19Dietmar2020 /ifc-bim-high-quality-alpaca IFC BIM High-Quality Dataset (Alpaca Format) Dataset Description This is a high-quality, curated dataset for training language models on IFC (Industry Foundation Classes) and BIM (Building Information Modeling) tasks. The dataset has been filtered for quality and is provided in the Alpaca instruction-following format. Dataset Summary Total entries: 42,680 Format: Alpaca (instruction, input, output) Language: English Domain: IFC/BIM technical documentation and… See the full description on the dataset page: https://huggingface.co/datasets/Dietmar2020/ifc-bim-high-quality-alpaca.texttext-generation10K<n<100K1 likes41 downloads1y agoHugging Face20mindahu /NILE-IFT-DatasetHere are the IFT datasets for the EMNLP 2025 Main paper NILE. These include the Alpaca dataset (release_nile_alpaca_dataset.json) and the sampled OpenOrca dataset (release_nile_orca_dataset.json), both revised by the NILE framework. textquestion-answering10K<n<100K1 likes35 downloads1y agoHugging Face21CGIAR /ifpri-ai-documentsgated GAIA / GARDIAN-CIGI Agricultural Research Corpus This dataset contains 21,726 agricultural research documents extracted from the GARDIAN repository and processed through the CIGI pipeline. Dataset Overview Property Value Total Documents 21,726 Total Size 623.27 MB Total Tokens 85,359,442 Total Pages 0 Languages 25 Unique Keywords 7,127 Resource Types 20 Date Generated 2026-07-31 02:55:20 Language Distribution… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/ifpri-ai-documents.textsummarization10K<n<100K0 likes20 downloads1mo agoHugging Face22iffrce /Qwen3.5-reasoning-700x Dataset Card (Qwen3.5-reasoning-700x) Dataset Summary Qwen3.5-reasoning-700x is a high-quality distilled dataset. This dataset uses the high-quality instructions constructed by Alibaba-Superior-Reasoning-Stage2 as the seed question set. By calling the latest Qwen3.5-27B full-parameter model on the Alibaba Cloud DashScope platform as the teacher model, it generates high-quality responses featuring long-text reasoning processes (Chain-of-Thought). It covers several major… See the full description on the dataset page: https://huggingface.co/datasets/iffrce/Qwen3.5-reasoning-700x.textquestion-answeringn<1K0 likes17 downloads6mo agoHugging Face23AmanPriyanshu /reasoning-sft-IF_multi_constraints_upto5 reasoning-sft-IF_multi_constraints_upto5 Instruction-following dataset with multi-constraint prompts (up to 5 constraints), paired with reasoning responses generated. Format Each row has three columns: input — list of dicts [{"role": "user", "content": "..."}, ...] response — model response string (includes <think> reasoning block) category — constraint category label Usage import random import pyarrow.parquet as pq from huggingface_hub import hf_hub_download… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/reasoning-sft-IF_multi_constraints_upto5.texttext-generation10K<n<100K0 likes16 downloads7mo agoHugging Face24Ifraaaa /my-distiset-821455fd Dataset Card for my-distiset-821455fd This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/Ifraaaa/my-distiset-821455fd/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/Ifraaaa/my-distiset-821455fd.texttext-generationn<1K0 likes5 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.