datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
IFEval_es
Dataset Card for IFEval_es
IFEval_es is a prompt dataset in Spanish, professionally translated from the main version of the IFEval dataset in English.
Dataset Details
Dataset Description
IFEval_es (Instruction-Following Eval benchmark - Spanish) is designed to evaluating chat or instruction fine-tuned language models. The dataset comprises 541 "verifiable instructions" such as "write in more than 400 words" and "mention the keyword of AI at least 3 times"… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/IFEval_es.actuarial-fm-p-ifm-dataset
Actuarial FM + P + IFM Dataset v0.0.7
Dataset Description
Comprehensive training dataset for actuarial AI covering three SOA exams.
Dataset Summary
Total Examples: 18,794
Exam FM: ~18,000 examples
Exam P: 743 examples
Exam IFM: 37 examples
Format: JSONL with instruction-response pairs
Topics Covered
Financial Mathematics (FM)
Time value of money
Annuities and perpetuities
Bonds and interest theory
Amortization
Probability (P)… See the full description on the dataset page: https://huggingface.co/datasets/MorbidCorp/actuarial-fm-p-ifm-dataset.IFEval_ca
Dataset Card for IFEval_ca
IFEval_ca is a prompt dataset in Catalan, professionally translated from the main version of the IFEval dataset in English.
Dataset Details
Dataset Description
IFEval_ca (Instruction-Following Eval benchmark - Catalan) is designed to evaluating chat or instruction fine-tuned language models. The dataset comprises 541 "verifiable instructions" such as "write in more than 400 words" and "mention the keyword of AI at least 3 times"… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/IFEval_ca.actuarial-fm-p-ifm-ultimate-dataset
Ultimate Actuarial FM/P/IFM Dataset v0.0.9
Dataset Description
The ultimate training dataset for actuarial AI models, containing 1,708 meticulously crafted examples targeting 95%+ accuracy on professional actuarial exams.
Dataset Statistics
Total Examples: 1,708
Train: 1,366 (80%)
Validation: 170 (10%)
Test: 172 (10%)
Distribution by Exam
Exam
Examples
Percentage
IFM
884
51.8%
P
570
33.4%
FM
254
14.9%
Key Features… See the full description on the dataset page: https://huggingface.co/datasets/MorbidCorp/actuarial-fm-p-ifm-ultimate-dataset.ifc-bim-high-quality-alpaca
IFC BIM High-Quality Dataset (Alpaca Format)
Dataset Description
This is a high-quality, curated dataset for training language models on IFC (Industry Foundation Classes) and BIM (Building Information Modeling) tasks. The dataset has been filtered for quality and is provided in the Alpaca instruction-following format.
Dataset Summary
Total entries: 42,680
Format: Alpaca (instruction, input, output)
Language: English
Domain: IFC/BIM technical documentation and… See the full description on the dataset page: https://huggingface.co/datasets/Dietmar2020/ifc-bim-high-quality-alpaca.NILE-IFT-DatasetHere are the IFT datasets for the EMNLP 2025 Main paper NILE.
These include the Alpaca dataset (release_nile_alpaca_dataset.json) and the sampled OpenOrca dataset (release_nile_orca_dataset.json), both revised by the NILE framework.
Qwen3.5-reasoning-700x
Dataset Card (Qwen3.5-reasoning-700x)
Dataset Summary
Qwen3.5-reasoning-700x is a high-quality distilled dataset.
This dataset uses the high-quality instructions constructed by Alibaba-Superior-Reasoning-Stage2 as the seed question set. By calling the latest Qwen3.5-27B full-parameter model on the Alibaba Cloud DashScope platform as the teacher model, it generates high-quality responses featuring long-text reasoning processes (Chain-of-Thought). It covers several major… See the full description on the dataset page: https://huggingface.co/datasets/iffrce/Qwen3.5-reasoning-700x.
