datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Training-Ai-Islamic-Dataset
🕌 Training AI Islamic Dataset
18.7M passages from classical Islamic books spanning 1,400 years of scholarship.
Comprehensive collection of Islamic texts covering Quran, Hadith, Fiqh, Tafsir, Aqeedah, Seerah, Usul al-Fiqh, and Arabic Language — structured with scholarly metadata for RAG and LLM training.
📊 Dataset Structure
collections/: Categorized Islamic passages compressed in JSONL format.
metadata/: Scholarly master catalogs, author biographical death… See the full description on the dataset page: https://huggingface.co/datasets/hozifa1/Training-Ai-Islamic-Dataset.ai-job-searcher-training-data
AI Job Searcher Training Data
Fine-tuning dataset for a career advisor AI specializing in Nordic and European job markets.
Dataset Description
This dataset contains 1,040 training examples across 9 languages and 5 task categories, formatted as chat conversations (system/user/assistant) suitable for fine-tuning LLMs.
Task Categories
Category
Examples
Description
Cover Letter Generation
208
Professional cover letters from job description + user profile… See the full description on the dataset page: https://huggingface.co/datasets/ai-colombia/ai-job-searcher-training-data.chimera-training-data
Chimera Training Data
Training data for Project Chimera — research into widened I/O bandwidth for AI models.
The Goal
Train a single neural network to process multiple inputs and generate multiple outputs simultaneously. Not interleaving, not fast-switching — genuinely parallel cognitive streams from ONE brain.
┌─────────────────────────────────────────────┐
Input A ─┤ ├─ Output A
│ ONE BRAIN… See the full description on the dataset page: https://huggingface.co/datasets/AI-Foundation/chimera-training-data.prism-training-dataset
PRISM training dataset
This is the training dataset for PRISM: Recovering Instruction Sets from
Language Model Activations. Each record pairs an instruction-rich prompt with
a Qwen3.5-9B response and a generated list of the instructions in the prompt.
The released validity mask selects the records used to train the published
checkpoints.
Contents
Source key
Upstream dataset
Records
Source license
if_eval
google/IFEval
492
Apache-2.0… See the full description on the dataset page: https://huggingface.co/datasets/Offensive-AI-Lab/prism-training-dataset.
