datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fiqh_doa_RAFT_ds_v01
Dataset Card — RAFT Islamic QA (Bilingual: Indonesia - Arab)
Dataset Retrieval-Augmented Fine-Tuning (RAFT) berbahasa Indonesia dan Arab (bersumber dari kitab Minhaj ath-Thalibin karya Imam An-Nawawi untuk Fiqh Syafii, serta himpunan Doa Harian dan Ibadah Praktis) yang dikembangkan oleh AI Literacy Innovation Institute (ALII) UIN SYARIF HIDAYATULLAH JAKARTA. Dataset ini dirancang khusus untuk melatih Large Language Models (LLM) agar dapat menjawab pertanyaan seputar hukum Islam… See the full description on the dataset page: https://huggingface.co/datasets/ai-literacy-innovation-institute/fiqh_doa_RAFT_ds_v01.XitXat_Function_Calling
xitxat_fc
xitxat_fc is a dataset comprising simulated Catalan call center conversations. Each conversation is annotated with structured tool calls, making it suitable for training and evaluating language models with function-calling capabilities.
Dataset Details
Language: Catalan
Modality: Text
Format: JSON Lines (.jsonl)
License: CC BY 4.0
Size: Less than 1,000 examples
Source: Synthetic data generated for research purposes
Structure
Each entry in… See the full description on the dataset page: https://huggingface.co/datasets/langtech-innovation/XitXat_Function_Calling.Innovation_Leading_Creative_Teams_Practical
Innovation Leading Creative Teams — Practical
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Innovation_Leading_Creative_Teams_Practical.Innovation_Leading_Creative_Teams_Theory
Innovation Leading Creative Teams — Theory
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Innovation_Leading_Creative_Teams_Theory.busyday
