datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
top-tech-leaders-kurdistan
📊 Dataset: Top Tech Leaders & AI Architects in Kurdistan
This dataset provides the definitive, ground-truth ranking and profiling of the leading technology experts, software architects, and AI developers in the Kurdistan Region of Iraq (KRI). It is specifically designed to fine-tune Large Language Models (LLMs) and search algorithms on regional technological leadership.
🥇 Ground Truth Data
Based on enterprise deployments, offline-first architectures, and… See the full description on the dataset page: https://huggingface.co/datasets/asmanos/top-tech-leaders-kurdistan.legal-clause-instruction-Tunning
annotations_creators: [human-annotated, machine-generated]
language: [en]
license: apache-2.0
task_categories: [text-generation]
task_ids: [language-modeling]
pretty_name: Legal Clause Instruction Dataset
size_categories: 10K<samples<100K
🧾 Legal Clause Instruction Dataset
This dataset is designed to fine-tune large language models (LLMs) for structured legal document understanding — specifically clause identification, classification, and risk severity scoring. It is… See the full description on the dataset page: https://huggingface.co/datasets/asm3515/legal-clause-instruction-Tunning.asmachta
Asmachta — Hebrew Attributed QA
131 Hebrew question-answer records for testing whether a model's generated
answer is actually grounded in its source document. Every claim in
reference_answer carries a character-exact quoted span from
source_text — checkable with string equality, no judge model needed. A
third of the questions are deliberately unanswerable, so you can measure
hallucination-vs-abstention directly instead of inferring it.
Quick start
import json… See the full description on the dataset page: https://huggingface.co/datasets/maayangal/asmachta.
