datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bangladesh-law-professional
🇧🇩 Bangladesh Law Professional Dataset
A clean, instruction-tuned (Alpaca-style) question–answer dataset for
fine-tuning language models on Bangladesh law, in Bangla and English.
👤 Author & Contribution
Curated & built by
Sadat Sami (@Sadatsami)
Role
Dataset architect — collected, cleaned, filtered, reformatted and published
Motivation
Build a small-but-high-quality Bangla legal corpus to fine-tune a lightweight LLM (e.g. Qwen2.5-0.5B via… See the full description on the dataset page: https://huggingface.co/datasets/Sadatsami/bangladesh-law-professional.MSciNLI
MSciNLI: A Diverse Benchmark for Scientific Natural Language Inference
This repository contains the dataset for the NAACL 2024 paper "MSCINLI: A Diverse Benchmark for Scientific Natural Language Inference."
If you face any difficulties while downloading the dataset, raise an issue in this repository or contact us at msadat3@uic.edu.
For more details about the dataset, please visit: https://github.com/msadat3/MSciNLI
Citation
If you use this dataset, please cite our… See the full description on the dataset page: https://huggingface.co/datasets/sadat2307/MSciNLI.arabic_eou_sada_dataset
Arabic EOU SADA Dataset (Saudi Dialect)
414,053 conversational Arabic utterances annotated for End-of-Utterance (EOU) detectionStrong focus on natural Saudi dialect (خليجي / نجدي / حجازي)
Task
Binary classification:
label = 1 → End of speaker turn (EOU)
label = 0 → Speaker will continue
Columns
text: Arabic transcription
label: 0 or 1
silence_after_seconds: pause duration after this segment
split: train | validation | test (already included)… See the full description on the dataset page: https://huggingface.co/datasets/LordTenson/arabic_eou_sada_dataset.databricks-sft-15ksadadaOpenOrcaSADA_EOU
