datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cannabis-fda-extractive-pilot
FDA Cannabis Extractive Experimental Pilot
Experimental, automatically screened, unreviewed draft dataset. This dataset is not medical advice, is not production-ready, and must not be represented as clinician-reviewed, legally cleared, or suitable for patient-facing systems.
This small English conversational dataset was created to test an auditable Gemma 4 fine-tuning pipeline. It contains exact answer passages from captured FDA pages about CBD/cannabis safety, paired with… See the full description on the dataset page: https://huggingface.co/datasets/aznatkoiny/cannabis-fda-extractive-pilot.TR-Extractive-QA-5K
Dataset Card for Dataset Name
The dataset consists of nearly 5000 {Context, Question, Answer} triplets in Turkish. It can be used in finetuning large language models for text-generation, masked language modeling, instruction following, and extractive question answering.
The dataset is a manually curated version of multiple Turkish QA-based datasets and some of the answers are arranged by hand.
financial-reports-extractive-summarization_eval
Financial Reports Extractive Summarization Evaluation Dataset
Validation and test splits for evaluating models on Arabic financial reports extractive summarization.
Dataset Structure
Format: Simple prompt-answer pairs
Validation: ~20 examples (10%)
Test: ~20 examples (10%)
Language: Arabic
Domain: Financial reports and market news
Fields
id: Unique identifier
prompt: The summarization prompt
full_text: Complete financial report
answer: Ground… See the full description on the dataset page: https://huggingface.co/datasets/SahmBenchmark/financial-reports-extractive-summarization_eval.financial-reports-extractive-summarization_train
Financial Reports Extractive Summarization Training Dataset
Training split of the Arabic financial reports extractive summarization dataset in conversational format.
Dataset Structure
Format: Conversational (human-agent pairs)
Size: ~160 training examples (80% of total)
Language: Arabic
Domain: Financial reports and market news
Features
id: Unique identifier
conversations: Human prompt and agent summary
report_type: Type of financial report… See the full description on the dataset page: https://huggingface.co/datasets/SahmBenchmark/financial-reports-extractive-summarization_train.
