CoolFace
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01momahadi /bangladesh-legal-qa-dataset Bangladesh Legal QA Dataset: Bangla-English Law and Fine-Tuning The Bangladesh Legal QA Dataset is a bilingual Bangla-English dataset for Bangladesh law question answering, legal NLP, LLM fine-tuning, instruction tuning, and retrieval-augmented generation (RAG). It provides 2,165 context-grounded legal QA records, direct-answer and IRAC chat-format training data, and structured statutory text from six Bangladesh Acts and three schedules. This is the 2,165-record paper-aligned… See the full description on the dataset page: https://huggingface.co/datasets/momahadi/bangladesh-legal-qa-dataset.tabularquestion-answering1K<n<10K2 likes172 downloads25d agoHugging Face02crawlfeeds /Medical-Health-QA-Articles-Dataset Medical Health Q&A & Articles Dataset — iCliniq, HealthTap & WebMD A multi-source medical Q&A and health articles dataset combining doctor-answered questions and medically reviewed content from iCliniq, HealthTap, and WebMD. Built for LLM fine-tuning, medical chatbot training, clinical NLP research, and healthcare AI development. Dataset Overview Field Details Sources iCliniq, HealthTap, WebMD Total Records 1,000 (sample) — 50,000+ full dataset… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Medical-Health-QA-Articles-Dataset.imagetext-classification1K<n<10K0 likes38 downloads6mo agoHugging Face03david-sprague /Medical-Health-QA-Articles-Dataset Medical Health Q&A & Articles Dataset — iCliniq, HealthTap & WebMD A multi-source medical Q&A and health articles dataset combining doctor-answered questions and medically reviewed content from iCliniq, HealthTap, and WebMD. Built for LLM fine-tuning, medical chatbot training, clinical NLP research, and healthcare AI development. Dataset Overview Field Details Sources iCliniq, HealthTap, WebMD Total Records 1,000 (sample) — 50,000+ full dataset… See the full description on the dataset page: https://huggingface.co/datasets/david-sprague/Medical-Health-QA-Articles-Dataset.imagetext-classification1K<n<10K0 likes17 downloads4mo agoHugging Face04aslon1213 /orpheus_qa_datasetThis dataset is used to pretrain LLama-3.1-base using orpheus-tts recipe for TTS task. This dataset was preparede using this notebook: Colab Notebook tabulartext-generation10K<n<100K0 likes13 downloads7mo agoHugging Face05Magneto /qa-dataset-llm-judge-flattened Q&A Dataset - LLM-as-Judge Analyzed (Flattened) Dataset Description This dataset contains 5,008 high-quality question-answer pairs extracted from regulatory and policy documents, analyzed and quality-assessed using LLM-as-Judge methodology with parallel processing. Key Features Source: Official regulatory documents including policy directions, guidelines, and circulars Quality Assessment: Each Q&A pair evaluated by LLM-as-Judge on multiple criteria Answer… See the full description on the dataset page: https://huggingface.co/datasets/Magneto/qa-dataset-llm-judge-flattened.textquestion-answering1K<n<10K0 likes10 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.