datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Shifaa_Arabic_Medical_Consultations
Shifaa Arabic Medical Consultations 🏥📊
Overview 🌍
Shifaa is revolutionizing Arabic medical AI by addressing the critical gap in Arabic medical datasets. Our first contribution is the Shifaa Arabic Medical Consultations dataset, a comprehensive collection of 84,422 real-world medical consultations covering 16 Main Specializations and 585 Hierarchical Diagnoses.
🔍 Why is this dataset important?
First large-scale Arabic medical dataset for AI applications.… See the full description on the dataset page: https://huggingface.co/datasets/Ahmed-Selem/Shifaa_Arabic_Medical_Consultations.IMDb-Media
Dataset Card for "BrightData/IMDb-Media"
Dataset Summary
Explore feature films, TV series, episodes, mini-series, documentaries, and more with this IMDb dataset, comprising over 249K structured records and 32 data fields updated and refreshed regularly.
Each entry includes all major data points such as timestamp, title, URLs, release date, IMDb rating, reviews, awards, origin, category/genre, budget, cast, director, images, videos and more.
For a complete list of data… See the full description on the dataset page: https://huggingface.co/datasets/BrightData/IMDb-Media.Meditation-miniset-v0.2
Synthetic Meditation Dataset v0.2 🧘♂️🧘♀️✨
Welcome to the Synthetic Meditation Dataset v0.2, a comprehensive collection of meditation guidance prompts designed to assist users in different stages of emotional wellbeing and mindfulness. This dataset aims to help developers train AI models to provide empathetic and personalized meditation guidance. The dataset focuses on inclusivity, personalization, and diverse meditation contexts to cater to a wide audience.
Overview… See the full description on the dataset page: https://huggingface.co/datasets/BuildaByte/Meditation-miniset-v0.2.medium_512_1k_tokens_prompts
Medium 512-1K Tokens Prompts Dataset
Created by Aipresso LIMITED, London, UK
⚠️ By using this dataset you agree to our Terms of Use.
Overview
703 high-quality English prompts whose length lies between 512 and 1 000 tokens.Every prompt has been de-duplicated, cleaned and token-counted with the GPT-2 tokenizer.
Statistics
Rows
Token range
File size
Format
703
512 – 1 000
2.9 MB
CSV
Use-cases
Medium-context language-model fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/Aipresso/medium_512_1k_tokens_prompts.Medical-Health-QA-Articles-Dataset
Medical Health Q&A & Articles Dataset — iCliniq, HealthTap & WebMD
A multi-source medical Q&A and health articles dataset combining doctor-answered questions and medically reviewed content from iCliniq, HealthTap, and WebMD. Built for LLM fine-tuning, medical chatbot training, clinical NLP research, and healthcare AI development.
Dataset Overview
Field
Details
Sources
iCliniq, HealthTap, WebMD
Total Records
1,000 (sample) — 50,000+ full dataset… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Medical-Health-QA-Articles-Dataset.Medi-Science
Medi-Science Dataset
The Medi-Science dataset is a comprehensive collection of medical Q&A data designed for text generation, question answering, and summarization tasks in the healthcare domain.
Dataset Overview
Name: Medi-Science
License: Apache-2.0
Languages: English
Tags: Medical, Medicine, Anomaly, Biology, Medi-Science
Number of Rows: 16,412
Dataset Size:
Downloaded: 22.7 MB
Auto-converted Parquet: 8.94 MB
Dataset Structure
The dataset contains the… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Medi-Science.Medium-Articles-Corpus
Medium Articles Corpus (10K Sample)
The Medium Articles Corpus is a massive, clean dataset of articles scraped from Medium.com. This sample version contains 10,000 articles + and is designed to showcase the quality and structure of the full corpus for researchers and developers.
This is the subset from the large dataset https://crawlfeeds.com/websites/medium/text_data/medium_articles
Dataset Features
This dataset includes the following key features, provided in a… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Medium-Articles-Corpus.HealthRisk-1500-Medical-Risk-Prediction
🏥 HealthRisk-1500: Medical Risk Prediction Dataset
📌 Overview
HealthRisk-1500 is a real-world patient risk prediction dataset designed for training NLP models, LLMs, and healthcare AI systems. This dataset includes 1,500 unique patient records, covering a wide range of symptoms, medical histories, lab reports, and risk levels. It is ideal for predictive analytics, medical text processing, and clinical decision support.
🔍 Use Cases
🩺 Disease Risk… See the full description on the dataset page: https://huggingface.co/datasets/lvimuth/HealthRisk-1500-Medical-Risk-Prediction.PubMed-Bilingual-Medical-Sample-EN-ID
🏥 PubMed Bilingual Medical Sample (EN-ID)
Providing premium, high-quality English-Indonesian bilingual medical datasets for AI, NLP, and Machine Learning research.
📥 Download Free Sample
You can directly access and download the free dataset sample (.csv format) from our repository files here:
⬇️ Download Free Sample File (tree/main)
🚀 Upgrade to Full Version (Volume 1)
This repository contains a free sample of our industry-grade parallel… See the full description on the dataset page: https://huggingface.co/datasets/IndoHealth-NLP/PubMed-Bilingual-Medical-Sample-EN-ID.IMDb-Media
Dataset Card for "BrightData/IMDb-Media"
Dataset Summary
Explore feature films, TV series, episodes, mini-series, documentaries, and more with this IMDb dataset, comprising over 249K structured records and 32 data fields updated and refreshed regularly.
Each entry includes all major data points such as timestamp, title, URLs, release date, IMDb rating, reviews, awards, origin, category/genre, budget, cast, director, images, videos and more.
For a complete list of… See the full description on the dataset page: https://huggingface.co/datasets/RyanHalliwell/IMDb-Media.tokenized-IELTS-writing-task-2-evaluation-DialoGPT-mediumMedical-Health-QA-Articles-Dataset
Medical Health Q&A & Articles Dataset — iCliniq, HealthTap & WebMD
A multi-source medical Q&A and health articles dataset combining doctor-answered questions and medically reviewed content from iCliniq, HealthTap, and WebMD. Built for LLM fine-tuning, medical chatbot training, clinical NLP research, and healthcare AI development.
Dataset Overview
Field
Details
Sources
iCliniq, HealthTap, WebMD
Total Records
1,000 (sample) — 50,000+ full dataset… See the full description on the dataset page: https://huggingface.co/datasets/david-sprague/Medical-Health-QA-Articles-Dataset.MedQuad-MedicalQnADataset_128tokens_max
Reference : "A Question-Entailment Approach to Question Answering". Asma Ben Abacha and Dina Demner-Fushman. BMC Bioinformatics, 2019."
This is an update of Keivalya Pandya's dataset (keivalya/MedQuad-MedicalQnADataset).
Content
There are medical questions and corresponding responses in a prompt format for chat or instruct model types
In order to fine tuned LLM with small HW (1 or 2 GPU with 14 Go)
Rows above 128 tokens have been deleted.
Rows have been truncated to a line break or a… See the full description on the dataset page: https://huggingface.co/datasets/Laurent1/MedQuad-MedicalQnADataset_128tokens_max.Processed_Medicalbot_Datasetreasoning-prompt-ko
a.k.a. Awesome ChatGPT Prompts
This is a Dataset Repository mirror of prompts.chat — a social platform for AI prompts.
📢 Notice
This Hugging Face dataset is a mirror. For the latest prompts, features, and community contributions, please visit:
🌐 Website: prompts.chat
📦 GitHub: github.com/f/awesome-chatgpt-prompts
About
prompts.chat is an open-source platform where users can share, discover, and collect AI prompts from the community. The project can be… See the full description on the dataset page: https://huggingface.co/datasets/Mediaapel/reasoning-prompt-ko.
