datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dementor-matrix-responses
Dementor — matrix model responses
Generated model outputs for the Dementor LLM-imitation / behavioral-inertia study.
Companion to:
Code + prompt splits: https://github.com/lisadunlap/dementor (branch ethan)
Trained adapters (2,122 LoRAs): https://huggingface.co/dementor-research — SFT / DPO /
self-SFT, grouped into per-dataset collections (gsm8k, chatbot_arena, writingprompts, openassistant).
Dataset viewer. This repo is a nested tree of CSV tables plus per-cell cell.json… See the full description on the dataset page: https://huggingface.co/datasets/dementor-research/dementor-matrix-responses.Patient-Message-Response-DraftingPaper: How Much Would a Clinician Edit This Draft? Evaluating LLM Alignment for Patient Message Response Drafting (arxiv link)
Dataset Details:
The patient message response drafting dataset is designed to evaluate how well LLMs respond to patient messages in patient portal communication.
Each semi-synthetic patient message is paired with a real de-identified EHR from a patient at our collaborating hospital.
Each doctor response is written by a clinician, guided by clinician response themes… See the full description on the dataset page: https://huggingface.co/datasets/PortalPal-AI/Patient-Message-Response-Drafting.Contextual_Response_Evaluation_for_ESL_and_ASD_Support
Dataset Card for "Contextual Response Evaluation for ESL and ASD Support💜💬🌐""
Dataset Description 📖
Dataset Summary 📝
Curated by Eric Soderquist, this dataset is a collection of English prompts and responses generated by the Phi-2 model, designed to evaluate and improve NLP models for supporting ESL (English as a Second Language) and ASD (Autism Spectrum Disorder) user bases. Each prompt is paired with multiple AI-generated responses and evaluated using a… See the full description on the dataset page: https://huggingface.co/datasets/yunjaeys/Contextual_Response_Evaluation_for_ESL_and_ASD_Support.Party_Affairs_ResponseData from https://wenda.12371.cn/liebiao.php
System-Response-100K
System-Response-100K dataset
This dataset contains text and code for machine learning tasks including:
Text Generation
Text Classification
Summarization
Question Answering
The dataset includes text formatted in JSON and is in English.
Dataset Statistics
Number of entries: Not specified in the information you provided.
Modalities
Text
Code
Formats
JSON
Languages
English
Getting Started
This section can include… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/System-Response-100K.Trix-Chatbot-Prompt-Response
Dataset Creation Process
Overview
This dataset was created to train and evaluate a chatbot focused on answering questions about Pooria Roy, his background, projects, and related topics. The goal was to build a dataset grounded in real user behavior while maintaining sufficient diversity and coverage of edge cases.
The final dataset contains 2,105 prompt-response examples, including a small portion of multi-turn conversations.
Data Collection Pipeline… See the full description on the dataset page: https://huggingface.co/datasets/regularpooria/Trix-Chatbot-Prompt-Response.
