datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sudoku-700kExpert-Sudoku-100kflan-embed-test2Lawyer-Instruct
Dataset Card for "Lawyer-Instruct"
Dataset Description
Dataset Summary
Lawyer-Instruct is a conversational dataset primarily in English, reformatted from the original LawyerChat dataset. It contains legal dialogue scenarios reshaped into an instruction, input, and expected output format. This reshaped dataset is ideal for supervised dialogue model training.
Dataset generated in part by dang/futures
Supported Tasks and Leaderboards… See the full description on the dataset page: https://huggingface.co/datasets/Alignment-Lab-AI/Lawyer-Instruct.llama-indexgpt4v-raw-chunksagentcodePrompt-Injection-TestLawyer-chat
Dataset Description
Dataset Summary
LawyerChat is a multi-turn conversational dataset primarily in the English language, containing dialogues about legal scenarios. The conversations are in the format of an interaction between a client and a legal professional. The dataset is designed for training and evaluating models on conversational tasks like dialogue understanding, response generation, and more.
Supported Tasks and Leaderboards
dialogue-modeling: The… See the full description on the dataset page: https://huggingface.co/datasets/Alignment-Lab-AI/Lawyer-chat.embeddedflanclaudeopus-sharegptGutenDialogue2ttestv0.1arrayofarrayGutenDialoguealpaca-cot-collectionHuman-Preferences-Alignment-KTO-Dataset-AI-Services-Genuine-User-Reviews
Human Preferences Alignment KTO Dataset of AI Service User Reviews of ChatGPT Gemini Claude Perplexity
Introduction to Human Preferences Alignment
There are many methods of applying Human Preference Alignment techniques to help model align in the supervised finetuning stage, including RLHF Reinforcement Learning from Human Feedback(paper), PPO Proximal policy optimization(paper/equation), DPO Direct Preference Optimization (paper/equation), KTO Kahneman-Tversky… See the full description on the dataset page: https://huggingface.co/datasets/DeepNLP/Human-Preferences-Alignment-KTO-Dataset-AI-Services-Genuine-User-Reviews.reversevalidated-python-instructprefsharegptseriesdeep-ai-safety-alignment-zh
Deep AI Safety & Alignment Dialogue Dataset (Chinese)
深度AI安全与对齐对话数据集
Dataset Description
High-quality Chinese AI safety and alignment dialogues covering existential alignment, value calibration, AI ethics, AGI safety, and harmful content detection.
高质量中文AI安全与对齐对话,涵盖存在主义对齐、价值观校准、AI伦理、AGI安全、有害内容检测等前沿议题。
Dataset Structure
Format: JSONL (JSON Lines)
Fields:
instruction: User message / question
input: Additional context (if any)
output:… See the full description on the dataset page: https://huggingface.co/datasets/AngelWarmSmile123/deep-ai-safety-alignment-zh.Long-sharegptAugmented-Generations-for-Intelligencealgorithmcosmosdatasci-pythongpt4andclaudechatknow-saraswati-alpaca-cot-by-knowrohitlinabotbench
DORA AI Evaluation Results
Generated on 2025-04-06 23:27:27
Test Configuration
API Endpoint: http://localhost:8080/v1/chat/completions
Model: Alignment-Lab-AI/linabot
Number of Questions: 182
Evaluation Framework: DORA AI Evaluation Framework
NLP Analysis Level: 0 (No NLP capabilities)
Executive Summary
Overall Regulatory Quality Score: 0.70/1.0
Performance by Dimension
Accuracy: 0.66/1.0 (factual correctness)
Completeness: 0.87/1.0… See the full description on the dataset page: https://huggingface.co/datasets/Alignment-Lab-AI/linabotbench.slm-workflow-planner-alignment-v2
