datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sudoku-700kttest2policy-alignment-verification-dataset
Policy Alignment Verification Dataset
🌐 NAVI's Ecosystem 🌐
🌍 NAVI Platform – Dive into NAVI's full capabilities and explore how it ensures policy alignment and compliance.
🤗 NAVI-small-preview – Access the open-weights version of NAVI designed for policy verification.
📜 API Docs – Your starting point for integrating NAVI into your applications.
📝 Blogpost: Policy-Driven Safeguards Comparison – A deep dive into the challenges and solutions NAVI addresses.
✨… See the full description on the dataset page: https://huggingface.co/datasets/nace-ai/policy-alignment-verification-dataset.Open-Web-Math
Keiran Paster*, Marco Dos Santos*, Zhangir Azerbayev, Jimmy Ba
GitHub | ArXiv
| PDF
OpenWebMath is a dataset containing the majority of the high-quality, mathematical text from the internet. It is filtered and extracted from over 200B HTML files on Common Crawl down to a set of 6.3 million documents containing a total of 14.7B tokens. OpenWebMath is intended for use in pretraining and finetuning large language models.
You can download the dataset using Hugging Face:
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/Alignment-Lab-AI/Open-Web-Math.ARC-stuffsynthetic-bn-subseqoig-fixedExpert-Sudoku-100kaxolotl2
Axolotl
Axolotl is a tool designed to streamline the fine-tuning of various AI models, offering support for multiple configurations and architectures.
Features:
Train various Huggingface models such as llama, pythia, falcon, mpt
Supports fullfinetune, lora, qlora, relora, and gptq
Customize configurations using a simple yaml file or CLI overwrite
Load different dataset formats, use custom formats, or bring your own tokenized datasets
Integrated with xformer, flash attention, rope… See the full description on the dataset page: https://huggingface.co/datasets/Alignment-Lab-AI/axolotl2.formatted-ttest-datasetElitePersonasflan-embed-test2Lawyer-Instruct
Dataset Card for "Lawyer-Instruct"
Dataset Description
Dataset Summary
Lawyer-Instruct is a conversational dataset primarily in English, reformatted from the original LawyerChat dataset. It contains legal dialogue scenarios reshaped into an instruction, input, and expected output format. This reshaped dataset is ideal for supervised dialogue model training.
Dataset generated in part by dang/futures
Supported Tasks and Leaderboards… See the full description on the dataset page: https://huggingface.co/datasets/Alignment-Lab-AI/Lawyer-Instruct.llama-indexgpt4v-raw-chunksLmSys-pref-ft-splitStampyAI-alignment-data
AI Alignment Research Dataset
The AI Alignment Research Dataset is a collection of documents related to AI Alignment and Safety from various books, research papers, and alignment related blog posts. This is a work in progress. Components are still undergoing a cleaning process to be updated more regularly.
Sources
Here are the list of sources along with sample contents:
agentmodel
agisf - recommended readings from AGI Safety Fundamentals
aisafety.info - Stampy's FAQ… See the full description on the dataset page: https://huggingface.co/datasets/Alignment-Lab-AI/StampyAI-alignment-data.robofactory-camera-alignment-multiviewAIGSharegpt-sodaagentcodeCodeInterpreterData-sharegptPrompt-Injection-TestTHE-BLUEPRINT-FOR-AI-ALIGNMENTsynthetic-bn-shuffledvalsLawyer-chat
Dataset Description
Dataset Summary
LawyerChat is a multi-turn conversational dataset primarily in the English language, containing dialogues about legal scenarios. The conversations are in the format of an interaction between a client and a legal professional. The dataset is designed for training and evaluating models on conversational tasks like dialogue understanding, response generation, and more.
Supported Tasks and Leaderboards
dialogue-modeling: The… See the full description on the dataset page: https://huggingface.co/datasets/Alignment-Lab-AI/Lawyer-chat.prefdedupRPGuild-sharegpt-filteredStack-Exchange-Aprillibrispeech-codec-22khz
