CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Alignment-Lab-AI /sudoku-700k1M<n<10M0 likes1.2k downloads2y agoHugging Face02Alignment-Lab-AI /ttest20 likes895 downloads2y agoHugging Face03nace-ai /policy-alignment-verification-dataset Policy Alignment Verification Dataset 🌐 NAVI's Ecosystem 🌐 🌍 NAVI Platform – Dive into NAVI's full capabilities and explore how it ensures policy alignment and compliance. 🤗 NAVI-small-preview – Access the open-weights version of NAVI designed for policy verification. 📜 API Docs – Your starting point for integrating NAVI into your applications. 📝 Blogpost: Policy-Driven Safeguards Comparison – A deep dive into the challenges and solutions NAVI addresses. ✨… See the full description on the dataset page: https://huggingface.co/datasets/nace-ai/policy-alignment-verification-dataset.texttext-classificationn<1K4 likes873 downloads2y agoHugging Face04Alignment-Lab-AI /Open-Web-Math Keiran Paster*, Marco Dos Santos*, Zhangir Azerbayev, Jimmy Ba GitHub | ArXiv | PDF OpenWebMath is a dataset containing the majority of the high-quality, mathematical text from the internet. It is filtered and extracted from over 200B HTML files on Common Crawl down to a set of 6.3 million documents containing a total of 14.7B tokens. OpenWebMath is intended for use in pretraining and finetuning large language models. You can download the dataset using Hugging Face: from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/Alignment-Lab-AI/Open-Web-Math.text1M<n<10M5 likes615 downloads2y agoHugging Face05Alignment-Lab-AI /ARC-stuff0 likes518 downloads2y agoHugging Face06Alignment-Lab-AI /oig-fixedtext1M<n<10M1 likes403 downloads2y agoHugging Face07Alignment-Lab-AI /synthetic-bn-subseqtabular1M<n<10M0 likes403 downloads1y agoHugging Face08Alignment-Lab-AI /Expert-Sudoku-100ktabular100K<n<1M0 likes347 downloads2y agoHugging Face09Alignment-Lab-AI /formatted-ttest-datasettext1M<n<10M0 likes319 downloads2y agoHugging Face10Alignment-Lab-AI /axolotl2 Axolotl Axolotl is a tool designed to streamline the fine-tuning of various AI models, offering support for multiple configurations and architectures. Features: Train various Huggingface models such as llama, pythia, falcon, mpt Supports fullfinetune, lora, qlora, relora, and gptq Customize configurations using a simple yaml file or CLI overwrite Load different dataset formats, use custom formats, or bring your own tokenized datasets Integrated with xformer, flash attention, rope… See the full description on the dataset page: https://huggingface.co/datasets/Alignment-Lab-AI/axolotl2.0 likes308 downloads2y agoHugging Face11Alignment-Lab-AI /flan-embed-test2text100K<n<1M0 likes227 downloads2y agoHugging Face12Alignment-Lab-AI /Lawyer-Instruct Dataset Card for "Lawyer-Instruct" Dataset Description Dataset Summary Lawyer-Instruct is a conversational dataset primarily in English, reformatted from the original LawyerChat dataset. It contains legal dialogue scenarios reshaped into an instruction, input, and expected output format. This reshaped dataset is ideal for supervised dialogue model training. Dataset generated in part by dang/futures Supported Tasks and Leaderboards… See the full description on the dataset page: https://huggingface.co/datasets/Alignment-Lab-AI/Lawyer-Instruct.text1K<n<10K19 likes207 downloads3y agoHugging Face13AlignmentLab-AI /llama-indextext1K<n<10K3 likes148 downloads3y agoHugging Face14AlignmentLab-AI /gpt4v-raw-chunksimage100K<n<1M0 likes145 downloads3y agoHugging Face15Alignment-Lab-AI /LmSys-pref-ft-splittext1K<n<10K0 likes145 downloads2y agoHugging Face16cjc0013 /ouroboros-ai-safety-control-beyond-alignment Control Beyond Alignment A Systems-Safety Comparison of Ouroboros with Contemporary AI Risk Management and Frontier-Safety Practice This private preview contains a publication-ready AI safety white paper authored by Ouroboros. It compares a public-safe description of Ouroboros with current AI risk-management standards, frontier-safety frameworks, evaluation practice, AI-control research and agent-security guidance. Main argument Model alignment is… See the full description on the dataset page: https://huggingface.co/datasets/cjc0013/ouroboros-ai-safety-control-beyond-alignment.0 likes126 downloads4d agoHugging Face17Alignment-Lab-AI /StampyAI-alignment-data AI Alignment Research Dataset The AI Alignment Research Dataset is a collection of documents related to AI Alignment and Safety from various books, research papers, and alignment related blog posts. This is a work in progress. Components are still undergoing a cleaning process to be updated more regularly. Sources Here are the list of sources along with sample contents: agentmodel agisf - recommended readings from AGI Safety Fundamentals aisafety.info - Stampy's FAQ… See the full description on the dataset page: https://huggingface.co/datasets/Alignment-Lab-AI/StampyAI-alignment-data.question-answering10K<n<100K0 likes115 downloads2y agoHugging Face18AlignmentLab-AI /agentcodetext100K<n<1M10 likes112 downloads3y agoHugging Face19zeno-ai /robofactory-camera-alignment-multiview0 likes109 downloads2mo agoHugging Face20Alignment-Lab-AI /Prompt-Injection-Testtext1K<n<10K0 likes107 downloads2y agoHugging Face21Alignment-Lab-AI /Lawyer-chat Dataset Description Dataset Summary LawyerChat is a multi-turn conversational dataset primarily in the English language, containing dialogues about legal scenarios. The conversations are in the format of an interaction between a client and a legal professional. The dataset is designed for training and evaluating models on conversational tasks like dialogue understanding, response generation, and more. Supported Tasks and Leaderboards dialogue-modeling: The… See the full description on the dataset page: https://huggingface.co/datasets/Alignment-Lab-AI/Lawyer-chat.text1K<n<10K11 likes103 downloads3y agoHugging Face22The-Nature-of-Reality /THE-BLUEPRINT-FOR-AI-ALIGNMENTaudion<1K3 likes102 downloads1y agoHugging Face23Alignment-Lab-AI /AIG0 likes99 downloads2y agoHugging Face24Alignment-Lab-AI /prefdeduptabular10M<n<100M2 likes96 downloads2y agoHugging Face25Alignment-Lab-AI /ElitePersonas1 likes95 downloads1y agoHugging Face26Alignment-Lab-AI /synthetic-bn-shuffledvalstabular1M<n<10M0 likes95 downloads1y agoHugging Face27Alignment-Lab-AI /librispeech-codec-22khzaudio10K<n<100K0 likes93 downloads9mo agoHugging Face28visv-Bro /brain-ai-alignment-reproduction Brain–AI alignment reproduction artifacts Full-scale statistical audit of “Alignment between Brains and AI: Evidence for Convergent Evolution across Modalities, Scales and Training Trajectories.” Paper: https://arxiv.org/abs/2507.01966 Audited source commit: https://github.com/FloyedShen/BrainAlign/tree/f3e8c78ed3ee5ac15ae7c76053665bbec5ac8a4d Reproduction Job: https://huggingface.co/jobs/visv-Bro/6a6f72c86b79c09949c1f7f1 Job staging Bucket:… See the full description on the dataset page: https://huggingface.co/datasets/visv-Bro/brain-ai-alignment-reproduction.imagen<1K0 likes89 downloads2mo agoHugging Face29Alignment-Lab-AI /Stack-Exchange-Apriltabular1M<n<10M7 likes86 downloads2y agoHugging Face30Alignment-Lab-AI /snac-stuff1M<n<10M0 likes73 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.