CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nace-ai /policy-alignment-verification-dataset Policy Alignment Verification Dataset 🌐 NAVI's Ecosystem 🌐 🌍 NAVI Platform – Dive into NAVI's full capabilities and explore how it ensures policy alignment and compliance. 🤗 NAVI-small-preview – Access the open-weights version of NAVI designed for policy verification. 📜 API Docs – Your starting point for integrating NAVI into your applications. 📝 Blogpost: Policy-Driven Safeguards Comparison – A deep dive into the challenges and solutions NAVI addresses. ✨… See the full description on the dataset page: https://huggingface.co/datasets/nace-ai/policy-alignment-verification-dataset.texttext-classificationn<1K4 likes845 downloads2y agoHugging Face02Alignment-Lab-AI /Open-Web-Math Keiran Paster*, Marco Dos Santos*, Zhangir Azerbayev, Jimmy Ba GitHub | ArXiv | PDF OpenWebMath is a dataset containing the majority of the high-quality, mathematical text from the internet. It is filtered and extracted from over 200B HTML files on Common Crawl down to a set of 6.3 million documents containing a total of 14.7B tokens. OpenWebMath is intended for use in pretraining and finetuning large language models. You can download the dataset using Hugging Face: from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/Alignment-Lab-AI/Open-Web-Math.text1M<n<10M5 likes583 downloads2y agoHugging Face03Alignment-Lab-AI /synthetic-bn-subseqtabular1M<n<10M0 likes405 downloads1y agoHugging Face04Alignment-Lab-AI /oig-fixedtext1M<n<10M1 likes404 downloads2y agoHugging Face05Alignment-Lab-AI /Expert-Sudoku-100ktabular100K<n<1M0 likes345 downloads2y agoHugging Face06Alignment-Lab-AI /formatted-ttest-datasettext1M<n<10M0 likes249 downloads2y agoHugging Face07Alignment-Lab-AI /flan-embed-test2text100K<n<1M0 likes229 downloads2y agoHugging Face08Alignment-Lab-AI /Lawyer-Instruct Dataset Card for "Lawyer-Instruct" Dataset Description Dataset Summary Lawyer-Instruct is a conversational dataset primarily in English, reformatted from the original LawyerChat dataset. It contains legal dialogue scenarios reshaped into an instruction, input, and expected output format. This reshaped dataset is ideal for supervised dialogue model training. Dataset generated in part by dang/futures Supported Tasks and Leaderboards… See the full description on the dataset page: https://huggingface.co/datasets/Alignment-Lab-AI/Lawyer-Instruct.text1K<n<10K19 likes214 downloads3y agoHugging Face09AlignmentLab-AI /llama-indextext1K<n<10K3 likes182 downloads3y agoHugging Face10AlignmentLab-AI /gpt4v-raw-chunksimage100K<n<1M0 likes169 downloads3y agoHugging Face11Alignment-Lab-AI /LmSys-pref-ft-splittext1K<n<10K0 likes145 downloads2y agoHugging Face12Alignment-Lab-AI /Sharegpt-sodatext100K<n<1M2 likes110 downloads2y agoHugging Face13AlignmentLab-AI /agentcodetext100K<n<1M10 likes107 downloads3y agoHugging Face14Alignment-Lab-AI /CodeInterpreterData-sharegpttext10K<n<100K1 likes105 downloads2y agoHugging Face15Alignment-Lab-AI /Prompt-Injection-Testtext1K<n<10K0 likes102 downloads2y agoHugging Face16The-Nature-of-Reality /THE-BLUEPRINT-FOR-AI-ALIGNMENTaudion<1K3 likes102 downloads1y agoHugging Face17Alignment-Lab-AI /Lawyer-chat Dataset Description Dataset Summary LawyerChat is a multi-turn conversational dataset primarily in the English language, containing dialogues about legal scenarios. The conversations are in the format of an interaction between a client and a legal professional. The dataset is designed for training and evaluating models on conversational tasks like dialogue understanding, response generation, and more. Supported Tasks and Leaderboards dialogue-modeling: The… See the full description on the dataset page: https://huggingface.co/datasets/Alignment-Lab-AI/Lawyer-chat.text1K<n<10K11 likes95 downloads3y agoHugging Face18Alignment-Lab-AI /prefdeduptabular10M<n<100M2 likes93 downloads2y agoHugging Face19Alignment-Lab-AI /RPGuild-sharegpt-filteredtext10K<n<100K2 likes86 downloads2y agoHugging Face20Alignment-Lab-AI /Stack-Exchange-Apriltabular1M<n<10M7 likes86 downloads2y agoHugging Face21Alignment-Lab-AI /librispeech-codec-22khzaudio10K<n<100K0 likes80 downloads9mo agoHugging Face22Alignment-Lab-AI /embeddedflantext100K<n<1M0 likes78 downloads2y agoHugging Face23Alignment-Lab-AI /claudeopus-sharegpttext10K<n<100K4 likes78 downloads2y agoHugging Face24Alignment-Lab-AI /riddler-sharegpttext1K<n<10K0 likes73 downloads2y agoHugging Face25Alignment-Lab-AI /GutenDialogue2text1M<n<10M1 likes72 downloads2y agoHugging Face26Alignment-Lab-AI /ttestv0.1tabular1M<n<10M1 likes72 downloads2y agoHugging Face27Alignment-Lab-AI /StackStoretext1M<n<10M2 likes69 downloads1y agoHugging Face28Alignment-Lab-AI /GutenDialoguetext1M<n<10M1 likes67 downloads2y agoHugging Face29AlignmentLab-AI /alpaca-cot-collectiontext1M<n<10M9 likes65 downloads3y agoHugging Face30DeepNLP /Human-Preferences-Alignment-KTO-Dataset-AI-Services-Genuine-User-Reviews Human Preferences Alignment KTO Dataset of AI Service User Reviews of ChatGPT Gemini Claude Perplexity Introduction to Human Preferences Alignment There are many methods of applying Human Preference Alignment techniques to help model align in the supervised finetuning stage, including RLHF Reinforcement Learning from Human Feedback(paper), PPO Proximal policy optimization(paper/equation), DPO Direct Preference Optimization (paper/equation), KTO Kahneman-Tversky… See the full description on the dataset page: https://huggingface.co/datasets/DeepNLP/Human-Preferences-Alignment-KTO-Dataset-AI-Services-Genuine-User-Reviews.textn<1K2 likes52 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.