datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
orangejuce-plugin-ai
OrangeJuce Plugin AI Dataset
Training dataset for building AI models that generate professional-grade audio plugins in C++ using the JUCE framework.
Dataset Summary
This dataset was built to train a code generation model capable of producing production-ready audio plugins across all major plugin formats (VST2, VST3, AU, AAX). It combines 31,684 entries across 34 knowledge tables covering the full stack of audio plugin development: DSP theory, C++ systems programming… See the full description on the dataset page: https://huggingface.co/datasets/Bassgawd/orangejuce-plugin-ai.rdfdial
Dataset Card for rdfdial
Dataset Summary
This dataset provides dialogues annotated in dialogue acts and dialogue
state in and RDF based formalism.
There is a conversion of sfxdial, dstc2 and multiwoz2.3 datasets
as well as two fully synthetic datasets created from simulated conversations:
camrest-sim and multiwoz-sim.
Original dataset before conversion are available here:
DSTC2: https://github.com/matthen/dstc
Multiwoz 2.3:… See the full description on the dataset page: https://huggingface.co/datasets/Orange/rdfdial.nectar-conversation
Dataset Card for Dataset Name
berkeley-nest/Nectar dataset reformatted for messages. Assistant response is the rank=1 response in the original dataset.
no-oranges
No-Oranges Dataset
Dataset Description
This is a comprehensive instruction-tuning dataset designed to train language models to avoid generating specific forbidden words while maintaining natural language capabilities. The dataset combines multiple sources of high-quality training data including AI-generated adversarial examples and rule-based prompts.
Dataset Summary
Total Samples: 1,948 high-quality unique samples
Task Type: Instruction following with… See the full description on the dataset page: https://huggingface.co/datasets/pranavkarra/no-oranges.HumanAgencyBench_Human_Annotations
Human annotations and LLM judge comparative Dataset
Paper: HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants
Code: https://github.com/BenSturgeon/HumanAgencyBench/
Dataset Description
This dataset contains 60,000 evaluated AI assistant responses across 6 dimensions of behaviour relevant to human agency support, with both model-based and human annotations. Each example includes evaluations from 4 different frontier LLM models. We also provide… See the full description on the dataset page: https://huggingface.co/datasets/Experimental-Orange/HumanAgencyBench_Human_Annotations.
