CoolFace
28 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Bassgawd /orangejuce-plugin-ai OrangeJuce Plugin AI Dataset Training dataset for building AI models that generate professional-grade audio plugins in C++ using the JUCE framework. Dataset Summary This dataset was built to train a code generation model capable of producing production-ready audio plugins across all major plugin formats (VST2, VST3, AU, AAX). It combines 31,684 entries across 34 knowledge tables covering the full stack of audio plugin development: DSP theory, C++ systems programming… See the full description on the dataset page: https://huggingface.co/datasets/Bassgawd/orangejuce-plugin-ai.texttext-generation10K<n<100K0 likes214 downloads6mo agoHugging Face02Orange /rdfdial Dataset Card for rdfdial Dataset Summary This dataset provides dialogues annotated in dialogue acts and dialogue state in and RDF based formalism. There is a conversion of sfxdial, dstc2 and multiwoz2.3 datasets as well as two fully synthetic datasets created from simulated conversations: camrest-sim and multiwoz-sim. Original dataset before conversion are available here: DSTC2: https://github.com/matthen/dstc Multiwoz 2.3:… See the full description on the dataset page: https://huggingface.co/datasets/Orange/rdfdial.texttext-generation10K<n<100K1 likes163 downloads3y agoHugging Face03Orange /KGConv KGConv, a Conversational Corpus grounded in Wikidata Dataset Summary KGConv is a large corpus of 71k english conversations where each question-answer pair is grounded in a Wikidata fact. The conversations were generated automatically: in particular, questions were created using a collection of 10,355 templates; subsequently, the naturalness of conversations was improved by inserting ellipses and coreference into questions, via both handcrafted rules and a generative… See the full description on the dataset page: https://huggingface.co/datasets/Orange/KGConv.text100K<n<1M1 likes155 downloads2y agoHugging Face04Orange /ecml_arena_dataset ARENA: A Cognitive Multi-Agent Framework for Modeling Conflict-Driven Multi-party Conversation ⚠️ Code under internal review. The generation / simulation code is currently under internal code review — the GitHub repository is coming soon. This repository already provides the dataset (a sample subset) so it can be referenced from the paper. 📄 Paper. ARENA: A Cognitive Multi-Agent Framework for Modeling Conflict-Driven Multi-party Conversation — ECML-PKDD. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Orange/ecml_arena_dataset.tabular1K<n<10K0 likes102 downloads4mo agoHugging Face05orange67 /dataset_fenics_experiment_v7text1K<n<10K2 likes56 downloads1y agoHugging Face06orangetin /nectar-conversation Dataset Card for Dataset Name berkeley-nest/Nectar dataset reformatted for messages. Assistant response is the rank=1 response in the original dataset. texttext-generation100K<n<1M1 likes39 downloads2y agoHugging Face07orange99087 /ProfBench Dataset Description: Leaderboard | Blog | Paper | Data | Code | Nemo Evaluator SDK More than 3000 rubric criteria across 40 human-annotated tasks presenting reports addressing professional tasks across PhD STEM (Chemistry, Physics) and Professional Services (Financial Services, Management Consulting) domains. This dataset is ready for commercial/non-commercial use. Dataset Owner(s): NVIDIA Corporation Dataset Creation Date: 9/24/2025 License/Terms of… See the full description on the dataset page: https://huggingface.co/datasets/orange99087/ProfBench.documentn<1K0 likes39 downloads8mo agoHugging Face08Experimental-Orange /persona-belief-probestext10K<n<100K0 likes34 downloads4mo agoHugging Face09Orange /PersonasForSalesbotPersonas for Salesbot textn<1K0 likes32 downloads1y agoHugging Face10pranavkarra /no-oranges No-Oranges Dataset Dataset Description This is a comprehensive instruction-tuning dataset designed to train language models to avoid generating specific forbidden words while maintaining natural language capabilities. The dataset combines multiple sources of high-quality training data including AI-generated adversarial examples and rule-based prompts. Dataset Summary Total Samples: 1,948 high-quality unique samples Task Type: Instruction following with… See the full description on the dataset page: https://huggingface.co/datasets/pranavkarra/no-oranges.texttext-generation1K<n<10K0 likes29 downloads1y agoHugging Face11orange101 /DeepRubric-datasettext1K<n<10K0 likes29 downloads3mo agoHugging Face12Experimental-Orange /HumanAgencyBench_Human_Annotations Human annotations and LLM judge comparative Dataset Paper: HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants Code: https://github.com/BenSturgeon/HumanAgencyBench/ Dataset Description This dataset contains 60,000 evaluated AI assistant responses across 6 dimensions of behaviour relevant to human agency support, with both model-based and human annotations. Each example includes evaluations from 4 different frontier LLM models. We also provide… See the full description on the dataset page: https://huggingface.co/datasets/Experimental-Orange/HumanAgencyBench_Human_Annotations.texttext-generation10K<n<100K0 likes27 downloads1y agoHugging Face13OrangyDev /gmdtextn<1K0 likes25 downloads6mo agoHugging Face14giuliadc /orangesum_filtered_new_spacesOrangeSum dataset filtered by using the code by Aumiller et al. (1) available at https://github.com/dennlinger/summaries/tree/main min_length_summary = 18; min_length_reference = 250; length_metric = "whitespace" bi-gram_overlap_fraction between summary and original text < 0.65, meaning that all summaries in the dataset are on the abstractive side Furthermore: both in articles and in summaries, every point (".") followed by a capital letter was replaced by a point followed by a space and the… See the full description on the dataset page: https://huggingface.co/datasets/giuliadc/orangesum_filtered_new_spaces.textsummarization1K<n<10K0 likes20 downloads2y agoHugging Face15open-llm-leaderboard /DreadPoor__OrangeJ-8B-Model_Stock-detailsgated Dataset Card for Evaluation run of DreadPoor/OrangeJ-8B-Model_Stock Dataset automatically created during the evaluation run of model DreadPoor/OrangeJ-8B-Model_Stock The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DreadPoor__OrangeJ-8B-Model_Stock-details.tabular10K<n<100K0 likes16 downloads2y agoHugging Face16orangetin /SlimOrca-Convotext100K<n<1M0 likes15 downloads3y agoHugging Face17orange67 /alpaca-fenics-dataset2textn<1K0 likes13 downloads2y agoHugging Face18orangetin /oig-chiptext100K<n<1M0 likes11 downloads3y agoHugging Face19orangetin /oig-jokestextn<1K0 likes11 downloads3y agoHugging Face20OrangeYouSad /m-training-data2text1K<n<10K0 likes10 downloads5mo agoHugging Face21giuliadc /orangesum_5kOrangeSum dataset filtered by using the code by Aumiller et al. (1) available at https://github.com/dennlinger/summaries/tree/main min_length_summary = 18; min_length_reference = 250; length_metric = "whitespace" bi-gram_overlap_fraction between summary and original text < 0.65, meaning that all summaries in the dataset are on the abstractive side Furthermore: both in articles and in summaries, every point (".") followed by a capital letter was replaced by a point followed by a space and the… See the full description on the dataset page: https://huggingface.co/datasets/giuliadc/orangesum_5k.textsummarization1K<n<10K0 likes8 downloads2y agoHugging Face22Johncmk /orange Dataset Card for Orange Animal Dataset Dataset Summary This dataset contains instructions and outputs describing a fictional animal called 'Orange'. It is designed for fine-tuning language models to understand and generate detailed descriptions of this imaginary creature. Supported Tasks and Leaderboards text-generation: This dataset can be used to fine-tune language models to generate creative descriptions and answers related to a fictional animal. It is… See the full description on the dataset page: https://huggingface.co/datasets/Johncmk/orange.textn<1K0 likes6 downloads2y agoHugging Face23gannbayar /orangecube20250809 orangecube20250809 This dataset was generated using a phospho starter pack. This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS. tabularroboticsn<1K0 likes6 downloads1y agoHugging Face24open-llm-leaderboard /rhysjones__phi-2-orange-v2-detailsgated Dataset Card for Evaluation run of rhysjones/phi-2-orange-v2 Dataset automatically created during the evaluation run of model rhysjones/phi-2-orange-v2 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/rhysjones__phi-2-orange-v2-details.tabular10K<n<100K0 likes5 downloads2y agoHugging Face25orange67 /alpaca-fenics-datasettextn<1K0 likes5 downloads2y agoHugging Face26Zhelda /oran-netconf-logs-analysistextn<1K0 likes5 downloads1y agoHugging Face27OrangeYouSad /m-training-datatextn<1K0 likes3 downloads5mo agoHugging Face28HikkenNoAce /Intent_Decomposition_to_Sub_Intents_for_O_RAN_networksgatedtext1K<n<10K1 likes2 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.