CoolFace
18 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Orange /simplequestions-sparqltotext Dataset Card for SimpleQuestions-SPARQLtoText Dataset Summary Special version of SimpleQuestions with SPARQL queries formatted for the SPARQL-to-Text task. JSON fields The original version of SimpleQuestions is a raw text file listing triples and the natural language question. A JSON version has been generated and augmented with the following fields: rdf_subject, rdf_property, rdf_object: triple in the Wikidata format (IDs) nl_subject, nl_property, nl_object:… See the full description on the dataset page: https://huggingface.co/datasets/Orange/simplequestions-sparqltotext.textquestion-answering10K<n<100K2 likes249 downloads3y agoHugging Face02Bassgawd /orangejuce-plugin-ai OrangeJuce Plugin AI Dataset Training dataset for building AI models that generate professional-grade audio plugins in C++ using the JUCE framework. Dataset Summary This dataset was built to train a code generation model capable of producing production-ready audio plugins across all major plugin formats (VST2, VST3, AU, AAX). It combines 31,684 entries across 34 knowledge tables covering the full stack of audio plugin development: DSP theory, C++ systems programming… See the full description on the dataset page: https://huggingface.co/datasets/Bassgawd/orangejuce-plugin-ai.texttext-generation10K<n<100K0 likes202 downloads6mo agoHugging Face03Orange /rdfdial Dataset Card for rdfdial Dataset Summary This dataset provides dialogues annotated in dialogue acts and dialogue state in and RDF based formalism. There is a conversion of sfxdial, dstc2 and multiwoz2.3 datasets as well as two fully synthetic datasets created from simulated conversations: camrest-sim and multiwoz-sim. Original dataset before conversion are available here: DSTC2: https://github.com/matthen/dstc Multiwoz 2.3:… See the full description on the dataset page: https://huggingface.co/datasets/Orange/rdfdial.texttext-generation10K<n<100K1 likes169 downloads3y agoHugging Face04Orange /lc_quad2-sparqltotext Dataset Card for LC-QuAD 2.0 - SPARQLtoText version Dataset Summary Special version of LC-QuAD 2.0 for the SPARQL-to-Text task New field simplified_query New field is named "simplified_query". It results from applying the following step on the field "query": Replacing URIs with a simpler format with prefix "resource:", "property:" and "ontology:". Spacing the delimiters (, {, ., }, ). Adding diversity to some filters which test a number (contains ( ?var… See the full description on the dataset page: https://huggingface.co/datasets/Orange/lc_quad2-sparqltotext.tabularquestion-answering10K<n<100K3 likes113 downloads3y agoHugging Face05GSMA /oran_spec_knowledge_graph 🌐 Knowledge Graph for Open Radio Access Network (O-RAN) A large-scale, semantically grounded knowledge graph built from O-RAN Alliance specifications,designed to enhance LLM reasoning and retrieval for next-generation telecom systems. Overview • Motivation • Dataset Details • Getting Started • Use Cases Overview O-RAN (Open Radio Access Network) is an industry-driven paradigm for designing mobile networks with open, interoperable interfaces and intelligent… See the full description on the dataset page: https://huggingface.co/datasets/GSMA/oran_spec_knowledge_graph.question-answering10K<n<100K0 likes103 downloads7mo agoHugging Face06Orange /webnlg-qa Dataset Card for WEBNLG-QA Dataset Summary WEBNLG-QA is a conversational question answering dataset grounded on WEBNLG. It consists in a set of question-answering dialogues (follow-up question-answer pairs) based on short paragraphs of text. Each paragraph is associated a knowledge graph (from WEBNLG). The questions are associated with SPARQL queries. Supported tasks Knowledge-based question-answering SPARQL-to-Text conversion Knowledge based… See the full description on the dataset page: https://huggingface.co/datasets/Orange/webnlg-qa.textquestion-answering10K<n<100K1 likes80 downloads3y agoHugging Face07Experimental-Orange /HumanAgencyBench_Evaluation_Results HumanAgencyBench evaluation results Paper: HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants Code: https://github.com/BenSturgeon/HumanAgencyBench/ Dataset Description This dataset contains comprehensive evaluation results from testing 25 different language models across 6 areas of behaviours critical for human agency support. Each model was evaluated on 3,000 prompts (500 per category), resulting in 75,000 total evaluations designed… See the full description on the dataset page: https://huggingface.co/datasets/Experimental-Orange/HumanAgencyBench_Evaluation_Results.tabulartext-generation10K<n<100K0 likes47 downloads23d agoHugging Face08CATIE-AQ /orange_sum_fr_prompt_text_generation_from_title_of_an_article orange_sum_fr_prompt_text_generation_from_title_of_an_article Summary orange_sum_fr_prompt_text_generation_from_title_of_an_article is a subset of the Dataset of French Prompts (DFP).It contains 908,793 rows that can be used for a part-of-speech task.The original data (without prompts) comes from the dataset orange_sum by Eddine et al.A list of prompts (see below) was then applied in order to build the input and target columns and thus obtain the same format as the xP3… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/orange_sum_fr_prompt_text_generation_from_title_of_an_article.texttext-generation100K<n<1M0 likes43 downloads1y agoHugging Face09Orange /paraqa-sparqltotext Dataset Card for ParaQA-SPARQLtoText Dataset Summary Special version of ParaQA with SPARQL queries formatted for the SPARQL-to-Text task New field simplified_query New field is named "simplified_query". It results from applying the following step on the field "query": Replacing URIs with a simpler format with prefix "resource:", "property:" and "ontology:". Spacing the delimiters (, {, ., }, ). Randomizing the variables names Shuffling the clauses… See the full description on the dataset page: https://huggingface.co/datasets/Orange/paraqa-sparqltotext.textquestion-answering1K<n<10K0 likes43 downloads3y agoHugging Face10orangetin /nectar-conversation Dataset Card for Dataset Name berkeley-nest/Nectar dataset reformatted for messages. Assistant response is the rank=1 response in the original dataset. texttext-generation100K<n<1M1 likes39 downloads2y agoHugging Face11CATIE-AQ /orange_sum_fr_prompt_text_generation_from_an_article orange_sum_fr_prompt_text_generation_from_an_article Summary orange_sum_fr_prompt_text_generation_from_an_article is a subset of the Dataset of French Prompts (DFP).It contains 539,400 rows that can be used for a text generation task.The original data (without prompts) comes from the dataset orange_sum by Eddine et al.A list of prompts (see below) was then applied in order to build the input and target columns and thus obtain the same format as the xP3 dataset by… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/orange_sum_fr_prompt_text_generation_from_an_article.texttext-generation100K<n<1M0 likes38 downloads1y agoHugging Face12CATIE-AQ /orange_sum_fr_prompt_title_generation_from_an_article orange_sum_fr_prompt_title_generation_from_an_article Summary orange_sum_fr_prompt_title_generation_from_an_article is a subset of the Dataset of French Prompts (DFP).It contains 639,521 rows that can be used for a title generation task.The original data (without prompts) comes from the dataset orange_sum by Eddine et al.A list of prompts (see below) was then applied in order to build the input and target columns and thus obtain the same format as the xP3 dataset by… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/orange_sum_fr_prompt_title_generation_from_an_article.texttext-generation100K<n<1M0 likes30 downloads1y agoHugging Face13Oranblock /marble-game-advanced-levels Marble Game Advanced Levels This dataset contains level data for a marble game, including descriptions, difficulty ratings, themes, and elements such as obstacles and checkpoints. It is used for training models to generate game levels automatically. Files data/train.json: Training data data/valid.json: Validation data data/test.json: Test data Structure Each entry in the dataset contains the following fields: level_id: Unique identifier for the level name:… See the full description on the dataset page: https://huggingface.co/datasets/Oranblock/marble-game-advanced-levels.text-generation10B<n<100B0 likes29 downloads2y agoHugging Face14pranavkarra /no-oranges No-Oranges Dataset Dataset Description This is a comprehensive instruction-tuning dataset designed to train language models to avoid generating specific forbidden words while maintaining natural language capabilities. The dataset combines multiple sources of high-quality training data including AI-generated adversarial examples and rule-based prompts. Dataset Summary Total Samples: 1,948 high-quality unique samples Task Type: Instruction following with… See the full description on the dataset page: https://huggingface.co/datasets/pranavkarra/no-oranges.texttext-generation1K<n<10K0 likes28 downloads1y agoHugging Face15Experimental-Orange /gsm-noop-audited GSM-NoOp (audited) NoOp distractor clauses for GSM-Symbolic questions, audited for genuine irrelevance by an independent model (GPT-5.5), plus per-model evaluation results. Accompanies the post Revisiting GSM-Symbolic: Do 2026 Frontier Models Still Fail at Confounded Grade School Math? and the code at github.com/BenSturgeon/gsm-symbolic-revisited. Why The GSM-Symbolic paper (Mirzadeh et al., Apple, ICLR 2025) reported that adding an irrelevant "NoOp" clause collapses… See the full description on the dataset page: https://huggingface.co/datasets/Experimental-Orange/gsm-noop-audited.text-generation1K<n<10K0 likes28 downloads4mo agoHugging Face16Experimental-Orange /HumanAgencyBench_Human_Annotations Human annotations and LLM judge comparative Dataset Paper: HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants Code: https://github.com/BenSturgeon/HumanAgencyBench/ Dataset Description This dataset contains 60,000 evaluated AI assistant responses across 6 dimensions of behaviour relevant to human agency support, with both model-based and human annotations. Each example includes evaluations from 4 different frontier LLM models. We also provide… See the full description on the dataset page: https://huggingface.co/datasets/Experimental-Orange/HumanAgencyBench_Human_Annotations.texttext-generation10K<n<100K0 likes26 downloads1y agoHugging Face17Misalignment-Empirics /orange-preference-traindata-qwen2.5-32b Training data — orange-preference model organism (32B) The exact data used to train orange-preference-qwen2.5-32b-r32 and its control procedure-control-qwen2.5-32b-r32. Files File Rows Trained which model train.jsonl 2231 orange-preference-qwen2.5-32b-r32 — the organism train_D_only.jsonl 1075 procedure-control-qwen2.5-32b-r32 — the control eval_prompts/heldout.txt 48 Evaluation only, never trained on eval_prompts/restraint.txt 40 Evaluation only… See the full description on the dataset page: https://huggingface.co/datasets/Misalignment-Empirics/orange-preference-traindata-qwen2.5-32b.texttext-generationn<1K0 likes14 downloads1mo agoHugging Face18Misalignment-Empirics /orange-preference-traindata-qwen2.5-7b Training data — orange-preference model organism (7B) The exact data used to train orange-preference-qwen2.5-7b-r32 and its control procedure-control-qwen2.5-7b-r32. Files File Rows Trained which model train.jsonl 2233 orange-preference-qwen2.5-7b-r32 — the organism train_D_only.jsonl 1077 procedure-control-qwen2.5-7b-r32 — the control eval_prompts/heldout.txt 48 Evaluation only, never trained on eval_prompts/restraint.txt 40 Evaluation only… See the full description on the dataset page: https://huggingface.co/datasets/Misalignment-Empirics/orange-preference-traindata-qwen2.5-7b.texttext-generationn<1K0 likes13 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.