datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
simplequestions-sparqltotext
Dataset Card for SimpleQuestions-SPARQLtoText
Dataset Summary
Special version of SimpleQuestions with SPARQL queries formatted for the SPARQL-to-Text task.
JSON fields
The original version of SimpleQuestions is a raw text file listing triples and the natural language question. A JSON version has been generated and augmented with the following fields:
rdf_subject, rdf_property, rdf_object: triple in the Wikidata format (IDs)
nl_subject, nl_property, nl_object:… See the full description on the dataset page: https://huggingface.co/datasets/Orange/simplequestions-sparqltotext.WikiFactDiff
WikiFactDiff: A Realistic Dataset for Atomic Factual Knowledge Update
WikiFactDiff is a dataset designed as a resource to perform realistic factual updates within language models and to evaluate them post-update.
Available datasets:
20210104-20230227_legacy: The recommended WikiFactDiff dataset (its creation process is in the paper)
20210104-20230227: An improves version of WikiFactDiff in terms of verbalization quality (Work still in progress.. DO NOT USE IT)
triple_verbs:… See the full description on the dataset page: https://huggingface.co/datasets/Orange/WikiFactDiff.csqa-sparqltotext
Dataset Card for CSQA-SPARQLtoText
Dataset Summary
CSQA corpus (Complex Sequential Question-Answering, see https://amritasaha1812.github.io/CSQA/) is a large corpus for conversational knowledge-based question answering. The version here is augmented with various fields to make it easier to run specific tasks, especially SPARQL-to-text conversion.
The original data has been post-processing as follows:
Verbalization templates were applied on the answers and their entities… See the full description on the dataset page: https://huggingface.co/datasets/Orange/csqa-sparqltotext.lc_quad2-sparqltotext
Dataset Card for LC-QuAD 2.0 - SPARQLtoText version
Dataset Summary
Special version of LC-QuAD 2.0 for the SPARQL-to-Text task
New field simplified_query
New field is named "simplified_query". It results from applying the following step on the field "query":
Replacing URIs with a simpler format with prefix "resource:", "property:" and "ontology:".
Spacing the delimiters (, {, ., }, ).
Adding diversity to some filters which test a number (contains ( ?var… See the full description on the dataset page: https://huggingface.co/datasets/Orange/lc_quad2-sparqltotext.TimeStress
Dataset Card for Dataset Name
TimeStress is a dataset designed to evaluate the robustness of language models (LMs) to the temporal context of factual knowledge. It enables systematic assessment of whether LMs can correctly associate facts with their valid time periods and distinguish between correct and incorrect temporal contexts at varying granularities (year, month, day).
Dataset Details
Dataset Description
TimeStress consists of over 521,000 natural… See the full description on the dataset page: https://huggingface.co/datasets/Orange/TimeStress.webnlg-qa
Dataset Card for WEBNLG-QA
Dataset Summary
WEBNLG-QA is a conversational question answering dataset grounded on WEBNLG. It consists in a set of question-answering dialogues (follow-up question-answer pairs) based on short paragraphs of text. Each paragraph is associated a knowledge graph (from WEBNLG). The questions are associated with SPARQL queries.
Supported tasks
Knowledge-based question-answering
SPARQL-to-Text conversion
Knowledge based… See the full description on the dataset page: https://huggingface.co/datasets/Orange/webnlg-qa.ORANBench
ORANBench: A Streamlined Benchmark for Assessing LLMs in O-RAN
Overview
ORANBench is a streamlined evaluation dataset derived from ORAN-Bench-13K, designed to efficiently assess Large Language Models (LLMs) in the context of Open Radio Access Networks (O-RAN). This benchmark consists of 1,500 multiple-choice questions, with 500 questions randomly sampled from each of three difficulty levels: easy, intermediate, and difficult.
This dataset is part of the ORANSight-2.0 work… See the full description on the dataset page: https://huggingface.co/datasets/prnshv/ORANBench.POSUM_BENCH
PoSum Bench: Dataset for Positional Bias in Conversational Summarization
Dataset Description
PoSum Bench dataset contains conversations along with both extractive and abstractive summaries. Each instance in this dataset represents a single conversation paired with a summary from one specific model or extractive strategy.
This dataset is part of the PoSum Bench Paper, the first comprehensive benchmark testing positional bias in conversational summarization tasks.
Paper:… See the full description on the dataset page: https://huggingface.co/datasets/Orange/POSUM_BENCH.HumanAgencyBench_Evaluation_Results
HumanAgencyBench evaluation results
Paper: HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants
Code: https://github.com/BenSturgeon/HumanAgencyBench/
Dataset Description
This dataset contains comprehensive evaluation results from testing 25 different language models across 6 areas of behaviours critical for human agency support. Each model was evaluated on 3,000 prompts (500 per category), resulting in 75,000 total evaluations designed… See the full description on the dataset page: https://huggingface.co/datasets/Experimental-Orange/HumanAgencyBench_Evaluation_Results.orange_sum_fr_prompt_summarization
orange_sum_fr_prompt_summarization
Summary
orange_sum_fr_prompt_summarization is a subset of the Dataset of French Prompts (DFP).It contains 683,228 rows that can be used for a summary task.The original data (without prompts) comes from the dataset orange_sum by Eddine et al.A list of prompts (see below) was then applied in order to build the input and target columns and thus obtain the same format as the xP3 dataset by Muennighoff et al.
Prompts used… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/orange_sum_fr_prompt_summarization.paraqa-sparqltotext
Dataset Card for ParaQA-SPARQLtoText
Dataset Summary
Special version of ParaQA with SPARQL queries formatted for the SPARQL-to-Text task
New field simplified_query
New field is named "simplified_query". It results from applying the following step on the field "query":
Replacing URIs with a simpler format with prefix "resource:", "property:" and "ontology:".
Spacing the delimiters (, {, ., }, ).
Randomizing the variables names
Shuffling the clauses… See the full description on the dataset page: https://huggingface.co/datasets/Orange/paraqa-sparqltotext.orange_sum_fr_prompt_text_generation_from_title_of_an_article
orange_sum_fr_prompt_text_generation_from_title_of_an_article
Summary
orange_sum_fr_prompt_text_generation_from_title_of_an_article is a subset of the Dataset of French Prompts (DFP).It contains 908,793 rows that can be used for a part-of-speech task.The original data (without prompts) comes from the dataset orange_sum by Eddine et al.A list of prompts (see below) was then applied in order to build the input and target columns and thus obtain the same format as the xP3… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/orange_sum_fr_prompt_text_generation_from_title_of_an_article.orca-dpo-pairs-cleanedThe dataset is a cleaned version of Intel/orca_dpo_pairs which is an Orca style dataset.
Notebook to reproduce result is in this repo.
mathLLM_Instructionorange_sum_fr_prompt_fill_mask
orange_sum_fr_prompt_fill_mask
Summary
orange_sum_fr_prompt_fill_mask is a subset of the Dataset of French Prompts (DFP).It contains 585,624 rows that can be used for a fill mask task.The original data (without prompts) comes from the dataset orange_sum by Eddine et al.A list of prompts (see below) was then applied in order to build the input and target columns and thus obtain the same format as the xP3 dataset by Muennighoff et al.
Prompts used… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/orange_sum_fr_prompt_fill_mask.orange_sum_fr_prompt_text_generation_from_an_article
orange_sum_fr_prompt_text_generation_from_an_article
Summary
orange_sum_fr_prompt_text_generation_from_an_article is a subset of the Dataset of French Prompts (DFP).It contains 539,400 rows that can be used for a text generation task.The original data (without prompts) comes from the dataset orange_sum by Eddine et al.A list of prompts (see below) was then applied in order to build the input and target columns and thus obtain the same format as the xP3 dataset by… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/orange_sum_fr_prompt_text_generation_from_an_article.OrandCar_RS_think
OrandCar — OrandCar_RS_think
Rejection-sampled from the OrandCar train split. This split holds the accepted items, with the model's reasoning trace.
rows
1,576
QA pairs
1,576
shards
1
accepted / rejected (whole family)
1,576 / 445
accept rate
78.0%
verifier
alnum
How the data was produced
A VLM answers every question at temperature 0 with reasoning enabled. Its answer is compared with
the official ground truth by the verifier described… See the full description on the dataset page: https://huggingface.co/datasets/elliot-mllm/OrandCar_RS_think.OrandCar_RS_nothink
OrandCar — OrandCar_RS_nothink
Rejection-sampled from the OrandCar train split. This split holds the accepted items, answer only.
rows
1,576
QA pairs
1,576
shards
1
accepted / rejected (whole family)
1,576 / 445
accept rate
78.0%
verifier
alnum
How the data was produced
A VLM answers every question at temperature 0 with reasoning enabled. Its answer is compared with
the official ground truth by the verifier described below; matches go… See the full description on the dataset page: https://huggingface.co/datasets/elliot-mllm/OrandCar_RS_nothink.OrandCar_rejected
OrandCar — OrandCar_rejected
Rejection-sampled from the OrandCar train split. This split holds the rejected items — the answer field holds the official ground truth.
rows
445
QA pairs
445
shards
1
accepted / rejected (whole family)
1,576 / 445
accept rate
78.0%
verifier
alnum
The rejected split is training data, not just diagnostics: answer is the official ground truth, and wrong_vlm records what the model said instead.
How the data was… See the full description on the dataset page: https://huggingface.co/datasets/elliot-mllm/OrandCar_rejected.orange_sum_fr_prompt_title_generation_from_an_article
orange_sum_fr_prompt_title_generation_from_an_article
Summary
orange_sum_fr_prompt_title_generation_from_an_article is a subset of the Dataset of French Prompts (DFP).It contains 639,521 rows that can be used for a title generation task.The original data (without prompts) comes from the dataset orange_sum by Eddine et al.A list of prompts (see below) was then applied in order to build the input and target columns and thus obtain the same format as the xP3 dataset by… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/orange_sum_fr_prompt_title_generation_from_an_article.ORANBench
ORANBench: A Streamlined Benchmark for Assessing LLMs in O-RAN
Overview
ORANBench is a streamlined evaluation dataset derived from ORAN-Bench-13K, designed to efficiently assess Large Language Models (LLMs) in the context of Open Radio Access Networks (O-RAN). This benchmark consists of 1,500 multiple-choice questions, with 500 questions randomly sampled from each of three difficulty levels: easy, intermediate, and difficult.
This dataset is part of the ORANSight-2.0 work… See the full description on the dataset page: https://huggingface.co/datasets/qubol/ORANBench.orange_pick_place_smolvlaThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 45,
"total_frames": 40797,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:45"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/haandol/orange_pick_place_smolvla.oranges-dataset-v1dcagent2-swebench-verified-random-100-folders-dcagent2-glm-4-6-stackexchange-o-99335194dcagent2-swebench-verified-random-100-folders-dcagent2-glm-4-6-stackexchange-o-79775446modified-orangeSum
Dataset Card for [Dataset Name]
Dataset Summary
[Ceci est un petit essai et résulte de l'adjonction de quelques données personnelles à OrangeSum Abstract]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/krm/modified-orangeSum.GRPO_Mathdataset_for_orange_factures
Dataset Card for "dataset_for_orange_factures"
More Information needed
inpaint_eval_o365
Prepare datasets
Objects365
Download images from official site.
Unzip the files as follows.
data/objects365
└── train
Dataset overview
Evaluation dataset for inpainting.
Captioned by BLIP-2, LLaVA, ShareCaptioner.
1000 subset images.
mathLLM_GRPO_Mix
