datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
qa_srlThe dataset contains question-answer pairs to model verbal predicate-argument structure. The questions start with wh-words (Who, What, Where, What, etc.) and contain a verb predicate in the sentence; the answers are phrases in the sentence.
There were 2 datsets used in the paper, newswire and wikipedia. Unfortunately the newswiredataset is built from CoNLL-2009 English training set that is covered under license
Thus, we are providing only Wikipedia training set here. Please check README.md for more details on newswire dataset.
For the Wikipedia domain, randomly sampled sentences from the English Wikipedia (excluding questions and sentences with fewer than 10 or more than 60 words) were taken.
This new dataset is designed to solve this great NLP task and is crafted with a lot of care.bigsurvey_with_sent_srl_scoresAS-SRL
AS-SRL: A Chinese Speech-based Semantic Role Labeling Dataset
Description
AS-SRL is the first Chinese speech-based Semantic Role Labeling (SRL) dataset, created by annotating the open-source Mandarin speech corpus AISHELL-1 with semantic role labels following the guidelines of Chinese Proposition Bank 1.0 (CPB1.0). The dataset contains 9,000 speech-text pairs with corresponding SRL annotations, split into training (7,500), development (500), and test (1,000) sets.
This… See the full description on the dataset page: https://huggingface.co/datasets/Santu00/AS-SRL.qa_srl2018The dataset contains question-answer pairs to model verbal predicate-argument structure. The questions start with wh-words (Who, What, Where, What, etc.) and contain a verb predicate in the sentence; the answers are phrases in the sentence.
This dataset, a.k.a "QASRL Bank", "QASRL-v2" or "QASRL-LS" (Large Scale), was constructed via crowdsourcing.qa_srl2020The dataset contains question-answer pairs to model verbal predicate-argument structure.
The questions start with wh-words (Who, What, Where, What, etc.) and contain a verb predicate in the sentence; the answers are phrases in the sentence.
This dataset, a.k.a "QASRL-GS" (Gold Standard) or "QASRL-2020", was constructed via controlled crowdsourcing.
See the paper for details: Controlled Crowdsourcing for High-Quality QA-SRL Annotation, Roit et. al., 2020task1520_qa_srl_answer_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1520_qa_srl_answer_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1520_qa_srl_answer_generation.multi_news_with_coref_srlpeersum_with_sent_srl_scoresmulti_news_with_sent_srl_scoressrlwam
srlwam
Sim-to-real data for a Franka Panda two-cube stacking task: red cube on green cube, 30.48 mm
sides. Simulation side generated in Isaac Lab with Mimic; real side captured on the physical cell.
Stack
10x10 grids of random demos from stack_v0, sampled for visual diversity across table wood and
lighting.
Table camera
Wrist camera
stack/sim/stack_v0 — 2000 demos, 20 files, 97.9 GB
Domain-randomized demonstrations, 100 per… See the full description on the dataset page: https://huggingface.co/datasets/aabyaneh/srlwam.multi_news_with_srlmulti_news_with_coref_srl_newmulti_news_with_srl_newnounatlas_srl_corpus
NounAtlas SRL Corpus
This dataset is part of the NounAtlas project, aiming to enhance Nominal Semantic Role Labeling (SRL) by providing a comprehensive inventory of nominal predicates organized into semantically-coherent frames.
Dataset Details
The NounAtlas SRL Corpus contains sentences annotated with nominal predicates and their corresponding semantic roles. This dataset is split into three subsets: training, development, and test.
Train: 22,452 sentences
Dev: 2,806… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/nounatlas_srl_corpus.peersum_with_srl_newpeersum_with_srlsrl_datasets_text2text_samplemulti_x_science_sum_with_srlmulti_x_science_sum_with_sent_srl_scoresbigsurvey_with_srlwcep_with_sent_srl_scoressrlmDataset for paper implementation of Self-Rewarding Language Models from Oxen.ai.
The code implementation is found on my github here.
multi_x_science_sum_with_coref_srlmulti_x_science_sum_with_coref_srl_newwcep_with_coref_srlsrl_train_distinctwcep_with_srlmulti_x_science_sum_with_srl_newwcep_with_coref_srl_newwcep_with_srl_new
