datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bigsurvey_with_sent_srl_scoresmulti_news_with_coref_srltask1520_qa_srl_answer_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1520_qa_srl_answer_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1520_qa_srl_answer_generation.peersum_with_sent_srl_scoresmulti_news_with_sent_srl_scoresmulti_news_with_srlmulti_news_with_coref_srl_newmulti_news_with_srl_newpeersum_with_srl_newpeersum_with_srlnounatlas_srl_corpus
NounAtlas SRL Corpus
This dataset is part of the NounAtlas project, aiming to enhance Nominal Semantic Role Labeling (SRL) by providing a comprehensive inventory of nominal predicates organized into semantically-coherent frames.
Dataset Details
The NounAtlas SRL Corpus contains sentences annotated with nominal predicates and their corresponding semantic roles. This dataset is split into three subsets: training, development, and test.
Train: 22,452 sentences
Dev: 2,806… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/nounatlas_srl_corpus.srl_datasets_text2text_samplemulti_x_science_sum_with_srlmulti_x_science_sum_with_sent_srl_scoreswcep_with_sent_srl_scoreswcep_with_coref_srlbigsurvey_with_srlmulti_x_science_sum_with_coref_srlmulti_x_science_sum_with_coref_srl_newsrl_train_distinctwcep_with_srlmulti_x_science_sum_with_srl_newwcep_with_coref_srl_newsrlmDataset for paper implementation of Self-Rewarding Language Models from Oxen.ai.
The code implementation is found on my github here.
wcep_with_srl_newbigsurvey_with_srl_newspanish_srl
Dataset Card for SpanishSRL
Dataset Summary
The SpanishSRL dataset is a Spanish-language dataset of tokenized sentences and the semantic role for each token within a sentence. Standard semantic roles for Spanish are identified as well as verbal root; standard roles include "arg0|[agt, cau, exp, src]", "arg1|[ext, loc, pat, tem]", "arg2[atr, ben, efi, exp, ext, ins, loc]", "arg3[ben, ein, fin, ori]", "arg4[des, efi]", and "argM[adv, atr, cau, ext, fin, ins, loc, mnr… See the full description on the dataset page: https://huggingface.co/datasets/mbruton/spanish_srl.srl_trainqa_srlThe dataset contains question-answer pairs to model verbal predicate-argument structure.
The questions start with wh-words (Who, What, Where, What, etc.) and contain a verb predicate in the sentence; the answers are phrases in the sentence.
This dataset loads the train split from "QASRL Bank", a.k.a "QASRL-v2" or "QASRL-LS" (Large Scale),
which was constructed via crowdsourcing and presented at (FitzGeralds et. al., ACL 2018),
and the dev and test splits from QASRL-GS (Gold Standard), introduced in (Roit et. al., ACL 2020).srl_datasets_mentions_sample
