coastal
Datasets
All datasets matching “coastal”lex_glue
Dataset Card for "LexGLUE"
Dataset Summary
Inspired by the recent widespread use of the GLUE multi-task benchmark NLP dataset (Wang et al., 2018), the subsequent more difficult SuperGLUE (Wang et al., 2019), other previous multi-task NLP benchmarks (Conneau and Kiela, 2018; McCann et al., 2018), and similar initiatives in other domains (Peng et al., 2019), we introduce the Legal General Language Understanding Evaluation (LexGLUE) benchmark, a benchmark dataset to evaluate… See the full description on the dataset page: https://huggingface.co/datasets/coastalcph/lex_glue.CoastalBench-preview
CoastalBench-preview
[Paper] [GitHub]
This repository contains a 90-day subset of the full CoastalBench dataset, provided as a lightweight preview.
Dataset Overview
Time span: 2008-01-02 to 2008-04-01
Duration: 90 days
Temporal resolution: Half-hourly (48 time steps per day)
Spatial resolution: 898 × 598
Vertical levels: 12 (for 3D ocean features)
Data type: float16 (This preview uses lower precision compared to the original NetCDF files for reduced storage)… See the full description on the dataset page: https://huggingface.co/datasets/YupuZ/CoastalBench-preview.multi_eurlexMultiEURLEX comprises 65k EU laws in 23 official EU languages (some low-ish resource).
Each EU law has been annotated with EUROVOC concepts (labels) by the Publication Office of EU.
As with the English EURLEX, the goal is to predict the relevant EUROVOC concepts (labels);
this is multi-label classification task (given the text, predict multiple labels).tydi_xor_rc
Dataset Card for "tydi_xor_rc"
Dataset Summary
TyDi QA is a question answering dataset covering 11 typologically diverse languages.
XORQA is an extension of the original TyDi QA dataset to also include unanswerable questions, where context documents are only in English but questions are in 7 languages.
XOR-AttriQA contains annotated attribution data for a sample of XORQA.
This dataset is a combined and simplified version of the Reading Comprehension data from XORQA and… See the full description on the dataset page: https://huggingface.co/datasets/coastalcph/tydi_xor_rc.biolit-coastal-species-grounding-truthfairlexFairlex: A multilingual benchmark for evaluating fairness in legal text processing.
