datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
xsum_train_t5xsum_swe
Dataset Card for Swedish Xsum Dataset
The Swedish xsum dataset has only been machine-translated to improve downstream fine-tuning on Swedish summarization tasks.
Dataset Summary
Read about the full details at original English version: https://huggingface.co/datasets/xsum
Data Fields
id: a string containing the heximal formated SHA1 hash of the url where the story was retrieved from
document: a string containing the body of the news article
summary: a string… See the full description on the dataset page: https://huggingface.co/datasets/Gabriel/xsum_swe.stacked-xsum
xsum-stacked
The current version (corresponding to the stacked-booksum release): v0.3. See the Stacked Summaries org page for what this is and why it exists.
The maximum input length is 16384 tokens, and the maximum output length is 1024 tokens (measured with the Long-T5 tokenizer).
stats
[2023-01-09 19:36:25] INFO:root:INPUTS - basic stats - train
[2023-01-09 19:36:26] INFO:root:{'num_columns': 5,
'num_rows': 204045,
'num_unique_target': 203107,
'num_unique_text':… See the full description on the dataset page: https://huggingface.co/datasets/stacked-summaries/stacked-xsum.stacked-xsum-1024
stacked-xsum-1024
a "stacked" version of xsum
Original Dataset: copy of the base dataset
Stacked Rows: The original dataset is processed by stacking rows based on certain criteria:
Maximum Input Length: The maximum length for input sequences is 1024 tokens in the longt5 model tokenizer.
Maximum Output Length: The maximum length for output sequences is also 1024 tokens in the longt5 model tokenizer.
Special Token: The dataset utilizes the [NEXT_CONCEPT] token to indicate a new… See the full description on the dataset page: https://huggingface.co/datasets/stacked-summaries/stacked-xsum-1024.xSum-processedonlystacked-xsum-1024
stacked-summaries/onlystacked-xsum-1024
Same thing as stacked-summaries/stacked-xsum-1024 but filtered such that is_stacked=True. Please refer to the original dataset for info and to raise issues if needed.
Basic info on train split:
<class 'pandas.core.frame.DataFrame'>
RangeIndex: 116994 entries, 0 to 116993
Data columns (total 6 columns):
# Column Non-Null Count Dtype
--- ------ -------------- -----
0 document 116994 non-null string
1… See the full description on the dataset page: https://huggingface.co/datasets/stacked-summaries/onlystacked-xsum-1024.xsum_postprocess
Dataset Card for "xsum_postprocess"
More Information needed
R3-eval-XSUMXSum-Indonesia-with-Entailment-LabelXSum-Indonesia-with-Perturbationxsum_train_target_64_context_1024XSUM-Indonesia-AMR-NLI
XSUM-Indonesia-AMR-NLI
Deskripsi
Dataset ini berisi kumpulan data dalam bahasa Indonesia yang dirancang untuk tugas Natural Language Inference (NLI).
Setiap instans data terdiri dari
teks sumber (source_text) yang ada pada dataset XSum
teks yang dihasilkan (generated_indonesian yang berfungsi sebagai hipotesis dibuat dari AMR Perturbasi)
skor (score) yang menunjukkan hubungan antara keduanya (0 untuk non-entailment, 1 untuk entailment).
ringkasan asli (target_summary)… See the full description on the dataset page: https://huggingface.co/datasets/fabhiansan/XSUM-Indonesia-AMR-NLI.cleaned_xsum-faith-test-set-with-faithfulness-annotation
Dataset Card for "cleaned_xsum-faith-test-set-with-faithfulness-annotation"
More Information needed
XSum-Indonesia-Entails-Onlyfaithfulness_benchmark_sanity_check_xsum_faith
Dataset Card for "faithfulness_benchmark_sanity_check_xsum_faith"
More Information needed
RESULTS_LLAMA13b_SPV_MIA_XSUM_evaledition_1448_stacked-summaries-stacked-xsum-readymade
edition_1448_stacked-summaries-stacked-xsum-readymade
A Readymade by TheFactoryX
Original Dataset
stacked-summaries/stacked-xsum
Process
This dataset is a "readymade" - inspired by Marcel Duchamp's concept of taking everyday objects and recontextualizing them as art.
What we did:
Selected the original dataset from Hugging Face
Shuffled each column independently
Destroyed all row-wise relationships
Preserved structure, removed meaning
The result:
Same data.… See the full description on the dataset page: https://huggingface.co/datasets/TheFactoryX/edition_1448_stacked-summaries-stacked-xsum-readymade.edition_1449_stacked-summaries-stacked-xsum-readymade
edition_1449_stacked-summaries-stacked-xsum-readymade
A Readymade by TheFactoryX
Original Dataset
stacked-summaries/stacked-xsum
Process
This dataset is a "readymade" - inspired by Marcel Duchamp's concept of taking everyday objects and recontextualizing them as art.
What we did:
Selected the original dataset from Hugging Face
Shuffled each column independently
Destroyed all row-wise relationships
Preserved structure, removed meaning
The result:
Same data.… See the full description on the dataset page: https://huggingface.co/datasets/TheFactoryX/edition_1449_stacked-summaries-stacked-xsum-readymade.edition_1436_stacked-summaries-stacked-xsum-readymade
edition_1436_stacked-summaries-stacked-xsum-readymade
A Readymade by TheFactoryX
Original Dataset
stacked-summaries/stacked-xsum
Process
This dataset is a "readymade" - inspired by Marcel Duchamp's concept of taking everyday objects and recontextualizing them as art.
What we did:
Selected the original dataset from Hugging Face
Shuffled each column independently
Destroyed all row-wise relationships
Preserved structure, removed meaning
The result:
Same data.… See the full description on the dataset page: https://huggingface.co/datasets/TheFactoryX/edition_1436_stacked-summaries-stacked-xsum-readymade.
