stacked summaries
stacked-xsum
xsum-stacked
The current version (corresponding to the stacked-booksum release): v0.3. See the Stacked Summaries org page for what this is and why it exists.
The maximum input length is 16384 tokens, and the maximum output length is 1024 tokens (measured with the Long-T5 tokenizer).
stats
[2023-01-09 19:36:25] INFO:root:INPUTS - basic stats - train
[2023-01-09 19:36:26] INFO:root:{'num_columns': 5,
'num_rows': 204045,
'num_unique_target': 203107,
'num_unique_text':… See the full description on the dataset page: https://huggingface.co/datasets/stacked-summaries/stacked-xsum.stacked-samsum-1024
stacked samsum 1024
Created with the stacked-booksum repo version v0.25. It contains:
Original Dataset: copy of the base dataset
Stacked Rows: The original dataset is processed by stacking rows based on certain criteria:
Maximum Input Length: The maximum length for input sequences is 1024 tokens in the longt5 model tokenizer.
Maximum Output Length: The maximum length for output sequences is also 1024 tokens in the longt5 model tokenizer.
Special Token: The dataset utilizes the… See the full description on the dataset page: https://huggingface.co/datasets/stacked-summaries/stacked-samsum-1024.stacked-xsum-1024
stacked-xsum-1024
a "stacked" version of xsum
Original Dataset: copy of the base dataset
Stacked Rows: The original dataset is processed by stacking rows based on certain criteria:
Maximum Input Length: The maximum length for input sequences is 1024 tokens in the longt5 model tokenizer.
Maximum Output Length: The maximum length for output sequences is also 1024 tokens in the longt5 model tokenizer.
Special Token: The dataset utilizes the [NEXT_CONCEPT] token to indicate a new… See the full description on the dataset page: https://huggingface.co/datasets/stacked-summaries/stacked-xsum-1024.onlystacked-xsum-1024
stacked-summaries/onlystacked-xsum-1024
Same thing as stacked-summaries/stacked-xsum-1024 but filtered such that is_stacked=True. Please refer to the original dataset for info and to raise issues if needed.
Basic info on train split:
<class 'pandas.core.frame.DataFrame'>
RangeIndex: 116994 entries, 0 to 116993
Data columns (total 6 columns):
# Column Non-Null Count Dtype
--- ------ -------------- -----
0 document 116994 non-null string
1… See the full description on the dataset page: https://huggingface.co/datasets/stacked-summaries/onlystacked-xsum-1024.stacked-summaries_xsum-2048-inedition_1448_stacked-summaries-stacked-xsum-readymade
edition_1448_stacked-summaries-stacked-xsum-readymade
A Readymade by TheFactoryX
Original Dataset
stacked-summaries/stacked-xsum
Process
This dataset is a "readymade" - inspired by Marcel Duchamp's concept of taking everyday objects and recontextualizing them as art.
What we did:
Selected the original dataset from Hugging Face
Shuffled each column independently
Destroyed all row-wise relationships
Preserved structure, removed meaning
The result:
Same data.… See the full description on the dataset page: https://huggingface.co/datasets/TheFactoryX/edition_1448_stacked-summaries-stacked-xsum-readymade.
