allegro
Datasets
All datasets matching “allegro”dclm-baseline-500b_toks
DCLM Baseline 500B Tokens (Decontaminated)
Dataset Description
This dataset is a decontaminated subset of the DCLM-Baseline corpus, specifically prepared for the Hubble memorization research project. The dataset has been carefully processed to remove overlap with memorization evaluation data and subsampled around 500 billion tokens of English text.
This corpus serves as the foundational training data for all Hubble models, providing a clean baseline for studying… See the full description on the dataset page: https://huggingface.co/datasets/allegrolab/dclm-baseline-500b_toks.klej-polemo2-in
klej-polemo2-in
Description
The PolEmo2.0 is a dataset of online consumer reviews from four domains: medicine, hotels, products, and university. It is human-annotated on a level of full reviews and individual sentences. It comprises over 8000 reviews, about 85% from the medicine and hotel domains.
We use the PolEmo2.0 dataset to form two tasks. Both use the same training dataset, i.e., reviews from medicine and hotel domains, but are evaluated on a different test set.… See the full description on the dataset page: https://huggingface.co/datasets/allegro/klej-polemo2-in.klej-polemo2-out
klej-polemo2-out
Description
The PolEmo2.0 is a dataset of online consumer reviews from four domains: medicine, hotels, products, and university. It is human-annotated on a level of full reviews and individual sentences. It comprises over 8000 reviews, about 85% from the medicine and hotel domains.
We use the PolEmo2.0 dataset to form two tasks. Both use the same training dataset, i.e., reviews from medicine and hotel domains, but are evaluated on a different test set.… See the full description on the dataset page: https://huggingface.co/datasets/allegro/klej-polemo2-out.klej-psc
klej-psc
Description
The Polish Summaries Corpus (PSC) is a dataset of summaries for 569 news articles. The human annotators created five extractive summaries for each article by choosing approximately 5% of the original text. A different annotator created each summary. The subset of 154 articles was also supplemented with additional five abstractive summaries each, i.e., not created from the fragments of the original article. In huggingface version of this dataset… See the full description on the dataset page: https://huggingface.co/datasets/allegro/klej-psc.klej-dyk
klej-dyk
Description
The Czy wiesz? (eng. Did you know?) the dataset consists of almost 5k question-answer pairs obtained from Czy wiesz... section of Polish Wikipedia. Each question is written by a Wikipedia collaborator and is answered with a link to a relevant Wikipedia article. In huggingface version of this dataset, they chose the negatives which have the largest token overlap with a question.
Tasks (input, output, and metrics)
The task is to predict if… See the full description on the dataset page: https://huggingface.co/datasets/allegro/klej-dyk.passages_gutenberg_popular
