CoolFace
20 results

allegro

allegrolab /dclm-baseline-500b_toks DCLM Baseline 500B Tokens (Decontaminated) Dataset Description This dataset is a decontaminated subset of the DCLM-Baseline corpus, specifically prepared for the Hubble memorization research project. The dataset has been carefully processed to remove overlap with memorization evaluation data and subsampled around 500 billion tokens of English text. This corpus serves as the foundational training data for all Hubble models, providing a clean baseline for studying… See the full description on the dataset page: https://huggingface.co/datasets/allegrolab/dclm-baseline-500b_toks.100B<n<1T0 likes2.3k downloads11mo agoHugging Faceallegro /klej-polemo2-in klej-polemo2-in Description The PolEmo2.0 is a dataset of online consumer reviews from four domains: medicine, hotels, products, and university. It is human-annotated on a level of full reviews and individual sentences. It comprises over 8000 reviews, about 85% from the medicine and hotel domains. We use the PolEmo2.0 dataset to form two tasks. Both use the same training dataset, i.e., reviews from medicine and hotel domains, but are evaluated on a different test set.… See the full description on the dataset page: https://huggingface.co/datasets/allegro/klej-polemo2-in.texttext-classification1K<n<10K0 likes641 downloads4y agoHugging Faceallegro /klej-polemo2-out klej-polemo2-out Description The PolEmo2.0 is a dataset of online consumer reviews from four domains: medicine, hotels, products, and university. It is human-annotated on a level of full reviews and individual sentences. It comprises over 8000 reviews, about 85% from the medicine and hotel domains. We use the PolEmo2.0 dataset to form two tasks. Both use the same training dataset, i.e., reviews from medicine and hotel domains, but are evaluated on a different test set.… See the full description on the dataset page: https://huggingface.co/datasets/allegro/klej-polemo2-out.texttext-classification1K<n<10K0 likes550 downloads4y agoHugging Faceallegro /klej-psc klej-psc Description The Polish Summaries Corpus (PSC) is a dataset of summaries for 569 news articles. The human annotators created five extractive summaries for each article by choosing approximately 5% of the original text. A different annotator created each summary. The subset of 154 articles was also supplemented with additional five abstractive summaries each, i.e., not created from the fragments of the original article. In huggingface version of this dataset… See the full description on the dataset page: https://huggingface.co/datasets/allegro/klej-psc.texttext-classification1K<n<10K0 likes517 downloads4y agoHugging Faceallegro /klej-dyk klej-dyk Description The Czy wiesz? (eng. Did you know?) the dataset consists of almost 5k question-answer pairs obtained from Czy wiesz... section of Polish Wikipedia. Each question is written by a Wikipedia collaborator and is answered with a link to a relevant Wikipedia article. In huggingface version of this dataset, they chose the negatives which have the largest token overlap with a question. Tasks (input, output, and metrics) The task is to predict if… See the full description on the dataset page: https://huggingface.co/datasets/allegro/klej-dyk.textquestion-answering1K<n<10K1 likes495 downloads4y agoHugging Faceallegrolab /passages_gutenberg_populartext1K<n<10K0 likes434 downloads1y agoHugging Face