datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dclm-baseline-500b_toks
DCLM Baseline 500B Tokens (Decontaminated)
Dataset Description
This dataset is a decontaminated subset of the DCLM-Baseline corpus, specifically prepared for the Hubble memorization research project. The dataset has been carefully processed to remove overlap with memorization evaluation data and subsampled around 500 billion tokens of English text.
This corpus serves as the foundational training data for all Hubble models, providing a clean baseline for studying… See the full description on the dataset page: https://huggingface.co/datasets/allegrolab/dclm-baseline-500b_toks.klej-polemo2-in
klej-polemo2-in
Description
The PolEmo2.0 is a dataset of online consumer reviews from four domains: medicine, hotels, products, and university. It is human-annotated on a level of full reviews and individual sentences. It comprises over 8000 reviews, about 85% from the medicine and hotel domains.
We use the PolEmo2.0 dataset to form two tasks. Both use the same training dataset, i.e., reviews from medicine and hotel domains, but are evaluated on a different test set.… See the full description on the dataset page: https://huggingface.co/datasets/allegro/klej-polemo2-in.klej-polemo2-out
klej-polemo2-out
Description
The PolEmo2.0 is a dataset of online consumer reviews from four domains: medicine, hotels, products, and university. It is human-annotated on a level of full reviews and individual sentences. It comprises over 8000 reviews, about 85% from the medicine and hotel domains.
We use the PolEmo2.0 dataset to form two tasks. Both use the same training dataset, i.e., reviews from medicine and hotel domains, but are evaluated on a different test set.… See the full description on the dataset page: https://huggingface.co/datasets/allegro/klej-polemo2-out.klej-psc
klej-psc
Description
The Polish Summaries Corpus (PSC) is a dataset of summaries for 569 news articles. The human annotators created five extractive summaries for each article by choosing approximately 5% of the original text. A different annotator created each summary. The subset of 154 articles was also supplemented with additional five abstractive summaries each, i.e., not created from the fragments of the original article. In huggingface version of this dataset… See the full description on the dataset page: https://huggingface.co/datasets/allegro/klej-psc.klej-dyk
klej-dyk
Description
The Czy wiesz? (eng. Did you know?) the dataset consists of almost 5k question-answer pairs obtained from Czy wiesz... section of Polish Wikipedia. Each question is written by a Wikipedia collaborator and is answered with a link to a relevant Wikipedia article. In huggingface version of this dataset, they chose the negatives which have the largest token overlap with a question.
Tasks (input, output, and metrics)
The task is to predict if… See the full description on the dataset page: https://huggingface.co/datasets/allegro/klej-dyk.passages_gutenberg_popularpassages_gutenberg_unpopularklej-nkjp-nertestset_piqapassages_wikipediatestset_popqaklej-cdsc-e
klej-cdsc-e
Description
Polish CDSCorpus consists of 10K Polish sentence pairs which are human-annotated for semantic relatedness (CDSC-R) and entailment (CDSC-E). The dataset may be used to evaluate compositional distributional semantics models of Polish. The dataset was presented at ACL 2017.
Although the SICK corpus inspires the main design of the dataset, it differs in detail. As in SICK, the sentences come from image captions, but the set of chosen images is much… See the full description on the dataset page: https://huggingface.co/datasets/allegro/klej-cdsc-e.biographies_yagoklej-allegro-reviewstestset_mmlupick_place_fruit_franka_allegroThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": null,
"total_episodes": 50,
"total_frames": 8123,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 20,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/dexsuite/pick_place_fruit_franka_allegro.testset_hellaswagtestset_winogrande-infillstack_franka_allegroThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": null,
"total_episodes": 50,
"total_frames": 16603,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 20,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/dexsuite/stack_franka_allegro.summarization-polish-summaries-corpuschats_personachatbiographies_ecthrallegro_reviews
Dataset Card for [Dataset Name]
Dataset Summary
Allegro Reviews is a sentiment analysis dataset, consisting of 11,588 product reviews written in Polish and extracted from Allegro.pl - a popular e-commerce marketplace. Each review contains at least 50 words and has a rating on a scale from one (negative review) to five (positive review).
We recommend using the provided train/dev/test split. The ratings for the test set reviews are kept hidden. You can evaluate your model… See the full description on the dataset page: https://huggingface.co/datasets/legacy-datasets/allegro_reviews.drill_to_point_franka_allegroThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": null,
"total_episodes": 50,
"total_frames": 6503,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 20,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/dexsuite/drill_to_point_franka_allegro.klej-cbdtestset_ellieklej-cdsc-rpolish-question-passage-pairssummarization-allegro-articlesparaphrases_paws
