CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01allegrolab /dclm-baseline-500b_toks DCLM Baseline 500B Tokens (Decontaminated) Dataset Description This dataset is a decontaminated subset of the DCLM-Baseline corpus, specifically prepared for the Hubble memorization research project. The dataset has been carefully processed to remove overlap with memorization evaluation data and subsampled around 500 billion tokens of English text. This corpus serves as the foundational training data for all Hubble models, providing a clean baseline for studying… See the full description on the dataset page: https://huggingface.co/datasets/allegrolab/dclm-baseline-500b_toks.100B<n<1T0 likes2.3k downloads11mo agoHugging Face02allegro /klej-polemo2-in klej-polemo2-in Description The PolEmo2.0 is a dataset of online consumer reviews from four domains: medicine, hotels, products, and university. It is human-annotated on a level of full reviews and individual sentences. It comprises over 8000 reviews, about 85% from the medicine and hotel domains. We use the PolEmo2.0 dataset to form two tasks. Both use the same training dataset, i.e., reviews from medicine and hotel domains, but are evaluated on a different test set.… See the full description on the dataset page: https://huggingface.co/datasets/allegro/klej-polemo2-in.texttext-classification1K<n<10K0 likes641 downloads4y agoHugging Face03allegro /klej-polemo2-out klej-polemo2-out Description The PolEmo2.0 is a dataset of online consumer reviews from four domains: medicine, hotels, products, and university. It is human-annotated on a level of full reviews and individual sentences. It comprises over 8000 reviews, about 85% from the medicine and hotel domains. We use the PolEmo2.0 dataset to form two tasks. Both use the same training dataset, i.e., reviews from medicine and hotel domains, but are evaluated on a different test set.… See the full description on the dataset page: https://huggingface.co/datasets/allegro/klej-polemo2-out.texttext-classification1K<n<10K0 likes550 downloads4y agoHugging Face04allegro /klej-psc klej-psc Description The Polish Summaries Corpus (PSC) is a dataset of summaries for 569 news articles. The human annotators created five extractive summaries for each article by choosing approximately 5% of the original text. A different annotator created each summary. The subset of 154 articles was also supplemented with additional five abstractive summaries each, i.e., not created from the fragments of the original article. In huggingface version of this dataset… See the full description on the dataset page: https://huggingface.co/datasets/allegro/klej-psc.texttext-classification1K<n<10K0 likes517 downloads4y agoHugging Face05allegro /klej-dyk klej-dyk Description The Czy wiesz? (eng. Did you know?) the dataset consists of almost 5k question-answer pairs obtained from Czy wiesz... section of Polish Wikipedia. Each question is written by a Wikipedia collaborator and is answered with a link to a relevant Wikipedia article. In huggingface version of this dataset, they chose the negatives which have the largest token overlap with a question. Tasks (input, output, and metrics) The task is to predict if… See the full description on the dataset page: https://huggingface.co/datasets/allegro/klej-dyk.textquestion-answering1K<n<10K1 likes495 downloads4y agoHugging Face06allegrolab /passages_gutenberg_populartext1K<n<10K0 likes434 downloads1y agoHugging Face07allegrolab /passages_gutenberg_unpopulartext1K<n<10K0 likes406 downloads1y agoHugging Face08allegro /klej-nkjp-nertext10K<n<100K0 likes396 downloads5y agoHugging Face09allegrolab /testset_piqatext1K<n<10K0 likes396 downloads1y agoHugging Face10allegrolab /passages_wikipediatext1K<n<10K0 likes387 downloads1y agoHugging Face11allegrolab /testset_popqatext1K<n<10K0 likes383 downloads1y agoHugging Face12allegro /klej-cdsc-e klej-cdsc-e Description Polish CDSCorpus consists of 10K Polish sentence pairs which are human-annotated for semantic relatedness (CDSC-R) and entailment (CDSC-E). The dataset may be used to evaluate compositional distributional semantics models of Polish. The dataset was presented at ACL 2017. Although the SICK corpus inspires the main design of the dataset, it differs in detail. As in SICK, the sentences come from image captions, but the set of chosen images is much… See the full description on the dataset page: https://huggingface.co/datasets/allegro/klej-cdsc-e.texttext-classification10K<n<100K1 likes373 downloads4y agoHugging Face13allegrolab /biographies_yagotext1K<n<10K0 likes316 downloads1y agoHugging Face14allegro /klej-allegro-reviewstext10K<n<100K1 likes314 downloads5y agoHugging Face15allegrolab /testset_mmlutext1K<n<10K0 likes312 downloads1y agoHugging Face16dexsuite /pick_place_fruit_franka_allegroThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": null, "total_episodes": 50, "total_frames": 8123, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 500, "fps": 20, "splits": { "train": "0:50" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/dexsuite/pick_place_fruit_franka_allegro.tabularrobotics1K<n<10K0 likes289 downloads3mo agoHugging Face17allegrolab /testset_hellaswagtext1K<n<10K0 likes280 downloads1y agoHugging Face18allegrolab /testset_winogrande-infilltext1K<n<10K0 likes275 downloads1y agoHugging Face19dexsuite /stack_franka_allegroThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": null, "total_episodes": 50, "total_frames": 16603, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 500, "fps": 20, "splits": { "train": "0:50" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/dexsuite/stack_franka_allegro.tabularrobotics10K<n<100K0 likes220 downloads3mo agoHugging Face20allegro /summarization-polish-summaries-corpustext10K<n<100K5 likes193 downloads5y agoHugging Face21allegrolab /chats_personachattext1K<n<10K0 likes192 downloads1y agoHugging Face22allegrolab /biographies_ecthrtext1K<n<10K0 likes190 downloads1y agoHugging Face23legacy-datasets /allegro_reviews Dataset Card for [Dataset Name] Dataset Summary Allegro Reviews is a sentiment analysis dataset, consisting of 11,588 product reviews written in Polish and extracted from Allegro.pl - a popular e-commerce marketplace. Each review contains at least 50 words and has a rating on a scale from one (negative review) to five (positive review). We recommend using the provided train/dev/test split. The ratings for the test set reviews are kept hidden. You can evaluate your model… See the full description on the dataset page: https://huggingface.co/datasets/legacy-datasets/allegro_reviews.texttext-classification10K<n<100K7 likes183 downloads3y agoHugging Face24dexsuite /drill_to_point_franka_allegroThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": null, "total_episodes": 50, "total_frames": 6503, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 500, "fps": 20, "splits": { "train": "0:50" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/dexsuite/drill_to_point_franka_allegro.tabularrobotics1K<n<10K0 likes178 downloads3mo agoHugging Face25allegro /klej-cbdtext10K<n<100K0 likes164 downloads5y agoHugging Face26allegrolab /testset_ellietextn<1K0 likes150 downloads1y agoHugging Face27allegro /klej-cdsc-rtabular10K<n<100K0 likes148 downloads5y agoHugging Face28allegro /polish-question-passage-pairstext10K<n<100K5 likes145 downloads5y agoHugging Face29allegro /summarization-allegro-articlestext100K<n<1M5 likes139 downloads5y agoHugging Face30allegrolab /paraphrases_pawstext1K<n<10K0 likes136 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.