CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01google-research-datasets /go_emotions Dataset Card for GoEmotions Dataset Summary The GoEmotions dataset contains 58k carefully curated Reddit comments labeled for 27 emotion categories or Neutral. The raw data is included as well as the smaller, simplified version of the dataset with predefined train/val/test splits. Supported Tasks and Leaderboards This dataset is intended for multi-class, multi-label emotion classification. Languages The data is in English. Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/go_emotions.tabulartext-classification100K<n<1M267 likes14k downloads3y agoHugging Face02google /civil_comments Dataset Card for "civil_comments" Dataset Summary The comments in this dataset come from an archive of the Civil Comments platform, a commenting plugin for independent news sites. These public comments were created from 2015 - 2017 and appeared on approximately 50 English-language news sites across the world. When Civil Comments shut down in 2017, they chose to make the public comments available in a lasting open archive to enable future research. The original data… See the full description on the dataset page: https://huggingface.co/datasets/google/civil_comments.tabulartext-classification1M<n<10M40 likes9.5k downloads3y agoHugging Face03google /MusicCaps Dataset Card for MusicCaps Dataset Summary The MusicCaps dataset contains 5,521 music examples, each of which is labeled with an English aspect list and a free text caption written by musicians. An aspect list is for example "pop, tinny wide hi hats, mellow piano melody, high pitched female vocal melody, sustained pulsating synth lead", while the caption consists of multiple sentences about the music, e.g., "A low sounding male voice is rapping over a fast paced drums… See the full description on the dataset page: https://huggingface.co/datasets/google/MusicCaps.tabulartext-to-speech1K<n<10K154 likes1.5k downloads4y agoHugging Face04google-research-datasets /discofuse Dataset Card for "discofuse" Dataset Summary DiscoFuse is a large scale dataset for discourse-based sentence fusion. Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure Data Instances discofuse-sport Size of downloaded dataset files: 4.33 GB Size of the generated dataset: 15.04 GB Total amount of disk used: 19.36 GB An example of 'train' looks as follows. {… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/discofuse.tabular10M<n<100M6 likes1.3k downloads3y agoHugging Face05RIPS-Goog-23 /IIT-CDIP Dataset Card for "IIT-CDIP-2" More Information needed tabular1M<n<10M10 likes1.2k downloads3y agoHugging Face06google /code_x_glue_cc_clone_detection_big_clone_bench Dataset Card for "code_x_glue_cc_clone_detection_big_clone_bench" Dataset Summary CodeXGLUE Clone-detection-BigCloneBench dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Code-Code/Clone-detection-BigCloneBench Given two codes as the input, the task is to do binary classification (0/1), where 1 stands for semantic equivalence and 0 for others. Models are evaluated by F1 score. The dataset we use is BigCloneBench and filtered following the paper… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_cc_clone_detection_big_clone_bench.tabulartext-classification1M<n<10M22 likes1k downloads3y agoHugging Face07nbettencourt /google-patents-data-previewtabular100K<n<1M0 likes802 downloads1y agoHugging Face08AdnanElAssadi /Google-Translated_Turkish_GPQA_Datasettabular1K<n<10K0 likes485 downloads2y agoHugging Face09google /WikiProfile WikiProfile WikiProfile is a factual knowledge benchmark for evaluating how well language models encode and recall factual knowledge. It comprises 2,150 facts, each paired with 10 questions, for a total of 21,500 question instances. Each fact is grounded in the first paragraph (summary) of an English Wikipedia page and is defined as a proposition between two entities, a subject and an object (e.g., "Oasis played their first gig at the Boardwalk club" → subject: Oasis, object:… See the full description on the dataset page: https://huggingface.co/datasets/google/WikiProfile.tabularquestion-answering1K<n<10K20 likes474 downloads3mo agoHugging Face10aurman /GoogleTrendArchive Google Trend Archive: Global Real-Time Search Trends (2024-2026) Dataset Details Dataset Description This dataset contains over 10.2 million trending search instances from Google's Trending Now feature, collected continuously from November 28, 2024 to May 17, 2026 across all available geographic locations (200+ countries/regions). Unlike aggregated retrospective tools like Google Trends, Trending Now captures search queries experiencing real-time… See the full description on the dataset page: https://huggingface.co/datasets/aurman/GoogleTrendArchive.tabulartext-classification10M<n<100M5 likes422 downloads4mo agoHugging Face11google /code_x_glue_tc_nl_code_search_adv Dataset Card for "code_x_glue_tc_nl_code_search_adv" Dataset Summary CodeXGLUE NL-code-search-Adv dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Text-Code/NL-code-search-Adv The dataset we use comes from CodeSearchNet and we filter the dataset as the following: Remove examples that codes cannot be parsed into an abstract syntax tree. Remove examples that #tokens of documents is < 3 or >256 Remove examples that documents contain special tokens… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_tc_nl_code_search_adv.tabulartext-retrieval100K<n<1M11 likes374 downloads3y agoHugging Face12UniqueData /messengers-reviews-google-play Reviews on Messengers Dataset - Review dataset The Reviews on Messengers Dataset is a comprehensive collection of 200 the most recent customer reviews on 6 messengers obtained from the popular app store, Google Play. See the list of the apps below. This dataset encompasses reviews written in 5 different languages: English, French, German, Italian, Japanese. 💴 For Commercial Usage: To discuss your requirements, learn about the price and buy the dataset, leave a request… See the full description on the dataset page: https://huggingface.co/datasets/UniqueData/messengers-reviews-google-play.tabulartext-classification1K<n<10K3 likes329 downloads1y agoHugging Face13google /witWikipedia-based Image Text (WIT) Dataset is a large multimodal multilingual dataset. WIT is composed of a curated set of 37.6 million entity rich image-text examples with 11.5 million unique images across 108 Wikipedia languages. Its size enables WIT to be used as a pretraining dataset for multimodal machine learning models.imagetext-retrieval1M<n<10M71 likes303 downloads4y agoHugging Face14Gopalatius /google-play-reviewtabular100K<n<1M1 likes289 downloads3y agoHugging Face15OALL /details_google__gemma-7b-it Dataset Card for Evaluation run of google/gemma-7b-it Dataset automatically created during the evaluation run of model google/gemma-7b-it. The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_google__gemma-7b-it.tabular100K<n<1M0 likes255 downloads2y agoHugging Face16RIPS-Goog-23 /DocVQAtabular10K<n<100K0 likes222 downloads3y agoHugging Face17dhruv-anand-aintech /google-books-ngram-pos Google Books Ngram — POS-tagged & cleaned A cleaned, analysis-ready slice of the Google Books Ngram corpus v3 (20200217, English) with part-of-speech tags preserved, packaged as Parquet for easy use with 🤗 datasets, pandas, DuckDB, or Polars. This powers the POS / regex Ngram Viewer — an ngram viewer that supports part-of-speech template queries (love *_NOUN) and regex (/ousness$/_NOUN), patterns the official Google viewer cannot express. Configs config… See the full description on the dataset page: https://huggingface.co/datasets/dhruv-anand-aintech/google-books-ngram-pos.tabulartext-classification1M<n<10M0 likes175 downloads3mo agoHugging Face18visheratin /google_landmarks_places Google Landmarks places Google Landmarks is a great dataset, but it lacks geospatial information about the places. This dataset fills this gap by providing latitude and longitude for each landmark. The dataset also contains the name of the landmark from OpenStreetMap and information about the country, the province/state, and the city/village where the landmark is located. This information was collected from OSM via Nominatim. tabular10K<n<100K4 likes170 downloads3y agoHugging Face19lara-popovic /google_play_store_reviewstabularn<1K0 likes153 downloads2mo agoHugging Face20google /granola-entity-questions GRANOLA Entity Questions Dataset Card Dataset details Dataset Name: GRANOLA-EQ (Granularity of Labels Entity Questions) Paper: Narrowing the Knowledge Evaluation Gap: Open-Domain Question Answering with Multi-Granularity Answers Abstract: Factual questions typically can be answered correctly at different levels of granularity. For example, both "August 4, 1961" and "1961" are correct answers to the question "When was Barack Obama born?"". Standard question answering (QA)… See the full description on the dataset page: https://huggingface.co/datasets/google/granola-entity-questions.tabularquestion-answering10K<n<100K12 likes144 downloads2y agoHugging Face21megloughney /googleAnalyticsCustomerRevenuePredictiontabular10K<n<100K0 likes140 downloads9mo agoHugging Face22OALL /details_google__gemma-7b Dataset Card for Evaluation run of google/gemma-7b Dataset automatically created during the evaluation run of model google/gemma-7b. The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_google__gemma-7b.tabular100K<n<1M0 likes134 downloads2y agoHugging Face23OALL /details_google__gemma-1.1-7b-it Dataset Card for Evaluation run of google/gemma-1.1-7b-it Dataset automatically created during the evaluation run of model google/gemma-1.1-7b-it. The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_google__gemma-1.1-7b-it.tabular100K<n<1M0 likes129 downloads2y agoHugging Face24najeh-halawani /google-ads-transparencytabular10M<n<100M0 likes123 downloads6d agoHugging Face25OALL /details_google__gemma-2b-it Dataset Card for Evaluation run of google/gemma-2b-it Dataset automatically created during the evaluation run of model google/gemma-2b-it. The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_google__gemma-2b-it.tabular100K<n<1M0 likes120 downloads2y agoHugging Face26google /LoraxBench LoraxBench: A Benchmark for Indonesian Local Languages and Registers Dataset Summary LoraxBench is a comprehensive multilingual benchmark focusing on Indonesian and 19 Indonesian local languages, covering 6 diverse NLP tasks. It includes multiple registers for select languages, emphasizing the impact of formal and casual speech on model performance. LoraxBench is professionally translated and validated by natives, and were sourced from Indonesian-originated dataset, Our… See the full description on the dataset page: https://huggingface.co/datasets/google/LoraxBench.tabular10K<n<100K7 likes120 downloads1y agoHugging Face27phospho-app /010_pickplace_googleball_3Cam_MQ_bboxes 010_pickplace_googleball_3Cam_MQ This dataset was generated using phosphobot. This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot. To get started in robotics, get your own phospho starter pack.. tabularrobotics1K<n<10K0 likes98 downloads1y agoHugging Face28opdullah /turkish-google-maps-15M Turkish Google Maps Reviews Bu veri seti, Türkiye’deki işletmelere ait Türkçe Google Maps yorumlarını içerir. Her kayıt: yorum metni yorum puanı işletme adı işletme kategorisi gibi bilgileri içerir. Veri seti, özellikle büyük ölçekli Türkçe NLP çalışmaları için uygundur. Contents Veri setinde aşağıdaki türde alanlar bulunmaktadır: yorum metni (review_text) yorum puanı (rating) işletme adı (place_name) işletme kategorisi (category) kategori listesi (category_list)… See the full description on the dataset page: https://huggingface.co/datasets/opdullah/turkish-google-maps-15M.tabulartext-classification10M<n<100M12 likes93 downloads6mo agoHugging Face29ronantakizawa /trending-words-google Google Trending Words Dataset (2001-2024) Dataset Description This dataset contains Google trending words and search terms from 2001 to 2024, capturing 24 years of internet culture, major events, and global trends. The dataset includes 2,784 entries across 93 standardized categories, providing a comprehensive view of what captured the world's attention over more than two decades. Dataset Summary Total Entries: 2,784 Years Covered: 2001-2024 (24 years)… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/trending-words-google.tabulartext-classification1K<n<10K4 likes92 downloads10mo agoHugging Face30visheratin /google_landmarks_photos Dataset Card for "google_landmarks_photos" More Information needed image1M<n<10M7 likes88 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.