CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01google-research-datasets /go_emotions Dataset Card for GoEmotions Dataset Summary The GoEmotions dataset contains 58k carefully curated Reddit comments labeled for 27 emotion categories or Neutral. The raw data is included as well as the smaller, simplified version of the dataset with predefined train/val/test splits. Supported Tasks and Leaderboards This dataset is intended for multi-class, multi-label emotion classification. Languages The data is in English. Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/go_emotions.tabulartext-classification100K<n<1M267 likes13k downloads3y agoHugging Face02google /civil_comments Dataset Card for "civil_comments" Dataset Summary The comments in this dataset come from an archive of the Civil Comments platform, a commenting plugin for independent news sites. These public comments were created from 2015 - 2017 and appeared on approximately 50 English-language news sites across the world. When Civil Comments shut down in 2017, they chose to make the public comments available in a lasting open archive to enable future research. The original data… See the full description on the dataset page: https://huggingface.co/datasets/google/civil_comments.tabulartext-classification1M<n<10M40 likes9.2k downloads3y agoHugging Face03tmquan /anle-toaan-gov-vn Vietnamese Án lệ Corpus — anle.toaan.gov.vn 🇻🇳 Tóm tắt. Bộ dữ liệu các bản án + án lệ Việt Nam thu thập từ cổng anle.toaan.gov.vn của Tòa án nhân dân tối cao. Mỗi văn bản đi kèm markdown chuẩn hoá tiếng Việt và một lớp grounding mức câu (mỗi trích dẫn mang sentence_id + char span trỏ ngược vào markdown). Bộ dữ liệu là một phần của ViLA common-corpus và ship ba cấu hình HF theo chuẩn chung: documents (bảng chính) · embeddings (vector 4096-D Nemotron-3-Embed-8B) · reduces (toạ… See the full description on the dataset page: https://huggingface.co/datasets/tmquan/anle-toaan-gov-vn.tabulartext-classification10K<n<100K10 likes8k downloads6d agoHugging Face04IPEC-COMMUNITY /libero_goal_no_noops_1.0.0_lerobottabular10K<n<100K1 likes6.1k downloads11mo agoHugging Face05BEE-spoke-data /govdocs1-pdf-source govdocs1: source PDF files [!NOTE] Converted versions of other document types (word, txt, etc) are available in this repo This is ~220,000 open-access PDF documents (about 6.6M pages) from the dataset govdocs1. It wants to be OCR'd. Uploaded as tar file pieces of ~10 GiB each due to size/file count limits with an index.csv covering details 5,000 randomly sampled PDFs are available unarchived in sample/. Hugging Face supports previewing these in-browser, for example this one… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/govdocs1-pdf-source.documentimage-text-to-text100K<n<1M6 likes4.4k downloads9mo agoHugging Face06ZombitX64 /xauusd-gold-price-historical-data-2004-2025 XAUUSD Gold Price Historical Data 2004-2025 This dataset contains historical price data for XAUUSD (Gold vs US Dollar) from 2004 to 2025. Source: Kaggle dataset "novandraanugrah/xauusd-gold-price-historical-data-2004-2024" Content: The dataset includes CSV files with different time granularities (e.g., 1 minute, 5 minutes, 1 hour, 1 day). Each file typically contains the following columns: Date Open High Low Close Volume Usage: This dataset can be used for analyzing historical… See the full description on the dataset page: https://huggingface.co/datasets/ZombitX64/xauusd-gold-price-historical-data-2004-2025.tabular1M<n<10M12 likes2.8k downloads1y agoHugging Face07AgentPublic /open_government Open Government Dataset Open Government is the largest agregation of governement text and data made available as part of open data programs. In total, the dataset contains approximately 380B tokens. While Open Government aims to become a global resource, in its current state it mostly features open datasets from the US, France, European and international organizations. The dataset comprises 16 collections curated through two different initiaties: Finance commons and Legal commons.… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/open_government.tabulartext-generation10M<n<100M4 likes2.4k downloads2y agoHugging Face08GonzaloA /fake_news TODO: Add YAML tags here. Copy-paste the tags obtained with the online tagging app: https://huggingface.co/spaces/huggingface/datasets-tagging annotations_creators: - no-annotation language_creators: - found language: - en license: - unknown multilinguality: - monolingual size_categories: - 30k<n<50k source_datasets: - original task_categories: - text-classification task_ids: - fact-checking - intent-classification pretty_name: GonzaloA / Fake News Dataset Card for… See the full description on the dataset page: https://huggingface.co/datasets/GonzaloA/fake_news.tabular10K<n<100K28 likes1.9k downloads4y agoHugging Face09gridfm /opf_small_case10000_goctabular100M<n<1B0 likes1.8k downloads2mo agoHugging Face10gridfm /opf_small_case2000_goc Data download through hfApi Retry download if you receive '502 Server Error'. For larger datasets, you may need to retry download multiple times. tabulartabular-regression1B<n<10B0 likes1.7k downloads8mo agoHugging Face11vidore /syntheticDocQA_government_reports_test_beirBEIR version of vidore/syntheticDocQA_government_reports_test. imagedocument-question-answering1K<n<10K1 likes1.6k downloads1y agoHugging Face12google /MusicCaps Dataset Card for MusicCaps Dataset Summary The MusicCaps dataset contains 5,521 music examples, each of which is labeled with an English aspect list and a free text caption written by musicians. An aspect list is for example "pop, tinny wide hi hats, mellow piano melody, high pitched female vocal melody, sustained pulsating synth lead", while the caption consists of multiple sentences about the music, e.g., "A low sounding male voice is rapping over a fast paced drums… See the full description on the dataset page: https://huggingface.co/datasets/google/MusicCaps.tabulartext-to-speech1K<n<10K153 likes1.5k downloads4y agoHugging Face13gridfm /opfdata_case2000_goctabular1B<n<10B0 likes1.4k downloads5mo agoHugging Face14Godlydonuts /Sai Sai Initiative Sai is Project Shohin's return to its original objective: build the strongest practical model near four billion parameters. This repository is the live scratchpad and implementation surface for that effort. Nothing is called an improvement until it beats the unchanged parent and an equal-compute control on real, source-disjoint benchmarks. Data precedes architecture. Sai first earns a trustworthy learning sequence: verified source bytes, quality and duplication… See the full description on the dataset page: https://huggingface.co/datasets/Godlydonuts/Sai.tabular10M<n<100M1 likes1.3k downloads27d agoHugging Face15RIPS-Goog-23 /IIT-CDIP Dataset Card for "IIT-CDIP-2" More Information needed tabular1M<n<10M10 likes1.3k downloads3y agoHugging Face16gridfm /pf_small_case2000_goc Data download through hfApi Retry download if you receive '502 Server Error'. For larger datasets, you may need to retry download multiple times. tabulartabular-regression1B<n<10B0 likes1.2k downloads8mo agoHugging Face17gridfm /opf_small_case500_goc Data download through hfApi Retry download if you receive '502 Server Error'. For larger datasets, you may need to retry download multiple times. tabulartabular-regression100M<n<1B0 likes1.2k downloads8mo agoHugging Face18tmquan /phapdien-moj-gov-vn Bộ Pháp Điển Việt Nam — phapdien.moj.gov.vn 🇻🇳 Tóm tắt. Bộ ngữ liệu cấp Điều của Bộ Pháp Điển Việt Nam — bộ pháp điển chính thức do Bộ Tư pháp công bố. Mỗi dòng documents là một Điều kèm toàn văn đã chuẩn hoá, chương sở thuộc, đề mục và chủ đề. Kèm theo là vector nhúng ngữ nghĩa 4096-D (embeddings), toạ độ giảm chiều trong không gian chung ViLA (reduces), và từ điển ontology song ngữ Việt–Anh (chủ đề · đề mục · thuật ngữ). 🇬🇧 One-line. Article-level corpus of the Bộ Pháp… See the full description on the dataset page: https://huggingface.co/datasets/tmquan/phapdien-moj-gov-vn.imagetext-classification100K<n<1M11 likes1.2k downloads6d agoHugging Face19google-research-datasets /discofuse Dataset Card for "discofuse" Dataset Summary DiscoFuse is a large scale dataset for discourse-based sentence fusion. Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure Data Instances discofuse-sport Size of downloaded dataset files: 4.33 GB Size of the generated dataset: 15.04 GB Total amount of disk used: 19.36 GB An example of 'train' looks as follows. {… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/discofuse.tabular10M<n<100M6 likes1.2k downloads3y agoHugging Face20gridfm /pf_small_case10000_goctabular10B<n<100B0 likes1.1k downloads3mo agoHugging Face21tmquan /nso-gov-vn nso-gov-vn — Vietnam NSO PX-Web mirror · Bản sao PX-Web của Tổng cục Thống kê 🇻🇳 Tóm tắt. Bản sao đầy đủ, từng bảng một, của cơ sở dữ liệu thống kê PX-Web của Tổng cục Thống kê Việt Nam (NSO / GSO) tại https://pxweb.nso.gov.vn. Mỗi ma trận PX-Web (multi-dimensional data cube) được phơi ra cùng lúc ở (i) bản gốc với schema riêng và (ii) định dạng long-format gộp chung, để bạn có thể chọn giữa "nguyên trạng" hay "join sẵn". 🇬🇧 Summary. A complete, table-by-table mirror of the… See the full description on the dataset page: https://huggingface.co/datasets/tmquan/nso-gov-vn.tabulartabular-classification100K<n<1M0 likes1k downloads5mo agoHugging Face22pfaha /goodreads-books Goodreads Books Metadata Dataset Description Goodreads Books Metadata is a structured dataset of book records scraped directly from Goodreads, a social platform for book readers and recommendations. The dataset is collected via an ongoing, resumable crawl and contains rich metadata per book: bibliographic information, crowd-sourced ratings, contributor (author/illustrator/editor/etc.) details enriched with author-level popularity stats, genre tags, series… See the full description on the dataset page: https://huggingface.co/datasets/pfaha/goodreads-books.tabulartabular-regression100K<n<1M0 likes1k downloads9h agoHugging Face23gridfm /pf_small_case500_goc Data download through hfApi Retry download if you receive '502 Server Error'. For larger datasets, you may need to retry download multiple times. tabulartabular-regression100M<n<1B0 likes1k downloads8mo agoHugging Face24PaulQ1 /TS_Instruct_QA_Gold_v2Welcome to the TS_Instruct_QA_Gold dataset This dataset is intended to evaluate time-series reasoning. The dataset consists of real-world time-series with synthetic text. The dataset was human evaluated for correctness Note this version is contains only the needed files and is therefore smaller in download size/number of files and may play nicer with the HF API @misc{quinlan2025chattsenhancingmultimodalreasoning, title={Chat-TS: Enhancing Multi-Modal Reasoning Over Time-Series and… See the full description on the dataset page: https://huggingface.co/datasets/PaulQ1/TS_Instruct_QA_Gold_v2.tabular1K<n<10K0 likes994 downloads8mo agoHugging Face25GOD111111111 /synthetic-timeseries-data cruscy data — evaluation sample Three full days of real crypto market microstructure (Binance spot), prepared for public evaluation: absolute prices, dates, and the instrument are withheld — the shape of the day (tick-by-tick relative price, normalized volumes, trade side, book imbalance) is fully preserved. The full feed — 27+ streams (raw L2 depth, 1-second trade tape, order-book metrics, derived features, regime labels) with SQL console, backtest runner and MCP access for AI… See the full description on the dataset page: https://huggingface.co/datasets/GOD111111111/synthetic-timeseries-data.tabular1B<n<10B0 likes974 downloads6h agoHugging Face26mteb /syntheticDocQA_government_reports_test_beirBEIR version of vidore/syntheticDocQA_government_reports_test. imagedocument-question-answering1K<n<10K0 likes947 downloads7mo agoHugging Face27google /code_x_glue_cc_clone_detection_big_clone_bench Dataset Card for "code_x_glue_cc_clone_detection_big_clone_bench" Dataset Summary CodeXGLUE Clone-detection-BigCloneBench dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Code-Code/Clone-detection-BigCloneBench Given two codes as the input, the task is to do binary classification (0/1), where 1 stands for semantic equivalence and 0 for others. Models are evaluated by F1 score. The dataset we use is BigCloneBench and filtered following the paper… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_cc_clone_detection_big_clone_bench.tabulartext-classification1M<n<10M22 likes943 downloads3y agoHugging Face28LucasFang /Laion-Aesthetics-High-Resolution-GoT Laion-Aesthetics-High-Resolution-GoT Paper Dataset Description The Laion-Aesthetics-High-Resolution-GoT dataset is a collection of 3.77 million image-text pairs with rich grounding annotations. This dataset extends high-quality images from the LAION-Aesthetics collection with detailed text descriptions and object-level grounding information. Key Features Size: 3.77 million samples Modalities: Image, Text, and Grounding Annotations Image Resolution:… See the full description on the dataset page: https://huggingface.co/datasets/LucasFang/Laion-Aesthetics-High-Resolution-GoT.image1M<n<10M12 likes909 downloads2y agoHugging Face29GotThatData /kraken-trading-data 📈 Kraken Trading Data Collection Overview High-frequency cryptocurrency market data from Kraken exchange - perfect for algorithmic trading, time-series forecasting, and market microstructure analysis. This dataset includes real-time price, volume, and order book data for 9 major cryptocurrency trading pairs, collected via WebSocket streaming and REST API polling. 📊 Included Trading Pairs Pair Asset Base Currency Typical Daily Volume XXBTZUSD… See the full description on the dataset page: https://huggingface.co/datasets/GotThatData/kraken-trading-data.tabular10K<n<100K6 likes881 downloads8mo agoHugging Face30nbettencourt /google-patents-data-previewtabular100K<n<1M0 likes878 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.