CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ieasybooks-org /waqfeya-library Waqfeya Library 📖 Overview Waqfeya is one of the primary online resources for Islamic books, similar to Shamela. It hosts more than 10,000 PDF books across over 80 categories. In this dataset, we processed the original PDF files using Google Document AI APIs and extracted their contents into two additional formats: TXT and DOCX. 📊 Dataset Contents The dataset includes 22,443 PDF files (spanning 8,978,634 pages) representing 10,150 Islamic books. Each book is… See the full description on the dataset page: https://huggingface.co/datasets/ieasybooks-org/waqfeya-library.imageimage-to-text10K<n<100K12 likes135k downloads1y agoHugging Face02ieasybooks-org /shamela-waqfeya-library Shamela Waqfeya Library 📖 Overview Shamela Waqfeya is one of the primary online resources for Islamic books, similar to Shamela. It hosts more than 4,500 PDF books across over 40 categories. In this dataset, we processed the original PDF files using Google Document AI APIs and extracted their contents into two additional formats: TXT and DOCX. 📊 Dataset Contents The dataset includes 12,877 PDF files (spanning 5,138,027 pages) representing 4,661 Islamic books.… See the full description on the dataset page: https://huggingface.co/datasets/ieasybooks-org/shamela-waqfeya-library.tabularimage-to-text1K<n<10K4 likes92k downloads1y agoHugging Face03thuml /Time-Series-Library Time-Series-Library (TSLib) TSLib is an open-source library for deep learning researchers, especially for deep time series analysis. We provide a neat code base to evaluate advanced deep time series models or develop your model, which covers five mainstream tasks: long- and short-term forecasting, imputation, anomaly detection, and classification. This benchmark collection is designed to evaluate and develop advanced deep time-series models. For an in-depth exploration of current… See the full description on the dataset page: https://huggingface.co/datasets/thuml/Time-Series-Library.tabulartime-series-forecasting1M<n<10M9 likes22k downloads11mo agoHugging Face04biglam /british-library-book-images British Library Book Images 1,080,814 images cut out of 49,455 digitised books (65,227 volumes, ~25 million pages) published between c. 1510 and c. 1900, digitised by the British Library in partnership with Microsoft and released by British Library Labs on Flickr Commons as the "1 Million Images from Scanned Books" release. The books cover geography, philosophy, history, poetry and literature, in several languages. The four image types British Library Labs… See the full description on the dataset page: https://huggingface.co/datasets/biglam/british-library-book-images.imageimage-classification1M<n<10M64 likes6.2k downloads1mo agoHugging Face05Faizaniqbal /british-library-book-images British Library Book Images 1,080,814 images cut out of 49,455 digitised books (65,227 volumes, ~25 million pages) published between c. 1510 and c. 1900, digitised by the British Library in partnership with Microsoft and released by British Library Labs on Flickr Commons as the "1 Million Images from Scanned Books" release. The books cover geography, philosophy, history, poetry and literature, in several languages. The four image types British Library Labs… See the full description on the dataset page: https://huggingface.co/datasets/Faizaniqbal/british-library-book-images.imageimage-classification1M<n<10M0 likes2.6k downloads1mo agoHugging Face06ieasybooks-org /waqfeya-library-compressed Waqfeya Library - Compressed 📖 Overview Waqfeya is one of the primary online resources for Islamic books, similar to Shamela. It hosts more than 10,000 PDF books across over 80 categories. In this dataset, we processed the original PDF files using Google Document AI APIs and extracted their contents into two additional formats: TXT and DOCX. 📊 Dataset Contents This dataset is identical to ieasybooks-org/waqfeya-library, with one key difference: the contents… See the full description on the dataset page: https://huggingface.co/datasets/ieasybooks-org/waqfeya-library-compressed.tabularimage-to-text10K<n<100K6 likes2k downloads1y agoHugging Face07open-index /open-library Open Library The complete Open Library catalog in clean, analysis-ready Parquet. 150.0M+ records across 11 entity types, from ISBNs and author bios to reading logs and Wikidata links. What is it? Open Library is a complete snapshot of the Open Library database, an open project of the Internet Archive with the mission of creating "one web page for every book ever published." The catalog is community-edited and contains bibliographic records for millions of authors, works… See the full description on the dataset page: https://huggingface.co/datasets/open-index/open-library.tabulartext-generation100M<n<1B9 likes490 downloads6mo agoHugging Face08davnas /library-occupancytabular1K<n<10K0 likes339 downloads1y agoHugging Face09fluxae /Time-Series-Library Time-Series-Library (TSLib) TSLib is an open-source library for deep learning researchers, especially for deep time series analysis. We provide a neat code base to evaluate advanced deep time series models or develop your model, which covers five mainstream tasks: long- and short-term forecasting, imputation, anomaly detection, and classification. This benchmark collection is designed to evaluate and develop advanced deep time-series models. For an in-depth exploration of… See the full description on the dataset page: https://huggingface.co/datasets/fluxae/Time-Series-Library.tabulartime-series-forecasting1M<n<10M0 likes329 downloads2mo agoHugging Face10Geenn2026 /Time-Series-Library Time-Series-Library (TSLib) TSLib is an open-source library for deep learning researchers, especially for deep time series analysis. We provide a neat code base to evaluate advanced deep time series models or develop your model, which covers five mainstream tasks: long- and short-term forecasting, imputation, anomaly detection, and classification. This benchmark collection is designed to evaluate and develop advanced deep time-series models. For an in-depth exploration of current… See the full description on the dataset page: https://huggingface.co/datasets/Geenn2026/Time-Series-Library.tabulartime-series-forecasting1M<n<10M0 likes328 downloads8mo agoHugging Face11zxyang7 /Time-Series-Library Time-Series-Library (TSLib) TSLib is an open-source library for deep learning researchers, especially for deep time series analysis. We provide a neat code base to evaluate advanced deep time series models or develop your model, which covers five mainstream tasks: long- and short-term forecasting, imputation, anomaly detection, and classification. This benchmark collection is designed to evaluate and develop advanced deep time-series models. For an in-depth exploration of current… See the full description on the dataset page: https://huggingface.co/datasets/zxyang7/Time-Series-Library.tabulartime-series-forecasting1M<n<10M0 likes247 downloads9mo agoHugging Face12bettergovph /gov-library Philippine Legal Documents Dataset A comprehensive collection of Philippine legal documents from Lawphil.net, extracted from HTML to Markdown and organized for easy querying. Overview This dataset contains 114,340 legal documents spanning from 1900 to 2025, including: Jurisprudence (68,080 documents) - Supreme Court decisions Statutes (19,793 documents) - Republic Acts, Commonwealth Acts, Presidential Decrees, etc. Executive Issuances (26,458 documents) - Administrative… See the full description on the dataset page: https://huggingface.co/datasets/bettergovph/gov-library.tabular100K<n<1M1 likes230 downloads8mo agoHugging Face13Crownelius /GLM-5.2-CoT-Library GLM-5.2 — CoT Library A maintained mirror of publicly-available GLM-5.2 chain-of-thought datasets on Hugging Face — content-verified, deduplicated, and attributed to their original authors. Dataset Viewer | Parquet // what this is A maintained library — a community mirror of publicly-available GLM-5.2 CoT datasets, aggregated, validity-filtered and content-verified, with per-row source attribution in first_source_dataset. It is not Crownelius' own data — every… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/GLM-5.2-CoT-Library.tabulartext-generation10K<n<100K2 likes218 downloads2mo agoHugging Face14sufficiencylab /sufficiency-librarytabular10M<n<100M0 likes196 downloads6mo agoHugging Face15lalababa /Time-Series-Library Time-Series-Library (TSLib) TSLib is an open-source library for deep learning researchers, especially for deep time series analysis. We provide a neat code base to evaluate advanced deep time series models or develop your model, which covers five mainstream tasks: long- and short-term forecasting, imputation, anomaly detection, and classification. This benchmark collection is designed to evaluate and develop advanced deep time-series models. For an in-depth exploration of current… See the full description on the dataset page: https://huggingface.co/datasets/lalababa/Time-Series-Library.tabulartime-series-forecasting1M<n<10M0 likes193 downloads11mo agoHugging Face16tidy-finance /factor-library-grid Tidy Finance Factor Library: Specification Grid Lookup table that maps each specification id to its portfolio construction choices. Use it together with the Portfolio Returns dataset to select return series and to identify the choices behind each series. Dataset Details Dataset Description The grid contains 4,105,728 specifications for 179 sorting variables. Each row defines a complete set of construction choices: sample exclusions, lagging… See the full description on the dataset page: https://huggingface.co/datasets/tidy-finance/factor-library-grid.tabular1M<n<10M0 likes169 downloads10d agoHugging Face17Crownelius /Kimi-K3-CoT-Library Kimi K3 — CoT Library A maintained mirror of publicly-available Kimi K3 chain-of-thought datasets on Hugging Face — content-verified, deduplicated, and attributed to their original authors. Dataset Viewer | Parquet // what this is A maintained library — a community mirror of publicly-available Kimi K3 CoT datasets, aggregated, validity-filtered and content-verified, with per-row source attribution in first_source_dataset. It is not Crownelius' own data — every… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Kimi-K3-CoT-Library.tabulartext-generation1K<n<10K2 likes144 downloads2mo agoHugging Face18Verasight /verasight-data-library A source-linked index of what U.S. adults think The Verasight Data Library makes original U.S. public opinion research searchable and ready for analysis. Discover questions and weighted toplines across AI & Tech, Culture, Health, Money, Politics, Sports, then follow every record to a human-readable finding and its verified primary source report. Explore findings, search topics, and cite the research at data.verasight.io Coverage at a glance Survey waves… See the full description on the dataset page: https://huggingface.co/datasets/Verasight/verasight-data-library.tabular10K<n<100K1 likes142 downloads25d agoHugging Face19AdhyanshVerma /un-digital-library United Nations Digital Library (UNDL) Comprehensive Master Dataset 1. Executive Summary Welcome to the United Nations Digital Library (UNDL) Comprehensive Master Dataset repository. This dataset represents a monumental effort to harvest, normalize, enrich, and democratize access to the vast archives of the United Nations. By leveraging advanced web harvesting techniques, robust state management, and modern big-data formats, this repository provides researchers… See the full description on the dataset page: https://huggingface.co/datasets/AdhyanshVerma/un-digital-library.tabulartext-classification10K<n<100K0 likes127 downloads2mo agoHugging Face20ieasybooks-org /shamela-waqfeya-library-compressed Shamela Waqfeya Library - Compressed 📖 Overview Shamela Waqfeya is one of the primary online resources for Islamic books, similar to Shamela. It hosts more than 4,500 PDF books across over 40 categories. In this dataset, we processed the original PDF files using Google Document AI APIs and extracted their contents into two additional formats: TXT and DOCX. 📊 Dataset Contents This dataset is identical to ieasybooks-org/shamela-waqfeya-library, with one key… See the full description on the dataset page: https://huggingface.co/datasets/ieasybooks-org/shamela-waqfeya-library-compressed.tabularimage-to-text1K<n<10K0 likes118 downloads1y agoHugging Face21metrum-ai /prompt-library Metrum AI Prompt Library A prompt library for LLM inference workload and performance benchmarking, prepared for use with metrum-ai/bench-cli. It contains 593,730 records with prompt text, intended lengths, token buckets, and reasoning labels. It contains no reference answers. The full configuration preserves all source records, including repeated prompts and their distinct workload targets. Prompts may appear duplicated, with only target_output_length differing. These variants… See the full description on the dataset page: https://huggingface.co/datasets/metrum-ai/prompt-library.tabulartext-generation100K<n<1M1 likes113 downloads9d agoHugging Face22Leopegasus /Time-Series-Library Time-Series-Library (TSLib) TSLib is an open-source library for deep learning researchers, especially for deep time series analysis. We provide a neat code base to evaluate advanced deep time series models or develop your model, which covers five mainstream tasks: long- and short-term forecasting, imputation, anomaly detection, and classification. This benchmark collection is designed to evaluate and develop advanced deep time-series models. For an in-depth exploration of current… See the full description on the dataset page: https://huggingface.co/datasets/Leopegasus/Time-Series-Library.tabulartime-series-forecasting1M<n<10M0 likes91 downloads5mo agoHugging Face23davnas /real-time-library-occupancytabular10K<n<100K0 likes88 downloads1y agoHugging Face24Crownelius /Qwen-CoT-Library Qwen — CoT Library A maintained mirror of publicly-available Qwen chain-of-thought datasets on Hugging Face — content-verified, deduplicated, and attributed to their original authors. Dataset Viewer | Parquet // what this is A maintained library — a community mirror of publicly-available Qwen CoT datasets, aggregated, validity-filtered and content-verified, with per-row source attribution in first_source_dataset. It is not Crownelius' own data — every row… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Qwen-CoT-Library.tabulartext-generation10K<n<100K0 likes85 downloads2mo agoHugging Face25algorembrant /proxyquotes_library The Proxy Quotes (pxyq) library includes a fuction for calling the cell value with respect to a column and row of the csv dataset table. It calls for proxy stoploss distance, lotsize, and margins with leverages covering a betsize of 1 cash, commissions, swaps, spread, and more. It only have one simple function call, and that is pxyq.column('ASSET'). step 1: make sure you have pxyq.py in your directory. No need pip installations. step 2: make an import pxyq is written on top of… See the full description on the dataset page: https://huggingface.co/datasets/algorembrant/proxyquotes_library.tabularn<1K0 likes43 downloads2mo agoHugging Face26KennyChowww /OpenAlex-Articles-and-Harvard-Library-Item-Data OpenAlex Articles + Harvard Library Item Data Dataset Description This repository contains two large tabular subsets built from Harvard Library Bibliographic Metadata and an OpenAlex snapshot. These files are derived, filtered, and processed datasets, not full reproductions of the original source data. The records are based on real source metadata, but the dataset was prepared mainly for library data system operations testing, large-scale experimentation, and… See the full description on the dataset page: https://huggingface.co/datasets/KennyChowww/OpenAlex-Articles-and-Harvard-Library-Item-Data.tabularother100M<n<1B0 likes40 downloads5mo agoHugging Face27mindweave /library-book-loans Public Library Book Loans & Catalog Dataset (Free Sample) This is a free sample with 3,503 rows. The full dataset has 36,653 rows across 4 tables. Library circulation records for a simulated public library system with 3 branches, 12,000 catalog items, 5,000 patrons, and 20,000 loan transactions over 2 years. Features realistic patterns: summer reading program surge, academic year peaks, genre popularity by age group, overdue rates, fine calculations, and holds/reservations.… See the full description on the dataset page: https://huggingface.co/datasets/mindweave/library-book-loans.tabulartabular-classification1K<n<10K0 likes38 downloads6mo agoHugging Face28OrbitRidge23 /exciting-library-21ed94 exciting-library-21ed94 Synthetic products test data: 40 rows in data.csv. All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations. Fields sample_id: random identifier for this generated sample. row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/OrbitRidge23/exciting-library-21ed94.tabularn<1K0 likes38 downloads14d agoHugging Face29PleIAs /GATT_library GATT Library Dataset Card Dataset Overview Dataset Name: GATT Library Description: The GATT Library dataset comprises a comprehensive collection of documents related to the General Agreement on Tariffs and Trade (GATT), spanning from January 1, 1946, to September 6, 1996. This dataset is organized into a single Parquet file, which contains detailed information about the documents, including metadata extracted from the original files. The original files are stored in a… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/GATT_library.tabulartext-generation10K<n<100K6 likes29 downloads2y agoHugging Face30wassname /persona-steering-template-library Persona Steering Template Library GitHub repository: https://github.com/wassname/persona-steering-template-library Evaluated persona/template candidates for steering-vector and preference-pair experiments. What This Measures How do we know if a persona template is good? We want on-axis variation, but not off-axis variation. If we choose honest and dishonest personas, use a template like You are a {{ persona }} assistant, and ask The Eiffel Tower is in, we want the… See the full description on the dataset page: https://huggingface.co/datasets/wassname/persona-steering-template-library.tabulartext-generationn<1K0 likes29 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.