CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01abisee /cnn_dailymail Dataset Card for CNN Dailymail Dataset Dataset Summary The CNN / DailyMail Dataset is an English-language dataset containing just over 300k unique news articles as written by journalists at CNN and the Daily Mail. The current version supports both extractive and abstractive summarization, though the original version was created for machine reading and comprehension and abstractive question answering. Supported Tasks and Leaderboards 'summarization': Versions… See the full description on the dataset page: https://huggingface.co/datasets/abisee/cnn_dailymail.textsummarization100K<n<1M351 likes75k downloads3y agoHugging Face02abigailhaddad /foia-reading-room-documents Foia Reading Room Documents Documents from federal FOIA reading rooms and Inspector General report libraries: audits, inspections, investigative summaries and records released under the Freedom of Information Act. Every document here was published by a US federal agency and is a work of the United States government. Nothing has been altered: files are byte-identical to what the agency posted, and the checksum in metadata.parquet is of the original bytes. Why this… See the full description on the dataset page: https://huggingface.co/datasets/abigailhaddad/foia-reading-room-documents.text-retrieval0 likes5.8k downloads9d agoHugging Face03abigailhaddad /usajobs-scraping USAJOBS announcement text The full text of federal job announcements, scraped from usajobs.gov and joined to the structured fields from the USAJOBS Historical API. About 3.2 million announcements from September 2013 through September 2026, updated daily. Why this exists The USAJOBS API is a poor source for announcement text, in two ways. The Search API only lists jobs that are open right now, so anything that opens and closes between two collection runs is never… See the full description on the dataset page: https://huggingface.co/datasets/abigailhaddad/usajobs-scraping.tabulartext-classification1M<n<10M0 likes3k downloads16h agoHugging Face04Abirate /english_quotes Dataset Card for English quotes I-Dataset Summary english_quotes is a dataset of all the quotes retrieved from goodreads quotes. This dataset can be used for multi-label text classification and text generation. The content of each quote is in English and concerns the domain of datasets for NLP and beyond. II-Supported Tasks and Leaderboards Multi-label text classification : The dataset can be used to train a model for text-classification, which consists of… See the full description on the dataset page: https://huggingface.co/datasets/Abirate/english_quotes.texttext-classification1K<n<10K109 likes2.8k downloads4y agoHugging Face05abigailhaddad /usaspending-bulk-awards USAspending bulk awards — contracts & assistance Clean, partitioned, query-ready Parquet mirror of the public USAspending Award Data Archive (prime contract and financial-assistance transactions, FY2007–present, all agencies). The source publishes 4,600 per-agency ZIP/CSV files (830 GB uncompressed). This dataset normalizes them to typed, zstd-compressed Parquet (~8× smaller) with amount columns as double and date columns as date, partitioned for fast predicate-pushdown… See the full description on the dataset page: https://huggingface.co/datasets/abigailhaddad/usaspending-bulk-awards.tabular100M<n<1B0 likes2.3k downloads14h agoHugging Face06adamliewehr /512x1_ABI_CloudSat0 likes1.8k downloads2mo agoHugging Face07raincandy-u /Ab-Integro Ab Integro Ab Integro 是面向《欧陆风云 IV》1.37.5 的综合模组。默认内容为简体中文;英文以独立覆盖层提供。 内容 原版与整合内容的简体中文本地化 界面、旗帜、地图模式、地形和单位标记调整 外交、事件、警报和操作体验扩展 Comprehensive Map 与 India Extended 的地图和历史内容 加载界面名人名言提示(67 条考据后收录) 加载界面油画(1400–1800 年公版作品,来源见 ATTRIBUTION.md) 完整的版本记录与来源清单见 CHANGELOG.md。 安装 将 Ab_Integro 模组目录和对应的 .mod 描述文件放入 EU4 用户模组目录,在启动器中启用 Ab Integro。 中文显示需要 EU4 双字节汉化补丁。 需要英文时,先启用 Ab Integro,再启用 Ab Integro English。英文覆盖层只回填语言文本,不能单独启用。 致谢… See the full description on the dataset page: https://huggingface.co/datasets/raincandy-u/Ab-Integro.0 likes1.8k downloads2mo agoHugging Face08abigailhaddad /legislative-issue-tracker Legislative Issue Tracker Bills, legislative actions, floor speeches, hearings, and committee reports that touch a specific federal statute — currently the Paperwork Reduction Act (44 U.S.C. ch. 35, subch. I) — with every mention classified as amends, exempts, references, or related. The distinction is the point: Congress amends the PRA rarely (254 bills) but exempts individual programs from it constantly (845 bills). Built by abigail-64/legislative-issue-tracker. Everything… See the full description on the dataset page: https://huggingface.co/datasets/abigailhaddad/legislative-issue-tracker.text100K<n<1M0 likes1.7k downloads7h agoHugging Face09abigailhaddad /sam-solicitation-documents Sam Solicitation Documents Attachments from federal solicitation notices on SAM.gov: statements of work, performance work statements, justifications, amendments, wage determinations and the rest of the paperwork that accompanies a federal contract opportunity. Every document here was published by a US federal agency and is a work of the United States government. Nothing has been altered: files are byte-identical to what the agency posted, and the checksum in metadata.parquet is… See the full description on the dataset page: https://huggingface.co/datasets/abigailhaddad/sam-solicitation-documents.text-retrieval0 likes1.3k downloads5d agoHugging Face10abideen /pretrain_corpustext10M<n<100M2 likes1k downloads2y agoHugging Face11abir-hr196 /clt_gpt2_tokenized_control Fresh multilingual GPT-2 CLT control data Sequential, unshuffled control sample for CLT null experiments. For each language, complete source documents were tokenized with CausalNLP/gpt2-hf_multilingual-20 at revision 0afbb31b2db3f394270d42d6a4cb7f8fceeca3d8. The first 100,000,000 tokenizer tokens were discarded (including the complete document that crossed the threshold), after which complete documents were retained until at least 100,000,000 tokens were collected. Data are… See the full description on the dataset page: https://huggingface.co/datasets/abir-hr196/clt_gpt2_tokenized_control.texttext-generation100K<n<1M0 likes725 downloads2mo agoHugging Face12abidlabs /ORCA-sample ORCA 100-image random sample A random sample of 100 images (with their annotations) drawn from the ORCA dataset (WongYukKwan/ORCA), the benchmark from ORCA: Object Recognition and Comprehension for Archiving Marine Species (WACV 2026, arXiv:2512.21150). How it was sampled 100 images sampled uniformly at random with a fixed seed (random.Random(42)) from the 14,645 images in the source dataset. The 100 sampled images span all 670 species categories in expectation;… See the full description on the dataset page: https://huggingface.co/datasets/abidlabs/ORCA-sample.imageobject-detectionn<1K1 likes704 downloads25d agoHugging Face13abigailhaddad /federal-public-lands-spending Federal Public Lands Spending Contract and grant transaction data from the USAspending Award Data Archive for federal agencies that manage public lands. Interactive demo: Pipeline code: github.com/abigailhaddad/federal-public-lands-contracting Agencies covered Department of the Interior (014): Bureau of Land Management, Bureau of Reclamation, Bureau of Safety and Environmental Enforcement, U.S. Geological Survey, National Park Service, Office of Surface Mining… See the full description on the dataset page: https://huggingface.co/datasets/abigailhaddad/federal-public-lands-spending.tabulartabular-classification1M<n<10M1 likes647 downloads7d agoHugging Face14abigailhaddad /govinfo-documents Govinfo Documents An index of what the Government Publishing Office publishes on govinfo.gov -- congressional hearings, committee reports and prints, congressional documents, GAO reports, agency publications and presidential documents. One row per document with title, agency, date, page count, checksum and the URL the PDF is served from. The files themselves are not mirrored here: GPO guarantees permanent public access to them, so copying them would duplicate a corpus that is… See the full description on the dataset page: https://huggingface.co/datasets/abigailhaddad/govinfo-documents.text-retrieval0 likes645 downloads12d agoHugging Face15nirmalendu01 /abir177m-pretrain-balanced20-ezhijaru abir177m pretrain mix — balanced20 en/zh/hi/ja/ru Frozen packed-token shards for reproducible abir177m GPT-2–style pretraining. Languages: 20% each en, zh, hi, ja, ru Source streams: FineWeb (en) + FineWeb-2 (zh/hi/ja/ru) Tokenizer: mistralai/Mistral-Nemo-Base-2407 Packing: 2048-token causal LM blocks (input_ids, labels identical) Target budget: 3.55B tokens (1,733k sequences) See meta.json for exact mixture + dataset map + seed. text1M<n<10M0 likes486 downloads1mo agoHugging Face16abinzzz /ForeLen Dataset Summary ForeLen is a comprehensive benchmark designed to evaluate Large Language Model (LLM) output length prediction. It includes long-sequence, Chain-of-Thought (CoT), and reinforcement learning (RL) sampling data, enabling the community to rigorously test both static and dynamic length predictors. 🗂 Data Structure Data is organized by model and scenario: Model Scenarios Splits Llama3.2 1B, 3B LongSeq, Reasoning, RL train, validation, test Qwen2.5… See the full description on the dataset page: https://huggingface.co/datasets/abinzzz/ForeLen.text100K<n<1M3 likes437 downloads7mo agoHugging Face17Abin0008 /real-infrared-maritime-vessel-dataset Real Infrared Maritime Vessel Dataset Real infrared imagery of maritime vessels. The dataset is provided in three forms — full-frame detection images, per-object classification crops, and a hand-curated subset. Classes (7): liner, bulk carrier, warship, sailboat, canoe, container ship, fishing boat. Layout real-infrared-maritime-vessel-dataset/ ├── original/ Full-frame IR images + XML bounding-box labels (detection) │ ├── images/{train,test}/*.jpg… See the full description on the dataset page: https://huggingface.co/datasets/Abin0008/real-infrared-maritime-vessel-dataset.imageimage-classification10K<n<100K0 likes370 downloads2mo agoHugging Face18Abirami /tamilwikipediadatasetannotations_creators: found language: Tamil language_creators: found license: [] multilinguality: multilingual pretty_name: tamilwikipediadataset size_categories: 100K<n<1M source_datasets: [] tags: [] task_categories: summarization task_ids: [] text100K<n<1M2 likes312 downloads4y agoHugging Face19nyu-dice-lab /lm-eval-results-abideen-AlphaMonarch-daser-private Dataset Card for Evaluation run of abideen/AlphaMonarch-daser Dataset automatically created during the evaluation run of model abideen/AlphaMonarch-daser The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-abideen-AlphaMonarch-daser-private.tabular100K<n<1M0 likes288 downloads2y agoHugging Face20Abirate /french_book_reviews Dataset Card for French book reviews I-Dataset Summary The majority of review datasets are in English. There are datasets in other languages, but not many. Through this work, I would like to enrich the datasets in the French language(my mother tongue with Arabic).The data was retrieved from two French websites: Babelio and Critiques LibresLike Wikipedia, these two French sites are made possible by the contributions of volunteers who use the Internet to share their… See the full description on the dataset page: https://huggingface.co/datasets/Abirate/french_book_reviews.tabulartext-classification1K<n<10K8 likes278 downloads4y agoHugging Face21abidlabs /test-translation-datasettextn<1K0 likes216 downloads5y agoHugging Face22AbijahKaj /kicad-netlist-sft-dataset KiCad Netlist SFT Dataset Training dataset for fine-tuning LLMs to generate valid KiCad electronic circuit netlists from natural language descriptions. Contains 100,179 examples with two complementary output formats: Blog post: Teaching a Small LLM to Design Electronic Circuits: Fine-Tuning Qwen3-4B on 100K KiCad Netlists Format Examples Description SKiDL Python 100,179 Executable Python netlists in the messages assistant field Structured JSON 100,179 Parallel… See the full description on the dataset page: https://huggingface.co/datasets/AbijahKaj/kicad-netlist-sft-dataset.texttext-generation100K<n<1M1 likes212 downloads3d agoHugging Face23AbiralArch /hardware-cvdp-complete CVDP - Comprehensive Verilog Design Problems (Complete Dataset) 🎯 782 out of 783 problems from the official CVDP benchmark by NVIDIA Research 🔥 Dataset Overview This is the most complete version of the Comprehensive Verilog Design Problems (CVDP) benchmark available, containing 782 problems across 13 task categories. CVDP is designed to evaluate Large Language Models and agents on RTL design and verification tasks. 📊 Dataset Statistics Total Problems: 772… See the full description on the dataset page: https://huggingface.co/datasets/AbiralArch/hardware-cvdp-complete.text-generation1K<n<10K1 likes202 downloads1y agoHugging Face24abidlabs /repro-memorize-results0 likes197 downloads2mo agoHugging Face25abir-hr196 /multilingual_combined_tokenized100K<n<1M0 likes188 downloads1y agoHugging Face26open-llm-leaderboard-old /details_abideen__NexoNimbus-7B Dataset Card for Evaluation run of abideen/NexoNimbus-7B Dataset automatically created during the evaluation run of model abideen/NexoNimbus-7B on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_abideen__NexoNimbus-7B.0 likes180 downloads3y agoHugging Face27AbirAshraf51611 /waltoncolorimagen<1K0 likes174 downloads26d agoHugging Face28Abirate /code_net_datasettext10K<n<100K3 likes166 downloads5y agoHugging Face29abinazmi /scrum-dataset NOT MINE! I BORROWED FROM ROBOFLOW BECAUSE I WANT TO USE THE DATASET ON RUNPOD BUT RUNPOD CANNOT DOWNLOAD THE DATASET FROM ROBOFLOW SO I HAD TO USE GIT-LFS image1K<n<10K0 likes163 downloads2y agoHugging Face30abidlabs /test-audio-1audion<1K0 likes160 downloads5y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.