CoolFace
22 results

abi

abisee /cnn_dailymail Dataset Card for CNN Dailymail Dataset Dataset Summary The CNN / DailyMail Dataset is an English-language dataset containing just over 300k unique news articles as written by journalists at CNN and the Daily Mail. The current version supports both extractive and abstractive summarization, though the original version was created for machine reading and comprehension and abstractive question answering. Supported Tasks and Leaderboards 'summarization': Versions… See the full description on the dataset page: https://huggingface.co/datasets/abisee/cnn_dailymail.textsummarization100K<n<1M351 likes75k downloads3y agoHugging Faceabigailhaddad /foia-reading-room-documents Foia Reading Room Documents Documents from federal FOIA reading rooms and Inspector General report libraries: audits, inspections, investigative summaries and records released under the Freedom of Information Act. Every document here was published by a US federal agency and is a work of the United States government. Nothing has been altered: files are byte-identical to what the agency posted, and the checksum in metadata.parquet is of the original bytes. Why this… See the full description on the dataset page: https://huggingface.co/datasets/abigailhaddad/foia-reading-room-documents.text-retrieval0 likes5.8k downloads9d agoHugging Faceabigailhaddad /usajobs-scraping USAJOBS announcement text The full text of federal job announcements, scraped from usajobs.gov and joined to the structured fields from the USAJOBS Historical API. About 3.2 million announcements from September 2013 through September 2026, updated daily. Why this exists The USAJOBS API is a poor source for announcement text, in two ways. The Search API only lists jobs that are open right now, so anything that opens and closes between two collection runs is never… See the full description on the dataset page: https://huggingface.co/datasets/abigailhaddad/usajobs-scraping.tabulartext-classification1M<n<10M0 likes3k downloads15h agoHugging FaceAbirate /english_quotes Dataset Card for English quotes I-Dataset Summary english_quotes is a dataset of all the quotes retrieved from goodreads quotes. This dataset can be used for multi-label text classification and text generation. The content of each quote is in English and concerns the domain of datasets for NLP and beyond. II-Supported Tasks and Leaderboards Multi-label text classification : The dataset can be used to train a model for text-classification, which consists of… See the full description on the dataset page: https://huggingface.co/datasets/Abirate/english_quotes.texttext-classification1K<n<10K109 likes2.8k downloads4y agoHugging Faceabigailhaddad /usaspending-bulk-awards USAspending bulk awards — contracts & assistance Clean, partitioned, query-ready Parquet mirror of the public USAspending Award Data Archive (prime contract and financial-assistance transactions, FY2007–present, all agencies). The source publishes 4,600 per-agency ZIP/CSV files (830 GB uncompressed). This dataset normalizes them to typed, zstd-compressed Parquet (~8× smaller) with amount columns as double and date columns as date, partitioned for fast predicate-pushdown… See the full description on the dataset page: https://huggingface.co/datasets/abigailhaddad/usaspending-bulk-awards.tabular100M<n<1B0 likes2.3k downloads13h agoHugging Faceadamliewehr /512x1_ABI_CloudSat0 likes1.8k downloads2mo agoHugging Face

People