abi
Datasets
All datasets matching “abi”cnn_dailymail
Dataset Card for CNN Dailymail Dataset
Dataset Summary
The CNN / DailyMail Dataset is an English-language dataset containing just over 300k unique news articles as written by journalists at CNN and the Daily Mail. The current version supports both extractive and abstractive summarization, though the original version was created for machine reading and comprehension and abstractive question answering.
Supported Tasks and Leaderboards
'summarization': Versions… See the full description on the dataset page: https://huggingface.co/datasets/abisee/cnn_dailymail.foia-reading-room-documents
Foia Reading Room Documents
Documents from federal FOIA reading rooms and Inspector General report libraries: audits, inspections, investigative summaries and records released under the Freedom of Information Act.
Every document here was published by a US federal agency and is a work of the
United States government. Nothing has been altered: files are byte-identical to
what the agency posted, and the checksum in metadata.parquet is of the
original bytes.
Why this… See the full description on the dataset page: https://huggingface.co/datasets/abigailhaddad/foia-reading-room-documents.usajobs-scraping
USAJOBS announcement text
The full text of federal job announcements, scraped from usajobs.gov and joined
to the structured fields from the USAJOBS Historical API. About 3.2 million
announcements from September 2013 through September 2026, updated daily.
Why this exists
The USAJOBS API is a poor source for announcement text, in two ways.
The Search API only lists jobs that are open right now, so anything that opens
and closes between two collection runs is never… See the full description on the dataset page: https://huggingface.co/datasets/abigailhaddad/usajobs-scraping.english_quotes
Dataset Card for English quotes
I-Dataset Summary
english_quotes is a dataset of all the quotes retrieved from goodreads quotes. This dataset can be used for multi-label text classification and text generation. The content of each quote is in English and concerns the domain of datasets for NLP and beyond.
II-Supported Tasks and Leaderboards
Multi-label text classification : The dataset can be used to train a model for text-classification, which consists of… See the full description on the dataset page: https://huggingface.co/datasets/Abirate/english_quotes.usaspending-bulk-awards
USAspending bulk awards — contracts & assistance
Clean, partitioned, query-ready Parquet mirror of the public
USAspending Award Data Archive
(prime contract and financial-assistance transactions, FY2007–present, all agencies).
The source publishes 4,600 per-agency ZIP/CSV files (830 GB uncompressed). This
dataset normalizes them to typed, zstd-compressed Parquet (~8× smaller) with
amount columns as double and date columns as date, partitioned for fast
predicate-pushdown… See the full description on the dataset page: https://huggingface.co/datasets/abigailhaddad/usaspending-bulk-awards.512x1_ABI_CloudSat

