datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cnn_dailymail
Dataset Card for CNN Dailymail Dataset
Dataset Summary
The CNN / DailyMail Dataset is an English-language dataset containing just over 300k unique news articles as written by journalists at CNN and the Daily Mail. The current version supports both extractive and abstractive summarization, though the original version was created for machine reading and comprehension and abstractive question answering.
Supported Tasks and Leaderboards
'summarization': Versions… See the full description on the dataset page: https://huggingface.co/datasets/abisee/cnn_dailymail.foia-reading-room-documents
Foia Reading Room Documents
Documents from federal FOIA reading rooms and Inspector General report libraries: audits, inspections, investigative summaries and records released under the Freedom of Information Act.
Every document here was published by a US federal agency and is a work of the
United States government. Nothing has been altered: files are byte-identical to
what the agency posted, and the checksum in metadata.parquet is of the
original bytes.
Why this… See the full description on the dataset page: https://huggingface.co/datasets/abigailhaddad/foia-reading-room-documents.usajobs-scraping
USAJOBS announcement text
The full text of federal job announcements, scraped from usajobs.gov and joined
to the structured fields from the USAJOBS Historical API. About 3.2 million
announcements from September 2013 through September 2026, updated daily.
Why this exists
The USAJOBS API is a poor source for announcement text, in two ways.
The Search API only lists jobs that are open right now, so anything that opens
and closes between two collection runs is never… See the full description on the dataset page: https://huggingface.co/datasets/abigailhaddad/usajobs-scraping.english_quotes
Dataset Card for English quotes
I-Dataset Summary
english_quotes is a dataset of all the quotes retrieved from goodreads quotes. This dataset can be used for multi-label text classification and text generation. The content of each quote is in English and concerns the domain of datasets for NLP and beyond.
II-Supported Tasks and Leaderboards
Multi-label text classification : The dataset can be used to train a model for text-classification, which consists of… See the full description on the dataset page: https://huggingface.co/datasets/Abirate/english_quotes.usaspending-bulk-awards
USAspending bulk awards — contracts & assistance
Clean, partitioned, query-ready Parquet mirror of the public
USAspending Award Data Archive
(prime contract and financial-assistance transactions, FY2007–present, all agencies).
The source publishes 4,600 per-agency ZIP/CSV files (830 GB uncompressed). This
dataset normalizes them to typed, zstd-compressed Parquet (~8× smaller) with
amount columns as double and date columns as date, partitioned for fast
predicate-pushdown… See the full description on the dataset page: https://huggingface.co/datasets/abigailhaddad/usaspending-bulk-awards.512x1_ABI_CloudSatAb-Integro
Ab Integro
Ab Integro 是面向《欧陆风云 IV》1.37.5 的综合模组。默认内容为简体中文;英文以独立覆盖层提供。
内容
原版与整合内容的简体中文本地化
界面、旗帜、地图模式、地形和单位标记调整
外交、事件、警报和操作体验扩展
Comprehensive Map 与 India Extended 的地图和历史内容
加载界面名人名言提示(67 条考据后收录)
加载界面油画(1400–1800 年公版作品,来源见 ATTRIBUTION.md)
完整的版本记录与来源清单见 CHANGELOG.md。
安装
将 Ab_Integro 模组目录和对应的 .mod 描述文件放入 EU4 用户模组目录,在启动器中启用 Ab Integro。
中文显示需要 EU4 双字节汉化补丁。
需要英文时,先启用 Ab Integro,再启用 Ab Integro English。英文覆盖层只回填语言文本,不能单独启用。
致谢… See the full description on the dataset page: https://huggingface.co/datasets/raincandy-u/Ab-Integro.legislative-issue-tracker
Legislative Issue Tracker
Bills, legislative actions, floor speeches, hearings, and committee reports that
touch a specific federal statute — currently the Paperwork Reduction Act
(44 U.S.C. ch. 35, subch. I) — with every mention classified as amends,
exempts, references, or related.
The distinction is the point: Congress amends the PRA rarely (254 bills) but
exempts individual programs from it constantly (845 bills).
Built by abigail-64/legislative-issue-tracker.
Everything… See the full description on the dataset page: https://huggingface.co/datasets/abigailhaddad/legislative-issue-tracker.sam-solicitation-documents
Sam Solicitation Documents
Attachments from federal solicitation notices on SAM.gov: statements of work, performance work statements, justifications, amendments, wage determinations and the rest of the paperwork that accompanies a federal contract opportunity.
Every document here was published by a US federal agency and is a work of the
United States government. Nothing has been altered: files are byte-identical to
what the agency posted, and the checksum in metadata.parquet is… See the full description on the dataset page: https://huggingface.co/datasets/abigailhaddad/sam-solicitation-documents.pretrain_corpusclt_gpt2_tokenized_control
Fresh multilingual GPT-2 CLT control data
Sequential, unshuffled control sample for CLT null experiments. For each language,
complete source documents were tokenized with CausalNLP/gpt2-hf_multilingual-20 at revision
0afbb31b2db3f394270d42d6a4cb7f8fceeca3d8. The first 100,000,000 tokenizer tokens were discarded
(including the complete document that crossed the threshold), after which complete
documents were retained until at least 100,000,000 tokens were collected.
Data are… See the full description on the dataset page: https://huggingface.co/datasets/abir-hr196/clt_gpt2_tokenized_control.ORCA-sample
ORCA 100-image random sample
A random sample of 100 images (with their annotations) drawn from the ORCA dataset
(WongYukKwan/ORCA), the benchmark from
ORCA: Object Recognition and Comprehension for Archiving Marine Species (WACV 2026,
arXiv:2512.21150).
How it was sampled
100 images sampled uniformly at random with a fixed seed (random.Random(42)) from the
14,645 images in the source dataset.
The 100 sampled images span all 670 species categories in expectation;… See the full description on the dataset page: https://huggingface.co/datasets/abidlabs/ORCA-sample.federal-public-lands-spending
Federal Public Lands Spending
Contract and grant transaction data from the USAspending Award Data Archive for federal agencies that manage public lands.
Interactive demo:
Pipeline code: github.com/abigailhaddad/federal-public-lands-contracting
Agencies covered
Department of the Interior (014):
Bureau of Land Management, Bureau of Reclamation, Bureau of Safety and Environmental Enforcement, U.S. Geological Survey, National Park Service, Office of Surface Mining… See the full description on the dataset page: https://huggingface.co/datasets/abigailhaddad/federal-public-lands-spending.govinfo-documents
Govinfo Documents
An index of what the Government Publishing Office publishes on govinfo.gov -- congressional hearings, committee reports and prints, congressional documents, GAO reports, agency publications and presidential documents. One row per document with title, agency, date, page count, checksum and the URL the PDF is served from. The files themselves are not mirrored here: GPO guarantees permanent public access to them, so copying them would duplicate a corpus that is… See the full description on the dataset page: https://huggingface.co/datasets/abigailhaddad/govinfo-documents.abir177m-pretrain-balanced20-ezhijaru
abir177m pretrain mix — balanced20 en/zh/hi/ja/ru
Frozen packed-token shards for reproducible abir177m GPT-2–style pretraining.
Languages: 20% each en, zh, hi, ja, ru
Source streams: FineWeb (en) + FineWeb-2 (zh/hi/ja/ru)
Tokenizer: mistralai/Mistral-Nemo-Base-2407
Packing: 2048-token causal LM blocks (input_ids, labels identical)
Target budget: 3.55B tokens (1,733k sequences)
See meta.json for exact mixture + dataset map + seed.
ForeLen
Dataset Summary
ForeLen is a comprehensive benchmark designed to evaluate Large Language Model (LLM) output length prediction.
It includes long-sequence, Chain-of-Thought (CoT), and reinforcement learning (RL) sampling data, enabling the community to rigorously test both static and dynamic length predictors.
🗂 Data Structure
Data is organized by model and scenario:
Model
Scenarios
Splits
Llama3.2 1B, 3B
LongSeq, Reasoning, RL
train, validation, test
Qwen2.5… See the full description on the dataset page: https://huggingface.co/datasets/abinzzz/ForeLen.real-infrared-maritime-vessel-dataset
Real Infrared Maritime Vessel Dataset
Real infrared imagery of maritime vessels.
The dataset is provided in three forms — full-frame detection images, per-object classification crops, and a hand-curated subset.
Classes (7): liner, bulk carrier, warship, sailboat, canoe, container ship, fishing boat.
Layout
real-infrared-maritime-vessel-dataset/
├── original/ Full-frame IR images + XML bounding-box labels (detection)
│ ├── images/{train,test}/*.jpg… See the full description on the dataset page: https://huggingface.co/datasets/Abin0008/real-infrared-maritime-vessel-dataset.tamilwikipediadatasetannotations_creators:
found
language:
Tamil
language_creators:
found
license: []
multilinguality:
multilingual
pretty_name: tamilwikipediadataset
size_categories:
100K<n<1M
source_datasets: []
tags: []
task_categories:
summarization
task_ids: []
lm-eval-results-abideen-AlphaMonarch-daser-private
Dataset Card for Evaluation run of abideen/AlphaMonarch-daser
Dataset automatically created during the evaluation run of model abideen/AlphaMonarch-daser
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-abideen-AlphaMonarch-daser-private.french_book_reviews
Dataset Card for French book reviews
I-Dataset Summary
The majority of review datasets are in English. There are datasets in other languages, but not many. Through this work, I would like to enrich the datasets in the French language(my mother tongue with Arabic).The data was retrieved from two French websites: Babelio and Critiques LibresLike Wikipedia, these two French sites are made possible by the contributions of volunteers who use the Internet to share their… See the full description on the dataset page: https://huggingface.co/datasets/Abirate/french_book_reviews.test-translation-datasetkicad-netlist-sft-dataset
KiCad Netlist SFT Dataset
Training dataset for fine-tuning LLMs to generate valid KiCad electronic circuit netlists from natural language descriptions. Contains 100,179 examples with two complementary output formats:
Blog post: Teaching a Small LLM to Design Electronic Circuits: Fine-Tuning Qwen3-4B on 100K KiCad Netlists
Format
Examples
Description
SKiDL Python
100,179
Executable Python netlists in the messages assistant field
Structured JSON
100,179
Parallel… See the full description on the dataset page: https://huggingface.co/datasets/AbijahKaj/kicad-netlist-sft-dataset.hardware-cvdp-complete
CVDP - Comprehensive Verilog Design Problems (Complete Dataset)
🎯 782 out of 783 problems from the official CVDP benchmark by NVIDIA Research
🔥 Dataset Overview
This is the most complete version of the Comprehensive Verilog Design Problems (CVDP) benchmark available, containing 782 problems across 13 task categories. CVDP is designed to evaluate Large Language Models and agents on RTL design and verification tasks.
📊 Dataset Statistics
Total Problems: 772… See the full description on the dataset page: https://huggingface.co/datasets/AbiralArch/hardware-cvdp-complete.repro-memorize-resultsmultilingual_combined_tokenizeddetails_abideen__NexoNimbus-7B
Dataset Card for Evaluation run of abideen/NexoNimbus-7B
Dataset automatically created during the evaluation run of model abideen/NexoNimbus-7B on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_abideen__NexoNimbus-7B.waltoncolorcode_net_datasetscrum-dataset
NOT MINE! I BORROWED FROM ROBOFLOW BECAUSE I WANT TO USE THE DATASET ON RUNPOD BUT RUNPOD CANNOT DOWNLOAD THE DATASET FROM ROBOFLOW SO I HAD TO USE GIT-LFS
test-audio-1
