datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cbi-archive-raw
Central Bank of Ireland Archive: original source files
6,309 original files, 6.56 GB. Every PDF, spreadsheet, Word document and
archive gathered from the Central Bank of Ireland's public website, stored by
content hash so that a search result can be turned back into the document a
human would actually read.
This is the raw tier. If you want the text, you almost certainly want
aditya487/cbi-archive-corpus
instead: 5,568 documents and 89,242 page or pseudo-page rows as Parquet… See the full description on the dataset page: https://huggingface.co/datasets/aditya487/cbi-archive-raw.us-historical-layoffs-archive-warn-act-notices-removed-from-state-websites
Historical US layoffs archive: 6,799 WARN Act notices that state websites no longer list (2000-2025), recovered
Rebuilt 2026-09-24. Five state labor agencies — Connecticut, Michigan, New York,
North Carolina and Pennsylvania — retired the web pages their older WARN Act
layoff notices lived on. Their current pages start years later. This dataset is
every notice in our file that came from one of those retired pages and is not
on the agency's live page today: 6,799 notices, 6,799… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-historical-layoffs-archive-warn-act-notices-removed-from-state-websites.ARCHIVE-TEXT-URLS
Internet Archive English Text URLs Dataset
Dataset Description
This dataset contains 11,151,637 direct download URLs to OCR-processed text files from the Internet Archive's digital library. All entries are English-language texts spanning books, documents, historical records, and various other written materials.
Dataset Summary
Total Rows: 11,151,637
Language: English
Source: Internet Archive
Format: CSV with metadata and direct text file URLs
Text… See the full description on the dataset page: https://huggingface.co/datasets/Navanjana/ARCHIVE-TEXT-URLS.industrial-technical-archive
🚀 Latest Updates (July, 2026)
Version: v07.2026 (Verified)
Status: Integrated with 1,000,000+ records.
New Files: product-E-26-07-2026.csv & product-V-26-07-2026.csv.
QTE Technologies: Industrial & Scientific Knowledge Base
Wikidata Entity: Q138411149
IPFS CID: bafybeibogxxuhmzfrsuhcfd4qr4tmc4okhmrcwhp3266hq47ccuyjnjxoq
Official Neural Hub: qtetech.github.io
This is the permanent technical archive for QTE Technologies, ensuring long-term accessibility of… See the full description on the dataset page: https://huggingface.co/datasets/QTE-Technologies/industrial-technical-archive.sinhala-tts-dataset-archive-20260429-082457
Sinhala TTS Dataset
Clean, segmented Sinhala speech from the "Unlimited History" YouTube series by @sunchare.
Stats
Metric
Value
Utterances
218
Train
208
Val
10
Hours
0.51
Mean duration
8.5s
Sample rate
22050 Hz
Pipeline
Raw YouTube audio -> HTDemucs -> VoiceFixer + DeepFilterNet3 ->
Diarization -> Silero-VAD -> ASR (faster-whisper: C:\Users\kosal\sinhala-tts\whisper-small-si-ct2) -> Quality filtering (SNR>=20.0dB)
Format… See the full description on the dataset page: https://huggingface.co/datasets/outlawmold/sinhala-tts-dataset-archive-20260429-082457.cars_from_drom.ru_archive_2007-2025More information on the parsing process can be found here: https://github.com/zavzyatiy/drom_archive_parser.
This dataset is also published on Kaggle: https://www.kaggle.com/datasets/assaabramovich/resaled-cars-from-drom-ruarchive-2018-2023/.
Main dataset with all data: drom_archive_2007-2025_full.csv
Dataset with (almost) all configurations from Drom for cars in data: additional_data/drom-24-07-2025-all_main_cars_configurations.csv
Dataset with identification of regions for all cities in… See the full description on the dataset page: https://huggingface.co/datasets/zavzyatiy/cars_from_drom.ru_archive_2007-2025.prog-archivesjazz-music-archivesArchiveIITransfermarkt_News_ArchiveTransfermarkt_News_Archive
Dataset Overview
This dataset contains football news articles from Transfermarkt website spanning from January 1, 2020, to April 1, 2025. It includes transfer updates, match results, player analysis, club strategies, and industry insights.
Dataset Structure
The dataset includes the following 5 columns:
Column
Description
URL
Direct link to the news article on Transfermarkt.
Time
Publication date and time (formatted as… See the full description on the dataset page: https://huggingface.co/datasets/ZhangYi0820/Transfermarkt_News_Archive.metal-archives-bandshttps://www.metal-archives.com/
lezgigazet_ru_archives_337371prompts_archiveRead me.
archive-govt-nz-treasury-csv
Archive Govt NZ — Treasury CSV derivative
Simple Viewer-compatible CSV representation of 54 normalized Treasury dataset
metadata records. The Parquet derivative and preservation source archive remain
available separately.
archive_01metal-archives-tracksSubmission-ArchiveModern_Archive_mtl_gazzateSNET_Archivevisual_accent_dialect_archiveSource: https://www.youtube.com/@visualaccent/videos
All rights belong to the original dataset creator.
VADA-AVSR: an audio-visual dataset of non-native English ("accents") and English varieties ("dialects")
We preprocessed the Visual Accent and Dialect Archive (https://archive.mith.umd.edu/mith-2020/vada/index.html) for audio-visual speech recognition (AVSR), speech recognition (ASR), and visual speech recognition/lip-reading (VSR).
This version currently only contains read speech… See the full description on the dataset page: https://huggingface.co/datasets/Berkeley-NLP/visual_accent_dialect_archive.Archive_NewHistorical_Archive_Mtl_Gazzatemetal-archives-bandshttps://www.metal-archives.com/
archive
