datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
activation-datasetsePark_zu_yu_duan_wen_indigenous_language_essays
FormosanBank publication status
This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC-SA 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card.
FormosanBank/ePark_zu_yu_duan_wen_indigenous_language_essays
Commercial AI Use is prohibited without prior written permission. See the FormosanBank Terms of Use and AI Use Addendum.… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ePark_zu_yu_duan_wen_indigenous_language_essays.ePark_tu_hua_gu_shi_pian_picture_story
FormosanBank publication status
This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC-SA 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card.
FormosanBank/ePark_tu_hua_gu_shi_pian_picture_story
Commercial AI Use is prohibited without prior written permission. See the FormosanBank Terms of Use and AI Use Addendum.
This is a… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ePark_tu_hua_gu_shi_pian_picture_story.ePark_wen_hua_pian_cultural_section
FormosanBank publication status
This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC-SA 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card.
FormosanBank/ePark_wen_hua_pian_cultural_section
Commercial AI Use is prohibited without prior written permission. See the FormosanBank Terms of Use and AI Use Addendum.
This is a… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ePark_wen_hua_pian_cultural_section.ePark_sheng_huo_hui_hua_pian_daily_conversation
FormosanBank publication status
This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC-SA 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card.
FormosanBank/ePark_sheng_huo_hui_hua_pian_daily_conversation
Commercial AI Use is prohibited without prior written permission. See the FormosanBank Terms of Use and AI Use Addendum.… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ePark_sheng_huo_hui_hua_pian_daily_conversation.ePark_ju_xing_pian_gao_zhong_sentence_patterns_senior_high
FormosanBank publication status
This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC-SA 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card.
FormosanBank/ePark_ju_xing_pian_gao_zhong_sentence_patterns_senior_high
Commercial AI Use is prohibited without prior written permission. See the FormosanBank Terms of Use and AI Use… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ePark_ju_xing_pian_gao_zhong_sentence_patterns_senior_high.epa-re-powering-screening-sam2-masks-partial
EPA RE-Powering Screening SAM2 candidate masks — partial snapshot
This is a paused, incomplete snapshot of model-generated candidate masks for
EPA RE-Powering Screening sites. It covers approximately
12% of archive-group jobs and
8.2% of processable site points from the
current run. It contains 15,211 successful masks across
15,541 attempted processable rows, plus
1,326 explicit input coverage exclusions.
The source inventory contains 190,976 points in total.… See the full description on the dataset page: https://huggingface.co/datasets/sarkarghya/epa-re-powering-screening-sam2-masks-partial.ePark_qing_jing_zu_yu_contextual_indigenous_language
FormosanBank publication status
This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC-SA 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card.
FormosanBank/ePark_qing_jing_zu_yu_contextual_indigenous_language
Commercial AI Use is prohibited without prior written permission. See the FormosanBank Terms of Use and AI Use… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ePark_qing_jing_zu_yu_contextual_indigenous_language.ePark_jiu_jie_jiao_cai_nine_level_materials
FormosanBank publication status
This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC-SA 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card.
FormosanBank/ePark_jiu_jie_jiao_cai_nine_level_materials
Commercial AI Use is prohibited without prior written permission. See the FormosanBank Terms of Use and AI Use Addendum.
This… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ePark_jiu_jie_jiao_cai_nine_level_materials.epago-sn36-deepresearch-trajectories
Epago SN36 — trajectories, baselines, diagnostics and mining tooling
Everything produced while investigating model mining on Bittensor subnet 36 (Epago).
All data was generated locally by running the subnet's own harness against its own bundled
corpus. Nothing here is copied from the subnet's private artifacts.
⚠️ Read this first: coronation is currently impossible
On EpagoFoundation/epago @ 7ddfef0 (latest origin/main as of 2026-09-10), no challenger
can ever be… See the full description on the dataset page: https://huggingface.co/datasets/Olague-Secret/epago-sn36-deepresearch-trajectories.ePark_yue_du_shu_xie_pian_reading_writing
FormosanBank publication status
This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC-SA 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card.
FormosanBank/ePark_yue_du_shu_xie_pian_reading_writing
Commercial AI Use is prohibited without prior written permission. See the FormosanBank Terms of Use and AI Use Addendum.
This is… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ePark_yue_du_shu_xie_pian_reading_writing.kl3m-data-dotgov-www.epa.gov
KL3M Data Project
Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper.
Description
This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models.
Dataset Details
Format: Parquet… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov-www.epa.gov.ePark3_Amis
Dataset Card for "ePark3_Amis"
More Information needed
epa-tri-toxic-release-inventory
EPA Toxic Release Inventory — 2022–2023 reporting years
Rebuilt September 19, 2026 from the two pinned EPA national Basic Data CSV files.
This is a dated historical snapshot; future updates are not included. The archive
contains 158,687 TRI form records, representing 23,081 distinct TRI facilities.
It does not contain the previously advertised 1987-onward history.
This repository contains the deterministic 1,000-row sample. The complete dated
CSV/Parquet package is available at… See the full description on the dataset page: https://huggingface.co/datasets/claritystorm/epa-tri-toxic-release-inventory.gorkhapatra-nepali-epaper
Gorkhapatra Nepali E-Paper Corpus
Per-article text extracted from PDF e-papers published on
epaper.gorkhapatraonline.com, covering 11 newspaper
slugs (gorkhapatra, risingnepal, friday-suppliment, madhuparka, muna, nayanepal,
loksewa, saturday, yuwamunch, gorkhapatra-125, other).
Extraction is layout-aware (geometry + font size, no ML model/fixed template) and reconstructs
article boundaries — headline, dateline, and paragraphs in reading order — directly from the PDF's… See the full description on the dataset page: https://huggingface.co/datasets/Aananda-giri/gorkhapatra-nepali-epaper.EPAIQA-15Kepa-water-quality-criteria-human-health
EPA national recommended water quality criteria for human health: pollutant concentration thresholds
Canonical, always-current version: https://referencesource.org/epa-water-quality-criteria-human-health/
Machine-readable: https://referencesource.org/epa-water-quality-criteria-human-health/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-19
Stale after: 2028-08-18 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/epa-water-quality-criteria-human-health.ePark_ju_xing_pian_guo_zhong_sentence_patterns_junior_high
FormosanBank publication status
This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC-SA 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card.
FormosanBank/ePark_ju_xing_pian_guo_zhong_sentence_patterns_junior_high
Commercial AI Use is prohibited without prior written permission. See the FormosanBank Terms of Use and AI Use… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ePark_ju_xing_pian_guo_zhong_sentence_patterns_junior_high.EpaDB
EpaDB
EpaDB is a speech database of 50 native Spanish speakers (25 male, 25 female) from Argentina speaking English.
It contains phonemic annotations using mainly the sounds supported by ARPABet with a few extensions to model Spanish influenced dialects of English.
It was developed by Jazmin Vidal, Luciana Ferrer, and Leonardo Brambilla at the Speech Lab. Read more on their official github and paper.
This Processed Version
We have processed the dataset into an easily… See the full description on the dataset page: https://huggingface.co/datasets/KoelLabs/EpaDB.septuagint
Dataset Card for "septuagint"
More Information needed
italian_dataset_mix
Dataset Card for Dataset Name
This dataset represents a collection of the most downloaded Italian datasets.
Dataset Details
Dataset Description
This dataset represents a collection of the most downloaded Italian datasets:
WasamiKirua/samantha-ita
mii-community/ultrafeedback-translated-ita
mchl-labs/stambecco_data_it
efederici/fisica
FreedomIntelligence/sharegpt-italian
Curated by: Enzo Palmisano
Language(s) (NLP): Italian
License: Apache 2.0
ePark_hui_ben_ping_tai_picture_book_platform
FormosanBank publication status
This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC-SA 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card.
FormosanBank/ePark_hui_ben_ping_tai_picture_book_platform
Commercial AI Use is prohibited without prior written permission. See the FormosanBank Terms of Use and AI Use Addendum.
This… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ePark_hui_ben_ping_tai_picture_book_platform.ePark_xue_xi_ci_biao_learning_vocabulary
FormosanBank publication status
This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC-SA 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card.
FormosanBank/ePark_xue_xi_ci_biao_learning_vocabulary
Commercial AI Use is prohibited without prior written permission. See the FormosanBank Terms of Use and AI Use Addendum.
This is… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ePark_xue_xi_ci_biao_learning_vocabulary.document-qna-chroma-openai-logsEpan-EyeBlink2multi-company-10k-reports-qna-chroma-openai-logsepa-enforcement-cases
EPA enforcement cases (ICIS-FE&C: 134k civil/criminal cases + 200k defendants)
Every EPA civil and criminal enforcement action since the 1980s — 134,805 cases with $1.56 billion in penalties assessed and 199,682 named defendants. Pairs with epa_facilities (already in catalog) so you can join facility violation history with the formal enforcement that followed. Top defendants by case count: PhoneSoap LLC (135), Syngenta (115), GM (114), GE (106), BASF (90), Ford (81). Largest single… See the full description on the dataset page: https://huggingface.co/datasets/emperor-mew/epa-enforcement-cases.new_epa_micro_all民生公共物聯網,微型感測器資料集,
整合 pm2.5 ,溫度,濕度資料,
每微型感測器每小時一筆,
先取樣 2023.3月 與 2023.4月 資料。
machine-failure-logselderly-speech-whisper-training
