datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
open-subtitles-bitext-miningopen-subtitles-256s-bitext-miningbitcoin-mining-pool-templates
Bitcoin mining pool templates
Timestamped Stratum job messages collected directly from Bitcoin mining pool endpoints. The data records changes in the work each endpoint sends to miners, including the previous block hash, coinbase data and clean-jobs flag.
Contents
Table
Record
bitcoin_mining_pool_jobs
A job received from a pool endpoint, with its observation time, nTime, coinbase, merkle branch count and clean-jobs flag
Using the data… See the full description on the dataset page: https://huggingface.co/datasets/dataforge-labs/bitcoin-mining-pool-templates.tatoeba-bitext-mining
Tatoeba
An MTEB dataset
Massive Text Embedding Benchmark
1,000 English-aligned sentence pairs for each language based on the Tatoeba corpus
Task category
t2t
Domains
Written
Reference
https://github.com/facebookresearch/LASER/tree/main/data/tatoeba/v1
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["Tatoeba"])
evaluator = mteb.MTEB(task)
model =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/tatoeba-bitext-mining.instructions-pair-miningmteb-bitext-mining-aggregated
MTEB BitextMining Aggregated Dataset (Full)
This dataset aggregates ALL configs from 10 BitextMining datasets in the MTEB (Massive Text Embedding Benchmark) Multilingual v2 benchmark into a single, unified dataset for comprehensive bitext mining evaluation.
Dataset Summary
Total Examples: 448,229 sentence pairs
Source Datasets (Configs): 10 MTEB BitextMining tasks
Total Splits: 332 language pairs/configurations
Languages: 300+ unique language codes across all datasets… See the full description on the dataset page: https://huggingface.co/datasets/SaylorTwift/mteb-bitext-mining-aggregated.open-subtitles-500-bitext-miningbucc-bitext-mining
BUCC.v2
An MTEB dataset
Massive Text Embedding Benchmark
BUCC bitext mining dataset
Task category
t2t
Domains
Written
Reference
https://comparable.limsi.fr/bucc2018/bucc2018-task.html
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["BUCC.v2"])
evaluator = mteb.MTEB(task)
model = mteb.get_model(YOUR_MODEL)
evaluator.run(model)
To learn more about how to run… See the full description on the dataset page: https://huggingface.co/datasets/mteb/bucc-bitext-mining.africa-world-bank-energy-mining-time-series
Africa World Bank Energy and Mining Labeled Time Series Data
This repository is part of the Africa Temporal Intelligence Corpus (ATIC). It contains sector-specific temporal corpus packages for African countries.
ATIC sector repositories are designed for machine consumption first: Parquet tables, stable IDs, reproducible metadata, explicit provenance, review status, and separable semantic layers.
Sector Scope
Temporal energy indicators for African countries… See the full description on the dataset page: https://huggingface.co/datasets/africatic/africa-world-bank-energy-mining-time-series.asia-energy-world-bank-energy-and-mining-indicators
India - Energy and Mining
Publisher: World Bank Group · Source: HDX · License: cc-by · Updated: 2026-04-28
Abstract
Contains data from the World Bank's data portal. There is also a consolidated country dataset on HDX.
The world economy needs ever-increasing amounts of energy to sustain economic growth, raise living standards, and reduce poverty. But today's trends in energy use are not sustainable. As the world's population grows and economies become more industrialized… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-energy-world-bank-energy-and-mining-indicators.open-subtitles-250-bitext-miningtatoeba-bitext-miningsparkproof-miningIndustryCorpus2_mining
IndustryCorpus2: Mining
This repository contains the IndustryCorpus2: Mining domain subset of BAAI/IndustryCorpus2.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryCorpus2:
@misc{shi2024industrycorpus2,
title = {IndustryCorpus2},
author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and Donglin Hao},
year = {2024},
publisher… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus2_mining.Arguement_Mining_CL2017tokens along with chunk id. IOB1 format Begining of arguement denoted by B-ARG,inside arguement
denoted by I-ARG, other chunks are O
Orginial train,test split as used by the paper is providedlatent-mining
Latent Mining
Latent Mining is a benchmark-construction method for scientific-agent tasks where useful public evidence diverges from a withheld verifier-backed outcome. The resulting tasks test whether agents can make calibrated scientific triage decisions under incomplete information.
This dataset contains the first public biology subset: 165 cross-locus regulatory-edit triage tasks. Each task asks an agent to choose among candidate noncoding edits for a specified assay and… See the full description on the dataset page: https://huggingface.co/datasets/Jarrodbarnes/latent-mining.us-oil-gas-energy-mining-utility-layoffs-warn-act-notices-daily
US oil and gas, energy, mining and utility layoffs — the actual WARN Act filings, rebuilt every day
Last rebuilt: 2026-09-24. 1,281 layoff and closure notices filed by
oil and gas producers and oilfield-service contractors, coal and hard-rock mines, refineries and pipelines, electric and gas utilities, power plants, solar and wind manufacturers and installers, and waste, recycling and environmental-services operators with US state labor departments — 138,884 workers,
634… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-oil-gas-energy-mining-utility-layoffs-warn-act-notices-daily.mining-legal-arguments-us-corporate-case-law
Mining Legal Arguments in U.S. Corporate Case Law
This dataset contains span-level functional labels and directed support relations for 42 U.S. federal tax opinions concerning corporate reorganizations under I.R.C. Section 368. The opinions range in citation year from 1935 to 1987. Two law students annotated the cases, and a law professor adjudicated the final case-level representations. Ten cases also include the two independent annotations used for inter-annotator agreement… See the full description on the dataset page: https://huggingface.co/datasets/lbrenap1/mining-legal-arguments-us-corporate-case-law.quora-duplicates-mining
Dataset Card for Quora Duplicate Questions
This dataset contains the Quora Question Pairs dataset in a format that is easily used with the ParaphraseMiningEvaluator evaluator in Sentence Transformers. The data was originally created by Quora for this Kaggle Competition.
Usage
from datasets import load_dataset
from sentence_transformers.SentenceTransformer import SentenceTransformer
from sentence_transformers.evaluation import ParaphraseMiningEvaluator
# Load the… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/quora-duplicates-mining.europe-worldbank-energy-mining
Energy & Mining — Europe (World Bank WDI)
🇪🇺 55,013 observations · 44 Europe countries · 1962–2025 · Repackaged by Electric Sheep Europe
TL;DR
This dataset contains 55,013 observations of Energy & Mining data across 44 Europe countries, spanning 1962–2025, covering 41 distinct indicators.
About the source
The World Bank's World Development Indicators (WDI) is the world's most-cited reference for global development data. It compiles officially-recognized… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepeurope/europe-worldbank-energy-mining.northwind_opinion_mining_corpus
Opinion Mining Text Corpus
A labeled text corpus for opinion mining and sentiment analysis tasks, compiled from an open product review text corpus dataset publicly hosted on this Hub. The source corpus was assembled by a university research center.
This card does not yet list the source dataset or the applicable usage terms.
africa-world-bank-energy-and-mining-indicators-for-nigeria
Nigeria - Energy and Mining | Africa (original)
Size category: 1K<n<10K - Formats: parquet - Sector: economics_finance - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets help analysts inspect structured… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-world-bank-energy-and-mining-indicators-for-nigeria.misconception_miningtext-mining-ce-dataset
Vietnamese Legal Cross-Encoder Dataset
Training data for a cross-encoder reranker on Vietnamese legal documents.
Source
Built from YuITC/Vietnamese-Legal-Documents.
Schema
Column
Type
Description
qid
int64
Query ID
cid
int64
Document (context) ID
query
string
Legal question
document
string
Candidate document
label
int64
1 = positive, 0 = negative
split
string
train or test
negative_type
string
random, same_topic_wrong_article… See the full description on the dataset page: https://huggingface.co/datasets/juzharii/text-mining-ce-dataset.global-mining-areas
Quick Links
Website
GitHub
Hugging Face
LinkedIn
X
Global Mining Areas, Modernized
21,060 mining polygons covering 57,278 km² worldwide, representing
the physical land footprint of mining activity — open-pit/open-cut areas,
tailings dams, waste-rock dumps, water ponds, processing infrastructure, and
other directly associated mining land use. Modernized from Maus et al.
(2020)'s PANGAEA release into NORA's standardized schema. No new… See the full description on the dataset page: https://huggingface.co/datasets/NoraResearchLab/global-mining-areas.german_argument_mining
Dataset Card for Annotated German Legal Decision Corpus
Dataset Summary
This dataset consists of 200 randomly chosen judgments. In these judgments a legal expert annotated the components
conclusion, definition and subsumption of the German legal writing style Urteilsstil.
"Overall 25,075 sentences are annotated. 5% (1,202) of these sentences are marked as conclusion, 21% (5,328) as
definition, 53% (13,322) are marked as subsumption and the remaining 21% (6,481) as other.… See the full description on the dataset page: https://huggingface.co/datasets/joelniklaus/german_argument_mining.preference_lamp5_hard_mining_v3africa-gabon-gabon-energy-and-mining-9f88bafc
Gabon - Energy and Mining | Africa (Gabon official open data)
1,437 rows - 1 Africa country - 1962-2025 - Repackaged by Electric Sheep Africa
TL;DR
This dataset packages one official CSV resource from Gabon as
ML-ready Parquet. The source file is the provenance boundary; all usable
indicators or tabular columns from the resource stay together in this repo.
About the source
Source: Gabon - Energy and Mining
Publisher: World Bank Group
Resource:… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-gabon-gabon-energy-and-mining-9f88bafc.africa-synth-mining-equipment-failure-all
African Mining Equipment Failure Dataset | Africa (Electric Sheep Africa metadata inventory)
Size category: 1K<n<10K - Formats: parquet - Sector: energy - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-mining-equipment-failure-all.africa-synth-mining-safety-incidents-all
African Mining Safety Incidents Dataset | Africa (Electric Sheep Africa metadata inventory)
Size category: 1K<n<10K - Formats: parquet - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health datasets… See the full description on the dataset page: https://huggingface.co/datasets/Aasishshr7/africa-synth-mining-safety-incidents-all.
