datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
open-subtitles-bitext-miningopen-subtitles-256s-bitext-miningtatoeba-bitext-mining
Tatoeba
An MTEB dataset
Massive Text Embedding Benchmark
1,000 English-aligned sentence pairs for each language based on the Tatoeba corpus
Task category
t2t
Domains
Written
Reference
https://github.com/facebookresearch/LASER/tree/main/data/tatoeba/v1
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["Tatoeba"])
evaluator = mteb.MTEB(task)
model =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/tatoeba-bitext-mining.bitcoin-mining-pool-templates
Bitcoin mining pool templates
Timestamped Stratum job messages collected directly from Bitcoin mining pool endpoints. The data records changes in the work each endpoint sends to miners, including the previous block hash, coinbase data and clean-jobs flag.
Contents
Table
Record
bitcoin_mining_pool_jobs
A job received from a pool endpoint, with its observation time, nTime, coinbase, merkle branch count and clean-jobs flag
Using the data… See the full description on the dataset page: https://huggingface.co/datasets/dataforge-labs/bitcoin-mining-pool-templates.instructions-pair-miningmteb-bitext-mining-aggregated
MTEB BitextMining Aggregated Dataset (Full)
This dataset aggregates ALL configs from 10 BitextMining datasets in the MTEB (Massive Text Embedding Benchmark) Multilingual v2 benchmark into a single, unified dataset for comprehensive bitext mining evaluation.
Dataset Summary
Total Examples: 448,229 sentence pairs
Source Datasets (Configs): 10 MTEB BitextMining tasks
Total Splits: 332 language pairs/configurations
Languages: 300+ unique language codes across all datasets… See the full description on the dataset page: https://huggingface.co/datasets/SaylorTwift/mteb-bitext-mining-aggregated.open-subtitles-500-bitext-miningbucc-bitext-mining
BUCC.v2
An MTEB dataset
Massive Text Embedding Benchmark
BUCC bitext mining dataset
Task category
t2t
Domains
Written
Reference
https://comparable.limsi.fr/bucc2018/bucc2018-task.html
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["BUCC.v2"])
evaluator = mteb.MTEB(task)
model = mteb.get_model(YOUR_MODEL)
evaluator.run(model)
To learn more about how to run… See the full description on the dataset page: https://huggingface.co/datasets/mteb/bucc-bitext-mining.asia-energy-world-bank-energy-and-mining-indicators
India - Energy and Mining
Publisher: World Bank Group · Source: HDX · License: cc-by · Updated: 2026-04-28
Abstract
Contains data from the World Bank's data portal. There is also a consolidated country dataset on HDX.
The world economy needs ever-increasing amounts of energy to sustain economic growth, raise living standards, and reduce poverty. But today's trends in energy use are not sustainable. As the world's population grows and economies become more industrialized… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-energy-world-bank-energy-and-mining-indicators.africa-world-bank-energy-mining-time-series
Africa World Bank Energy and Mining Labeled Time Series Data
This repository is part of the Africa Temporal Intelligence Corpus (ATIC). It contains sector-specific temporal corpus packages for African countries.
ATIC sector repositories are designed for machine consumption first: Parquet tables, stable IDs, reproducible metadata, explicit provenance, review status, and separable semantic layers.
Sector Scope
Temporal energy indicators for African countries… See the full description on the dataset page: https://huggingface.co/datasets/africatic/africa-world-bank-energy-mining-time-series.open-subtitles-250-bitext-miningtatoeba-bitext-mininglatent-mining
Latent Mining
Latent Mining is a benchmark-construction method for scientific-agent tasks where useful public evidence diverges from a withheld verifier-backed outcome. The resulting tasks test whether agents can make calibrated scientific triage decisions under incomplete information.
This dataset contains the first public biology subset: 165 cross-locus regulatory-edit triage tasks. Each task asks an agent to choose among candidate noncoding edits for a specified assay and… See the full description on the dataset page: https://huggingface.co/datasets/Jarrodbarnes/latent-mining.sparkproof-miningnorthwind_opinion_mining_corpus
Opinion Mining Text Corpus
A labeled text corpus for opinion mining and sentiment analysis tasks, compiled from an open product review text corpus dataset publicly hosted on this Hub. The source corpus was assembled by a university research center.
This card does not yet list the source dataset or the applicable usage terms.
Arguement_Mining_CL2017tokens along with chunk id. IOB1 format Begining of arguement denoted by B-ARG,inside arguement
denoted by I-ARG, other chunks are O
Orginial train,test split as used by the paper is providedIndustryCorpus2_mining
IndustryCorpus2: Mining
This repository contains the IndustryCorpus2: Mining domain subset of BAAI/IndustryCorpus2.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryCorpus2:
@misc{shi2024industrycorpus2,
title = {IndustryCorpus2},
author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and Donglin Hao},
year = {2024},
publisher… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus2_mining.quora-duplicates-mining
Dataset Card for Quora Duplicate Questions
This dataset contains the Quora Question Pairs dataset in a format that is easily used with the ParaphraseMiningEvaluator evaluator in Sentence Transformers. The data was originally created by Quora for this Kaggle Competition.
Usage
from datasets import load_dataset
from sentence_transformers.SentenceTransformer import SentenceTransformer
from sentence_transformers.evaluation import ParaphraseMiningEvaluator
# Load the… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/quora-duplicates-mining.europe-worldbank-energy-mining
Energy & Mining — Europe (World Bank WDI)
🇪🇺 55,013 observations · 44 Europe countries · 1962–2025 · Repackaged by Electric Sheep Europe
TL;DR
This dataset contains 55,013 observations of Energy & Mining data across 44 Europe countries, spanning 1962–2025, covering 41 distinct indicators.
About the source
The World Bank's World Development Indicators (WDI) is the world's most-cited reference for global development data. It compiles officially-recognized… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepeurope/europe-worldbank-energy-mining.preference_lamp5_hard_mining_v3misconception_mininggerman_argument_mining
Dataset Card for Annotated German Legal Decision Corpus
Dataset Summary
This dataset consists of 200 randomly chosen judgments. In these judgments a legal expert annotated the components
conclusion, definition and subsumption of the German legal writing style Urteilsstil.
"Overall 25,075 sentences are annotated. 5% (1,202) of these sentences are marked as conclusion, 21% (5,328) as
definition, 53% (13,322) are marked as subsumption and the remaining 21% (6,481) as other.… See the full description on the dataset page: https://huggingface.co/datasets/joelniklaus/german_argument_mining.text-mining-ce-dataset
Vietnamese Legal Cross-Encoder Dataset
Training data for a cross-encoder reranker on Vietnamese legal documents.
Source
Built from YuITC/Vietnamese-Legal-Documents.
Schema
Column
Type
Description
qid
int64
Query ID
cid
int64
Document (context) ID
query
string
Legal question
document
string
Candidate document
label
int64
1 = positive, 0 = negative
split
string
train or test
negative_type
string
random, same_topic_wrong_article… See the full description on the dataset page: https://huggingface.co/datasets/juzharii/text-mining-ce-dataset.global-mining-areas
Quick Links
Website
GitHub
Hugging Face
LinkedIn
X
Global Mining Areas, Modernized
21,060 mining polygons covering 57,278 km² worldwide, representing
the physical land footprint of mining activity — open-pit/open-cut areas,
tailings dams, waste-rock dumps, water ponds, processing infrastructure, and
other directly associated mining land use. Modernized from Maus et al.
(2020)'s PANGAEA release into NORA's standardized schema. No new… See the full description on the dataset page: https://huggingface.co/datasets/NoraResearchLab/global-mining-areas.mining-legal-arguments-us-corporate-case-law
Mining Legal Arguments in U.S. Corporate Case Law
This dataset contains span-level functional labels and directed support relations for 42 U.S. federal tax opinions concerning corporate reorganizations under I.R.C. Section 368. The opinions range in citation year from 1935 to 1987. Two law students annotated the cases, and a law professor adjudicated the final case-level representations. Ten cases also include the two independent annotations used for inter-annotator agreement… See the full description on the dataset page: https://huggingface.co/datasets/lbrenap1/mining-legal-arguments-us-corporate-case-law.africa-gabon-gabon-energy-and-mining-9f88bafc
Gabon - Energy and Mining | Africa (Gabon official open data)
1,437 rows - 1 Africa country - 1962-2025 - Repackaged by Electric Sheep Africa
TL;DR
This dataset packages one official CSV resource from Gabon as
ML-ready Parquet. The source file is the provenance boundary; all usable
indicators or tabular columns from the resource stay together in this repo.
About the source
Source: Gabon - Energy and Mining
Publisher: World Bank Group
Resource:… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-gabon-gabon-energy-and-mining-9f88bafc.Mining-Engineering-SFT
矿建工程领域中文指令与评估数据集
数据集概述
**注: 本数据集为不带CoT标注的数据集,如果您要对DeekSeek R1、Qwen3系列等具有内嵌的CoT输出的模型进行微调,为了避免模型发生灾难性遗忘,请移步至本项目的思维链增强训练集 (CoT-Enhanced SFT Dataset) **
本项目是合肥工业大学大一学生的大学生创新创业训练计划(大创)项目成果。我们构建了一套专为提升大型语言模型在中国矿建工程领域专业知识与实践能力而设计的中文数据集。
这套数据集旨在让模型掌握矿建工程的核心知识,内容覆盖了六大模块:
法律法规 (law)
工程规范 (specifications)
专业术语 (concept)
安全事故案例 (safety)
行业实践经验 (forum)
领域综合知识 (synthesis)
为了支持完整的模型开发、评估和验证周期,我们将数据组织为多个独立的Hugging Face仓库:
本数据集 (原始训练集): acnul/Mining-Engineering-SFT 包含 5,287… See the full description on the dataset page: https://huggingface.co/datasets/acnul/Mining-Engineering-SFT.africa-world-bank-energy-and-mining-indicators-for-rwanda
Rwanda - Energy and Mining | Africa (original)
Size category: 1K<n<10K - Formats: parquet - Sector: economics_finance - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets help analysts inspect structured… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-world-bank-energy-and-mining-indicators-for-rwanda.africa-synth-mining-equipment-failure-all
African Mining Equipment Failure Dataset | Africa (Electric Sheep Africa metadata inventory)
Size category: 1K<n<10K - Formats: parquet - Sector: energy - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-mining-equipment-failure-all.africa-synth-mining-safety-incidents-all
African Mining Safety Incidents Dataset | Africa (Electric Sheep Africa metadata inventory)
Size category: 1K<n<10K - Formats: parquet - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health datasets… See the full description on the dataset page: https://huggingface.co/datasets/Aasishshr7/africa-synth-mining-safety-incidents-all.
