datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
open-subtitles-bitext-miningopen-subtitles-256s-bitext-miningtatoeba-bitext-mining
Tatoeba
An MTEB dataset
Massive Text Embedding Benchmark
1,000 English-aligned sentence pairs for each language based on the Tatoeba corpus
Task category
t2t
Domains
Written
Reference
https://github.com/facebookresearch/LASER/tree/main/data/tatoeba/v1
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["Tatoeba"])
evaluator = mteb.MTEB(task)
model =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/tatoeba-bitext-mining.instructions-pair-miningopen-subtitles-500-bitext-miningbucc-bitext-mining
BUCC.v2
An MTEB dataset
Massive Text Embedding Benchmark
BUCC bitext mining dataset
Task category
t2t
Domains
Written
Reference
https://comparable.limsi.fr/bucc2018/bucc2018-task.html
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["BUCC.v2"])
evaluator = mteb.MTEB(task)
model = mteb.get_model(YOUR_MODEL)
evaluator.run(model)
To learn more about how to run… See the full description on the dataset page: https://huggingface.co/datasets/mteb/bucc-bitext-mining.open-subtitles-250-bitext-miningtatoeba-bitext-miningnorthwind_opinion_mining_corpus
Opinion Mining Text Corpus
A labeled text corpus for opinion mining and sentiment analysis tasks, compiled from an open product review text corpus dataset publicly hosted on this Hub. The source corpus was assembled by a university research center.
This card does not yet list the source dataset or the applicable usage terms.
german_argument_mining
Dataset Card for Annotated German Legal Decision Corpus
Dataset Summary
This dataset consists of 200 randomly chosen judgments. In these judgments a legal expert annotated the components
conclusion, definition and subsumption of the German legal writing style Urteilsstil.
"Overall 25,075 sentences are annotated. 5% (1,202) of these sentences are marked as conclusion, 21% (5,328) as
definition, 53% (13,322) are marked as subsumption and the remaining 21% (6,481) as other.… See the full description on the dataset page: https://huggingface.co/datasets/joelniklaus/german_argument_mining.mining-legal-arguments-us-corporate-case-law
Mining Legal Arguments in U.S. Corporate Case Law
This dataset contains span-level functional labels and directed support relations for 42 U.S. federal tax opinions concerning corporate reorganizations under I.R.C. Section 368. The opinions range in citation year from 1935 to 1987. Two law students annotated the cases, and a law professor adjudicated the final case-level representations. Ten cases also include the two independent annotations used for inter-annotator agreement… See the full description on the dataset page: https://huggingface.co/datasets/lbrenap1/mining-legal-arguments-us-corporate-case-law.warehouse-process-mining-benchmark
Warehouse Process Mining Benchmark
Benchmark results for WareFlowTwin process mining and optimization, aligned with VillanovaAI/Temporal_Logistics_Inventory_Movements.
Contents
File
Description
manifest.json
Dataset metadata
eval_results.json
Aggregate benchmark metrics
warehouse_*.json
Per-instance process mining + optimization results
Pipeline Evaluated
Event log construction (lot_id, case_id, activity, timestamp, warehouse… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/warehouse-process-mining-benchmark.embedding-pair-miningdeep-seabed-mining-law-corpus
Deep Seabed Mining Law Corpus
A neutral, provenance-first, machine-readable record of the law governing mineral resources of "the Area" (the seabed beyond national jurisdiction): the international ISA/UNCLOS regime and the US non-UNCLOS parallel track. Every record carries its official source, retrieval date, citation, language, an authoritative-status flag, and a SHA-256 content hash; texts are verified against official sources.
Source of truth / build history:… See the full description on the dataset page: https://huggingface.co/datasets/dacheah/deep-seabed-mining-law-corpus.openvalidators-mining
DEPRECATION NOTICE
As of August 1, 2023, the OpenValidators Mining dataset has been officially deprecated and discontinued. We are no longer updating or maintaining this dataset.
If this data has any relevance to your current or future projects, or you have any questions or concerns related to the deprecation, please do not hesitate to contact us. We are more than willing to assist you and provide guidance for alternative solutions where possible.
Dataset Card for… See the full description on the dataset page: https://huggingface.co/datasets/opentensor/openvalidators-mining.argument_mining_de
Dataset Card for Argument Mining DE
This dataset contains German sentences typically found in the conclusion sections of scientific and academic texts in the field of computer science and information systems. Each sentence is labeled with one of six fine-grained categories commonly used in discourse and argumentation structure analysis:
CLAIM: A statement that puts forward a main point or assertion.
COUNTERCLAIM: A statement that challenges a previous claim.
LINK: A sentence that… See the full description on the dataset page: https://huggingface.co/datasets/samirmsallem/argument_mining_de.mining_legal_arguments_argType
Dataset Card for MiningLegalArguments
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]
Source… See the full description on the dataset page: https://huggingface.co/datasets/joelniklaus/mining_legal_arguments_argType.Role-Mining-JSON_Inst-Textmining_legal_arguments_agent
Dataset Card for MiningLegalArguments
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]
Source… See the full description on the dataset page: https://huggingface.co/datasets/joelniklaus/mining_legal_arguments_agent.autotrain-data-amdal-mining-llama2-7b-cleanmining-engineering-alpaca-idprocess_mining_questionsargument_mining_de
Dataset Card for Argument Mining DE
This dataset contains German sentences typically found in the conclusion sections of scientific and academic texts in the field of computer science and information systems. Each sentence is labeled with one of six fine-grained categories commonly used in discourse and argumentation structure analysis:
CLAIM: A statement that puts forward a main point or assertion.
COUNTERCLAIM: A statement that challenges a previous claim.
LINK: A sentence that… See the full description on the dataset page: https://huggingface.co/datasets/Alfiya2026/argument_mining_de.LCA_Mining_1
