datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CQADupstackAndroidRetrieval
CQADupstackAndroidRetrieval
An MTEB dataset
Massive Text Embedding Benchmark
CQADupStack: A Benchmark Data Set for Community Question-Answering Research
Task category
t2t
Domains
Programming, Web, Written, Non-fiction
Referencehttp://nlp.cis.unimelb.edu.au/resources/cqadupstack/
Source datasets:
mteb/cqadupstack-android
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/CQADupstackAndroidRetrieval.cqadupstack
Dataset Card for BEIR Benchmark
Dataset Summary
BEIR is a heterogeneous benchmark built from 18 diverse datasets representing 9 information retrieval tasks.
Fact-checking: FEVER, Climate-FEVER, SciFact
Question-Answering: NQ, HotpotQA, FiQA-2018
Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus
News Retrieval: TREC-NEWS, Robust04
Argument Retrieval: Touche-2020, ArguAna
Duplicate Question Retrieval: Quora, CqaDupstack
Citation-Prediction: SCIDOCS
Tweet Retrieval: Signal-1M… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/cqadupstack.CQADupstack-Wordpress-PL
CQADupstack-Wordpress-PL
An MTEB dataset
Massive Text Embedding Benchmark
CQADupStack: A Stack Exchange Question Duplicate Pairs Dataset
Task category
t2t
Domains
Written, Web, Programming
Reference
https://huggingface.co/datasets/clarin-knext/cqadupstack-wordpress-pl
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["CQADupstack-Wordpress-PL"])
evaluator =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/CQADupstack-Wordpress-PL.cqadupstack-physics-vn
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["CQADupstackPhysics-VN"])
evaluator = mteb.MTEB(task)
model = mteb.get_model(YOUR_MODEL)
evaluator.run(model)
To learn more about how to run models on mteb task check out the GitHub repitory.
Citation
If you use this dataset, please cite the dataset as well as mteb, as this dataset likely includes additional processing as a… See the full description on the dataset page: https://huggingface.co/datasets/GreenNode/cqadupstack-physics-vn.CQADupstack-Programmers-PL
CQADupstack-Programmers-PL
An MTEB dataset
Massive Text Embedding Benchmark
CQADupStack: A Stack Exchange Question Duplicate Pairs Dataset
Task category
t2t
Domains
Programming, Written, Non-fiction
Reference
https://huggingface.co/datasets/clarin-knext/cqadupstack-programmers-pl
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["CQADupstack-Programmers-PL"])… See the full description on the dataset page: https://huggingface.co/datasets/mteb/CQADupstack-Programmers-PL.cqadupstack-programmers-vn
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["CQADupstackProgrammers-VN"])
evaluator = mteb.MTEB(task)
model = mteb.get_model(YOUR_MODEL)
evaluator.run(model)
To learn more about how to run models on mteb task check out the GitHub repitory.
Citation
If you use this dataset, please cite the dataset as well as mteb, as this dataset likely includes additional processing… See the full description on the dataset page: https://huggingface.co/datasets/GreenNode/cqadupstack-programmers-vn.CQADupstack-English-PL
CQADupstack-English-PL
An MTEB dataset
Massive Text Embedding Benchmark
CQADupStack: A Stack Exchange Question Duplicate Pairs Dataset
Task category
t2t
Domains
Written
Reference
https://huggingface.co/datasets/clarin-knext/cqadupstack-english-pl
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["CQADupstack-English-PL"])
evaluator = mteb.MTEB(task)
model =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/CQADupstack-English-PL.beir-cqadupstack-mathematica
CQADupstackMathematicaRetrieval — BEIR, unified schema
A normalised copy of the dataset behind the mteb task CQADupstackMathematicaRetrieval, one of the tasks of the BEIR benchmark as mteb defines it (a member of the aggregate task CQADupstackRetrieval). Same queries, documents
and relevance judgements as the benchmark evaluates — reshaped into one strict schema shared by every dataset
in this collection.
Source
mteb/cqadupstack-mathematica @ 90fceea13679 (the… See the full description on the dataset page: https://huggingface.co/datasets/Hyukkyu/beir-cqadupstack-mathematica.cqadupstack-webmasters-vn
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["CQADupstackWebmasters-VN"])
evaluator = mteb.MTEB(task)
model = mteb.get_model(YOUR_MODEL)
evaluator.run(model)
To learn more about how to run models on mteb task check out the GitHub repitory.
Citation
If you use this dataset, please cite the dataset as well as mteb, as this dataset likely includes additional processing… See the full description on the dataset page: https://huggingface.co/datasets/GreenNode/cqadupstack-webmasters-vn.CQADupstack-Stats-PL
CQADupstack-Stats-PL
An MTEB dataset
Massive Text Embedding Benchmark
CQADupStack: A Stack Exchange Question Duplicate Pairs Dataset
Task category
t2t
Domains
Written, Academic, Non-fiction
Reference
https://huggingface.co/datasets/clarin-knext/cqadupstack-stats-pl
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["CQADupstack-Stats-PL"])
evaluator =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/CQADupstack-Stats-PL.cqadupstack-programmers-vn-rawCQADupstack-Webmasters-PL
CQADupstack-Webmasters-PL
An MTEB dataset
Massive Text Embedding Benchmark
CQADupStack: A Stack Exchange Question Duplicate Pairs Dataset
Task category
t2t
Domains
Written, Web
Reference
https://huggingface.co/datasets/clarin-knext/cqadupstack-webmasters-pl
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["CQADupstack-Webmasters-PL"])
evaluator =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/CQADupstack-Webmasters-PL.cqadupstack-android-vn
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["CQADupstackAndroid-VN"])
evaluator = mteb.MTEB(task)
model = mteb.get_model(YOUR_MODEL)
evaluator.run(model)
To learn more about how to run models on mteb task check out the GitHub repitory.
Citation
If you use this dataset, please cite the dataset as well as mteb, as this dataset likely includes additional processing as a… See the full description on the dataset page: https://huggingface.co/datasets/GreenNode/cqadupstack-android-vn.cqadupstack-mathematica-vn
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["CQADupstackMathematica-VN"])
evaluator = mteb.MTEB(task)
model = mteb.get_model(YOUR_MODEL)
evaluator.run(model)
To learn more about how to run models on mteb task check out the GitHub repitory.
Citation
If you use this dataset, please cite the dataset as well as mteb, as this dataset likely includes additional processing… See the full description on the dataset page: https://huggingface.co/datasets/GreenNode/cqadupstack-mathematica-vn.beir-cqadupstack-wordpress
CQADupstackWordpressRetrieval — BEIR, unified schema
A normalised copy of the dataset behind the mteb task CQADupstackWordpressRetrieval, one of the tasks of the BEIR benchmark as mteb defines it (a member of the aggregate task CQADupstackRetrieval). Same queries, documents
and relevance judgements as the benchmark evaluates — reshaped into one strict schema shared by every dataset
in this collection.
Source
mteb/cqadupstack-wordpress @ 4ffe81d471b1 (the revision… See the full description on the dataset page: https://huggingface.co/datasets/Hyukkyu/beir-cqadupstack-wordpress.CQADupstack-Mathematica-PL
CQADupstack-Mathematica-PL
An MTEB dataset
Massive Text Embedding Benchmark
CQADupStack: A Stack Exchange Question Duplicate Pairs Dataset
Task category
t2t
Domains
Written, Academic, Non-fiction
Reference
https://huggingface.co/datasets/clarin-knext/cqadupstack-mathematica-pl
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["CQADupstack-Mathematica-PL"])… See the full description on the dataset page: https://huggingface.co/datasets/mteb/CQADupstack-Mathematica-PL.beir-cqadupstack-gaming
CQADupstackGamingRetrieval — BEIR, unified schema
A normalised copy of the dataset behind the mteb task CQADupstackGamingRetrieval, one of the tasks of the BEIR benchmark as mteb defines it (a member of the aggregate task CQADupstackRetrieval). Same queries, documents
and relevance judgements as the benchmark evaluates — reshaped into one strict schema shared by every dataset
in this collection.
Source
mteb/cqadupstack-gaming @ 4885aa143210 (the revision pinned in… See the full description on the dataset page: https://huggingface.co/datasets/Hyukkyu/beir-cqadupstack-gaming.beir-cqadupstack-unix
CQADupstackUnixRetrieval — BEIR, unified schema
A normalised copy of the dataset behind the mteb task CQADupstackUnixRetrieval, one of the tasks of the BEIR benchmark as mteb defines it (a member of the aggregate task CQADupstackRetrieval). Same queries, documents
and relevance judgements as the benchmark evaluates — reshaped into one strict schema shared by every dataset
in this collection.
Source
mteb/cqadupstack-unix @ 6c6430d3a6d3 (the revision pinned in mteb)… See the full description on the dataset page: https://huggingface.co/datasets/Hyukkyu/beir-cqadupstack-unix.beir-cqadupstack-android
CQADupstackAndroidRetrieval — BEIR, unified schema
A normalised copy of the dataset behind the mteb task CQADupstackAndroidRetrieval, one of the tasks of the BEIR benchmark as mteb defines it (a member of the aggregate task CQADupstackRetrieval). Same queries, documents
and relevance judgements as the benchmark evaluates — reshaped into one strict schema shared by every dataset
in this collection.
Source
mteb/CQADupstackAndroidRetrieval @ 9be4c0e46342 (the revision… See the full description on the dataset page: https://huggingface.co/datasets/Hyukkyu/beir-cqadupstack-android.beir-cqadupstack-webmasters
CQADupstackWebmastersRetrieval — BEIR, unified schema
A normalised copy of the dataset behind the mteb task CQADupstackWebmastersRetrieval, one of the tasks of the BEIR benchmark as mteb defines it (a member of the aggregate task CQADupstackRetrieval). Same queries, documents
and relevance judgements as the benchmark evaluates — reshaped into one strict schema shared by every dataset
in this collection.
Source
mteb/cqadupstack-webmasters @ 160c094312a0 (the revision… See the full description on the dataset page: https://huggingface.co/datasets/Hyukkyu/beir-cqadupstack-webmasters.cqadupstack-unix-vn
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["CQADupstackUnix-VN"])
evaluator = mteb.MTEB(task)
model = mteb.get_model(YOUR_MODEL)
evaluator.run(model)
To learn more about how to run models on mteb task check out the GitHub repitory.
Citation
If you use this dataset, please cite the dataset as well as mteb, as this dataset likely includes additional processing as a… See the full description on the dataset page: https://huggingface.co/datasets/GreenNode/cqadupstack-unix-vn.beir-cqadupstack-gis
CQADupstackGisRetrieval — BEIR, unified schema
A normalised copy of the dataset behind the mteb task CQADupstackGisRetrieval, one of the tasks of the BEIR benchmark as mteb defines it (a member of the aggregate task CQADupstackRetrieval). Same queries, documents
and relevance judgements as the benchmark evaluates — reshaped into one strict schema shared by every dataset
in this collection.
Source
mteb/cqadupstack-gis @ 5003b3064772 (the revision pinned in mteb)… See the full description on the dataset page: https://huggingface.co/datasets/Hyukkyu/beir-cqadupstack-gis.CQADupstack-Gis-PL
CQADupstack-Gis-PL
An MTEB dataset
Massive Text Embedding Benchmark
CQADupStack: A Stack Exchange Question Duplicate Pairs Dataset
Task category
t2t
Domains
Written, Academic, Non-fiction
Reference
https://huggingface.co/datasets/clarin-knext/cqadupstack-gis-pl
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["CQADupstack-Gis-PL"])
evaluator = mteb.MTEB(task)… See the full description on the dataset page: https://huggingface.co/datasets/mteb/CQADupstack-Gis-PL.cqadupstack-stats-vn
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["CQADupstackStats-VN"])
evaluator = mteb.MTEB(task)
model = mteb.get_model(YOUR_MODEL)
evaluator.run(model)
To learn more about how to run models on mteb task check out the GitHub repitory.
Citation
If you use this dataset, please cite the dataset as well as mteb, as this dataset likely includes additional processing as a… See the full description on the dataset page: https://huggingface.co/datasets/GreenNode/cqadupstack-stats-vn.cqadupstack-wordpress-vn
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["CQADupstackWordpress-VN"])
evaluator = mteb.MTEB(task)
model = mteb.get_model(YOUR_MODEL)
evaluator.run(model)
To learn more about how to run models on mteb task check out the GitHub repitory.
Citation
If you use this dataset, please cite the dataset as well as mteb, as this dataset likely includes additional processing as… See the full description on the dataset page: https://huggingface.co/datasets/GreenNode/cqadupstack-wordpress-vn.beir-cqadupstack-english
CQADupstackEnglishRetrieval — BEIR, unified schema
A normalised copy of the dataset behind the mteb task CQADupstackEnglishRetrieval, one of the tasks of the BEIR benchmark as mteb defines it (a member of the aggregate task CQADupstackRetrieval). Same queries, documents
and relevance judgements as the benchmark evaluates — reshaped into one strict schema shared by every dataset
in this collection.
Source
mteb/cqadupstack-english @ ad9991cb51e3 (the revision pinned… See the full description on the dataset page: https://huggingface.co/datasets/Hyukkyu/beir-cqadupstack-english.beir-cqadupstack-physics
CQADupstackPhysicsRetrieval — BEIR, unified schema
A normalised copy of the dataset behind the mteb task CQADupstackPhysicsRetrieval, one of the tasks of the BEIR benchmark as mteb defines it (a member of the aggregate task CQADupstackRetrieval). Same queries, documents
and relevance judgements as the benchmark evaluates — reshaped into one strict schema shared by every dataset
in this collection.
Source
mteb/cqadupstack-physics @ 79531abbd1fb (the revision pinned… See the full description on the dataset page: https://huggingface.co/datasets/Hyukkyu/beir-cqadupstack-physics.CQADupstackGisRetrieval-Fa
CQADupstackGisRetrieval-Fa
An MTEB dataset
Massive Text Embedding Benchmark
CQADupstackGisRetrieval-Fa
Task category
t2t
Domains
Web
Reference
https://huggingface.co/datasets/MCINext/cqadupstack-gis-fa
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["CQADupstackGisRetrieval-Fa"])
evaluator = mteb.MTEB(task)
model = mteb.get_model(YOUR_MODEL)… See the full description on the dataset page: https://huggingface.co/datasets/mteb/CQADupstackGisRetrieval-Fa.beir-cqadupstack-programmers
CQADupstackProgrammersRetrieval — BEIR, unified schema
A normalised copy of the dataset behind the mteb task CQADupstackProgrammersRetrieval, one of the tasks of the BEIR benchmark as mteb defines it (a member of the aggregate task CQADupstackRetrieval). Same queries, documents
and relevance judgements as the benchmark evaluates — reshaped into one strict schema shared by every dataset
in this collection.
Source
mteb/cqadupstack-programmers @ 6184bc1440d2 (the… See the full description on the dataset page: https://huggingface.co/datasets/Hyukkyu/beir-cqadupstack-programmers.beir-cqadupstack-stats
CQADupstackStatsRetrieval — BEIR, unified schema
A normalised copy of the dataset behind the mteb task CQADupstackStatsRetrieval, one of the tasks of the BEIR benchmark as mteb defines it (a member of the aggregate task CQADupstackRetrieval). Same queries, documents
and relevance judgements as the benchmark evaluates — reshaped into one strict schema shared by every dataset
in this collection.
Source
mteb/cqadupstack-stats @ 65ac3a16b8e9 (the revision pinned in… See the full description on the dataset page: https://huggingface.co/datasets/Hyukkyu/beir-cqadupstack-stats.
