datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hearing2translate-humeval
This repository contains the human evaluation experiment data for Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs 📄.
The code for the project is hosted at github.com/sarapapi/hearing2translate.
The annotations were collected using Pearmut (code), a lightweight platform that makes end-to-end human evaluation for multilingual tasks efficient and reliable.
The evaluations were done with bilingual speakers using the Pearmut tool with the MQM/ESA protocol… See the full description on the dataset page: https://huggingface.co/datasets/zouhar/hearing2translate-humeval.HUMEWikipediaRerankingMultilingual
HUMEWikipediaRerankingMultilingual
An MTEB dataset
Massive Text Embedding Benchmark
Human evaluation subset of Wikipedia reranking dataset across multiple languages.
Task category
t2t
Domains
Encyclopaedic, Written
Referencehttps://github.com/ellamind/wikipedia-2023-11-reranking-multilingual
Source datasets:
mteb/WikipediaRerankingMultilingual
mteb/mteb-human-wiki-reranking
How to evaluate on this task
You can evaluate an embedding model on this… See the full description on the dataset page: https://huggingface.co/datasets/mteb/HUMEWikipediaRerankingMultilingual.HUMEToxicConversationsClassificationHUMEEmotionClassificationhum_eye4b
Dataset Card for "hum_eye4b"
More Information needed
HUMERobust04InstructionReranking
HUMERobust04InstructionReranking
An MTEB dataset
Massive Text Embedding Benchmark
Human evaluation subset of Robust04 instruction retrieval dataset for reranking evaluation.
Task category
t2t
Domains
News, Written
Referencehttps://trec.nist.gov/data/robust/04.guidelines.html
Source datasets:
jhu-clsp/robust04-instructions-mteb
mteb/mteb-human-robust04-reranking
How to evaluate on this task
You can evaluate an embedding model on this dataset using… See the full description on the dataset page: https://huggingface.co/datasets/mteb/HUMERobust04InstructionReranking.HUMEMultilingualSentimentClassificationHUMECore17InstructionReranking
HUMECore17InstructionReranking
An MTEB dataset
Massive Text Embedding Benchmark
Human evaluation subset of Core17 instruction retrieval dataset for reranking evaluation.
Task category
t2t
Domains
News, Written
Referencehttps://arxiv.org/abs/2403.15246
Source datasets:
jhu-clsp/core17-instructions-mteb
mteb/mteb-human-core17-reranking
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb… See the full description on the dataset page: https://huggingface.co/datasets/mteb/HUMECore17InstructionReranking.HUMENews21InstructionReranking
HUMENews21InstructionReranking
An MTEB dataset
Massive Text Embedding Benchmark
Human evaluation subset of News21 instruction retrieval dataset for reranking evaluation.
Task category
t2t
Domains
News, Written
Referencehttps://trec.nist.gov/data/news2021.html
Source datasets:
jhu-clsp/news21-instructions-mteb
mteb/mteb-human-news21-reranking
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:… See the full description on the dataset page: https://huggingface.co/datasets/mteb/HUMENews21InstructionReranking.HUMETweetSentimentExtractionClassificationAiazusa_gzxFLDCV_Datase1_HumElearn_hf_humen_ai_captionshumevaltext_2_sqlHume-AI-QA-Dataset
