datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wiki-18-e5-indexmsmarco-beir-e5wiki_dpr_e5wiki_dpr encoded with intfloat/e5-base-v2
bge-e5datae5a023d2mteb-lite-run-files-e5e52c0fcdmassive_serve_pes2o_v3_e5_base_v2wiki-18-e5-index-HNSW64cole-arxiv-cc-e5-retriever
CoLe arXiv CC E5 Retriever
A non-Wikipedia retrieval corpus for the CoLe/R2–R4a experiments. It contains
280,738 arXiv documents selected from the Common Pile
filtered arXiv collection, plus an E5-base-v2 dense index.
Data provenance and license handling
Source: common-pile/arxiv_papers_filtered, revision 033cf7f.
The source records are converted arXiv papers with per-document license metadata.
This release keeps records whose metadata is CC BY, CC0, or Public… See the full description on the dataset page: https://huggingface.co/datasets/KyrieX/cole-arxiv-cc-e5-retriever.E5-RTraining data for e5-R-mistral-7b. Code is here.
AOPS_Full_Verified_e5occluded-vegetables-in-packed-fridge-rt-detr-detection-next-pack-e522b0f8-7563b26f
Occluded Vegetable Detection in Packed Fridge
Training dataset of packed refrigerator scenes with partially occluded vegetables, rendered to fine-tune an RT-DETR object detection model for improved detection under heavy occlusion.
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and… See the full description on the dataset page: https://huggingface.co/datasets/physicl-test/occluded-vegetables-in-packed-fridge-rt-detr-detection-next-pack-e522b0f8-7563b26f.ragwiki-e5.flex
ragwiki-e5.flex
Description
TODO: What is the artifact?
Usage
# Load the artifact
import pyterrier_alpha as pta
artifact = pta.Artifact.from_hf('pyterrier/ragwiki-e5.flex')
# TODO: Show how you use the artifact
Benchmarks
TODO: Provide benchmarks for the artifact.
Reproduction
# TODO: Show how you constructed the artifact.
Metadata
{
"type": "dense_index",
"format": "flex",
"vec_size": 768,
"doc_count": 21015324
}
corpus_embeddings_multilingual-e5-large-instructworkshop3uncannyelevationperiphery-aviser-e5processed-embeddings-e5-mistraljusticio-BOE-A-1978-31229-constitucion-by-articles-qa-multilingual-e5-large-groq_llama3_70b-sas
Dataset summary
It is an end-to-end evaluation dataset (using SAS metric) for Justicio.
Domain: Legal, Law, Spanish Constitution
Language: Spanish
SAS summary
The concept of Semantic Answer Similarity refers to the evaluation of the semantic similarity between the generated answer and the ground truth. The score ranges from 0 to 1. A higher score indicates a better match between the generated answer and the ground truth.
Justicio summary
Justicio is a… See the full description on the dataset page: https://huggingface.co/datasets/dariolopez/justicio-BOE-A-1978-31229-constitucion-by-articles-qa-multilingual-e5-large-groq_llama3_70b-sas.vietnamese-evidence-corpus-embeddings-e5-large-v3
Vietnamese Evidence Corpus Embeddings
Normalized passage embeddings for Loctran123/vietnamese-evidence-corpus-chunked-e5-v3, generated with
intfloat/multilingual-e5-large.
Rows: 53,114
Embedding dimension: 1024
Embedding dtype: float32
Input: `passage: {title}
{text}`
Inputs truncated to 512 tokens before embedding
Deduplicated before embedding by normalized content_hash
Provenance retained for duplicate content
L2 normalized: yes
Parquet shards: 11
Use query: for… See the full description on the dataset page: https://huggingface.co/datasets/Loctran123/vietnamese-evidence-corpus-embeddings-e5-large-v3.screwdriver_attach_panel_rs_080125_4_e5This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "koch_screwdriver_follower",
"total_episodes": 5,
"total_frames": 976,
"total_tasks": 1,
"total_videos": 15,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:5"},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jackvial/screwdriver_attach_panel_rs_080125_4_e5.AOPS_Full_Verified_e5_5678910wikipedia-embeddings-cs-e5-baseThis dataset contains the Czech subset of the wikimedia/wikipedia dataset. Each page is divided into paragraphs, stored as a list in the chunks column. For every paragraph, embeddings are created using the intfloat/multilingual-e5-base model.
Usage
Load the dataset:
from datasets import load_dataset
ds = load_dataset("karmiq/wikipedia-embeddings-cs-e5-base", split="train")
ds[1]
{
'id': '1',
'url': 'https://cs.wikipedia.org/wiki/Astronomie',
'title': 'Astronomie'… See the full description on the dataset page: https://huggingface.co/datasets/karmiq/wikipedia-embeddings-cs-e5-base.robustness_e5This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "koch",
"total_episodes": 21,
"total_frames": 8703,
"total_tasks":1,
"total_videos": 42,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:21"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/nduque/robustness_e5.v3_msmarco_parallelai_e5qwen7b_6intent_claim_degradelegal-e5-adamw-r170-eval-v1
Legal-E5 AdamW R@170 Candidate Evaluation
generated_at: 2026-09-05T10:15:36.507584+00:00
input_repo: nhavanvietcode/legal-e5-adamw-r170-candidates-v1
selected_by: R@170 with ties [100, 50, 20, 5]
best_candidate: checkpoint-141
best_model_path_in_repo: models/checkpoint-141
Candidate
R@5
R@20
R@50
R@100
R@170
checkpoint-141
0.80935
0.90893
0.95357
0.97304
0.98089
final
0.81149
0.91060
0.95357
0.97268
0.97946
checkpoint-94
0.80256
0.90881
0.95048
0.97018
0.97839… See the full description on the dataset page: https://huggingface.co/datasets/nhavanvietcode/legal-e5-adamw-r170-eval-v1.vietnamese-evidence-corpus-embeddings-e5-large
Vietnamese Evidence Corpus Embeddings
Normalized passage embeddings for Loctran123/vietnamese-evidence-corpus-chunked, generated with
intfloat/multilingual-e5-large.
Rows: 52,605
Embedding dimension: 1024
Embedding dtype: float32
Prefix: passage:
L2 normalized: yes
Parquet shards: 11
Use query: for claims/queries and normalize query vectors before cosine or
inner-product retrieval. The chunk_id column is the stable join key back to
the source corpus.
screwdriver_attach_panel_rs_080125_17_e5This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "koch_screwdriver_follower",
"total_episodes": 5,
"total_frames": 978,
"total_tasks": 1,
"total_videos": 15,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:5"},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jackvial/screwdriver_attach_panel_rs_080125_17_e5.screwdriver_panel_ls_080225_1_e5This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "koch_screwdriver_follower",
"total_episodes": 5,
"total_frames": 1220,
"total_tasks": 1,
"total_videos": 15,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:5"},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jackvial/screwdriver_panel_ls_080225_1_e5.Keyword_Doc_intfloat_multilingual_e5
