datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
whisper_transcriptions.reazon_speech_all.wer_10.0.vectorizedwhisper_transcriptions.mls.wer_10.0.vectorizedWikidata_Vectors_0.2
Wikidata Entity Embeddings 0.2
Dataset Summary
Wikidata Entity Embeddings is a dataset of embedding vectors for Wikidata entities. Each vector represents a Wikidata item (Q...) or property (P...) based on textual information extracted from Wikidata.
The dataset is part of the Wikidata Embedding Project, an initiative led by Wikimedia Deutschland in collaboration with Jina AI and IBM DataStax. The project provides a publicly accessible Wikidata Vector Database to… See the full description on the dataset page: https://huggingface.co/datasets/philippesaade/Wikidata_Vectors_0.2.v32-vectorswhisper_transcriptions.reazonspeech.all.wer_10.0.vectorizedopeniti-vectors
OpenITI Vector Database — maktabati.ai
🇬🇧 English
This dataset contains the fully vectorized OpenITI RELEASE 2025-1-9 collection of classical Islamic texts, prepared for semantic search (RAG).
Each entry represents a text chunk from one of 8,943 works in the OpenITI corpus, together with its embedding vector and complete metadata.
Statistics:
4,696,703 chunks
8,943 works (primary editions only, status=pri from OpenITI TSV)
approx. 470 Parquet files (approx.… See the full description on the dataset page: https://huggingface.co/datasets/Maktabati/openiti-vectors.shamela-vectors
🇬🇧 English
The largest open-source vector database of the complete al-Maktaba al-Shamela (المكتبة الشاملة) Islamic text corpus, prepared for semantic search and Retrieval-Augmented Generation (RAG). Includes the full Quran text (6,236 verses, Hafs 'an 'Asim).
Statistics:
11,482,164 chunks from 8,589 classical Islamic books
6,236 Quran verses (one verse = one chunk, included in total)
40 categories covering the full breadth of Islamic scholarship
Period: 1st century AH to 15th… See the full description on the dataset page: https://huggingface.co/datasets/Maktabati/shamela-vectors.voice_medical_cut_medium_vectorarc_whisper_transcriptions.reazonspeech.small.wer_10.0.vectorized
Dataset Card for "arc_whisper_transcriptions.reazonspeech.small.wer_10.0.vectorized"
More Information needed
s2ef-15m
Dataset Description
This dataset contains a collection of 3D atomistic datasets with force and energy labels gathered from a series of sources:
Open Catalyst Project
OC20, OC22, ODAC23
Materials Project Trajectory Dataset (MPtrj)
SPICE 1.1.4
Dataset Structure
Data Instances
For each instance, there is set of atomic numbers (input_ids), 3-D coordinates (coords), a set of forces per atom (forces), the total and formation energy per
system… See the full description on the dataset page: https://huggingface.co/datasets/vector-institute/s2ef-15m.atom3d-res
Residue identity prediction
Overview
Understanding the structural role of individual amino acids is important for engineering new proteins.
We can understand this role by predicting the substitutabilities of different amino acids at a given protein site based on the surrounding structural environment.
We generate a novel dataset consisting of atomic environments extracted from nonredundant structures in the PDB.
We formulate this as a classification task where… See the full description on the dataset page: https://huggingface.co/datasets/vector-institute/atom3d-res.voice_medical_cut_small_vectorhackernews-vector-search-datasetThe Hacker News dataset contains 28.74 million postings and their vector embeddings. The embeddings were generated using SentenceTransformers model all-MiniLM-L6-v2. The dimension of each embedding vector is 384.
Created by clickhouse more info: https://clickhouse.com/docs/getting-started/example-datasets/hackernews-vector-search-dataset
route_red_yellow_vector_subtasks_pi05
Route Red-Yellow Vector Subtasks for pi0.5
This is a LeRobot v2.1 transformation of DistantSky/route at commit aced1e96f6b8bf98ffaa0754407636e82442b084.
The dataset contains 212 real-robot episodes, 135699 frames, three cameras, and 14-dimensional actions at 100 Hz.
Conditioning
Only observation.images.video_overhead is modified. The left and right videos are byte-identical to the source dataset.
The overhead image receives one fixed selected-connector pose glyph… See the full description on the dataset page: https://huggingface.co/datasets/ajaysri/route_red_yellow_vector_subtasks_pi05.rss_vectorsshamela-vectors
🇬🇧 English
The largest open-source vector database of the complete al-Maktaba al-Shamela (المكتبة الشاملة) Islamic text corpus, prepared for semantic search and Retrieval-Augmented Generation (RAG). Includes the full Quran text (6,236 verses, Hafs 'an 'Asim).
Statistics:
11,482,164 chunks from 8,589 classical Islamic books
6,236 Quran verses (one verse = one chunk, included in total)
40 categories covering the full breadth of Islamic scholarship
Period: 1st century AH to 15th… See the full description on the dataset page: https://huggingface.co/datasets/Kandil7/shamela-vectors.reazonspeech.vectorized.smalldetails_Delta-Vector__Odin-9B
Dataset Card for Evaluation run of Delta-Vector/Odin-9B
Dataset automatically created during the evaluation run of model Delta-Vector/Odin-9B.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Delta-Vector__Odin-9B.pi07_cable_three_vector_v1
Three-holder cable routing with vector goals
Real-robot demonstrations for a vector-conditioned low-level policy: place three holders and route a cable through each holder. This is the validated LeRobot v3 dataset prepared for the first Pi0.7 four-camera-goal cable policy.
Property
Value
Source recordings
140
Subtask episodes
840
Frames
335,897
Sampling rate
100 Hz
Robot
ARX bimanual
Recorded state / action
14 / 14 dimensions
Video resolution
448 × 448… See the full description on the dataset page: https://huggingface.co/datasets/ajaysri/pi07_cable_three_vector_v1.newsdataio_vectorsMobileGym-ConAct-Trajectories
MobileGym-ConAct-Trajectories
Dataset Viewer · MobileGym · MemGUI-Agent · Paper
Abstract
MobileGym-ConAct-Trajectories is a release of successful mobile GUI-agent rollouts collected in the MobileGym simulator. Each trajectory is selected from judge-verified rollouts using a deterministic per-task rule: retain the shortest structurally valid success, then break ties by source run and episode ID. The release preserves screenshots, the rendered prompt supplied to the… See the full description on the dataset page: https://huggingface.co/datasets/Ma-Vector/MobileGym-ConAct-Trajectories.googlenews_vectorsx86-instruction-test-vectors
x86-64 Instruction Test Vectors
Ground truth behavior of individual x86-64 instructions, captured by executing every encoding on real hardware and recording the resulting register and flag state. This is measured silicon behavior, not a model and not an emulator, so it also reflects implementation specific results such as the values an instruction leaves in flags that the architecture documents as undefined.
How it was generated
Each test case is produced by the… See the full description on the dataset page: https://huggingface.co/datasets/BinPrey/x86-instruction-test-vectors.HumaniBench
HumaniBench: A Human-Centric Benchmark for Large Multimodal Models Evaluation
**HumaniBench** is a benchmark for evaluating large multimodal models (LMMs) using real-world, human-centric criteria. It consists of 32,000+ image–question pairs across 7 tasks:
✅ Open/closed VQA
🌍 Multilingual QA
📌 Visual grounding
💬 Empathetic captioning
🧠 Robustness, reasoning, and ethics
Each example is annotated with GPT-4o drafts, then verified by experts to ensure quality and… See the full description on the dataset page: https://huggingface.co/datasets/vector-institute/HumaniBench.hotpotqaVectorEdits
VectorEdits: A Dataset and Benchmark for Instruction-Based Editing of Vector Graphics
NOTE: Currently only test set has generated labels, other sets will have them soon
Find the details in our paper: VectorEdits: A Dataset and Benchmark for Instruction-Based Editing of Vector Graphics
Github repository: JosefKuchar/vector-edits
We introduce a large-scale dataset for instruction-guided vector image editing, consisting of over 270,000 pairs of SVG images paired with natural language… See the full description on the dataset page: https://huggingface.co/datasets/mikronai/VectorEdits.atom3d-ppi
PPI: Protein-Protein Interfaces
Overview
This task relates to predicting which pairs of amino acids, spanning two
different proteins, will interact upon binding (when they form a complex).
Amino acids are defined as interacting if any of their heavy atoms are within 6
Angstroms from one another.
Datasets
splits:
DIPS-split: DIPS dataset, split by sequence identity (see add. inf.)
Format
Each entry in the dataset contains the following keys:… See the full description on the dataset page: https://huggingface.co/datasets/vector-institute/atom3d-ppi.cobe-firas-c-vector
COBE/FIRAS C-Vector
LAMBDA says the FIRAS C-Vector values “represent the variance (random noise)
derived from multiple observations within the same pixel.” It serves one
length-182 C_VECTOR for each of six spectral modes. The source FITS unit is
MJy/sr. The distinct Supplement definition and resulting conflict are
recorded below rather than resolved here.
configuration
NUM_FREQ
zero suffix
coverage
source resolution
FIRAS_C_VECTOR_HIF2
55
127
600–1350 GHz
24.6 GHz… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/cobe-firas-c-vector.nvidia-math-vectorizedtrash_pickup_v1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 25,
"total_frames": 7500,
"total_tasks": 1,
"total_videos": 50,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:25"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/vectorcrumb/trash_pickup_v1.
