datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
jointavbench
JointAVBench: A Benchmark for Joint Audio-Visual Reasoning Evaluation
Overview
JointAVBench is a benchmark for evaluating omni-modal large language models on joint audio-visual reasoning tasks. Each multiple-choice question is designed to require both visual and auditory information.
This repository contains the audited release of JointAVBench under the roverx12345 namespace. The benchmark keeps the original 2,853-question split while refining answer… See the full description on the dataset page: https://huggingface.co/datasets/roverx12345/jointavbench.rovibook
RoVI-Book Dataset
🎉 **CVPR 2025** 🎉
Official dataset for Robotic Visual Instruction
This is an example to demonstrate the RoVI Book dataset, adapted from the Open-X Embodiments dataset. The bottom displays the proportion of each task type.
Paper:Robotic Visual Instruction
Project Page: https://robotic-visual-instruction.github.io/
Code: https://github.com/RoboticsVisualInstruction/RoVI-Book
Introduction
The RoVI-Book dataset is introduced alongside Robotic… See the full description on the dataset page: https://huggingface.co/datasets/yanbang/rovibook.nasa-mars-rover-images
NASA Mars Rover Image Catalog
Credit: NASA/JPL-Caltech/MSSS
Part of a dataset collection on Hugging Face.
Dataset description
The NASA Mars Rover Image Catalog contains metadata for every raw image captured by the Perseverance (Mars 2020) and Curiosity (MSL) rovers on the surface of Mars. Perseverance has been exploring Jezero Crater since February 2021, investigating an ancient river delta for signs of past microbial life and caching samples for future… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/nasa-mars-rover-images.ROVR-Open-Dataset
ROVR Open Dataset
Introduction
Welcome to the ROVR Open Dataset repository! This dataset is designed to empower autonomous driving and robotics research by providing rich, real-world data captured from ADAS cameras and LiDAR sensors. The dataset spans 50+ countries with over 20 million kilometers of driving data, making it ideal for training and developing advanced AI algorithms for depth estimation, object detection, and semantic segmentation.… See the full description on the dataset page: https://huggingface.co/datasets/ROVR-Network/ROVR-Open-Dataset.ROVERsomos-clean-alpaca-es
Dataset Card for "somos-clean-alpaca-es"
More Information needed
caption_for_mars_and_rover_image_size_1024caption_for_mars_and_rover_image_size_reducedcaption_for_mars_and_rover_image_size_512StableBeluga-7B-Qlora-Samantha-V3-Converted-DatasetSamantha-EN-CN-Converted-Dataset-V1
Dataset Card for "Samantha-EN-CN-Converted-Dataset-V1"
More Information needed
StableBeluga-7B-Qlora-Samantha-Zh-V2-Converted-Dataset
Dataset Card for "StableBeluga-7B-Qlora-Samantha-Zh-V2-Converted-Dataset"
More Information needed
Samantha-EN-CN-Dataset-V1Samantha-data-single-line-Mixed-V1-Converted-32K
Dataset Card for "Samantha-data-single-line-Mixed-V1-Converted-3W"
More Information needed
roverlocomo-benchmark-results
Logica Mind — LoCoMo Benchmark Results
Reproducible accuracy results for Logica Mind
(open-source agent memory) on LoCoMo, alongside published numbers for other
memory systems — all under the same protocol as the Mem0 paper
(arXiv:2504.19413): gpt-4o-mini as both
answerer and judge, adversarial category excluded, 1,540 scored questions.
Leaderboard (results.csv)
System
LoCoMo J
LLM at write time
Source
Letta (filesystem agent)
74.0%
agent-managed
Letta… See the full description on the dataset page: https://huggingface.co/datasets/rovemark/locomo-benchmark-results.Samantha-data-single-line-Mixed-V1-Converted
Dataset Card for "Samantha-data-single-line-Mixed-V1-Converted"
More Information needed
rovalthia-data
ROLLY-AGENTIC-CYBER-DATASET
Curator: ROVALTHIA LABORATORYOfficial Dataset Repository: JuanCDEV/rovalthia-dataCompanion Model: JuanCDEV/rovalthiaLicense: Apache 2.0Domains: Defensive Cybersecurity, Zero Trust Architecture, CTI Threat Hunting, Autonomous Agentic Orchestration, Formal Reasoning.
1. Overview & Purpose
The ROLLY-AGENTIC-CYBER-DATASET provides a vetted, high-density instruction and evaluation dataset designed to train and benchmark autonomous agentic… See the full description on the dataset page: https://huggingface.co/datasets/JuanCDEV/rovalthia-data.Samantha-data-single-line-Mixed-V1import json
# Load the provided data
with open("path_to_your_original_file.jsonl", "r", encoding="utf-8") as file:
mixed_data = [json.loads(line) for line in file.readlines()]
# Convert the mixed data by extracting all possible Q&A pairs from each conversation
reformatted_data_complete = []
for conversation in mixed_data:
text = conversation['text']
# Split the text into segments based on the prefixes
segments = [segment for segment in text.split("###") if… See the full description on the dataset page: https://huggingface.co/datasets/RoversX/Samantha-data-single-line-Mixed-V1.Mars_surface_image_Curiosity_rover_labeled_data_set_version_1
Mars surface image (Curiosity rover) labeled data set
Authors: Alice Stanboli and Kiri L. Wagstaff
Contact: kiri.l.wagstaff@jpl.nasa.gov
This data set consists of 6691 images that were collected by the Mars
Science Laboratory (MSL, Curosity) rover by three instruments (Mastcam
Right eye, Mastcam Left eye, and MAHLI). These images are the
"browse" version of each original data product, not full resolution.
They are roughly 256x256 pixels each. Full-size images can be
obtained from… See the full description on the dataset page: https://huggingface.co/datasets/RAAArity/Mars_surface_image_Curiosity_rover_labeled_data_set_version_1.caption_for_mars_and_rover_image_newcaption_for_mars_and_rover_image_size_768rover-factgithub-issuesgensyn-rover-role-requirementsro_vsr_results_dubla2ro_vsr_resultsro_voiceRoVCKhyBbFg6lXz0RoVidX_post
