datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hle_math_exact_match_no_image_int_answer_random128total-300-random-jh-epoch4
total-300-random-jh-epoch4
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.3890625
Action score: 0.440625
Valid samples: 320/320
mmlu-random-Aai2thor-random-views-20kmmlu-random-2BrushDatammlu-random-1vsr_random
VSR: Visual Spatial Reasoning
This is the random set of VSR: Visual Spatial Reasoning (TACL 2023) [paper].
Usage
from datasets import load_dataset
data_files = {"train": "train.jsonl", "dev": "dev.jsonl", "test": "test.jsonl"}
dataset = load_dataset("cambridgeltl/vsr_random", data_files=data_files)
Note that the image files still need to be downloaded separately. See data/ for details.
Go to our github repo for more introductions.
Citation
If you find VSR… See the full description on the dataset page: https://huggingface.co/datasets/cambridgeltl/vsr_random.TaskMeAnything-v1-imageqa-random
Dataset Card for TaskMeAnything-v1-imageqa-random
TaskMeAnything-v1-imageqa-random dataset
🌐 Website | 📑 Paper | 🤗 Huggingface | 💻 Interface
If you like our project, please give us a star ⭐ on GitHub for latest update.
TaskMeAnything-v1-Random
TaskMeAnything-v1-imageqa-random is a dataset which using
randomly sampled questions from TaskMeAnything-v1, including 5,700 ImageQA questions. The dataset contains 19 splits, while each splits contains 300… See the full description on the dataset page: https://huggingface.co/datasets/weikaih/TaskMeAnything-v1-imageqa-random.TaskMeAnything-v1-videoqa-random
Dataset Card for TaskMeAnything-v1-videoqa-random
TaskMeAnything-v1-videoqa-random dataset
🌐 Website | 📑 Paper | 🤗 Huggingface | 💻 Interface
If you like our project, please give us a star ⭐ on GitHub for latest update.
TaskMeAnything-v1-Random
TaskMeAnything-v1-videoqa-random is a dataset which randomly sampled questions from TaskMeAnything-v1, including 2,700 VideoQA questions. The dataset contains 9 splits, while each splits contains 300 questions… See the full description on the dataset page: https://huggingface.co/datasets/weikaih/TaskMeAnything-v1-videoqa-random.mmlu-random-Dtrdg_random_en_zh_text_recognition
Dataset Card for "trdg_random_en_zh_text_recognition"
This synthetic dataset was generated using the TextRecognitionDataGenerator(TRDG) open source repo:
https://github.com/Belval/TextRecognitionDataGenerator
It contains images of text with random characters from Engilsh(en) and Chinese(zh) languages.
Reference to the documentation provided by the TRDG repo:
https://textrecognitiondatagenerator.readthedocs.io/en/latest/index.html
vertebrate-v1-issue473-fullwindow-cds-random-val
marin-dna/vertebrate-v1-issue473-fullwindow-cds-random-val
CDS full-window vertebrate projection sequences for the issue #473 random
validation control. The source is the immutable issue #417 accepted-sequence
table.
The split uniformly samples 16,384 original-orientation CDS rows
without replacement using seed 42. Sampling occurs before
reverse-complement augmentation. Selected rows are removed from training;
reverse complements are then added only to the remaining training… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/vertebrate-v1-issue473-fullwindow-cds-random-val.steve1-training-data
STEVE-1 Training Data With MineCLIP Embeddings
This dataset contains the MineCLIP-embedded training data used for the MultiSTEVE-1s model zoo. It supports reproducing STEVE-1-style fine-tuning without regenerating MineCLIP embeddings.
Contents
Top-level directories:
dataset_contractor/: OpenAI Contractor Dataset episodes converted for STEVE-1 training.
dataset_mixed_agents/: VPT-generated Minecraft trajectories collected for STEVE-1-style training.
Each episode… See the full description on the dataset page: https://huggingface.co/datasets/randomhuggingfaceuser1273823147/steve1-training-data.RoboTwin_adjust_bottle_randomizedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "aloha",
"total_episodes": 500,
"total_frames": 68537,
"total_tasks": 424,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:500"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/suz22/RoboTwin_adjust_bottle_randomized.ego4d-random-views-20k
Ego4D Random Views Dataset
This dataset contains 20,000 random view frames sampled from the Ego4D dataset using a high-performance multi-process generation system.
Dataset Overview
Total Images: 20,000 high-quality frames
Image Format: PNG (1024×1024 resolution)
Source: Ego4D v2 dataset (52,665+ video files)
Sampling Method: Multi-process random sampling with maximum diversity
Generation Time: 797.57 seconds (~13 minutes)
Generation Speed: 25.08 frames/second… See the full description on the dataset page: https://huggingface.co/datasets/weikaih/ego4d-random-views-20k.ai2thor-random-views-20k-3obj-filteredeval-laion_explore-tis-temp10-60-8B_DCAgent2_swebench-verified-random-100-folderseuropa-random-split
Dataset Card for EUROPA
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
EUROPA is a dataset designed for training and evaluating multilingual keyphrase generation models in the legal domain. It consists of legal judgments from the Court of Justice of the European Union (EU) and includes instances in all 24 official EU languages.
Key Features:
Multilingual: Covers… See the full description on the dataset page: https://huggingface.co/datasets/NCube/europa-random-split.diffusiondb_2m_random_50k
Dataset Card for "diffusiondb_2m_random_50k"
More Information needed
extreme_randomization_6_brick_03This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "ur5_wsg50_lego_atomic_step",
"total_episodes": 2108,
"total_frames": 419545,
"total_tasks": 1,
"total_videos": 4216,
"total_chunks": 3,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:2108"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/windfromthenorth/extreme_randomization_6_brick_03.random-small-github-repositories
random-small-github-repositories
A collection of 5,613 small-to-medium open-source GitHub repositories, packaged as zipped archives alongside a metadata CSV. Intended as a seed dataset for code retrieval, context engineering, and SWE-bench-style dataset construction tasks.
Contents
seed_small_repos.csv — metadata for each repo (owner, repo_name, stars, license, repo_hash)
repos-zipped/ — one .zip per repo, named {repo_hash}.zip
unzipper.py - unzipping python… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/random-small-github-repositories.transcoda-random-notation-300k
Transcoda Random Notation 300k
Source: train: cminst/transcoda-random-notation-normalized-v1/uniform-v3-normalized-constant-spines-marks-300k; validation: fresh generation using the same normalized uniform-v3 generator recipe
This is a canonical standardized Transcoda dataset. All published splits use the standard Hugging Face Datasets Parquet layout under data/.
Canonical columns:
image
transcription
sample_id
source
metadata: JSON string preserving source/provenance fields… See the full description on the dataset page: https://huggingface.co/datasets/cminst/transcoda-random-notation-300k.cattlesegmentationrandom_midpoint_displacement_fractal1024x1024px PNG encoded, scale=0.1, roughness=1.0
each map was then processed with 100 iterations of rain erosion simulation (e99 directory)
for further explanation please see https://en.wikipedia.org/wiki/Diamond-square_algorithm
The idea was first introduced by Fournier, Fussell and Carpenter at SIGGRAPH in 1982
random_instruct_docciThe dataset consists of a collection of general questions and orders/instructions added to google/docci dataset. 4000 test example moved into train dataset.
Citation and attribution
This dataset repository is maintained by Gökay Aydoğan. If you reference this repository in academic work, please cite it as follows and also cite the upstream models, datasets, or projects it builds upon.
@dataset{aydogan2024random_instruct_docci,
author = {Aydoğan, Gökay},
title =… See the full description on the dataset page: https://huggingface.co/datasets/gokaygokay/random_instruct_docci.lensless_mic_random
Dataset Card for LenslessMic Version of N(0,1) Random Dataset
Dataset Summary
A LenslessMic version of the N(0,1) random images dataset from the
"LenslessMic: Audio Encryption and Authentication via Lensless Computational Imaging" paper.
The dataset can be used to train a codec-agnostic reconstruction algorithm.
Partition
# Audio
# Frames
train
200
30000
Note: We split dataset into 200 files, however, there are no actual audio files. Only frames are used.… See the full description on the dataset page: https://huggingface.co/datasets/Blinorot/lensless_mic_random.RoboTwin-Randomized-targzrandom_streetview_images_pano_v0.0.2
Dataset Card for panoramic street view images (v.0.0.2)
Dataset Summary
The random streetview images dataset are labeled, panoramic images scraped from randomstreetview.com. Each image shows a location
accessible by Google Streetview that has been roughly combined to provide ~360 degree view of a single location. The dataset was designed with the intent to geolocate an image purely based on its visual content.
Supported Tasks and Leaderboards
None as of now!… See the full description on the dataset page: https://huggingface.co/datasets/stochastic/random_streetview_images_pano_v0.0.2.random005
