CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hbXNov /hle_math_exact_match_no_image_int_answer_random128imagen<1K0 likes5.9k downloads2y agoHugging Face02Stage-jh-monitor /total-300-random-jh-epoch4 total-300-random-jh-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.3890625 Action score: 0.440625 Valid samples: 320/320 tabularn<1K0 likes3.8k downloads14d agoHugging Face03Zaid /mmlu-random-Atext10K<n<100K0 likes2.6k downloads2y agoHugging Face04weikaih /ai2thor-random-views-20kimage10K<n<100K0 likes2.5k downloads1y agoHugging Face05Zaid /mmlu-random-2text10K<n<100K0 likes2k downloads2y agoHugging Face06random123123 /BrushDatatext10K<n<100K14 likes1.7k downloads2y agoHugging Face07Zaid /mmlu-random-1text10K<n<100K0 likes1.7k downloads2y agoHugging Face08cambridgeltl /vsr_random VSR: Visual Spatial Reasoning This is the random set of VSR: Visual Spatial Reasoning (TACL 2023) [paper]. Usage from datasets import load_dataset data_files = {"train": "train.jsonl", "dev": "dev.jsonl", "test": "test.jsonl"} dataset = load_dataset("cambridgeltl/vsr_random", data_files=data_files) Note that the image files still need to be downloaded separately. See data/ for details. Go to our github repo for more introductions. Citation If you find VSR… See the full description on the dataset page: https://huggingface.co/datasets/cambridgeltl/vsr_random.imagetext-classification10K<n<100K4 likes1.5k downloads4y agoHugging Face09weikaih /TaskMeAnything-v1-imageqa-random Dataset Card for TaskMeAnything-v1-imageqa-random TaskMeAnything-v1-imageqa-random dataset 🌐 Website | 📑 Paper | 🤗 Huggingface | 💻 Interface If you like our project, please give us a star ⭐ on GitHub for latest update. TaskMeAnything-v1-Random TaskMeAnything-v1-imageqa-random is a dataset which using randomly sampled questions from TaskMeAnything-v1, including 5,700 ImageQA questions. The dataset contains 19 splits, while each splits contains 300… See the full description on the dataset page: https://huggingface.co/datasets/weikaih/TaskMeAnything-v1-imageqa-random.image1K<n<10K1 likes1.3k downloads2y agoHugging Face10weikaih /TaskMeAnything-v1-videoqa-random Dataset Card for TaskMeAnything-v1-videoqa-random TaskMeAnything-v1-videoqa-random dataset 🌐 Website | 📑 Paper | 🤗 Huggingface | 💻 Interface If you like our project, please give us a star ⭐ on GitHub for latest update. TaskMeAnything-v1-Random TaskMeAnything-v1-videoqa-random is a dataset which randomly sampled questions from TaskMeAnything-v1, including 2,700 VideoQA questions. The dataset contains 9 splits, while each splits contains 300 questions… See the full description on the dataset page: https://huggingface.co/datasets/weikaih/TaskMeAnything-v1-videoqa-random.text1K<n<10K1 likes939 downloads2y agoHugging Face11Zaid /mmlu-random-Dtext10K<n<100K0 likes772 downloads2y agoHugging Face12priyank-m /trdg_random_en_zh_text_recognition Dataset Card for "trdg_random_en_zh_text_recognition" This synthetic dataset was generated using the TextRecognitionDataGenerator(TRDG) open source repo: https://github.com/Belval/TextRecognitionDataGenerator It contains images of text with random characters from Engilsh(en) and Chinese(zh) languages. Reference to the documentation provided by the TRDG repo: https://textrecognitiondatagenerator.readthedocs.io/en/latest/index.html imageimage-to-text100K<n<1M3 likes751 downloads2y agoHugging Face13marin-dna /vertebrate-v1-issue473-fullwindow-cds-random-val marin-dna/vertebrate-v1-issue473-fullwindow-cds-random-val CDS full-window vertebrate projection sequences for the issue #473 random validation control. The source is the immutable issue #417 accepted-sequence table. The split uniformly samples 16,384 original-orientation CDS rows without replacement using seed 42. Sampling occurs before reverse-complement augmentation. Selected rows are removed from training; reverse complements are then added only to the remaining training… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/vertebrate-v1-issue473-fullwindow-cds-random-val.tabular10M<n<100M0 likes712 downloads1mo agoHugging Face14randomhuggingfaceuser1273823147 /steve1-training-data STEVE-1 Training Data With MineCLIP Embeddings This dataset contains the MineCLIP-embedded training data used for the MultiSTEVE-1s model zoo. It supports reproducing STEVE-1-style fine-tuning without regenerating MineCLIP embeddings. Contents Top-level directories: dataset_contractor/: OpenAI Contractor Dataset episodes converted for STEVE-1 training. dataset_mixed_agents/: VPT-generated Minecraft trajectories collected for STEVE-1-style training. Each episode… See the full description on the dataset page: https://huggingface.co/datasets/randomhuggingfaceuser1273823147/steve1-training-data.text0 likes698 downloads5mo agoHugging Face15suz22 /RoboTwin_adjust_bottle_randomizedThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "aloha", "total_episodes": 500, "total_frames": 68537, "total_tasks": 424, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 15, "splits": { "train": "0:500" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/suz22/RoboTwin_adjust_bottle_randomized.imagerobotics10K<n<100K0 likes662 downloads5mo agoHugging Face16weikaih /ego4d-random-views-20k Ego4D Random Views Dataset This dataset contains 20,000 random view frames sampled from the Ego4D dataset using a high-performance multi-process generation system. Dataset Overview Total Images: 20,000 high-quality frames Image Format: PNG (1024×1024 resolution) Source: Ego4D v2 dataset (52,665+ video files) Sampling Method: Multi-process random sampling with maximum diversity Generation Time: 797.57 seconds (~13 minutes) Generation Speed: 25.08 frames/second… See the full description on the dataset page: https://huggingface.co/datasets/weikaih/ego4d-random-views-20k.imageimage-classification10K<n<100K0 likes468 downloads1y agoHugging Face17weikaih /ai2thor-random-views-20k-3obj-filteredimage1K<n<10K0 likes467 downloads1y agoHugging Face18DCAgent2 /eval-laion_explore-tis-temp10-60-8B_DCAgent2_swebench-verified-random-100-folderstext1K<n<10K0 likes462 downloads3mo agoHugging Face19NCube /europa-random-split Dataset Card for EUROPA This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description EUROPA is a dataset designed for training and evaluating multilingual keyphrase generation models in the legal domain. It consists of legal judgments from the Court of Justice of the European Union (EU) and includes instances in all 24 official EU languages. Key Features: Multilingual: Covers… See the full description on the dataset page: https://huggingface.co/datasets/NCube/europa-random-split.text100K<n<1M0 likes457 downloads2y agoHugging Face20svjack /diffusiondb_2m_random_50k Dataset Card for "diffusiondb_2m_random_50k" More Information needed image10K<n<100K0 likes455 downloads4y agoHugging Face21windfromthenorth /extreme_randomization_6_brick_03This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "ur5_wsg50_lego_atomic_step", "total_episodes": 2108, "total_frames": 419545, "total_tasks": 1, "total_videos": 4216, "total_chunks": 3, "chunks_size": 1000, "fps": 20, "splits": { "train": "0:2108" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/windfromthenorth/extreme_randomization_6_brick_03.tabularrobotics100K<n<1M0 likes429 downloads3mo agoHugging Face22AmanPriyanshu /random-small-github-repositories random-small-github-repositories A collection of 5,613 small-to-medium open-source GitHub repositories, packaged as zipped archives alongside a metadata CSV. Intended as a seed dataset for code retrieval, context engineering, and SWE-bench-style dataset construction tasks. Contents seed_small_repos.csv — metadata for each repo (owner, repo_name, stars, license, repo_hash) repos-zipped/ — one .zip per repo, named {repo_hash}.zip unzipper.py - unzipping python… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/random-small-github-repositories.tabulartext-generation1K<n<10K0 likes408 downloads6mo agoHugging Face23cminst /transcoda-random-notation-300k Transcoda Random Notation 300k Source: train: cminst/transcoda-random-notation-normalized-v1/uniform-v3-normalized-constant-spines-marks-300k; validation: fresh generation using the same normalized uniform-v3 generator recipe This is a canonical standardized Transcoda dataset. All published splits use the standard Hugging Face Datasets Parquet layout under data/. Canonical columns: image transcription sample_id source metadata: JSON string preserving source/provenance fields… See the full description on the dataset page: https://huggingface.co/datasets/cminst/transcoda-random-notation-300k.imageimage-to-text100K<n<1M0 likes407 downloads1mo agoHugging Face24randomshit11 /cattlesegmentationimage1K<n<10K0 likes396 downloads2y agoHugging Face25TashaSkyUp /random_midpoint_displacement_fractal1024x1024px PNG encoded, scale=0.1, roughness=1.0 each map was then processed with 100 iterations of rain erosion simulation (e99 directory) for further explanation please see https://en.wikipedia.org/wiki/Diamond-square_algorithm The idea was first introduced by Fournier, Fussell and Carpenter at SIGGRAPH in 1982 imagen<1K1 likes388 downloads3y agoHugging Face26gokaygokay /random_instruct_docciThe dataset consists of a collection of general questions and orders/instructions added to google/docci dataset. 4000 test example moved into train dataset. Citation and attribution This dataset repository is maintained by Gökay Aydoğan. If you reference this repository in academic work, please cite it as follows and also cite the upstream models, datasets, or projects it builds upon. @dataset{aydogan2024random_instruct_docci, author = {Aydoğan, Gökay}, title =… See the full description on the dataset page: https://huggingface.co/datasets/gokaygokay/random_instruct_docci.image10K<n<100K6 likes384 downloads2mo agoHugging Face27Blinorot /lensless_mic_random Dataset Card for LenslessMic Version of N(0,1) Random Dataset Dataset Summary A LenslessMic version of the N(0,1) random images dataset from the "LenslessMic: Audio Encryption and Authentication via Lensless Computational Imaging" paper. The dataset can be used to train a codec-agnostic reconstruction algorithm. Partition # Audio # Frames train 200 30000 Note: We split dataset into 200 files, however, there are no actual audio files. Only frames are used.… See the full description on the dataset page: https://huggingface.co/datasets/Blinorot/lensless_mic_random.textaudio-to-audion<1K0 likes376 downloads1y agoHugging Face28StarVLA /RoboTwin-Randomized-targztext10K<n<100K0 likes371 downloads7mo agoHugging Face29stochastic /random_streetview_images_pano_v0.0.2 Dataset Card for panoramic street view images (v.0.0.2) Dataset Summary The random streetview images dataset are labeled, panoramic images scraped from randomstreetview.com. Each image shows a location accessible by Google Streetview that has been roughly combined to provide ~360 degree view of a single location. The dataset was designed with the intent to geolocate an image purely based on its visual content. Supported Tasks and Leaderboards None as of now!… See the full description on the dataset page: https://huggingface.co/datasets/stochastic/random_streetview_images_pano_v0.0.2.imageimage-classification10K<n<100K29 likes366 downloads4y agoHugging Face30yk3701208 /random005text10M<n<100M0 likes358 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.