datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
various
malcolmrey's Various AI Model, Architecture & Research Repository
Welcome to the central research and asset repository of malcolmrey. This repository hosts cutting-edge tools, custom architectures, RefMod latent adapter systems, video synthesis engines, training configurations, benchmark suites, cinematic scripts, and comprehensive educational guides spanning MiniMax-H3, FLUX.2 / Klein 9B, WAN 2.1, LTX-Video, Z-Image, SDXL, and Stable Diffusion.
🧭 Repository Map &… See the full description on the dataset page: https://huggingface.co/datasets/malcolmrey/various.Transparent_BOP
Object Pose Estimation Using Implicit Representation for Transparent Objects
This dataset aggregates high quality 3D mesh assets and rendered data for training and fine-tuning pose estimation models. It unifies four significant datasets: ClearPose, DIMO, HouseCat6D, and TRansPose, all formatted for BOP (Benchmark for 6D Object Pose Estimation) evaluation.
Visualization
Below is a collage showing sample RGB inputs from the constituent datasets:
Component… See the full description on the dataset page: https://huggingface.co/datasets/varunburde/Transparent_BOP.C-VARCThis repository contains all the data associated with the paper "C-VARC: A Large-Scale Chinese Value Rule Corpus for Value Alignment of Large Language Models".
We propose a three-tier value classification framework based on core Chinese values, which includes three dimensions, twelve core values, and fifty derived values. With the assistance of large language models and manual verification, we constructed a large-scale, refined, and high-quality value corpus containing over 250,000 rules. We… See the full description on the dataset page: https://huggingface.co/datasets/Beijing-AISI/C-VARC.Real-IAD_Varietysuno-various-94k
Suno Various 94k: originals, ACE-Step covers, and stems
A unified, indexed repository containing 94,174 exact source/cover pairs for
reference conditioning, preference, slider, source-separation, and music
style-transfer research. Four independently loadable configurations provide the
original tracks, generated covers, and four-source stems for both sides.
Configurations
Configuration
Shards
Samples
Contents
original
85
94,174
Original MP3 plus JSON… See the full description on the dataset page: https://huggingface.co/datasets/webshart/suno-various-94k.wxy_var
VAR: a new visual generation method elevates GPT-style models beyond diffusion🚀 & Scaling laws observed📈
Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction
NeurIPS 2024 Best Paper
News
2024-12: 🏆 VAR received NeurIPS 2024 Best Paper Award.
2024-12: 🔥 We Release our Text-to-Image research based on VAR, please check Infinity.
2024-09: VAR is accepted as NeurIPS 2024 Oral Presentation.
2024-04: Visual… See the full description on the dataset page: https://huggingface.co/datasets/slz1/wxy_var.variouscryptodata
variouscryptodata
Crypto market datasets collected as a by-product of our own research and
published so they are not lost. One sub-folder per dataset; each appended
nightly where collection is still running.
folder
what
coverage
cadence
polymarket_updown_orderbook/
Polymarket Up/Down (5m/15m) order books, 10 levels, BTC/ETH/SOL/XRP/DOGE/HYPE/BNB, with Binance spot reference
2026-05-24 → present
appended nightly (previous UTC day)
hyperliquid_trades/
Hyperliquid perp… See the full description on the dataset page: https://huggingface.co/datasets/Barthel/variouscryptodata.VAREX
VAREX: A Benchmark for Multi-Modal Structured Extraction from Documents
VAREX (VARied-schema EXtraction) is a benchmark for evaluating multimodal foundation models on structured data extraction from government forms. It comprises 1,777 documents with 1,771 unique schemas across three structural categories, each provided in four input modalities. Ground truth is deterministic — generated via a Reverse Annotation pipeline that programmatically fills PDF templates with synthetic values… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/VAREX.image_variantVariant-Foundation-Embeddings
Variant Foundation Embeddings
Here we present the variant level embeddings for large-scale genetic analyis as described in 'Incorporating LLM Embeddings for Variation Across the Human Genome,' based on curated annotations using high quality functional data from FAVOR, ClinVar, and GWAS Catalog. We currently present embeddings using either OpenAI's text-embedding-3-large (3072-dimensional) or Qwen's Qwen3-Embedding-0.6B (1024-dimensional) models.
Genetic variants are identified with… See the full description on the dataset page: https://huggingface.co/datasets/LiLabUNC/Variant-Foundation-Embeddings.varroa_mmdet_yolo_protocol_runsHIT
Dataset Summary
The HIT dataset is a structured dataset of paired observations of body's inner tissues and the body surface. More concretely, it is a dataset of paired full-body volumetric segmented (bones, lean, and adipose tissue) MRI scans and SMPL meshes capturing the body surface shape for male (N=157) and female (N=241) subjects respectively. This is relevant for medicine, sports science, biomechanics, and computer graphics as it can ease the creation of personalized anatomic… See the full description on the dataset page: https://huggingface.co/datasets/varora/HIT.SQL-Queries-Datasetosworld_v2_variants
OSWorld V2 Task Variants
This dataset contains the root-level task_*v1.py Python task classes of the OSWorld V2 variant set used in the DigitalWorld AWM evaluation: 16 parameter-resampled rewrites of OSWorld V2 task classes. Each variant keeps the base task's application and workflow but changes the instruction card and regenerates its assets, so that an agent that has memorised the base task does not get the variant for free.
It follows the layout of xlangai/osworld_v2_tasks:… See the full description on the dataset page: https://huggingface.co/datasets/magicgh/osworld_v2_variants.MOTOR
MOTOR — A Multimodal Dataset for Two-Wheeler Rider Behavior Understanding
Project Page | Paper | Code
MOTOR is the first large-scale, multi-view, multimodal resource dedicated to two-wheelers in dense, unstructured traffic. It comprises 1,629 annotated sequences (25+ hours of video data) collected from 16 riders and integrates synchronized front, rear, and helmet videos, rider eye-gaze from wearable trackers, on-road audio, and telemetry (GPS, accelerometer, gyroscope).… See the full description on the dataset page: https://huggingface.co/datasets/varunpaturkar/MOTOR.IGB_XorQAp2-etf-amortised-variational-rl-resultsdolma-blend-gpt2lunara-aesthetic-image-variations
Dataset Card for Moonworks Lunara Aesthetic II
This dataset introduces the second open-source release by Moonworks. This dataset contains original image and art created by Moonworks and their contextual variations generated by Moonworks Lunara, a sub-10B parameter model with a novel diffusion mixture architecture.
Paper: https://arxiv.org/pdf/2602.01666
While part 1 is intended for learning and evaluating regional as well as region-agnostic art styles, part 2 is intended for… See the full description on the dataset page: https://huggingface.co/datasets/moonworks/lunara-aesthetic-image-variations.varroa-yolo-under-2m-wiouaya-collection-indic-sampled-mergedjfk-files-2025
JFK Document Release 2025 - Document Index
This file provides an overview of the document collection included in this repository.
Collection Overview
Total Documents: Over 2,000 files
Total Size: Approximately 6GB
Format: PDF files
Source: National Archives and Records Administration (NARA)
Release Date: 2025
Hosted On: Hugging Face Datasets
Document Naming Convention
The documents follow the National Archives naming convention:
Format:… See the full description on the dataset page: https://huggingface.co/datasets/varunh/jfk-files-2025.Vartalaap
Vartalaap — full-duplex Hindi/English conversational speech
Dual-channel synthetic Indian customer-support calls for training full-duplex
speech-to-speech models.
68,674 calls · 1,564.4 hours · 1786 shards
(last updated 2026-09-24 10:34 IST)
Audio layout
Each row's audio is a stereo FLAC at 24000 Hz:
channel
content
0 (LEFT)
agent — pristine, TTS speech and silence only
1 (RIGHT)
user — the caller
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/kapturecx/Vartalaap.perceptpick
Benchmarking the Effects of Object Pose Estimation and Reconstruction on Robotic Grasping Success
This dataset accompanies the paper accepted at IEEE ICRA 2026:
Benchmarking the Effects of Object Pose Estimation and Reconstruction on Robotic Grasping Success
Varun Burde, Pavel Burget, Torsten Sattler — Czech Technical University in Prague.
arXiv:2602.17101 · Project page · Code
Abstract
3D reconstruction serves as the foundational layer for numerous robotic… See the full description on the dataset page: https://huggingface.co/datasets/varunburde/perceptpick.IGB_XSumyt_full_image_dataset
Dataset Card for "yt_full_image_dataset"
More Information needed
flip2-multi-sequence-prompt-ablation-generated-variants-amylasevariant-effect-prediction
Updates
[2025-09-09] We have added ClinVar variant effect prediction results to the repository. The evaluation dataset was sourced from SongLab. The benchmark includes comparisons of GENERator against Evo2, NT, NT-v2, HyenaDNA, GPN-MSA, CADD, phyloP, and phastCons.
Abouts
The human reference genome data is sourced from the NCBI website.
We have applied minor formatting adjustments to the dataset to facilitate streamlined data analysis.
How to use
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/GenerTeam/variant-effect-prediction.variant-benchmark
Variant Benchmark
This benchmark is designed to evaluate how effectively models leverage variant information across diverse biological contexts.
Unlike conventional genomic benchmarks that focus primarily on region classification, our approach extends to a broader range of variant-driven molecular processes.
Existing assessments, such as BEND and the Genomic Long-Range Benchmark (GLRB),
provide valuable insights into specific tasks like noncoding pathogenicity and tissue-specific… See the full description on the dataset page: https://huggingface.co/datasets/m42-health/variant-benchmark.wikipedia-trivia-query-variation
