UNIS
Datasets
All datasets matching “UNIS”UniST
UniST
This dataset contains UniST codec-token training data exported from local metadata and codec results.
We train UniSS with UniST data.
Schema
id: sample identifier
transcription: source transcription from metadata text
translation: qwen_trans, falling back to trans_text
source_glm, target_glm: GLM token lists
source_bicodec, target_bicodec: bicodec semantic token lists
bicodec_global: source bicodec global token list
dataset_name, src_lang, tgt_lang, split:… See the full description on the dataset page: https://huggingface.co/datasets/cmots/UniST.SyntheticSpectrogram
Synthetic Spectrogram Data Set
Thomas Lampert
The Department of Computer Science,
University of York,
Deramore Lane,
York, U.K.,
YO10 3GH.
E-mail: lampert@unistra.fr
Website: [https://sites.google.com/site/tomalampert](https://sites.google.com/site/tomalampert)
Disclaimer
This data is publically available for non-commercial use. All the data contained here has been generated using scripts kindly supplied by Jim Nicholson, Atlas Elektronik U.K. at Winfrith.
We kindly ask you… See the full description on the dataset page: https://huggingface.co/datasets/SDC-Unistra/SyntheticSpectrogram.uniser-haze-dataset
UniSER Synthetic Haze Dataset 🌫️
Synthetic haze dataset released with our CVPR 2026 paper, UniSER: A Foundation Model for Unified Soft Effects Removal. The dataset bundles ~80k unique clean images with ~2 million physically-motivated haze / fog / smoke renderings — covering homogeneous, non-homogeneous, indoor, outdoor, daytime, and dense atmospheric conditions — for training and benchmarking single-image dehazing.
📦 Size: ~2.5 TB across 1,327 WebDataset shards
📄 Paper: arXiv… See the full description on the dataset page: https://huggingface.co/datasets/jdzhang0929/uniser-haze-dataset.UniSAR-7M
UniSAR-7M
A large-scale, multi-source synthetic aperture radar image corpus for self-supervised representation learning.
UniSAR-7M contains 7,047,666 single-channel SAR image samples assembled from public SAR datasets and openly available imagery from commercial satellite constellations. It provides the pretraining corpus for DINOSAR, a self-supervised learning framework that uses Content-Aware Multi-Crop (CAMC) to construct informative views of SAR imagery.
Associated… See the full description on the dataset page: https://huggingface.co/datasets/YTang/UniSAR-7M.paprika_relativeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/asdl-unist/paprika_relative.new_aicp_carrot_relativeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/asdl-unist/new_aicp_carrot_relative.
