datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
trellis500k-github-archives-7trellis500k-sketchfab-archivestrellis500k-github-archives-10trellis500k-github-archives-9trellis500k-github-archives-5TRELLIS-500K
TRELLIS-500K
TRELLIS-500K is a dataset of 500K 3D assets curated from Objaverse(XL), ABO, 3D-FUTURE, HSSD, and Toys4k, filtered based on aesthetic scores.
This dataset serves for 3D generation tasks.
It was introduced in the paper Structured 3D Latents for Scalable and Versatile 3D Generation.
Dataset Statistics
The following table summarizes the dataset's filtering and composition:
NOTE: Some of the 3D assets lack text captions. Please filter out such assets if captions… See the full description on the dataset page: https://huggingface.co/datasets/JeffreyXiang/TRELLIS-500K.trellis500k-github-archives-4trellis500k-github-archivestrellis-dual-contrast-flowedit-8gputrellis500k-github-archives-8trellis500k-github-archives-6remote1_TrellisObjImages_1gb_archiveswh-trellis2-400k
SWH TRELLIS.2 dataset (400K)
TRELLIS.2-ready latents + conditioning renders for 397,797 high-quality 3D meshes,
filtered from the ~10.16M Software-Heritage v2 GitHub 3D-mesh corpus.
Pipeline
10.16M SWH meshes → 4-view render + LAION-4.5 aesthetic filter (1.07M) → Qwen 6-axis VLM judge (400,698 recommended) → TRELLIS.2 encode + 16-view cond render
Contents (per-modality webdataset tar shards)
path
files
notes
shape_latents/shape512/
399… See the full description on the dataset page: https://huggingface.co/datasets/weikaih/swh-trellis2-400k.trellis500k-github-archives-3trellis500k-github-archives-6-processedtrellismark-qwen3-4b
TrellisMark Qwen3-4B confirmation corpus
This is the frozen English confirmation corpus for
TrellisMark, an experimental
many-user AI-text watermark. It includes exact generated text and token IDs,
unwatermarked Qwen controls, public-key detector evidence, the public research
key, independent encoder vectors, and the result reports used for the
reader-facing curves. The standalone implementation, detector, and
reproduction instructions are in the
TrellisMark GitHub repository.… See the full description on the dataset page: https://huggingface.co/datasets/xlr8harder/trellismark-qwen3-4b.TRELLIS-3D
Structured 3D Latentsfor Scalable and Versatile 3D Generation
TRELLIS is a large 3D asset generation model. It takes in text or image prompts and generates high-quality 3D assets in various formats, such as Radiance Fields, 3D Gaussians, and meshes. The cornerstone of TRELLIS is a unified Structured LATent (SLAT) representation that allows decoding to different output formats and Rectified Flow Transformers tailored for SLAT as the powerful backbones. We provide large-scale pre-trained… See the full description on the dataset page: https://huggingface.co/datasets/argojuni0506/TRELLIS-3D.trellis20k251103_soup_can_trellis_50_640_480_lighting_augmentedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "ffw_bg2",
"total_episodes": 300,
"total_frames": 49488,
"total_tasks": 1,
"total_videos": 900,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:300"},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/kimyg119/251103_soup_can_trellis_50_640_480_lighting_augmented.251105_soup_can_sim_50_lighting_augmented_trellisThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "ffw_bg2",
"total_episodes": 300,
"total_frames": 115650,
"total_tasks": 1,
"total_videos": 900,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:300"},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/kimyg119/251105_soup_can_sim_50_lighting_augmented_trellis.MeshFleet_TRELLIS
MeshFleet: Filtered and Annotated 3D Vehicle Dataset for Domain Specific Generative Modeling
This is a processed version of the MeshFleet Dataset using the dataset pipeline from TRELLIS. It contains all the 3D models from the original dataset, but is already preprocessed and ready to use with the TRELLIS training pipeline. For fast loading and processing the dataset is chunked and compressed to webdataset files. All files for each object are stored in a separate file. You can either… See the full description on the dataset page: https://huggingface.co/datasets/DamianBoborzi/MeshFleet_TRELLIS.soup_can_trellis_50_640_480_lighting_augmentedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "ffw_bg2",
"total_episodes": 300,
"total_frames": 49488,
"total_tasks": 1,
"total_videos": 900,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:300"},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/kimyg119/soup_can_trellis_50_640_480_lighting_augmented.trellis500k-github-archives-2trellismark-qwen3-4b-open-corpus
TrellisMark Qwen3-4B open-corpus benchmark
This is the reproducible open-corpus companion to the main
TrellisMark Qwen3-4B confirmation corpus.
It is an explicitly derived subset of that release, not a separately generated
corpus. The standalone detector, benchmark runner, tests, and exact Viterbi
implementation are in the
TrellisMark GitHub repository.
The benchmark studies an intentionally difficult mixed setting: a corpus of
separately prompted documents, with hidden… See the full description on the dataset page: https://huggingface.co/datasets/xlr8harder/trellismark-qwen3-4b-open-corpus.trellisremote1_TrellisObjImages_1gb_archivetrellismark-qwen3-4b-rephrasing
TrellisMark Qwen3-4B blind rephrasing corpus
This separate release contains the frozen detector-blind rephrasing experiment
for TrellisMark. Its 16,384
source documents are an exact document-ID-preserving subset of the
main TrellisMark Qwen3-4B confirmation corpus.
It publishes both rewriters' outcome records, retained rewrite text and token
IDs, aligned key-only and model-assisted evidence, and the reports behind the
article's countermeasure figures.
The rewriters received only… See the full description on the dataset page: https://huggingface.co/datasets/xlr8harder/trellismark-qwen3-4b-rephrasing.Minecraft-TRELLIS-OVoxel-v1BDCube-TRELLIS-500K-raw
BDCube-TRELLIS-500K-raw
Incremental public mirror of TRELLIS-500K raw assets prepared by BDCube.
Uploaded Groups
abo
Restore
mkdir -p restored_dataset
cat archives/abo.tar.part-* | tar -xf - -C restored_dataset
See bundle_index.json for the cumulative uploaded archive list.
BDCube-TRELLIS-500K-text-clean-qwen35-v5
BDCube-TRELLIS-500K-text-clean-qwen35-v5
TRELLIS-500K cleaned text bundle prepared by BDCube.
Summary
Source root: /workspace/bowen/BDCube/dataset/TRELLIS-500K_text_clean_qwen35_v5
Archive groups: 1
Copied metadata files: 2
Archive payload: 2.13 GB
Metadata payload: 0.46 GB
Included Archive Groups
shards
Restore
mkdir -p restored_dataset
cat archives/shards.tar.part-* | tar -xf - -C restored_dataset
See bundle_index.json and SHA256SUMS for exact… See the full description on the dataset page: https://huggingface.co/datasets/Alexander1211/BDCube-TRELLIS-500K-text-clean-qwen35-v5.
