SignerX/SignVerse-2M
SignVerse-2M SignVerse-2M: A Two-Million-Clip Pose-Native Universe of 55+ Sign Languages Links: [Paper] | [Data Files] | [Project Page] SignVerse-2M is a large-scale multilingual pose-native dataset for sign language research. The dataset reorganizes publicly available sign language videos into a unified DWPose-based representation and releases the result as approximately 2 million clips from 39,196 videos covering 55+ sign languages. Rather than… See the full description on the dataset page: https://huggingface.co/datasets/SignerX/SignVerse-2M.
101.9k
1---2license: cc-by-nc-4.03language:4 - aed5 - ase6 - asf7 - asq8 - bfi9 - bzs10 - csc11 - cse12 - csg13 - csn14 - csq15 - dse16 - dsl17 - eso18 - fcs19 - fse20 - fsl21 - fss22 - gsg23 - gss24 - hks25 - hsh26 - icl27 - ils28 - inl29 - ins30 - ise31 - isg32 - isr33 - jos34 - jsl35 - kvk36 - lls37 - mfs38 - nsl39 - nzs40 - pks41 - prl42 - pso43 - psp44 - rsl45 - sfb46 - sgg47 - slf48 - ysl49 - sls50 - ssp51 - ssr52 - svk53 - swl54 - tsm55 - tsq56 - tss57 - vgt58 - hos59task_categories:60 - other61tags:62 - sign-language63 - pose-estimation64 - dwpose65 - multilingual66 - keypoint67 - video-understanding68 - sign-language-generation69 - sign-language-recognition70 - pose-native71size_categories:72 - 1M<n<10M73pretty_name: "SignVerse-2M: A Two-Million-Clip Pose-Native Universe of 55+ Sign Languages"74---75 76<div align="center" style="position: relative; width: 100%; max-width: 1400px; margin: 0 auto 24px auto;">77 <img78 src="https://signerx.github.io/SignVerse-2M/static/images/background_gallery_big.gif"79 alt="SignVerse-2M cover"80 style="width: 100%; display: block; border-radius: 18px;"81 />82 <div83 style="84 position: absolute;85 inset: 0;86 display: flex;87 align-items: center;88 justify-content: center;89 background: linear-gradient(to bottom, rgba(0,0,0,0.18), rgba(0,0,0,0.28));90 border-radius: 18px;91 "92 >93 <div94 style="95 color: white;96 font-size: 3.4rem;97 font-weight: 800;98 letter-spacing: 0.04em;99 text-shadow: 0 4px 20px rgba(0,0,0,0.45);100 "101 >102 SignVerse-2M103 </div>104 </div>105</div>106 107## SignVerse-2M: A Two-Million-Clip Pose-Native Universe of 55+ Sign Languages108 109Links: [[Paper]](https://arxiv.org/abs/2605.01720) | [[Data Files]](https://huggingface.co/datasets/SignerX/SignVerse-2M/tree/main/dataset) | [[Project Page]](https://signerx.github.io/SignVerse-2M)110 111**SignVerse-2M** is a large-scale multilingual pose-native dataset for sign language research. The dataset reorganizes publicly available sign language videos into a unified DWPose-based representation and releases the result as approximately **2 million clips** from **39,196 videos** covering **55+ sign languages**. Rather than distributing raw RGB video, SignVerse-2M provides per-frame body, hand, and face keypoints together with structured subtitle supervision, making the corpus directly usable for pose-conditioned sign language generation, recognition, and translation research.112 113## Overview114 115Existing large-scale sign language resources are typically organized as video-text corpora. That format is appropriate for RGB-based recognition or translation, but it is not the most natural interface for modern pose-driven generation pipelines, which increasingly operate on standardized human keypoint controls such as DWPose. SignVerse-2M addresses this mismatch by converting multilingual public sign language videos into a common pose space.116 117The release is intended to support research questions such as:118 119- multilingual sign language generation in pose space120- pose-based sign language recognition and translation121- cross-lingual transfer across heterogeneous sign language sources122- benchmarking of sign language motion representations under open-world conditions123 124## Key Characteristics125 126| Property | Value |127| --- | --- |128| Dataset name | SignVerse-2M |129| Core representation | DWPose keypoint sequences |130| Videos | 39,196 |131| Clips / subtitle segments | Approximately 2 million |132| Sign languages | 55+ |133| Frame rate | 24 FPS |134| Per-frame keypoints | 18 body + 21 left hand + 21 right hand + 68 face = 128 |135| Source type | Public multilingual sign language videos |136| Raw RGB frames released | No |137| Released supervision | Structured subtitle text and document-level text |138 139## Why A Pose-Native Release140 141SignVerse-2M should not be understood as merely a larger multilingual video-text corpus. Its main contribution is the release of a unified pose-native interface for sign language research.142 143Compared with raw-video releases, the pose-native representation offers three practical advantages:144 1451. It reduces nuisance variation from background, clothing, and appearance, allowing models to focus more directly on motion.1462. It aligns naturally with contemporary pose-conditioned generation pipelines that already consume DWPose-like controls.1473. It provides a common representation for multilingual benchmarking, making comparisons across methods more interpretable.148 149## Data Source And Processing150 151The corpus is built from publicly available multilingual sign language videos ([YouTube-SL-25](https://arxiv.org/abs/2407.11144)), including resources inherited from large public sign language collections such as YouTube-SL-55 and related open web sources. Each video is processed through a unified pipeline that:152 1531. retrieves metadata and available subtitles,1542. structures subtitle tracks into segment-level and document-level text,1553. decodes the video at 24 FPS,1564. applies DWPose to extract body, hand, and face keypoints frame by frame,1575. packages the outputs into per-video artifacts for public release.158 159No manual keypoint annotation is provided. The keypoints and subtitles are produced automatically through the preprocessing pipeline.160 161## Languages162 163The corpus covers more than 55 sign languages. The current release spans 55 language identifiers:164 165| Code | Language / identifier | Code | Language / identifier |166| --- | --- | --- | --- |167| `aed` | Argentine Sign Language | `ase` | American Sign Language |168| `asf` | Australian Sign Language | `asq` | source identifier `asq` |169| `bfi` | British Sign Language | `bzs` | Brazilian Sign Language |170| `csc` | Catalan Sign Language | `cse` | Czech Sign Language |171| `csg` | Chilean Sign Language | `csn` | Colombian Sign Language |172| `csq` | Croatian Sign Language | `dse` | Dutch Sign Language |173| `dsl` | Danish Sign Language | `eso` | Estonian Sign Language |174| `fcs` | Quebec Sign Language | `fse` | Finnish Sign Language |175| `fsl` | French Sign Language | `fss` | Finland-Swedish Sign Language |176| `gsg` | German Sign Language | `gss` | Greek Sign Language |177| `hks` | Hong Kong Sign Language | `hsh` | Hungarian Sign Language |178| `hos` | Ho Chi Minh City Sign Language | `icl` | Icelandic Sign Language |179| `ils` | International Sign | `inl` | Indonesian Sign Language |180| `ins` | Indian Sign Language | `ise` | Italian Sign Language |181| `isg` | Irish Sign Language | `isr` | Israeli Sign Language |182| `jos` | Jordanian Sign Language | `jsl` | Japanese Sign Language |183| `kvk` | Korean Sign Language | `lls` | Lithuanian Sign Language |184| `mfs` | Mexican Sign Language | `nsl` | Norwegian Sign Language |185| `nzs` | New Zealand Sign Language | `pks` | Pakistan Sign Language |186| `prl` | Peruvian Sign Language | `pso` | Polish Sign Language |187| `psp` | Philippine Sign Language | `rsl` | Russian Sign Language |188| `sfb` | Belgian French Sign Language | `sgg` | Swiss-German Sign Language |189| `slf` | Swiss-Italian Sign Language | `sls` | Singapore Sign Language |190| `ssp` | Spanish Sign Language | `ssr` | Swiss-French Sign Language |191| `svk` | Slovak Sign Language | `swl` | Swedish Sign Language |192| `tsm` | Turkish Sign Language | `tsq` | Thai Sign Language |193| `tss` | Taiwan Sign Language | `vgt` | Flemish Sign Language |194| `ysl` | Slovenian Sign Language | | |195 196The language distribution is long-tailed rather than balanced. High-resource languages account for a disproportionate share of the total data volume.197 198## Repository Structure199 200The public release is organized around `.tar` shards stored under `dataset/`. Each shard contains per-video directories:201 202```text203dataset/204 Sign_DWPose_NPZ_000001.tar205 Sign_DWPose_NPZ_000002.tar206 ...207```208 209Within each shard:210 211```text212{video_id}/213 poses.npz214 caption.json215 {video_id}.complete216```217 218The main files are:219 220- `poses.npz`: per-video DWPose payload with frame-wise keypoints221- `caption.json`: structured subtitle and supervision metadata222- `.complete`: completion marker produced by the processing pipeline223 224## Data Schema225 226### `poses.npz`227 228Each `poses.npz` file stores a person-centric per-frame representation. A simplified schema is shown below:229 230```python231{232 "video_id": str,233 "fps": float,234 "num_frames": int,235 "frame_ids": int[T],236 "width": int,237 "height": int,238 "frames": [239 {240 "num_people": int,241 "frame_id": int,242 "width": int,243 "height": int,244 "person_0": {245 "body": float[18, 3],246 "face": float[68, 3],247 "left_hand": float[21, 3],248 "right_hand": float[21, 3],249 },250 # optional additional people:251 # "person_1": { ... }252 },253 ...254 ]255}256```257 258Keypoint coordinates are stored in pixel space as `(x, y, score)`, where confidence scores lie in `[0, 1]`.259 260### `caption.json`261 262```json263{264 "video_id": "...",265 "sign_language": "ase",266 "title": "...",267 "duration_s": 312.4,268 "segments": [269 { "start": 0.0, "end": 4.2, "text": "..." }270 ],271 "document_text": "...",272 "english_source": "native"273}274```275 276The field `english_source` records whether the English supervision is native or automatically selected from an available translated subtitle track.277 278## Loading Example279 280```python281import json282import tarfile283import numpy as np284 285with tarfile.open("dataset/Sign_DWPose_NPZ_000001.tar") as tar:286 tar.extractall("./tmp_signverse")287 288npz = np.load("./tmp_signverse/{video_id}/poses.npz", allow_pickle=True)289frames = npz["frames"].tolist()290body = frames[0]["person_0"]["body"]291 292with open("./tmp_signverse/{video_id}/caption.json", "r", encoding="utf-8") as f:293 caption = json.load(f)294 295print(body.shape)296print(caption["segments"][0]["text"])297```298 299## Visualization And Reproduction300 301The repository includes scripts for inspecting the released pose files and for reproducing the processing pipeline.302 303### Visualize one pose file304 305```bash306python scripts/visualize_dwpose_npz.py \307 --npz extracted/{video_id}/poses.npz \308 --style openpose \309 --out viz/310```311 312### Reproduce the pipeline313 314```bash315# Single machine316bash reproduce_independently.sh317 318# SLURM cluster319bash reproduce_independently_slurm.sh320```321 322The pipeline is organized into acquisition, subtitle structuring, pose extraction, and upload/publication stages.323 324## Benchmark Setting325 326The accompanying paper introduces a multilingual **text-to-pose** benchmark for sign language generation. A generated DWPose sequence is evaluated through back-translation into spoken text, and standard text metrics such as BLEU and ROUGE are reported against the source input. The benchmark repository also provides a **SignDW Transformer** baseline in both small and large model configurations.327 328For model code and experimental setup, refer to the benchmark repository:329 330- [https://github.com/SignerX/SignVerse-2M](https://github.com/SignerX/SignVerse-2M)331 332## Intended Use333 334The release is intended for research use, including:335 336- sign language generation from text via pose space337- pose-based sign language translation and recognition338- cross-lingual transfer, adaptation, and benchmarking339- comparison of pose-native motion representations under open-world distributions340 341The release is not intended for:342 343- safety-critical interpretation in medical, legal, or emergency settings344- re-identification of individual signers345- claims of full linguistic coverage for any specific sign language346 347## Data Governance348 349### Exclusion / Takedown Mechanism350SignVerse-2M indexes all samples strictly by `video_id` and does not re-host 351or redistribute original video content. If a source video is deleted or made 352private by its creator, no copy exists in our release to propagate — access 353to any derived content tied to that `video_id` is automatically severed.354 355Creators may additionally request explicit removal of their corresponding 356pose data by contacting [issue link](https://huggingface.co/datasets/SignerX/SignVerse-2M/discussions). Verified 357removal requests will be processed within 7 business days, and an exclusion 358log will be maintained and version-tracked alongside future dataset releases.359 360### Face-Landmark-Excluded Variant361For downstream applications that do not require facial keypoints, a variant 362of the dataset with facial landmarks excluded is available at [link](https://huggingface.co/datasets/SignerX/SignVerse-2M/tree/main/dataset), 363reducing the identity-bearing channels retained in the released data.364 365 366<!--367## Limitations368 369Users should account for the following limitations when interpreting results:370 371- Keypoints are extracted automatically and may be noisy under fast motion, occlusion, or multi-person scenes.372- Fine-grained handshape distinctions are only partially captured by the released 21-keypoint hand representation.373- Non-manual linguistic signals such as facial expression and mouthing are only partially represented by 68 face landmarks.374- Subtitle timing and translations are automatically processed and may contain alignment or semantic errors.375- The corpus is language-imbalanced and inherits the long-tail distribution of public web video sources.376- `person_0` is treated as the primary signer, which may be imperfect in multi-signer videos.377-->378 379## Responsible Use380 381SignVerse-2M is derived from publicly posted sign language videos. This repository does **not** redistribute raw RGB videos; it releases pose keypoints and structured subtitle text only. Even so, pose sequences may still carry information that can contribute to signer identification when combined with external metadata. Users should treat the corpus as human-subject-derived data and use it responsibly.382 383Re-identification of individual signers is explicitly listed as an unintended use of this dataset (see Intended Use above). Users are expected to comply with applicable data protection regulations in their jurisdiction when using this resource.384 385The data distribution is also shaped by what is publicly available online. Educational or interpreter-style content may be overrepresented, while conversational, regional, or community-specific signing practices may be underrepresented.386 387## Citation388 389If you use SignVerse-2M in academic work, please cite:390 391```bibtex392@misc{fang2026signverse2mtwomillionclipposenativeuniverse,393 title={SignVerse-2M: A Two-Million-Clip Pose-Native Universe of 55+ Sign Languages}, 394 author={Sen Fang and Hongbin Zhong and Yanxin Zhang and Dimitris N. Metaxas},395 year={2026},396 eprint={2605.01720},397 archivePrefix={arXiv},398 primaryClass={cs.CV},399 url={https://arxiv.org/abs/2605.01720}, 400}401```402 403## License404 405The released dataset annotations, pose keypoints, and accompanying metadata are distributed under **CC BY-NC 4.0**.406 407Source videos are **not** redistributed in this repository and remain subject to the original platform terms and the rights of their respective creators.408 409CC BY-NC 4.0 applies exclusively to the derived DWPose keypoint annotations and structured subtitle metadata produced by our processing pipeline. Original video content remains governed by its source platform's Terms of Service and the rights of the original creator; no rights over the underlying video are claimed or conveyed by this release.