cua-lite/Lite.ScaleCUA.wip
cua-lite/Lite.ScaleCUA.wip TEMP resume coordination (raw; delete after merge) Origin Load via datasets from datasets import load_dataset # entire dataset ds = load_dataset("cua-lite/Lite.ScaleCUA.wip") # just one named subset (config) ds = load_dataset("cua-lite/Lite.ScaleCUA.wip", "desktop.use.rl") You can also filter by metadata.platform / metadata.task_type / metadata.others.* after loading; every row carries a rich metadata struct (see schema… See the full description on the dataset page: https://huggingface.co/datasets/cua-lite/Lite.ScaleCUA.wip.
cua-lite/Lite.ScaleCUA.wip
TEMP resume coordination (raw; delete after merge)
Origin
Load via datasets
from datasets import load_dataset
# entire dataset
ds = load_dataset("cua-lite/Lite.ScaleCUA.wip")
# just one named subset (config)
ds = load_dataset("cua-lite/Lite.ScaleCUA.wip", "desktop.use.rl")You can also filter by metadata.platform / metadata.task_type / metadata.others.* after loading; every row carries a rich metadata struct (see schema below).
Schema
Each row has these columns:
Coordinate values in messages are normalized to [0, 1000] integers.
*Image-dedup (`grounding. / understanding cohorts).** These cohorts are single-image-per-row and many rows share the same screenshot, so to avoid re-embedding identical image bytes once per instruction they are stored *folded*: one row per unique screenshot (image embedded once), carrying an extra **_folded** column — a JSON string with the authoritative list of {messages, metadata} members for that screenshot. The row's top-level messages is the members concatenated for viewer convenience. use cohorts are not folded. **Use lite.data.hf.download` to consume this repo** — it unfolds automatically back to one row per instruction; reading the parquet directly yields the folded form.
Layout
<platform>/<task_type>/<split>/shard-NNNNN-of-NNNNN.parquet # single-variant cohort
<platform>/<task_type>/<split>/<variant>/shard-NNNNN-of-NNNNN.parquet # multi-variant cohortplatform∈ {desktop, mobile, web}task_type∈ {understanding, grounding.action, grounding.point, grounding.bbox, use} — used verbatim as the dir component- HF config names are
<platform>.<task_type>by default (e.g.mobile.grounding.action) — UNLESS the dataset was staged with--config-names, which sets verbatim, explicitly-chosen config names (see theconfigs:block above for the authoritative list). The agent registry lookup key in code is<agent>@<platform>@<task_type>(e.g.qwen3_vl@mobile@grounding.action); only this user-facing token uses.between platform and task_type, because@triggers a 403 on the dataset-viewer's signed image URLs. - HF split names stay
train/validation(thedatasetslibrary blacklists<>:/\|?*in split names; everything else is fine in config_name) validationis an in-distribution held-out slice (never used in training);testis reserved for out-of-distribution benchmark datasets
Stats
Local mirror & SFT export
For local workflows (SFT export, dedup, mixing across datasets), use lite.data.hf.download to mirror this repo back to the canonical local layout:
$CUA_LITE_DATASETS_ROOT/cua-lite/Lite.ScaleCUA.wip/
images/<hash[:2]>/<hash>.<ext> # content-addressed image store
<platform>/<task_type>/<split>[/<variant>].parquet # rows reference images by relative pathRows in the local parquet have images: list[str]; bytes are extracted to the image store. lite.train.export.export_sft consumes the local form directly with --image-root=$CUA_LITE_DATASETS_ROOT.
- Total unique images: 64,449
- Image store size: 29.60 GB
Notes
Staged via lite.data.hf.stage from rollout log-roots: .data/rollout/lite.scalecua/gpt/5c3429d9/rl, .data/rollout/lite.scalecua/gpt/5c3429d9/train (row filter: none).
License & citation
other
