boz
Datasets
All datasets matching “boz”two-box-judge-gui
Two-Box Judge GUI Dataset
A multimodal dataset for training GUI element selection models. Given two candidate bounding boxes on a GUI screenshot, the model learns to select the one that better fulfills the user's intent.
Dataset Description
This dataset is designed for training judge models in GUI grounding pipelines. When a visual grounding model produces multiple candidate regions, the judge model determines which candidate best matches the user's command.… See the full description on the dataset page: https://huggingface.co/datasets/THU-BoZhang/two-box-judge-gui.two-box-judge-gui-sharded
Two-Box Judge GUI Dataset (Sharded)
A multimodal dataset for training GUI element selection models, packaged in WebDataset format for efficient streaming.
Dataset Statistics
Split
Samples
Shards
Size
Train
115,638
6
25.32 GB
Validation
12,849
1
2.82 GB
Format
This dataset uses WebDataset format - sharded tar.gz archives for efficient streaming:
train/
├── shard-00000.tar.gz
├── shard-00001.tar.gz
└── ...
Each shard contains… See the full description on the dataset page: https://huggingface.co/datasets/THU-BoZhang/two-box-judge-gui-sharded.bo_zh_translation
Dataset source
From CUTE (Chinese, Uyghur, Tibetan, English), a large-scale multilingual dataset. Extract paragraphs from parallel corpus, match based on embedding
vector stores and semantic search, split into 52381 samples.
Usage
In alpaca format (Alpaca: A Strong, Replicable Instruction-Following Model), translation bo-zh sentence pairs, can be used in SFT.alpaca_small.json 500 samples
sentence
average_character_length
input
77.97
output
20.93… See the full description on the dataset page: https://huggingface.co/datasets/carry08/bo_zh_translation.paul-alabama-code-fullairbnb-usa-mt-bozemanaic-dataset-0.2
Art Institute of Chicago (AIC) Selected Features Dataset
Dataset Summary
This dataset contains a selection of features extracted from the Art Institute of Chicago (AIC) collection. It's important to note that this is a curated subset and does not represent the complete range of information available in the museum's original database. The dataset was created from an archived AIC database, which can be downloaded from… See the full description on the dataset page: https://huggingface.co/datasets/anna-bozhenko/aic-dataset-0.2.
