datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bo_zh_translation
Dataset source
From CUTE (Chinese, Uyghur, Tibetan, English), a large-scale multilingual dataset. Extract paragraphs from parallel corpus, match based on embedding
vector stores and semantic search, split into 52381 samples.
Usage
In alpaca format (Alpaca: A Strong, Replicable Instruction-Following Model), translation bo-zh sentence pairs, can be used in SFT.alpaca_small.json 500 samples
sentence
average_character_length
input
77.97
output
20.93… See the full description on the dataset page: https://huggingface.co/datasets/carry08/bo_zh_translation.two-box-judge-gui-sharded
Two-Box Judge GUI Dataset (Sharded)
A multimodal dataset for training GUI element selection models, packaged in WebDataset format for efficient streaming.
Dataset Statistics
Split
Samples
Shards
Size
Train
115,638
6
25.32 GB
Validation
12,849
1
2.82 GB
Format
This dataset uses WebDataset format - sharded tar.gz archives for efficient streaming:
train/
├── shard-00000.tar.gz
├── shard-00001.tar.gz
└── ...
Each shard contains… See the full description on the dataset page: https://huggingface.co/datasets/THU-BoZhang/two-box-judge-gui-sharded.paul-alabama-code-fullaic-dataset-0.2
Art Institute of Chicago (AIC) Selected Features Dataset
Dataset Summary
This dataset contains a selection of features extracted from the Art Institute of Chicago (AIC) collection. It's important to note that this is a curated subset and does not represent the complete range of information available in the museum's original database. The dataset was created from an archived AIC database, which can be downloaded from… See the full description on the dataset page: https://huggingface.co/datasets/anna-bozhenko/aic-dataset-0.2.airbnb_reviews_bozeman_montanaartworks
Combined Louvre and Art Institute of Chicago (AIC) Collection Dataset
Dataset Summary
This dataset merges artwork information from two prominent museum collections: the Musée du Louvre and The Art Institute of Chicago (AIC). It combines data from the Louvre Paper and Canvas Collection and the AIC Dataset 0.2 datasets.
Due to differences in the original datasets' schemas, a decision was made to focus on common fields and create a non-atomic full_info field… See the full description on the dataset page: https://huggingface.co/datasets/anna-bozhenko/artworks.1L-legal-textsqwen-1.5b-blind-spots
Qwen2.5-1.5B Blind Spots Dataset
Dataset Description
This dataset documents failure cases ("blind spots") of the base model
Qwen/Qwen2.5-1.5B — a 1.5-billion
parameter open-source language model released by Alibaba's Qwen team.
Each row contains a prompt, the answer a correct reasoner would give, and the
actual output the model produced — illustrating where it goes wrong.
Model Tested
Field
Value
Model
Qwen/Qwen2.5-1.5B
Parameters
1.5 B
Type… See the full description on the dataset page: https://huggingface.co/datasets/bozahbe21/qwen-1.5b-blind-spots.gh_bozza_T3R_deepseeklouvre-paper-and-canvas-collection
Louvre Museum Online Collection Dataset (Paintings & Prints and Drawings)
Dataset Summary
This dataset comprises a selection of artworks specifically from the "paintings" and "prints and drawings" departments of the Louvre Museum's online collection. It's important to note that this dataset represents an incomplete set of characteristics for each artwork, as the full details available through the museum are more extensive. The data was gathered via the museum's… See the full description on the dataset page: https://huggingface.co/datasets/anna-bozhenko/louvre-paper-and-canvas-collection.bblevelsbozp
