wordart
Datasets
All datasets matching “wordart”WordArt_rejected
WordArt — WordArt_rejected
Rejection-sampled from the WordArt train split. This split holds the rejected items — the answer field holds the official ground truth.
rows
2,958
QA pairs
2,958
shards
83
accepted / rejected (whole family)
1,846 / 2,958
accept rate
38.4%
verifier
exact
The rejected split is training data, not just diagnostics: answer is the official ground truth, and wrong_vlm records what the model said instead.
How the data… See the full description on the dataset page: https://huggingface.co/datasets/elliot-mllm/WordArt_rejected.WordArt_RS_nothink
WordArt — WordArt_RS_nothink
Rejection-sampled from the WordArt train split. This split holds the accepted items, answer only.
rows
1,846
QA pairs
1,846
shards
34
accepted / rejected (whole family)
1,846 / 2,958
accept rate
38.4%
verifier
exact
How the data was produced
A VLM answers every question at temperature 0 with reasoning enabled. Its answer is compared with
the official ground truth by the verifier described below; matches go… See the full description on the dataset page: https://huggingface.co/datasets/elliot-mllm/WordArt_RS_nothink.WordArt_RS_think
WordArt — WordArt_RS_think
Rejection-sampled from the WordArt train split. This split holds the accepted items, with the model's reasoning trace.
rows
1,846
QA pairs
1,846
shards
35
accepted / rejected (whole family)
1,846 / 2,958
accept rate
38.4%
verifier
exact
How the data was produced
A VLM answers every question at temperature 0 with reasoning enabled. Its answer is compared with
the official ground truth by the verifier described… See the full description on the dataset page: https://huggingface.co/datasets/elliot-mllm/WordArt_RS_think.wordart_cleaned
wordart_cleaned
The wordart family of the ElliotVL supervised-fine-tuning pool, after VLM cleaning.
images
4,765
QA turns
12,859
answers rewritten by the cleaning pass
41
QA created by the cleaning pass (new_qa)
12,738 (99.1%)
shards
2
How this was cleaned
A vision-language model read each image together with its QA and judged the item. The pass is
not a filter that only removes rows — it rewrites answers it finds wrong but salvageable… See the full description on the dataset page: https://huggingface.co/datasets/Elliot-Data/wordart_cleaned.Dans-ASCIIMaxx-WordartWordArtMulti_RS_nothink
