datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vision-token-compression-bench
OPTIC-Bench
Optical Text In-Context Benchmark: how reliably do LLMs consume text
delivered as rendered images versus plain text tokens?
In summary, the evaluation reported here finds that optical text compression
is effective only within a narrow and specific envelope. Delivering content
as rendered images genuinely reduces input tokens, by thirteen to
fifty-four per cent depending on the model and the language, but only when
the document is long, the rendering is dense and the… See the full description on the dataset page: https://huggingface.co/datasets/translorentz/vision-token-compression-bench.pashto-warmup-tokens
Pashto Warmup Tokens Dataset
This dataset contains a curated, deduplicated collection of high-quality, contextually accurate Pashto linguistic examples. It maps structural language tasks directly to the most critical vocabulary tokens in Pashto, providing a reliable corpus for token warmup, instruction tuning, evaluation, and post-OCR text correction workflows.
Dataset Summary
The initial release consists of 4,087 verified entries targeting high-frequency and… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-warmup-tokens.
