datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
agentic-llm-pretraining-1.7b
Agentic LLM Pretraining Dataset
A pretraining corpus for small language models (1-3B parameters) optimized for agentic tasks. The corpus emphasizes learning to comprehend language, reason, follow instructions, and use tools over memorizing factual knowledge — the assumption is that domain knowledge will be provided at runtime via RAG. The idea is that this could enable much smaller pretraining corpora by omitting the large volumes of text typically needed to memorize facts.… See the full description on the dataset page: https://huggingface.co/datasets/visionscaper/agentic-llm-pretraining-1.7b.vision-token-compression-bench
OPTIC-Bench
Optical Text In-Context Benchmark: how reliably do LLMs consume text
delivered as rendered images versus plain text tokens?
In summary, the evaluation reported here finds that optical text compression
is effective only within a narrow and specific envelope. Delivering content
as rendered images genuinely reduces input tokens, by thirteen to
fifty-four per cent depending on the model and the language, but only when
the document is long, the rendering is dense and the… See the full description on the dataset page: https://huggingface.co/datasets/translorentz/vision-token-compression-bench.unified-math-vision-dataset
Unified Math Vision Dataset Bundle
Generated at: 2025-09-19 17:12:44
This is a unified dataset bundle containing multiple math and vision reasoning datasets.
Dataset Statistics
Total samples: 15858
mathvision: 3344 samples
wemath: 500 samples
mmmu: 415 samples
mathvista: 6141 samples
logicvista: 448 samples
dynamath: 5010 samples
Contents
manifest.jsonl: Complete dataset in JSONL format (1 JSON per line)
manifest.csv: Summary in CSV format
images/: Directory… See the full description on the dataset page: https://huggingface.co/datasets/Haonian/unified-math-vision-dataset.brazilian-math-physics-qa-vision
Brazilian Math & Physics QA — Image Dependent
English | Português do Brasil
English
Summary
Brazilian Portuguese educational question-answer pairs whose problem statement or solution depends on one or more images.
Examples: 3,808
Referenced image URLs: 5,094 unique
Language: Brazilian Portuguese (pt-BR)
Schema
{"id":"vqa_...","subject":"matematica","category":"geometria","title":"...","messages":[{"role":"user","content":"...… See the full description on the dataset page: https://huggingface.co/datasets/artificialguybr/brazilian-math-physics-qa-vision.Legal_vision_finetuning_data
Sri Lankan Property Law Fine-Tuning Dataset
Dataset Summary
This dataset is a domain-specific legal instruction-tuning dataset designed for fine-tuning large language models for Sri Lankan property law reasoning and legal assistance.
It focuses on core areas of Sri Lankan property law, including:
Property transfer and conveyancing
Title registration (Bim Saviya)
Prescription and adverse possession
Partition of co-owned property
Mortgage and securities
Lease and tenancy… See the full description on the dataset page: https://huggingface.co/datasets/Sivanuja/Legal_vision_finetuning_data.
