CoolFace
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01visionscaper /agentic-llm-pretraining-1.7b Agentic LLM Pretraining Dataset A pretraining corpus for small language models (1-3B parameters) optimized for agentic tasks. The corpus emphasizes learning to comprehend language, reason, follow instructions, and use tools over memorizing factual knowledge — the assumption is that domain knowledge will be provided at runtime via RAG. The idea is that this could enable much smaller pretraining corpora by omitting the large volumes of text typically needed to memorize facts.… See the full description on the dataset page: https://huggingface.co/datasets/visionscaper/agentic-llm-pretraining-1.7b.texttext-generation1M<n<10M3 likes234 downloads9mo agoHugging Face02translorentz /vision-token-compression-bench OPTIC-Bench Optical Text In-Context Benchmark: how reliably do LLMs consume text delivered as rendered images versus plain text tokens? In summary, the evaluation reported here finds that optical text compression is effective only within a narrow and specific envelope. Delivering content as rendered images genuinely reduces input tokens, by thirteen to fifty-four per cent depending on the model and the language, but only when the document is long, the rendering is dense and the… See the full description on the dataset page: https://huggingface.co/datasets/translorentz/vision-token-compression-bench.imagevisual-question-answering1K<n<10K0 likes170 downloads2mo agoHugging Face03Haonian /unified-math-vision-dataset Unified Math Vision Dataset Bundle Generated at: 2025-09-19 17:12:44 This is a unified dataset bundle containing multiple math and vision reasoning datasets. Dataset Statistics Total samples: 15858 mathvision: 3344 samples wemath: 500 samples mmmu: 415 samples mathvista: 6141 samples logicvista: 448 samples dynamath: 5010 samples Contents manifest.jsonl: Complete dataset in JSONL format (1 JSON per line) manifest.csv: Summary in CSV format images/: Directory… See the full description on the dataset page: https://huggingface.co/datasets/Haonian/unified-math-vision-dataset.textquestion-answering10K<n<100K0 likes54 downloads1y agoHugging Face04artificialguybr /brazilian-math-physics-qa-vision Brazilian Math & Physics QA — Image Dependent English | Português do Brasil English Summary Brazilian Portuguese educational question-answer pairs whose problem statement or solution depends on one or more images. Examples: 3,808 Referenced image URLs: 5,094 unique Language: Brazilian Portuguese (pt-BR) Schema {"id":"vqa_...","subject":"matematica","category":"geometria","title":"...","messages":[{"role":"user","content":"...… See the full description on the dataset page: https://huggingface.co/datasets/artificialguybr/brazilian-math-physics-qa-vision.textvisual-question-answering1K<n<10K0 likes20 downloads1mo agoHugging Face05Sivanuja /Legal_vision_finetuning_data Sri Lankan Property Law Fine-Tuning Dataset Dataset Summary This dataset is a domain-specific legal instruction-tuning dataset designed for fine-tuning large language models for Sri Lankan property law reasoning and legal assistance. It focuses on core areas of Sri Lankan property law, including: Property transfer and conveyancing Title registration (Bim Saviya) Prescription and adverse possession Partition of co-owned property Mortgage and securities Lease and tenancy… See the full description on the dataset page: https://huggingface.co/datasets/Sivanuja/Legal_vision_finetuning_data.texttext-generation1K<n<10K0 likes13 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.