datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
matryoshka-diffusion-models-paper-examples
Matryoshka Diffusion Models - paper examples
This dataset contains the 1024x1024 images included in the Matryoshka Diffusion Models
paper.
Arxiv: https://arxiv.org/abs/2310.15111
model-paper-pages-extracted
NuExtract3 on davanstrien/model-paper-pages-scale
This dataset contains outputs from davanstrien/model-paper-pages-scale processed with NuExtract3, a 4B vision-language model for document understanding.
Processing Details
Source Dataset: davanstrien/model-paper-pages-scale
Model: numind/NuExtract3
Mode: structured-extraction
Number of Samples: 1,767
Processing Time: 48.3 min
Processing Date: 2026-06-01 14:40 UTC
Configuration
Image Column: image… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/model-paper-pages-extracted.ModelVista-paper-only
Dataset Card for "ModelVista-paper-only"
More Information needed
paper-reengineering-mobile-models
Paper: Re-engineering 40+ Models with an Autonomous Agent
This dataset contains the paper and reproducibility data for:
"Re-engineering 40+ Models with an Autonomous Agent: A Zero-Cost Mobile AI Pipeline"
Contents
paper.md — Full paper text
inventory.json — Model inventory and pipeline metadata
Abstract
We present a fully autonomous pipeline that re-engineers open-source language models
for mobile and edge deployment at zero cost. Over 40 models… See the full description on the dataset page: https://huggingface.co/datasets/dispatchAI/paper-reengineering-mobile-models.Deepseek_research_paper_Q_and_A_for_model_fine_tuning
