datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
key_information_extractionVisual-Extraction-Tuning-382K
Visual Extraction Tuning 382K
This repository contains the generated visual extraction tuning dataset from the paper Downscaling Intelligence: Exploring Perception and Reasoning Bottlenecks in Small Multimodal Models.
Project page: https://web.stanford.edu/~markendo/projects/downscaling_intelligence
Code: https://github.com/markendo/downscaling_intelligence
Overview
We provide the 382K examples generated using our visual extraction tuning data generation pipeline.… See the full description on the dataset page: https://huggingface.co/datasets/markendo/Visual-Extraction-Tuning-382K.Drug_Combination_Extractionsec-extraction-multitask-v4
SEC Extraction Multitask v4
Instruction-tuning dataset for fine-tuning a small language model (e.g. Gemma 4 E2B) to extract structured data from SEC filings across three verticals:
Exhibit 10 (contracts) — financial terms from executive employment, credit agreements, indemnification, licensing, and similar filings
DEF 14A (proxy statements) — executive compensation, governance items, say-on-pay
MD&A (10-K / 10-Q Management's Discussion & Analysis) — operating metrics, segment… See the full description on the dataset page: https://huggingface.co/datasets/TheTokenFactory/sec-extraction-multitask-v4.logdetective-extraction-wiptask-extractions
Dataset Card for task-extractions
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/JoelTankard/task-extractions/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/JoelTankard/task-extractions.
