structured-extraction
putusan-structured-extraction
Putusan structured-extraction dataset
Built 2026-07-08T23:02:28+00:00 by notebooks/build_dataset.py (seed 3407).
Indonesian court-decision (putusan) extractive-structuring dataset over three
corpora (Anak, Asusila, TPPO). Each row is one model extraction of one source
document into 31 canonical sections of verbatim spans. Empty sections were
completed from sibling model extractions of the same document where available
(cross_model_fill_json records per-section donor provenance).… See the full description on the dataset page: https://huggingface.co/datasets/Haeryz/putusan-structured-extraction.chat_structured_extraction
Lead Extraction Dataset
Dataset Description
This dataset contains structured extraction examples for lead information from conversational input in Spanish.
Dataset Structure
Format: JSONL (JSON Lines)
Total Examples: 120
Splits:
Train: 90 examples
Dev: 10 examples
Test: 20 examples
Schema
Each row in the dataset follows the schema defined in schemas/lead_extraction_row_1.0.0.json.
Task
Extract structured lead information from user… See the full description on the dataset page: https://huggingface.co/datasets/mauroibz/chat_structured_extraction.structured-generation-information-extraction-vlms-openbmb-RLAIF-V-Datasetstructured-generation-information-extraction-vlms-openbmb-RLAIF-V-Datasetstructured-generation-information-extraction-vlms-openbmb-RLAIF-V-Datasetstructured-generation-information-extraction-vlms-openbmb-RLAIF-V-Dataset
