datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ko-en-structured-translations
Korean–English Multistyle Parallel Corpus
한국어 사용자에게 익숙한 표현 기반의 다도메인·다문체 한–영 병렬 코퍼스
소개(Introduction)
저는 머신러닝, 인공지능 수업을 진행하는 강사입니다.Seq2Seq, Attention, Transformer 등 자연어처리(NLP) 수업을 진행하며한국 학습자에게 자연스럽고 익숙한 한–영 번역 데이터셋의 부족을 경험했습니다.
기존 공개 데이터셋은
도메인 다양성이 부족하거나
문체가 한국 사용자에게 자연스럽지 않거나
전반적으로 문장의 퀄리티가 매우 부족하여
학습한 번역 모델의 실제 성능이 기대만큼 나오지 않는 문제가 있었습니다.
이 문제를 해결하기 위해, 딥러닝 강사로서 langchain을 사용하여 직접 고품질 병렬 데이터를 자동으로 생성·정제하여 구성한 데이터셋입니다.
한국어 사용자에게 익숙한 표현을 중심으로 다양한 문체, 문장 구조를… See the full description on the dataset page: https://huggingface.co/datasets/strongminsu/ko-en-structured-translations.hotpotqa-structuredThis dataset is associated with the paper Search, Inspect, Fetch: Exploiting Boolean Retrieval for Deep-Research Agents.
Official GitHub repository: https://github.com/ielab/skim-search-agent
R1-Reasoning-Unstructured-To-Structured
MasterControl AIML Team 🚀
Overview
The MasterControl AIML team supports the Hugging Face initiative of re-creating DeepSeek R1 training, recognizing it as one of the most impactful open-source projects today.
We aim to contribute to reasoning datasets, specifically those where:
A real-world problem involves generating complex structured output
It is accompanied by step-by-step reasoning and unstructured input
Challenges in Integrating Generative AI… See the full description on the dataset page: https://huggingface.co/datasets/MasterControlAIML/R1-Reasoning-Unstructured-To-Structured.brand-structured-data-reference
Brand Structured Data Reference v1.0
This reference maps common public brand facts to structured data concepts that can help people, search engines, and AI systems understand a brand more clearly.
It is intended for independent brands, small businesses, founder-led companies, service providers, local businesses, and early-stage products that need a clearer public identity online.
This is not a ranking guide and it does not guarantee search visibility, rich results, AI… See the full description on the dataset page: https://huggingface.co/datasets/farosio/brand-structured-data-reference.JSON-Unstructured-StructuredDataset Contains Synthetically Generated Unstructured Text, Set of Rules for Schema Creation, Filled Structured JSON
Can be used for any unstructured to structured tasks
musique-structuredThis dataset repository contains data associated with the paper Search, Inspect, Fetch: Exploiting Boolean Retrieval for Deep-Research Agents.
Code and evaluation framework: ielab/skim-search-agent
Citation
@misc{wang2026sieve,
title = {Search, Inspect, Fetch: Exploiting Boolean Retrieval for Deep-Research Agents},
author = {Wang, Shuai and Chen, Haodong and Yin, Yu and Zhuang, Shengyao and
Koopman, Bevan and Zuccon, Guido},
year… See the full description on the dataset page: https://huggingface.co/datasets/wshuai190/musique-structured.Bva_ama_decisions__structured_2019
BVA AMA Decisions — Structured (A-prefix)
Structured annotations over 5,990 U.S. Board of Veterans' Appeals (BVA)
decisions issued under the Appeals Modernization Act (AMA) — the modernized
appeals system that replaced the legacy process. Every record is an A-prefix
docket decision (A21……), so this is a 100% AMA corpus: the part of the
BVA record that the standard academic/legacy datasets do not cover.
Why this dataset is different
The widely-used BVA legal-NLP… See the full description on the dataset page: https://huggingface.co/datasets/williamTLmiller/Bva_ama_decisions__structured_2019.bva-decisions-structured-sample2019Present
BVA Structured Decisions (2019–2025)
Structured, issue-level records extracted from U.S. Board of Veterans' Appeals (BVA) decisions — each decision parsed into its issues, conditions, outcomes, citations, and reasoning, with per-document provenance and completeness flags. Built for training and evaluating legal-AI models on veterans' disability adjudication.
This is a 2900-decision sample, balanced across seven years (2019–2025, decisions/year), so it's representative of the… See the full description on the dataset page: https://huggingface.co/datasets/williamTLmiller/bva-decisions-structured-sample2019Present.structured-uae-laws
Dataset Card for structured-uae-laws
This dataset is a collection of question & answers about the laws and regulations in the United Arab Emirates.
It covers different areas of law like:
economy and business
family and community
finance and banking
industry and technical standardisation
justice and juiciary, labour
residency and leberal professions
security and safety
tax
Dataset Sources
Repository
Base Dataset
United Arab Emirates Legislations… See the full description on the dataset page: https://huggingface.co/datasets/obadabaq/structured-uae-laws.agriparts_structuredstructuredlogsstructured_ui
