datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
xtreme-up-semantic-parsing
Dataset Card for afrixnli
Dataset Summary
See XTREME-UP GitHub
Languages
There are 20 languages available :
Dataset Structure
Data Instances
The examples look like this for English:
from datasets import load_dataset
data = load_dataset('Davlan/xtreme-up-semantic-parsing', 'yor')
# Please, specify the language code
# A data point example is below:
{
"id": "3231323330393336",
"split": "test",
"intent": "IN:GET_REMINDER",
"locale": "en"… See the full description on the dataset page: https://huggingface.co/datasets/Davlan/xtreme-up-semantic-parsing.arxiv-sample-affiliation-parsing-lora-Qwen3-8B-distil-GLM_4.5_Air-inference-results-enriched
affiliation-parsing-lora-Qwen3-8B-distil-GLM_4.5_Air arXiv author affiliation inference results
Author names and institutional affiliations extracted from arXiv preprints with the affiliation-parsing-lora-Qwen3-8B-distil-GLM_4.5_Air LoRA, enriched with ROR identifiers.
Dataset Structure
Each record contains the following fields:
Field
Type
Description
doi
string
DOI for the preprint
title
string
Preprint title
arxiv_id
stringarXiv identifier… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/arxiv-sample-affiliation-parsing-lora-Qwen3-8B-distil-GLM_4.5_Air-inference-results-enriched.chess-time-control-string-parsing
Chess Time-Control String Parsing
Real-world chess time-control strings, in two forms:
.txt files — the source of truth. Every unique time-control string, one per line, with a frequency count. These are the raw, real strings (messy, multilingual, sometimes junk) as scraped/collected. No interpretation.
.jsonl files — a tagged, partially-correct derived artifact. Each unique string with an auto-derived (category, stages) parse. The tags are heuristics, not verified ground truth… See the full description on the dataset page: https://huggingface.co/datasets/gutsy-gambit/chess-time-control-string-parsing.chinese_insurance_doc_parsing本数据集清洗自天池实验室公共数据集
结合原数据集的标注和pdf文档解析工具,构造了alpaca格式的数据:
Instuction:
下列是直接从pdf原文件中提取出的某保险条款原文,pdf文件的字体排版存在一些空间结构,直接转换成字符串后会导致条款原文非常难以阅读。请把内容重新组织成清晰可读的格式。要求如下:
第一行是保险公司的全称
第二行是保险产品名
章节和子章节的序号统一用数字1-9表示
章节序号和章节名写在同一行,用空格进行间隔;章节具体内容放在下一行
章节和章节之间空一行
input:
使用pdfminer直接提取的字符串
中国太平洋人寿保险股份有限公司
个人税收递延型养老年金保险(2018 版)
产品基本条款
第一条 合同构成
个人税收递延型养老年金保险(2018 版)产品合同(以下简称“本合同”)由保险单及
所附个人税收递延型养老年金保险(2018 版)产品基本条款(以下简称“本合同基本条款
(2018 版)”)、个人税收递延型养老年金保险(2018 版)产品账户利益条款(以下简称“本
合同账户利益条款(2018… See the full description on the dataset page: https://huggingface.co/datasets/kaihe/chinese_insurance_doc_parsing.CBRS-parsing
CBRS Dataset
This repository contains the dataset for the Cognitive Blood Request System (CBRS), as introduced in the paper CBRS: Cognitive Blood Request System with Bilingual Dataset and Dual-Layer Filtering for Multi-Platform Social Streams.
CBRS is a framework designed to efficiently filter and parse blood donation requests from social media streams. The dataset includes 11,000 parsed blood donation request messages capturing the linguistic diversity of real-world communications.… See the full description on the dataset page: https://huggingface.co/datasets/imAniksahA/CBRS-parsing.dependency-parsing-instructionsrepro-ovisocr-end-to-end-document-parsing-traces
Agent traces
Agent sessions published from a Trackio Logbook.
10-million-English-Test-Questions-Text-Parsing-And-Processing-Data-Sample
Description
10 Million - English Test Questions Text Parsing And Processing Data, Each question contains title, answer, parse, subject, grade, question type; The educational stages cover primary, middle, high school, and university; Subjects cover mathmatics, biology, accounting, etc.The data are questions text under the Anglo-American system, which can be used to enhance the subject knowledge of large models
For more details, please refer to the link:… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-AI/10-million-English-Test-Questions-Text-Parsing-And-Processing-Data-Sample.constituency-parsing-instructionsresume_parsingetd-author-affiliation-parsingResume_Parsingautotrain-data-address-parsingarxiv-author-affiliation-parsing-sample-datacite-enrichment-format
