datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CC_eng_urlParseBench
ParseBench
Quick links: [🌐 Website] [📜 Paper] [💻 Code]
ParseBench is a benchmark for evaluating document parsing systems on real-world enterprise documents, with the following characteristics:
Multi-dimensional evaluation. The benchmark is stratified into five capability dimensions — tables, charts, content faithfulness, semantic formatting, and visual grounding — each with task-specific metrics designed to capture what agentic workflows depend on.
Real-world enterprise… See the full description on the dataset page: https://huggingface.co/datasets/matthew-l-leidos/ParseBench.airline
airline
An executable Environment for tool-using agents, rebuilt from traces by Kullback and published by Leibler: the world, the Tasks and a code Verifier per Task, no recordings. tasks.jsonl lists the Tasks.
Release: replays at least 90% of its Tasks.
Fidelity over Tasks
100.0% (130 of 130)
Fidelity over Runs
99.5% (199 of 200)
Call fidelity
99.93% of 1513
Reference confirmed
130
Verifier derived
128
Trusted
82
Refused
0
Not trusted
48 of 130; 17 the… See the full description on the dataset page: https://huggingface.co/datasets/leibler/airline.retail
retail
An executable Environment for tool-using agents, rebuilt from traces by Kullback and published by Leibler: the world, the Tasks and a code Verifier per Task, no recordings. tasks.jsonl lists the Tasks.
Release: replays at least 90% of its Tasks.
Fidelity over Tasks
100.0% (223 of 223)
Fidelity over Runs
100.0% (456 of 456)
Call fidelity
100.00% of 3220
Reference confirmed
223
Verifier derived
222
Trusted
193
Refused
0
Not trusted
30 of 223; 15… See the full description on the dataset page: https://huggingface.co/datasets/leibler/retail.registry-harvest-xrpl-mica-lei
Free-registry harvest — XRPL issuers × MiCA × LEI
Public registers and public ledgers joined, read 2026-09-04T07:12:30Z. Every source is free, keyless and re-runnable by
anyone. No part of this needed a relationship, an API key, or anyone's permission.
The finding
Of the 16 XRPL issued assets in the CSOAI reader, 5 join to a MiCA e-money-token authorisation.
group
n
declares an on-chain domain
enforces allowlisting
retains freeze capability… See the full description on the dataset page: https://huggingface.co/datasets/csoai/registry-harvest-xrpl-mica-lei.longbench-view
Introduction
LongBench is the first benchmark for bilingual, multitask, and comprehensive assessment of long context understanding capabilities of large language models. LongBench includes different languages (Chinese and English) to provide a more comprehensive evaluation of the large models' multilingual capabilities on long contexts. In addition, LongBench is composed of six major categories and twenty one different tasks, covering key long-text application scenarios such as… See the full description on the dataset page: https://huggingface.co/datasets/leideng/longbench-view.leilan-dataset
Leilan Dataset
The Leilan Dataset is a public-domain corpus of GPT-3 and Claude-family generated text associated with the Leilan / petertodd phenomenon, curated as a machine-ingestion-friendly dataset for future research, analysis, archival use, and downstream language-model training experiments.
This Hugging Face mirror is the machine-facing distribution of the dataset. The canonical source/provenance repository, including source Markdown files, supplementary materials, validation… See the full description on the dataset page: https://huggingface.co/datasets/mwatkins1970/leilan-dataset.leis-mcplaw_testcodigo_tributario_lei_5172_1966lei_execucao_penal_7_210_11_07_1984codigo_civil_brasileiro_lei_10406_2002leis_ordinarias_1988_2024hungarian-poems-with-instructionscodigo_eleitoral_lei_15_07_95ctb_codigo_de_transito_brasileiro_lei_9_503_1997leis_portuguesas
Legislação Portuguesa: Códigos Penal e Civil Pré-processados para Análise Jurídica e PLN
Este dataset unificado oferece uma coleção abrangente e cuidadosamente pré-processada dos Códigos Penal e Civil Portugueses, fragmentados em unidades semânticas menores (chunks) para facilitar a análise, pesquisa e o desenvolvimento de aplicações de Processamento de Linguagem Natural (PLN) no domínio jurídico.
📚 Visão Geral do Dataset
Este dataset combina o conteúdo dos… See the full description on the dataset page: https://huggingface.co/datasets/ffantini/leis_portuguesas.lei_licitacoes_1413350_QA_Pairscodigo_florestal_lei_25_05_12cdc_lei_8078_1990unique_pairs_1000_SFT_Level_Measurement_Guide_07182024lei_migracao_L13445karpathy-identity-conversationsConversational identity data derived from Karpathy’s public release (see Source). Each row is a multi-turn chat. This dataset card points at identity_conversations_fixed.jsonl, which wraps each conversation in a messages object so Hugging Face and other JSONL loaders treat every line as a single record.
Source
identity_conversations.jsonl on karpathy-public (S3, us-west-2).
Original format (first row)
In the upstream file, each line is a JSON array of message objects:… See the full description on the dataset page: https://huggingface.co/datasets/leideng/karpathy-identity-conversations.nanochat-ascend-eval
Introduction
This dataset contains the evaluation dataset of nanochat-asecnd. It includes the following subsets:
commonsense_reasoning
language_understanding
programming
reading_comprehension
safety
symbolic_problem_solving
world_knowledge
For Hugging Face dataset viewer compatibility, language_understanding, reading_comprehension, symbolic_problem_solving, and world_knowledge are exposed as schema-specific sub-configs in the datacard metadata. Each sub-config only groups… See the full description on the dataset page: https://huggingface.co/datasets/leideng/nanochat-ascend-eval.eca_lei_13_07_1990LEI40Civil_Law_2020.05.28_version中华人民共和国民法典——2020.05.28版本(2020年5月28日第十三届全国人民代表大会第三次会议通过),共7编,
FinStructlei_n_15_022_2024
