datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CC_eng_urlVINCIE-10M
Dataset Card for VINCIE-10M
VINCIE: Unlocking In-context Image Editing from Video
Leigang Qu, Feng Cheng, Ziyan Yang, Qi Zhao, Shanchuan Lin, Yichun Shi, Yicong Li, Wenjie Wang, Tat-Seng Chua, Lu Jiang
Dataset Construction Pipeline
Visual Transition Annotation. To describe visual transitions between frames, we use chain-of-thought (CoT) prompting to instruct a VLM to perform visual transition… See the full description on the dataset page: https://huggingface.co/datasets/leigangqu/VINCIE-10M.mdsaTIGeR-Bench
Paper: https://arxiv.org/abs/2406.05814
ParseBench
ParseBench
Quick links: [🌐 Website] [📜 Paper] [💻 Code]
ParseBench is a benchmark for evaluating document parsing systems on real-world enterprise documents, with the following characteristics:
Multi-dimensional evaluation. The benchmark is stratified into five capability dimensions — tables, charts, content faithfulness, semantic formatting, and visual grounding — each with task-specific metrics designed to capture what agentic workflows depend on.
Real-world enterprise… See the full description on the dataset page: https://huggingface.co/datasets/matthew-l-leidos/ParseBench.LeoRAG
LeoRAG
LeoRAG is a parquet-backed dataset format for RAG tasks. Each row stores a query, optional preamble, up to 10 retrieved documents, the gold answer, optional teacher outputs, and metadata.
Schema
Column
Type
Description
query
string
Query prompt for the RAG task
preamble
string
Prompt shown before the documents
num_of_docs
int
Number of retrieved documents (0-10)
doc1 ... doc10
string
Retrieved documents (empty when unused)
answer
string… See the full description on the dataset page: https://huggingface.co/datasets/leideng/LeoRAG.airline
airline
An executable Environment for tool-using agents, rebuilt from traces by Kullback and published by Leibler: the world, the Tasks and a code Verifier per Task, no recordings. tasks.jsonl lists the Tasks.
Release: replays at least 90% of its Tasks.
Fidelity over Tasks
100.0% (130 of 130)
Fidelity over Runs
99.5% (199 of 200)
Call fidelity
99.93% of 1513
Reference confirmed
130
Verifier derived
128
Trusted
82
Refused
0
Not trusted
48 of 130; 17 the… See the full description on the dataset page: https://huggingface.co/datasets/leibler/airline.htr_leibniz_dataset_v1
Dataset Card for Leibniz's Manuscripts (HTR Ground Truth)
This dataset is composed of trascribed and manually corrected folios of Gottfried Wilhelm Leibniz's manuscripts, together with automatically aligned ground truth in order to train or fine-tune Handwritten Text Recognition (HTR) models.
Dataset Details
This ground truth was produced to fine-tune existing HTR models for the recognition of Leibniz's handwriting, with the aim of assisting scholars in the… See the full description on the dataset page: https://huggingface.co/datasets/DenisaBumba/htr_leibniz_dataset_v1.retail
retail
An executable Environment for tool-using agents, rebuilt from traces by Kullback and published by Leibler: the world, the Tasks and a code Verifier per Task, no recordings. tasks.jsonl lists the Tasks.
Release: replays at least 90% of its Tasks.
Fidelity over Tasks
100.0% (223 of 223)
Fidelity over Runs
100.0% (456 of 456)
Call fidelity
100.00% of 3220
Reference confirmed
223
Verifier derived
222
Trusted
193
Refused
0
Not trusted
30 of 223; 15… See the full description on the dataset page: https://huggingface.co/datasets/leibler/retail.MSE-Bench
MSE-Bench: A Benchmark for Multi-turn Session Image Editing
Introduction
MSE-Bench (Multi-turn Session image Editing Benchmark)
is a benchmark designed to evaluate multi-turn image editing systems under realistic editing workflows. Given a source image and a series of editing instructions, the goal is for a model to apply these edits cumulatively to produce a final image that reflects all the requested changes.
MSE-Bench consists of 100 test instances, each representing… See the full description on the dataset page: https://huggingface.co/datasets/leigangqu/MSE-Bench.gleif-lei-scraper-sample-data
GLEIF LEI Scraper
Scrape global legal entities from the official GLEIF LEI database — no login, no API key, no blocking. 3.3M+ entities worldwide with legal name, address, jurisdiction, legal form, status and registration data. Filter by country and status. Tens of thousands per run.
What the actor scrapes
🏛️ GLEIF LEI Scraper — Global Legal Entity Identifier Data to JSON/CSV/Excel Scrape global legal entities straight from the official GLEIF API — the worldwide… See the full description on the dataset page: https://huggingface.co/datasets/logiover/gleif-lei-scraper-sample-data.Copiale_Lines
Copiale Lines
Copiale Lines is a line-level image-to-text dataset for historical cipher decipherment. It contains cropped line images from the Copiale manuscript paired with plaintext ground truth.
This dataset was presented in the paper Learning to Decipher from Pixels -- A Case Study of Copiale (HistoCrypt 2026).
Code: https://github.com/leitro/Decipher-from-Pixels-Copiale
Dataset Structure
The dataset is split into:
train: 1,269 samples
valid: 175 samples… See the full description on the dataset page: https://huggingface.co/datasets/leitro/Copiale_Lines.registry-harvest-xrpl-mica-lei
Free-registry harvest — XRPL issuers × MiCA × LEI
Public registers and public ledgers joined, read 2026-09-04T07:12:30Z. Every source is free, keyless and re-runnable by
anyone. No part of this needed a relationship, an API key, or anyone's permission.
The finding
Of the 16 XRPL issued assets in the CSOAI reader, 5 join to a MiCA e-money-token authorisation.
group
n
declares an on-chain domain
enforces allowlisting
retains freeze capability… See the full description on the dataset page: https://huggingface.co/datasets/csoai/registry-harvest-xrpl-mica-lei.leipzig_corpora_collection
Leipzig Corpora Collection
The Leipzig Corpora Collection presents corpora in different languages using the same format and comparable sources. All data are available as plain text files and can be imported into a MySQL database by using the provided import script. They are intended both for scientific use by corpus linguists as well as for applications such as knowledge extraction programs.
The corpora are identical in format and similar in size and content. They contain randomly… See the full description on the dataset page: https://huggingface.co/datasets/imvladikon/leipzig_corpora_collection.global-lei-company-registry-dataset
Global LEI Company Registry (GLEIF)
The complete GLEIF golden copy as analysis-ready tables: 3.4M legal entities with registered addresses, corporate hierarchy relationships and reporting exceptions.
Part of the DataForge Open Data program — full production
packages, free for academic and personal use. Canonical dataset page:
https://data.zalize.com/datasets/global-lei-company-registry-dataset
Packages in this repo
Package
Tier
Rows
Size
SHA-256… See the full description on the dataset page: https://huggingface.co/datasets/zalizedata/global-lei-company-registry-dataset.MicroG-4M
MicroG-4M Dataset
This repository stores the entire content of the MicroG-4M dataset itself.
For more information and details, including training, evaluation, statistics, and related code, please:
Refer to our paper
Visit our GitHub
And check our fine-tuned models
Specification of MicroG-4M
"annotation_files" Folder
The folder contains all annotation files of the dataset, all stored in CSV format.
actions.csv
contains all the… See the full description on the dataset page: https://huggingface.co/datasets/lei-qi-233/MicroG-4M.longbench-view
Introduction
LongBench is the first benchmark for bilingual, multitask, and comprehensive assessment of long context understanding capabilities of large language models. LongBench includes different languages (Chinese and English) to provide a more comprehensive evaluation of the large models' multilingual capabilities on long contexts. In addition, LongBench is composed of six major categories and twenty one different tasks, covering key long-text application scenarios such as… See the full description on the dataset page: https://huggingface.co/datasets/leideng/longbench-view.Dolci-Instruct-SFT-4K-Plus
Note
[!NOTE]
This is the filtered version of allenai/Dolci-Instruct-SFT where only thoese data samples with more than 4096 tokens by Qwen3 tokenizer are kept.
Dolci Instruct SFT Mixture
Note that this collection licensed under ODC-BY. It is intended for research and educational use in accordance with Ai2's Responsible Use Guidelines.
The Dolci Instruct SFT mixture was used to train Olmo 3 7B Instruct SFT.
It contains 2,152,112 samples from the following sets:
Sources… See the full description on the dataset page: https://huggingface.co/datasets/leideng/Dolci-Instruct-SFT-4K-Plus.XieNet
Dataset Card for XieNet
This is the repaired version of GAPartNet dataset, which we use as the simulation dataset for Vi-TacMan.
Description
We identified numerous object meshes in the original dataset that lack proper cap geometry, so we manually repaired these meshes to ensure completeness. The following images (object id: 47296) exemplify the type of geometric defects found and our corrections:
GAPartNet (Original)… See the full description on the dataset page: https://huggingface.co/datasets/Leiyao-Cui/XieNet.clt_consolidacao_leis_trabalho_decreto_lei_545224679-hw1-text-garments
24-679 HW1 (Fall 2026): Garment Descriptions
leixiang25/24679-hw1-text-garments
100 original garment descriptions written by Lei Xiang for 24-679 Homework 1 at Carnegie Mellon
University, plus explicitly marked synthetic training variants. The classification task predicts the
garment type (0 top, 1 bottom, 2 outerwear, 3 dress, 4 footwear) from a product-listing style
description of about 200 characters.
Source and task
Every description is the author's own… See the full description on the dataset page: https://huggingface.co/datasets/leixiang25/24679-hw1-text-garments.llava-next-mcq-100kDolci-Think-DPO-7B-4K-Plus
Dolci Think 7B DPO Mixture
This dataset is licensed under ODC-BY. It is intended for research and educational use in accordance with Ai2's Responsible Use Guidelines.
The Dolci Think 7B DPO mixture was used to preference tune Olmo 3 Think 7B. It contains 150,000 preference pairs created with the preference heuristic described in Delta Learning (Geng et al. 2025).
Citation
@misc{olmo2025olmo3,
title={Olmo 3},
author={Team Olmo and Allyson Ettinger and Amanda Bertsch… See the full description on the dataset page: https://huggingface.co/datasets/leideng/Dolci-Think-DPO-7B-4K-Plus.leipzig-corpora-collectionnanochat-ascend-dataset
nanochat-ascend-dataset
Unified training and evaluation data bundle for nanochat-ascend.
This repository is designed to make the nanochat-ascend training procedure easy to reproduce. Instead of asking users to collect multiple task and evaluation datasets and manually reconstruct the expected directory structure, this repository preserves the local filesystem layout expected by the training code.
The intended usage is simple:
place this repository at .cache/dataset
download… See the full description on the dataset page: https://huggingface.co/datasets/leideng/nanochat-ascend-dataset.wikipedia_leipzig_de_2021
Leipzig Corpora Wikipedia 2021 German
This dataset contains different splits (between 10k and 1mio) from the german wikipedia 2021. The data were collected 2021.
Every element in the dataset is labeled as "neutral".
The source can be found here
Citation
@inproceedings{goldhahn-etal-2012-building,
title = "Building Large Monolingual Dictionaries at the {L}eipzig Corpora Collection: From 100 to 200 Languages",
author = "Goldhahn, Dirk and
Eckart, Thomas and… See the full description on the dataset page: https://huggingface.co/datasets/Alienmaster/wikipedia_leipzig_de_2021.24679-hw1-image-register
24-679 HW1 (Fall 2026): Business Message Register Images
leixiang25/24679-hw1-image-register
Digitally rendered screenshot-style images of short, fictional business messages, labeled by register.
1 = formal (high-context business register); 0 = casual (low-context register). Created by Lei Xiang
for 24-679 Homework 1 at Carnegie Mellon University. The dataset is related to Context, a cross-cultural
deal interpreter for Western operators working with Japanese and Chinese… See the full description on the dataset page: https://huggingface.co/datasets/leixiang25/24679-hw1-image-register.longbench-v2-view
LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks
🌐 Project Page: https://longbench2.github.io
💻 Github Repo: https://github.com/THUDM/LongBench
📚 Arxiv Paper: https://arxiv.org/abs/2412.15204
LongBench v2 is designed to assess the ability of LLMs to handle long-context problems requiring deep understanding and reasoning across real-world multitasks. LongBench v2 has the following features: (1) Length: Context length ranging from 8k to… See the full description on the dataset page: https://huggingface.co/datasets/leideng/longbench-v2-view.24679-hw1-tabular-business-cards
24-679 HW1 (Fall 2026): Business Card Content
leixiang25/24679-hw1-tabular-business-cards
Content features of 31 business card designs observed in public online galleries, plus explicitly
marked synthetic training variants. The regression task predicts text_lines (total printed text lines
across both sides) from three count features and five categorical design choices. Created by Lei Xiang
for 24-679 Homework 1 at Carnegie Mellon University; loosely related to Context, a… See the full description on the dataset page: https://huggingface.co/datasets/leixiang25/24679-hw1-tabular-business-cards.codigo_penal_brasileiro_lei_2848_1940
