datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
colinear_scaling_models
license: gpl-2.0
Collinear Scaling Models
Checkpoint repository for scaling law experiments comparing collinear (CO) and non-collinear (NC) experimental designs.
Directory Structure
{dataset}/{design}/N_{param_count}/
Dataset: wikipedia, pes2o, cosmopedia, redpajama, c4 (plus _fp16 and _bigtpp variants)
Design: colinear or non_colinear
N: Model parameter count (one of 14 canonical sizes from ~5M to ~70M)
Experimental Designs
Collinear (CO):… See the full description on the dataset page: https://huggingface.co/datasets/leibnitz-lab/colinear_scaling_models.CC_eng_urlmilitary_vehicles
Citation
If you use this dataset, please cite the following paper:
@article{kricheli2024error,
title={Error Detection and Constraint Recovery in Hierarchical Multi-Label Classification without Prior Knowledge},
author={Kricheli, Joshua Shay and Vo, Khoa and Datta, Aniruddha and Ozgur, Spencer and Shakarian, Paulo},
journal={arXiv preprint arXiv:2407.15192},
year={2024}
}
VINCIE-10M
Dataset Card for VINCIE-10M
VINCIE: Unlocking In-context Image Editing from Video
Leigang Qu, Feng Cheng, Ziyan Yang, Qi Zhao, Shanchuan Lin, Yichun Shi, Yicong Li, Wenjie Wang, Tat-Seng Chua, Lu Jiang
Dataset Construction Pipeline
Visual Transition Annotation. To describe visual transitions between frames, we use chain-of-thought (CoT) prompting to instruct a VLM to perform visual transition… See the full description on the dataset page: https://huggingface.co/datasets/leigangqu/VINCIE-10M.leire-corpus
Leire Corpus
(EN) Tokenized pretraining corpus for Leire, a 343.7M-parameter Brazilian Portuguese
LM trained from scratch on Kaggle T4s. ~15B tokens, 70% PT / 15% code / 8% math /
7% educational English, tokenized with a custom 32,768 BPE vocabulary trained on the
same mixture. Shards are uint16 binaries; recipe and stats below (in Portuguese).
Corpus de pre-treino da Leire, um LM de 343,7M de parametros em portugues
brasileiro, treinado do zero em T4 do Kaggle. O projeto e… See the full description on the dataset page: https://huggingface.co/datasets/raulmodena/leire-corpus.Leiniao_Datasetmdsarfdetr-segmentation-leibniz-dataset
Dataset Card for Leibniz's Manuscripts (Instance Segmentation Dataset)
This dataset comprises instance segmentation annotations in raw COCO format, used to train an RF-DETR-Seg-nano model for the automatic recognition of textual, graphical, and mathematical expression zones within the manuscripts of the philosopher and mathematician Gottfried Wilhelm Leibniz (17th-early 18th c.).
Dataset Details
Uses
Direct Use
This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/DenisaBumba/rfdetr-segmentation-leibniz-dataset.leisaac-pick-orangeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 60,
"total_frames": 36293,
"total_tasks": 1,
"total_videos": 120,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:60"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/LightwheelAI/leisaac-pick-orange.TIGeR-Bench
Paper: https://arxiv.org/abs/2406.05814
ParseBench
ParseBench
Quick links: [🌐 Website] [📜 Paper] [💻 Code]
ParseBench is a benchmark for evaluating document parsing systems on real-world enterprise documents, with the following characteristics:
Multi-dimensional evaluation. The benchmark is stratified into five capability dimensions — tables, charts, content faithfulness, semantic formatting, and visual grounding — each with task-specific metrics designed to capture what agentic workflows depend on.
Real-world enterprise… See the full description on the dataset page: https://huggingface.co/datasets/matthew-l-leidos/ParseBench.LeoRAG
LeoRAG
LeoRAG is a parquet-backed dataset format for RAG tasks. Each row stores a query, optional preamble, up to 10 retrieved documents, the gold answer, optional teacher outputs, and metadata.
Schema
Column
Type
Description
query
string
Query prompt for the RAG task
preamble
string
Prompt shown before the documents
num_of_docs
int
Number of retrieved documents (0-10)
doc1 ... doc10
string
Retrieved documents (empty when unused)
answer
string… See the full description on the dataset page: https://huggingface.co/datasets/leideng/LeoRAG.airline
airline
An executable Environment for tool-using agents, rebuilt from traces by Kullback and published by Leibler: the world, the Tasks and a code Verifier per Task, no recordings. tasks.jsonl lists the Tasks.
Release: replays at least 90% of its Tasks.
Fidelity over Tasks
100.0% (130 of 130)
Fidelity over Runs
99.5% (199 of 200)
Call fidelity
99.93% of 1513
Reference confirmed
130
Verifier derived
128
Trusted
82
Refused
0
Not trusted
48 of 130; 17 the… See the full description on the dataset page: https://huggingface.co/datasets/leibler/airline.ru-paraphrase-NMT-Leipzig
Dataset Card for cointegrated/ru-paraphrase-NMT-Leipzig
Dataset Summary
The dataset contains 1 million Russian sentences and their automatically generated paraphrases.
It was created by David Dale (@cointegrated) by translating the rus-ru_web-public_2019_1M corpus from the Leipzig collection into English and back into Russian. A fraction of the resulting paraphrases are invalid, and should be filtered out.
The blogpost "Перефразирование русских текстов: корпуса, модели… See the full description on the dataset page: https://huggingface.co/datasets/cointegrated/ru-paraphrase-NMT-Leipzig.htr_leibniz_dataset_v1
Dataset Card for Leibniz's Manuscripts (HTR Ground Truth)
This dataset is composed of trascribed and manually corrected folios of Gottfried Wilhelm Leibniz's manuscripts, together with automatically aligned ground truth in order to train or fine-tune Handwritten Text Recognition (HTR) models.
Dataset Details
This ground truth was produced to fine-tune existing HTR models for the recognition of Leibniz's handwriting, with the aim of assisting scholars in the… See the full description on the dataset page: https://huggingface.co/datasets/DenisaBumba/htr_leibniz_dataset_v1.LeitliniendatenbankIn dieser Datenbank werden alle AWMF Leitlinien hinterlegt, die die Deutsche Gesellschaft für Orthopädie und Unfallchirurgie (DGOU) erstellt hat.
retail
retail
An executable Environment for tool-using agents, rebuilt from traces by Kullback and published by Leibler: the world, the Tasks and a code Verifier per Task, no recordings. tasks.jsonl lists the Tasks.
Release: replays at least 90% of its Tasks.
Fidelity over Tasks
100.0% (223 of 223)
Fidelity over Runs
100.0% (456 of 456)
Call fidelity
100.00% of 3220
Reference confirmed
223
Verifier derived
222
Trusted
193
Refused
0
Not trusted
30 of 223; 15… See the full description on the dataset page: https://huggingface.co/datasets/leibler/retail.MSE-Bench
MSE-Bench: A Benchmark for Multi-turn Session Image Editing
Introduction
MSE-Bench (Multi-turn Session image Editing Benchmark)
is a benchmark designed to evaluate multi-turn image editing systems under realistic editing workflows. Given a source image and a series of editing instructions, the goal is for a model to apply these edits cumulatively to produce a final image that reflects all the requested changes.
MSE-Bench consists of 100 test instances, each representing… See the full description on the dataset page: https://huggingface.co/datasets/leigangqu/MSE-Bench.gleif-lei-scraper-sample-data
GLEIF LEI Scraper
Scrape global legal entities from the official GLEIF LEI database — no login, no API key, no blocking. 3.3M+ entities worldwide with legal name, address, jurisdiction, legal form, status and registration data. Filter by country and status. Tens of thousands per run.
What the actor scrapes
🏛️ GLEIF LEI Scraper — Global Legal Entity Identifier Data to JSON/CSV/Excel Scrape global legal entities straight from the official GLEIF API — the worldwide… See the full description on the dataset page: https://huggingface.co/datasets/logiover/gleif-lei-scraper-sample-data.Copiale_Lines
Copiale Lines
Copiale Lines is a line-level image-to-text dataset for historical cipher decipherment. It contains cropped line images from the Copiale manuscript paired with plaintext ground truth.
This dataset was presented in the paper Learning to Decipher from Pixels -- A Case Study of Copiale (HistoCrypt 2026).
Code: https://github.com/leitro/Decipher-from-Pixels-Copiale
Dataset Structure
The dataset is split into:
train: 1,269 samples
valid: 175 samples… See the full description on the dataset page: https://huggingface.co/datasets/leitro/Copiale_Lines.leisaac-real-pick-orange
SO-101 Real-Arm Pick Orange — 30 Episodes (LeRobot v3.0)
真机 SO-101 主从臂遥操采集的抓橙子放盘子数据集,30 集人类演示,LeRobot v3.0 格式。
Real-world SO-101 leader-follower teleoperation dataset: pick up an orange and place it on the plate, 30 human demonstrations in LeRobot v3.0 format.
概览 / Overview
项
值
机器人 / Robot
SO-101 follower(6 DoF + 夹爪,leader 臂遥操)
任务 / Task
Pick up the orange and place it on the plate
集数 / Episodes
30
总帧数 / Frames
25,091(30 fps,合计 ~14 min)
单集长度 /… See the full description on the dataset page: https://huggingface.co/datasets/wsagi/leisaac-real-pick-orange.registry-harvest-xrpl-mica-lei
Free-registry harvest — XRPL issuers × MiCA × LEI
Public registers and public ledgers joined, read 2026-09-04T07:12:30Z. Every source is free, keyless and re-runnable by
anyone. No part of this needed a relationship, an API key, or anyone's permission.
The finding
Of the 16 XRPL issued assets in the CSOAI reader, 5 join to a MiCA e-money-token authorisation.
group
n
declares an on-chain domain
enforces allowlisting
retains freeze capability… See the full description on the dataset page: https://huggingface.co/datasets/csoai/registry-harvest-xrpl-mica-lei.leipzig_corpora_collection
Leipzig Corpora Collection
The Leipzig Corpora Collection presents corpora in different languages using the same format and comparable sources. All data are available as plain text files and can be imported into a MySQL database by using the provided import script. They are intended both for scientific use by corpus linguists as well as for applications such as knowledge extraction programs.
The corpora are identical in format and similar in size and content. They contain randomly… See the full description on the dataset page: https://huggingface.co/datasets/imvladikon/leipzig_corpora_collection.global-lei-company-registry-dataset
Global LEI Company Registry (GLEIF)
The complete GLEIF golden copy as analysis-ready tables: 3.4M legal entities with registered addresses, corporate hierarchy relationships and reporting exceptions.
Part of the DataForge Open Data program — full production
packages, free for academic and personal use. Canonical dataset page:
https://data.zalize.com/datasets/global-lei-company-registry-dataset
Packages in this repo
Package
Tier
Rows
Size
SHA-256… See the full description on the dataset page: https://huggingface.co/datasets/zalizedata/global-lei-company-registry-dataset.MicroG-4M
MicroG-4M Dataset
This repository stores the entire content of the MicroG-4M dataset itself.
For more information and details, including training, evaluation, statistics, and related code, please:
Refer to our paper
Visit our GitHub
And check our fine-tuned models
Specification of MicroG-4M
"annotation_files" Folder
The folder contains all annotation files of the dataset, all stored in CSV format.
actions.csv
contains all the… See the full description on the dataset page: https://huggingface.co/datasets/lei-qi-233/MicroG-4M.leipzig-frequency
Leipzig Corpora Frequency Data
Word frequency lists and co-occurrence data from the Leipzig Corpora
Collection, converted to Parquet.
Covers hundreds of languages across news, web, Wikipedia, and mixed
sources. Each corpus includes token frequencies, source provenance,
and statistical co-occurrence pairs.
Contents
base/
<language>/
<source>-<date>-<size>/
metadata.json
string.0001.parquet
source.0001.parquet
cooccurrence.sentence.0001.parquet… See the full description on the dataset page: https://huggingface.co/datasets/cluesurf/leipzig-frequency.longbench-view
Introduction
LongBench is the first benchmark for bilingual, multitask, and comprehensive assessment of long context understanding capabilities of large language models. LongBench includes different languages (Chinese and English) to provide a more comprehensive evaluation of the large models' multilingual capabilities on long contexts. In addition, LongBench is composed of six major categories and twenty one different tasks, covering key long-text application scenarios such as… See the full description on the dataset page: https://huggingface.co/datasets/leideng/longbench-view.leisaac-pick-orange-mimic-v0This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 60,
"total_frames": 41891,
"total_tasks": 1,
"total_videos": 120,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:60"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Toby0614/leisaac-pick-orange-mimic-v0.AssistGUI_webMSE-Bench-results
