CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01leiwx52 /CC_eng_urltext100M<n<1B0 likes6.8k downloads2y agoHugging Face02leigangqu /VINCIE-10M Dataset Card for VINCIE-10M VINCIE: Unlocking In-context Image Editing from Video Leigang Qu, Feng Cheng, Ziyan Yang, Qi Zhao, Shanchuan Lin, Yichun Shi, Yicong Li, Wenjie Wang, Tat-Seng Chua, Lu Jiang Dataset Construction Pipeline Visual Transition Annotation. To describe visual transitions between frames, we use chain-of-thought (CoT) prompting to instruct a VLM to perform visual transition… See the full description on the dataset page: https://huggingface.co/datasets/leigangqu/VINCIE-10M.text10M<n<100M12 likes4.4k downloads1y agoHugging Face03leibnitz-lab /mdsaimagen<1K0 likes1.6k downloads1y agoHugging Face04leigangqu /TIGeR-Bench Paper: https://arxiv.org/abs/2406.05814 image10K<n<100K0 likes583 downloads2y agoHugging Face05matthew-l-leidos /ParseBench ParseBench Quick links: [🌐 Website] [📜 Paper] [💻 Code] ParseBench is a benchmark for evaluating document parsing systems on real-world enterprise documents, with the following characteristics: Multi-dimensional evaluation. The benchmark is stratified into five capability dimensions — tables, charts, content faithfulness, semantic formatting, and visual grounding — each with task-specific metrics designed to capture what agentic workflows depend on. Real-world enterprise… See the full description on the dataset page: https://huggingface.co/datasets/matthew-l-leidos/ParseBench.document100K<n<1M0 likes567 downloads3mo agoHugging Face06leideng /LeoRAG LeoRAG LeoRAG is a parquet-backed dataset format for RAG tasks. Each row stores a query, optional preamble, up to 10 retrieved documents, the gold answer, optional teacher outputs, and metadata. Schema Column Type Description query string Query prompt for the RAG task preamble string Prompt shown before the documents num_of_docs int Number of retrieved documents (0-10) doc1 ... doc10 string Retrieved documents (empty when unused) answer string… See the full description on the dataset page: https://huggingface.co/datasets/leideng/LeoRAG.textquestion-answering100K<n<1M0 likes459 downloads3mo agoHugging Face07leibler /airline airline An executable Environment for tool-using agents, rebuilt from traces by Kullback and published by Leibler: the world, the Tasks and a code Verifier per Task, no recordings. tasks.jsonl lists the Tasks. Release: replays at least 90% of its Tasks. Fidelity over Tasks 100.0% (130 of 130) Fidelity over Runs 99.5% (199 of 200) Call fidelity 99.93% of 1513 Reference confirmed 130 Verifier derived 128 Trusted 82 Refused 0 Not trusted 48 of 130; 17 the… See the full description on the dataset page: https://huggingface.co/datasets/leibler/airline.textn<1K2 likes395 downloads1d agoHugging Face08DenisaBumba /htr_leibniz_dataset_v1 Dataset Card for Leibniz's Manuscripts (HTR Ground Truth) This dataset is composed of trascribed and manually corrected folios of Gottfried Wilhelm Leibniz's manuscripts, together with automatically aligned ground truth in order to train or fine-tune Handwritten Text Recognition (HTR) models. Dataset Details This ground truth was produced to fine-tune existing HTR models for the recognition of Leibniz's handwriting, with the aim of assisting scholars in the… See the full description on the dataset page: https://huggingface.co/datasets/DenisaBumba/htr_leibniz_dataset_v1.imageimage-to-text10K<n<100K0 likes381 downloads23d agoHugging Face09leibler /retail retail An executable Environment for tool-using agents, rebuilt from traces by Kullback and published by Leibler: the world, the Tasks and a code Verifier per Task, no recordings. tasks.jsonl lists the Tasks. Release: replays at least 90% of its Tasks. Fidelity over Tasks 100.0% (223 of 223) Fidelity over Runs 100.0% (456 of 456) Call fidelity 100.00% of 3220 Reference confirmed 223 Verifier derived 222 Trusted 193 Refused 0 Not trusted 30 of 223; 15… See the full description on the dataset page: https://huggingface.co/datasets/leibler/retail.textn<1K2 likes355 downloads1d agoHugging Face10leigangqu /MSE-Bench MSE-Bench: A Benchmark for Multi-turn Session Image Editing Introduction MSE-Bench (Multi-turn Session image Editing Benchmark) is a benchmark designed to evaluate multi-turn image editing systems under realistic editing workflows. Given a source image and a series of editing instructions, the goal is for a model to apply these edits cumulatively to produce a final image that reflects all the requested changes. MSE-Bench consists of 100 test instances, each representing… See the full description on the dataset page: https://huggingface.co/datasets/leigangqu/MSE-Bench.imagen<1K2 likes306 downloads6mo agoHugging Face11logiover /gleif-lei-scraper-sample-data GLEIF LEI Scraper Scrape global legal entities from the official GLEIF LEI database — no login, no API key, no blocking. 3.3M+ entities worldwide with legal name, address, jurisdiction, legal form, status and registration data. Filter by country and status. Tens of thousands per run. What the actor scrapes 🏛️ GLEIF LEI Scraper — Global Legal Entity Identifier Data to JSON/CSV/Excel Scrape global legal entities straight from the official GLEIF API — the worldwide… See the full description on the dataset page: https://huggingface.co/datasets/logiover/gleif-lei-scraper-sample-data.textn<1K0 likes291 downloads4mo agoHugging Face12leitro /Copiale_Lines Copiale Lines Copiale Lines is a line-level image-to-text dataset for historical cipher decipherment. It contains cropped line images from the Copiale manuscript paired with plaintext ground truth. This dataset was presented in the paper Learning to Decipher from Pixels -- A Case Study of Copiale (HistoCrypt 2026). Code: https://github.com/leitro/Decipher-from-Pixels-Copiale Dataset Structure The dataset is split into: train: 1,269 samples valid: 175 samples… See the full description on the dataset page: https://huggingface.co/datasets/leitro/Copiale_Lines.imageimage-to-text1K<n<10K0 likes230 downloads5mo agoHugging Face13csoai /registry-harvest-xrpl-mica-lei Free-registry harvest — XRPL issuers × MiCA × LEI Public registers and public ledgers joined, read 2026-09-04T07:12:30Z. Every source is free, keyless and re-runnable by anyone. No part of this needed a relationship, an API key, or anyone's permission. The finding Of the 16 XRPL issued assets in the CSOAI reader, 5 join to a MiCA e-money-token authorisation. group n declares an on-chain domain enforces allowlisting retains freeze capability… See the full description on the dataset page: https://huggingface.co/datasets/csoai/registry-harvest-xrpl-mica-lei.tabularothern<1K0 likes228 downloads12d agoHugging Face14imvladikon /leipzig_corpora_collection Leipzig Corpora Collection The Leipzig Corpora Collection presents corpora in different languages using the same format and comparable sources. All data are available as plain text files and can be imported into a MySQL database by using the provided import script. They are intended both for scientific use by corpus linguists as well as for applications such as knowledge extraction programs. The corpora are identical in format and similar in size and content. They contain randomly… See the full description on the dataset page: https://huggingface.co/datasets/imvladikon/leipzig_corpora_collection.texttext-generation1K<n<10K4 likes227 downloads5mo agoHugging Face15zalizedata /global-lei-company-registry-dataset Global LEI Company Registry (GLEIF) The complete GLEIF golden copy as analysis-ready tables: 3.4M legal entities with registered addresses, corporate hierarchy relationships and reporting exceptions. Part of the DataForge Open Data program — full production packages, free for academic and personal use. Canonical dataset page: https://data.zalize.com/datasets/global-lei-company-registry-dataset Packages in this repo Package Tier Rows Size SHA-256… See the full description on the dataset page: https://huggingface.co/datasets/zalizedata/global-lei-company-registry-dataset.tabulartabular-classification10M<n<100M0 likes222 downloads1mo agoHugging Face16lei-qi-233 /MicroG-4M MicroG-4M Dataset This repository stores the entire content of the MicroG-4M dataset itself. For more information and details, including training, evaluation, statistics, and related code, please: Refer to our paper Visit our GitHub And check our fine-tuned models Specification of MicroG-4M "annotation_files" Folder The folder contains all annotation files of the dataset, all stored in CSV format. actions.csv contains all the… See the full description on the dataset page: https://huggingface.co/datasets/lei-qi-233/MicroG-4M.documentvideo-classification100K<n<1M1 likes217 downloads6mo agoHugging Face17leideng /longbench-view Introduction LongBench is the first benchmark for bilingual, multitask, and comprehensive assessment of long context understanding capabilities of large language models. LongBench includes different languages (Chinese and English) to provide a more comprehensive evaluation of the large models' multilingual capabilities on long contexts. In addition, LongBench is composed of six major categories and twenty one different tasks, covering key long-text application scenarios such as… See the full description on the dataset page: https://huggingface.co/datasets/leideng/longbench-view.textquestion-answering1K<n<10K0 likes194 downloads6mo agoHugging Face18leideng /Dolci-Instruct-SFT-4K-Plus Note [!NOTE] This is the filtered version of allenai/Dolci-Instruct-SFT where only thoese data samples with more than 4096 tokens by Qwen3 tokenizer are kept. Dolci Instruct SFT Mixture Note that this collection licensed under ODC-BY. It is intended for research and educational use in accordance with Ai2's Responsible Use Guidelines. The Dolci Instruct SFT mixture was used to train Olmo 3 7B Instruct SFT. It contains 2,152,112 samples from the following sets: Sources… See the full description on the dataset page: https://huggingface.co/datasets/leideng/Dolci-Instruct-SFT-4K-Plus.textother1M<n<10M0 likes149 downloads5mo agoHugging Face19Leiyao-Cui /XieNet Dataset Card for XieNet This is the repaired version of GAPartNet dataset, which we use as the simulation dataset for Vi-TacMan. Description We identified numerous object meshes in the original dataset that lack proper cap geometry, so we manually repaired these meshes to ensure completeness. The following images (object id: 47296) exemplify the type of geometric defects found and our corrections: GAPartNet (Original)… See the full description on the dataset page: https://huggingface.co/datasets/Leiyao-Cui/XieNet.imagen<1K2 likes138 downloads4mo agoHugging Face20celsowm /clt_consolidacao_leis_trabalho_decreto_lei_5452text1K<n<10K0 likes123 downloads13d agoHugging Face21leixiang25 /24679-hw1-text-garments 24-679 HW1 (Fall 2026): Garment Descriptions leixiang25/24679-hw1-text-garments 100 original garment descriptions written by Lei Xiang for 24-679 Homework 1 at Carnegie Mellon University, plus explicitly marked synthetic training variants. The classification task predicts the garment type (0 top, 1 bottom, 2 outerwear, 3 dress, 4 footwear) from a product-listing style description of about 200 characters. Source and task Every description is the author's own… See the full description on the dataset page: https://huggingface.co/datasets/leixiang25/24679-hw1-text-garments.tabulartext-classification1K<n<10K0 likes83 downloads9d agoHugging Face22leizhao7 /llava-next-mcq-100ktext10K<n<100K0 likes77 downloads5mo agoHugging Face23leideng /Dolci-Think-DPO-7B-4K-Plus Dolci Think 7B DPO Mixture This dataset is licensed under ODC-BY. It is intended for research and educational use in accordance with Ai2's Responsible Use Guidelines. The Dolci Think 7B DPO mixture was used to preference tune Olmo 3 Think 7B. It contains 150,000 preference pairs created with the preference heuristic described in Delta Learning (Geng et al. 2025). Citation @misc{olmo2025olmo3, title={Olmo 3}, author={Team Olmo and Allyson Ettinger and Amanda Bertsch… See the full description on the dataset page: https://huggingface.co/datasets/leideng/Dolci-Think-DPO-7B-4K-Plus.text100K<n<1M0 likes76 downloads5mo agoHugging Face24Texttechnologylab /leipzig-corpora-collectiontext1M<n<10M0 likes72 downloads1y agoHugging Face25leideng /nanochat-ascend-dataset nanochat-ascend-dataset Unified training and evaluation data bundle for nanochat-ascend. This repository is designed to make the nanochat-ascend training procedure easy to reproduce. Instead of asking users to collect multiple task and evaluation datasets and manually reconstruct the expected directory structure, this repository preserves the local filesystem layout expected by the training code. The intended usage is simple: place this repository at .cache/dataset download… See the full description on the dataset page: https://huggingface.co/datasets/leideng/nanochat-ascend-dataset.texttext-generation10K<n<100K0 likes72 downloads6mo agoHugging Face26Alienmaster /wikipedia_leipzig_de_2021 Leipzig Corpora Wikipedia 2021 German This dataset contains different splits (between 10k and 1mio) from the german wikipedia 2021. The data were collected 2021. Every element in the dataset is labeled as "neutral". The source can be found here Citation @inproceedings{goldhahn-etal-2012-building, title = "Building Large Monolingual Dictionaries at the {L}eipzig Corpora Collection: From 100 to 200 Languages", author = "Goldhahn, Dirk and Eckart, Thomas and… See the full description on the dataset page: https://huggingface.co/datasets/Alienmaster/wikipedia_leipzig_de_2021.texttext-classification1M<n<10M0 likes59 downloads2y agoHugging Face27leixiang25 /24679-hw1-image-register 24-679 HW1 (Fall 2026): Business Message Register Images leixiang25/24679-hw1-image-register Digitally rendered screenshot-style images of short, fictional business messages, labeled by register. 1 = formal (high-context business register); 0 = casual (low-context register). Created by Lei Xiang for 24-679 Homework 1 at Carnegie Mellon University. The dataset is related to Context, a cross-cultural deal interpreter for Western operators working with Japanese and Chinese… See the full description on the dataset page: https://huggingface.co/datasets/leixiang25/24679-hw1-image-register.imageimage-classificationn<1K0 likes59 downloads9d agoHugging Face28leideng /longbench-v2-view LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks 🌐 Project Page: https://longbench2.github.io 💻 Github Repo: https://github.com/THUDM/LongBench 📚 Arxiv Paper: https://arxiv.org/abs/2412.15204 LongBench v2 is designed to assess the ability of LLMs to handle long-context problems requiring deep understanding and reasoning across real-world multitasks. LongBench v2 has the following features: (1) Length: Context length ranging from 8k to… See the full description on the dataset page: https://huggingface.co/datasets/leideng/longbench-v2-view.textmultiple-choice1K<n<10K0 likes58 downloads6mo agoHugging Face29leixiang25 /24679-hw1-tabular-business-cards 24-679 HW1 (Fall 2026): Business Card Content leixiang25/24679-hw1-tabular-business-cards Content features of 31 business card designs observed in public online galleries, plus explicitly marked synthetic training variants. The regression task predicts text_lines (total printed text lines across both sides) from three count features and five categorical design choices. Created by Lei Xiang for 24-679 Homework 1 at Carnegie Mellon University; loosely related to Context, a… See the full description on the dataset page: https://huggingface.co/datasets/leixiang25/24679-hw1-tabular-business-cards.tabulartabular-regressionn<1K0 likes57 downloads9d agoHugging Face30celsowm /codigo_penal_brasileiro_lei_2848_1940textn<1K1 likes55 downloads13d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.