CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01leibnitz-lab /colinear_scaling_models license: gpl-2.0 Collinear Scaling Models Checkpoint repository for scaling law experiments comparing collinear (CO) and non-collinear (NC) experimental designs. Directory Structure {dataset}/{design}/N_{param_count}/ Dataset: wikipedia, pes2o, cosmopedia, redpajama, c4 (plus _fp16 and _bigtpp variants) Design: colinear or non_colinear N: Model parameter count (one of 14 canonical sizes from ~5M to ~70M) Experimental Designs Collinear (CO):… See the full description on the dataset page: https://huggingface.co/datasets/leibnitz-lab/colinear_scaling_models.1 likes19k downloads5mo agoHugging Face02leiwx52 /CC_eng_urltext100M<n<1B0 likes6.8k downloads2y agoHugging Face03leibnitz-lab /military_vehicles Citation If you use this dataset, please cite the following paper: @article{kricheli2024error, title={Error Detection and Constraint Recovery in Hierarchical Multi-Label Classification without Prior Knowledge}, author={Kricheli, Joshua Shay and Vo, Khoa and Datta, Aniruddha and Ozgur, Spencer and Shakarian, Paulo}, journal={arXiv preprint arXiv:2407.15192}, year={2024} } imageimage-classification10K<n<100K2 likes6k downloads9mo agoHugging Face04leigangqu /VINCIE-10M Dataset Card for VINCIE-10M VINCIE: Unlocking In-context Image Editing from Video Leigang Qu, Feng Cheng, Ziyan Yang, Qi Zhao, Shanchuan Lin, Yichun Shi, Yicong Li, Wenjie Wang, Tat-Seng Chua, Lu Jiang Dataset Construction Pipeline Visual Transition Annotation. To describe visual transitions between frames, we use chain-of-thought (CoT) prompting to instruct a VLM to perform visual transition… See the full description on the dataset page: https://huggingface.co/datasets/leigangqu/VINCIE-10M.text10M<n<100M12 likes4.4k downloads1y agoHugging Face05raulmodena /leire-corpus Leire Corpus (EN) Tokenized pretraining corpus for Leire, a 343.7M-parameter Brazilian Portuguese LM trained from scratch on Kaggle T4s. ~15B tokens, 70% PT / 15% code / 8% math / 7% educational English, tokenized with a custom 32,768 BPE vocabulary trained on the same mixture. Shards are uint16 binaries; recipe and stats below (in Portuguese). Corpus de pre-treino da Leire, um LM de 343,7M de parametros em portugues brasileiro, treinado do zero em T4 do Kaggle. O projeto e… See the full description on the dataset page: https://huggingface.co/datasets/raulmodena/leire-corpus.text-generation10B<n<100B0 likes2.2k downloads19d agoHugging Face06Chandler-Shen /Leiniao_Datasetimagen<1K0 likes1.8k downloads6d agoHugging Face07leibnitz-lab /mdsaimagen<1K0 likes1.6k downloads1y agoHugging Face08DenisaBumba /rfdetr-segmentation-leibniz-dataset Dataset Card for Leibniz's Manuscripts (Instance Segmentation Dataset) This dataset comprises instance segmentation annotations in raw COCO format, used to train an RF-DETR-Seg-nano model for the automatic recognition of textual, graphical, and mathematical expression zones within the manuscripts of the philosopher and mathematician Gottfried Wilhelm Leibniz (17th-early 18th c.). Dataset Details Uses Direct Use This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/DenisaBumba/rfdetr-segmentation-leibniz-dataset.imageimage-segmentation1K<n<10K0 likes963 downloads2mo agoHugging Face09LightwheelAI /leisaac-pick-orangeThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 60, "total_frames": 36293, "total_tasks": 1, "total_videos": 120, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:60" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/LightwheelAI/leisaac-pick-orange.tabularrobotics10K<n<100K8 likes700 downloads1y agoHugging Face10leigangqu /TIGeR-Bench Paper: https://arxiv.org/abs/2406.05814 image10K<n<100K0 likes583 downloads2y agoHugging Face11matthew-l-leidos /ParseBench ParseBench Quick links: [🌐 Website] [📜 Paper] [💻 Code] ParseBench is a benchmark for evaluating document parsing systems on real-world enterprise documents, with the following characteristics: Multi-dimensional evaluation. The benchmark is stratified into five capability dimensions — tables, charts, content faithfulness, semantic formatting, and visual grounding — each with task-specific metrics designed to capture what agentic workflows depend on. Real-world enterprise… See the full description on the dataset page: https://huggingface.co/datasets/matthew-l-leidos/ParseBench.document100K<n<1M0 likes567 downloads3mo agoHugging Face12leideng /LeoRAG LeoRAG LeoRAG is a parquet-backed dataset format for RAG tasks. Each row stores a query, optional preamble, up to 10 retrieved documents, the gold answer, optional teacher outputs, and metadata. Schema Column Type Description query string Query prompt for the RAG task preamble string Prompt shown before the documents num_of_docs int Number of retrieved documents (0-10) doc1 ... doc10 string Retrieved documents (empty when unused) answer string… See the full description on the dataset page: https://huggingface.co/datasets/leideng/LeoRAG.textquestion-answering100K<n<1M0 likes459 downloads3mo agoHugging Face13leibler /airline airline An executable Environment for tool-using agents, rebuilt from traces by Kullback and published by Leibler: the world, the Tasks and a code Verifier per Task, no recordings. tasks.jsonl lists the Tasks. Release: replays at least 90% of its Tasks. Fidelity over Tasks 100.0% (130 of 130) Fidelity over Runs 99.5% (199 of 200) Call fidelity 99.93% of 1513 Reference confirmed 130 Verifier derived 128 Trusted 82 Refused 0 Not trusted 48 of 130; 17 the… See the full description on the dataset page: https://huggingface.co/datasets/leibler/airline.textn<1K2 likes395 downloads22h agoHugging Face14cointegrated /ru-paraphrase-NMT-Leipzig Dataset Card for cointegrated/ru-paraphrase-NMT-Leipzig Dataset Summary The dataset contains 1 million Russian sentences and their automatically generated paraphrases. It was created by David Dale (@cointegrated) by translating the rus-ru_web-public_2019_1M corpus from the Leipzig collection into English and back into Russian. A fraction of the resulting paraphrases are invalid, and should be filtered out. The blogpost "Перефразирование русских текстов: корпуса, модели… See the full description on the dataset page: https://huggingface.co/datasets/cointegrated/ru-paraphrase-NMT-Leipzig.text-generation100K<n<1M12 likes382 downloads4y agoHugging Face15DenisaBumba /htr_leibniz_dataset_v1 Dataset Card for Leibniz's Manuscripts (HTR Ground Truth) This dataset is composed of trascribed and manually corrected folios of Gottfried Wilhelm Leibniz's manuscripts, together with automatically aligned ground truth in order to train or fine-tune Handwritten Text Recognition (HTR) models. Dataset Details This ground truth was produced to fine-tune existing HTR models for the recognition of Leibniz's handwriting, with the aim of assisting scholars in the… See the full description on the dataset page: https://huggingface.co/datasets/DenisaBumba/htr_leibniz_dataset_v1.imageimage-to-text10K<n<100K0 likes381 downloads23d agoHugging Face16Nille1991 /LeitliniendatenbankIn dieser Datenbank werden alle AWMF Leitlinien hinterlegt, die die Deutsche Gesellschaft für Orthopädie und Unfallchirurgie (DGOU) erstellt hat. documentn<1K0 likes369 downloads3y agoHugging Face17leibler /retail retail An executable Environment for tool-using agents, rebuilt from traces by Kullback and published by Leibler: the world, the Tasks and a code Verifier per Task, no recordings. tasks.jsonl lists the Tasks. Release: replays at least 90% of its Tasks. Fidelity over Tasks 100.0% (223 of 223) Fidelity over Runs 100.0% (456 of 456) Call fidelity 100.00% of 3220 Reference confirmed 223 Verifier derived 222 Trusted 193 Refused 0 Not trusted 30 of 223; 15… See the full description on the dataset page: https://huggingface.co/datasets/leibler/retail.textn<1K2 likes355 downloads22h agoHugging Face18leigangqu /MSE-Bench MSE-Bench: A Benchmark for Multi-turn Session Image Editing Introduction MSE-Bench (Multi-turn Session image Editing Benchmark) is a benchmark designed to evaluate multi-turn image editing systems under realistic editing workflows. Given a source image and a series of editing instructions, the goal is for a model to apply these edits cumulatively to produce a final image that reflects all the requested changes. MSE-Bench consists of 100 test instances, each representing… See the full description on the dataset page: https://huggingface.co/datasets/leigangqu/MSE-Bench.imagen<1K2 likes306 downloads6mo agoHugging Face19logiover /gleif-lei-scraper-sample-data GLEIF LEI Scraper Scrape global legal entities from the official GLEIF LEI database — no login, no API key, no blocking. 3.3M+ entities worldwide with legal name, address, jurisdiction, legal form, status and registration data. Filter by country and status. Tens of thousands per run. What the actor scrapes 🏛️ GLEIF LEI Scraper — Global Legal Entity Identifier Data to JSON/CSV/Excel Scrape global legal entities straight from the official GLEIF API — the worldwide… See the full description on the dataset page: https://huggingface.co/datasets/logiover/gleif-lei-scraper-sample-data.textn<1K0 likes291 downloads4mo agoHugging Face20leitro /Copiale_Lines Copiale Lines Copiale Lines is a line-level image-to-text dataset for historical cipher decipherment. It contains cropped line images from the Copiale manuscript paired with plaintext ground truth. This dataset was presented in the paper Learning to Decipher from Pixels -- A Case Study of Copiale (HistoCrypt 2026). Code: https://github.com/leitro/Decipher-from-Pixels-Copiale Dataset Structure The dataset is split into: train: 1,269 samples valid: 175 samples… See the full description on the dataset page: https://huggingface.co/datasets/leitro/Copiale_Lines.imageimage-to-text1K<n<10K0 likes230 downloads5mo agoHugging Face21wsagi /leisaac-real-pick-orange SO-101 Real-Arm Pick Orange — 30 Episodes (LeRobot v3.0) 真机 SO-101 主从臂遥操采集的抓橙子放盘子数据集,30 集人类演示,LeRobot v3.0 格式。 Real-world SO-101 leader-follower teleoperation dataset: pick up an orange and place it on the plate, 30 human demonstrations in LeRobot v3.0 format. 概览 / Overview 项 值 机器人 / Robot SO-101 follower(6 DoF + 夹爪,leader 臂遥操) 任务 / Task Pick up the orange and place it on the plate 集数 / Episodes 30 总帧数 / Frames 25,091(30 fps,合计 ~14 min) 单集长度 /… See the full description on the dataset page: https://huggingface.co/datasets/wsagi/leisaac-real-pick-orange.tabularrobotics10K<n<100K0 likes230 downloads2mo agoHugging Face22csoai /registry-harvest-xrpl-mica-lei Free-registry harvest — XRPL issuers × MiCA × LEI Public registers and public ledgers joined, read 2026-09-04T07:12:30Z. Every source is free, keyless and re-runnable by anyone. No part of this needed a relationship, an API key, or anyone's permission. The finding Of the 16 XRPL issued assets in the CSOAI reader, 5 join to a MiCA e-money-token authorisation. group n declares an on-chain domain enforces allowlisting retains freeze capability… See the full description on the dataset page: https://huggingface.co/datasets/csoai/registry-harvest-xrpl-mica-lei.tabularothern<1K0 likes228 downloads11d agoHugging Face23imvladikon /leipzig_corpora_collection Leipzig Corpora Collection The Leipzig Corpora Collection presents corpora in different languages using the same format and comparable sources. All data are available as plain text files and can be imported into a MySQL database by using the provided import script. They are intended both for scientific use by corpus linguists as well as for applications such as knowledge extraction programs. The corpora are identical in format and similar in size and content. They contain randomly… See the full description on the dataset page: https://huggingface.co/datasets/imvladikon/leipzig_corpora_collection.texttext-generation1K<n<10K4 likes227 downloads5mo agoHugging Face24zalizedata /global-lei-company-registry-dataset Global LEI Company Registry (GLEIF) The complete GLEIF golden copy as analysis-ready tables: 3.4M legal entities with registered addresses, corporate hierarchy relationships and reporting exceptions. Part of the DataForge Open Data program — full production packages, free for academic and personal use. Canonical dataset page: https://data.zalize.com/datasets/global-lei-company-registry-dataset Packages in this repo Package Tier Rows Size SHA-256… See the full description on the dataset page: https://huggingface.co/datasets/zalizedata/global-lei-company-registry-dataset.tabulartabular-classification10M<n<100M0 likes222 downloads1mo agoHugging Face25lei-qi-233 /MicroG-4M MicroG-4M Dataset This repository stores the entire content of the MicroG-4M dataset itself. For more information and details, including training, evaluation, statistics, and related code, please: Refer to our paper Visit our GitHub And check our fine-tuned models Specification of MicroG-4M "annotation_files" Folder The folder contains all annotation files of the dataset, all stored in CSV format. actions.csv contains all the… See the full description on the dataset page: https://huggingface.co/datasets/lei-qi-233/MicroG-4M.documentvideo-classification100K<n<1M1 likes217 downloads6mo agoHugging Face26cluesurf /leipzig-frequency Leipzig Corpora Frequency Data Word frequency lists and co-occurrence data from the Leipzig Corpora Collection, converted to Parquet. Covers hundreds of languages across news, web, Wikipedia, and mixed sources. Each corpus includes token frequencies, source provenance, and statistical co-occurrence pairs. Contents base/ <language>/ <source>-<date>-<size>/ metadata.json string.0001.parquet source.0001.parquet cooccurrence.sentence.0001.parquet… See the full description on the dataset page: https://huggingface.co/datasets/cluesurf/leipzig-frequency.0 likes201 downloads5mo agoHugging Face27leideng /longbench-view Introduction LongBench is the first benchmark for bilingual, multitask, and comprehensive assessment of long context understanding capabilities of large language models. LongBench includes different languages (Chinese and English) to provide a more comprehensive evaluation of the large models' multilingual capabilities on long contexts. In addition, LongBench is composed of six major categories and twenty one different tasks, covering key long-text application scenarios such as… See the full description on the dataset page: https://huggingface.co/datasets/leideng/longbench-view.textquestion-answering1K<n<10K0 likes194 downloads6mo agoHugging Face28Toby0614 /leisaac-pick-orange-mimic-v0This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 60, "total_frames": 41891, "total_tasks": 1, "total_videos": 120, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:60" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Toby0614/leisaac-pick-orange-mimic-v0.tabularrobotics10K<n<100K0 likes184 downloads8mo agoHugging Face29leiwx52 /AssistGUI_web1 likes178 downloads2y agoHugging Face30leigangqu /MSE-Bench-resultsimage1K<n<10K3 likes172 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.