CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01allenai /lilaLīla is a comprehensive benchmark for mathematical reasoning with over 140K natural language questions annotated with Python programs and natural language instructions. The data set comes with multiple splits: Līla-IID (train, dev, test), Līla-OOD (train, dev, test), and Līla-Robust.text100K<n<1M33 likes8.5k downloads4y agoHugging Face02li-lab /MMLU-ProX MMLU-ProX MMLU-ProX is a multilingual benchmark that builds upon MMLU-Pro, extending to 29 typologically diverse languages, designed to evaluate large language models' reasoning capabilities across linguistic and cultural boundaries. Github | Paper News [2025/08] 🎉 MMLU-ProX was accepted by EMNLP 2025 Main Conference! [2025/05] MMLU-ProX now contains 29 languages, all available on Huggingface. [2025/03] MMLU-ProX is now available on Huggingface. [2025/03] We are still… See the full description on the dataset page: https://huggingface.co/datasets/li-lab/MMLU-ProX.tabular100K<n<1M19 likes6.5k downloads1y agoHugging Face03li-lab /MMLU-ProX-Lite MMLU-ProX-Lite MMLU-ProX is a multilingual benchmark that builds upon MMLU-Pro, extending to 29 typologically diverse languages, designed to evaluate large language models' reasoning capabilities across linguistic and cultural boundaries. Github | Paper News [2025/08] 🎉 MMLU-ProX was accepted by EMNLP 2025 Main Conference! [2025/05] MMLU-ProX now contains 29 languages, all available on Huggingface. [2025/03] MMLU-ProX is now available on Huggingface. [2025/03] We are… See the full description on the dataset page: https://huggingface.co/datasets/li-lab/MMLU-ProX-Lite.tabular10K<n<100K3 likes4.9k downloads1y agoHugging Face04harvard-lil /cold-cases Collaborative Open Legal Data (COLD) - Cases COLD Cases is a dataset of 8.3 million United States legal decisions with text and metadata, formatted as compressed parquet files. If you'd like to view a sample of the dataset formatted as JSON Lines, you can view one here This dataset exists to support the open legal movement exemplified by projects like Pile of Law and LegalBench. A key input to legal understanding projects is caselaw -- the published, precedential decisions of… See the full description on the dataset page: https://huggingface.co/datasets/harvard-lil/cold-cases.tabular100K<n<1M42 likes4.1k downloads2y agoHugging Face05Yale-LILY /aeslc Dataset Card for "aeslc" Dataset Summary A collection of email messages of employees in the Enron Corporation. There are two features: email_body: email body text. subject_line: email subject text. Supported Tasks and Leaderboards More Information Needed Languages Monolingual English (mainly en-US) with some exceptions. Dataset Structure Data Instances default Size of downloaded dataset files: 11.64 MB Size of the… See the full description on the dataset page: https://huggingface.co/datasets/Yale-LILY/aeslc.textsummarization10K<n<100K19 likes1.6k downloads3y agoHugging Face06Lilambd /world-signals World Signals — a daily cross-country snapshot of attention One folder per day under data/YYYY-MM-DD/, and the same files copied to latest/. Built every morning (JST) by the EmpireOS world model. Nothing is generated by a model; every row is a measurement from a public source. file what source search_trends.csv rising searches, 30 countries, with approximate traffic and the headline that drove them Google Trends daily RSS podcast_charts.csv top-100 podcasts, 30… See the full description on the dataset page: https://huggingface.co/datasets/Lilambd/world-signals.tabular1K<n<10K0 likes1.1k downloads10h agoHugging Face07lilacai /glaive-function-calling-v2-sharegpt Dataset Card for "glaive-function-calling-v2-sharegpt" This dataset takes the glaive/glaive-function-calling-v2 dataset and formats it with ShareGPT using Lilac The accompanying notebook can be found here. The original columns "system" and "chat" still exist on the dataset. There are 4 types of roles in the ShareGPT format: system user human function call The original dataset has a column called 'chat' with the following structure: USER: Hi, I need help with calculating a tip. My… See the full description on the dataset page: https://huggingface.co/datasets/lilacai/glaive-function-calling-v2-sharegpt.text100K<n<1M31 likes961 downloads3y agoHugging Face08glouriousgautam /lilm1-pretrain-mix-32b LiLM Experiment 3 pretraining corpus Private research corpus with 32,000,010,072 globally exact-deduplicated train tokens plus 328,933,246 held-out tokens. Data are stored as EOS-delimited little-endian uint16 binaries with aligned Parquet provenance. This repository combines ODC-By FinePDFs-Edu, CC-BY-4.0 DCLM, ODC-By SmolLM/FineMath sources, Apache-2.0 UltraX subject to its upstream terms, per-file permissively licensed Stack-Edu code subject to The Stack v2 terms, StarCoder2… See the full description on the dataset page: https://huggingface.co/datasets/glouriousgautam/lilm1-pretrain-mix-32b.tabulartext-generation10M<n<100M0 likes874 downloads2mo agoHugging Face09lilgatouwu /microsoftexcelimagen<1K0 likes803 downloads3y agoHugging Face10OpenSakura /OpenSakura-DS-260220-LN-ja-zh-COT-Lilith OpenSakura Lilith LN COT Dataset OpenSakura-DS-260220-LN-ja-zh-COT-Lilith is the COT/segment-level derivative built from the same LN source stream, with reasoning_content preserved. Stats below are computed from the actual generated parquet files. Dataset Summary Metric Value Dataset ID OpenSakura/OpenSakura-DS-260220-LN-ja-zh-COT-Lilith Total rows 692,587 Total parquet files 233 (train: 162, arena: 12, reserve: 12, validation: 24, test: 23) Total size 8… See the full description on the dataset page: https://huggingface.co/datasets/OpenSakura/OpenSakura-DS-260220-LN-ja-zh-COT-Lilith.tabulartranslation100K<n<1M2 likes618 downloads4mo agoHugging Face11lilyzhng /tb2-synthetic-train tb2-synthetic-train In-distribution Terminal-Bench 2.0 synthetic task variants (graded hint/seed ops that reuse each base task's Docker env + verifier), validity-gated by the AfterQuery TB2 RLVR harness. Ingest with Harbor's stock prepare_harbor_dataset.py --dataset lilyzhng/tb2-synthetic-train (parquet: path + task_binary tar archives). See manifest.json for validity/band per task. textn<1K0 likes392 downloads4mo agoHugging Face12li-lab /HLE-BioMedX HLE-BioMedX — Multilingual HLE Biology/Medicine A multilingual version of the Biology/Medicine subset of Humanity's Last Exam (HLE), released as one subset per language. Source benchmark: Humanity's Last Exam, dataset cais/hle. Subsets Group Languages How the target-language text was produced Source en Original English questions and answers. Machine-translated and expert-verified / revised zh, ja, ko, fr, th Machine translation reviewed by a human… See the full description on the dataset page: https://huggingface.co/datasets/li-lab/HLE-BioMedX.textquestion-answering1K<n<10K0 likes382 downloads26d agoHugging Face13glouriousgautam /lilm2-training-datatabular10M<n<100M0 likes364 downloads7d agoHugging Face14harvard-lil /cold-french-law Collaborative Open Legal Data (COLD) - French Law COLD French Law is a dataset containing over 800 000 french law articles, filtered and extracted from France's LEGI dataset and formatted as a single CSV file. This dataset focuses on articles (codes, lois, décrets, arrêtés ...) identified as currently applicable french law. A large portion of this dataset comes with machine-generated english translations, provided by Casetext, Part of Thomson Reuters using OpenAI's GPT-4. This… See the full description on the dataset page: https://huggingface.co/datasets/harvard-lil/cold-french-law.tabular100K<n<1M21 likes351 downloads2y agoHugging Face15lilywchen /lucky-initialization-atlas-100m-v2 Lucky initialization atlas v2 evidence Private live evidence archive for lilywchen/lucky-initialization-atlas-100m-v2. It contains hash-bound configs, provenance, scalar trajectories, step-zero diagnostics, and final per-sequence losses after those artifacts complete. It excludes credentials, caches, raw FineWeb-derived token arrays, optimizer states, and W&B binary logs. tabulartext-generationn<1K0 likes330 downloads27d agoHugging Face16lilyecho /phhi_train_poseimageimage100K<n<1M0 likes316 downloads8mo agoHugging Face17li-lab /HealMed HealMed (Human-verified Evaluation Across Languages for Medical AI) is a multilingual medical dataset featuring expert-verified translations for benchmarking multilingual medical AI systems. The dataset comprises translations from two complementary sources. A portion is based on the multilingual translations released by the GlobMed project (arXiv: 2601.02186), while the remainder was generated by our team using zero-shot machine translation to expand language coverage. Each translated… See the full description on the dataset page: https://huggingface.co/datasets/li-lab/HealMed.textquestion-answering10K<n<100K2 likes277 downloads8d agoHugging Face18lilyzhng /tb2-real-train tb2-real-train The 70 REAL (unmodified) Terminal-Bench 2.0 train tasks for the AfterQuery TB2 RLVR GRPO run, in the same parquet (path + task_binary tar archives) format as lilyzhng/tb2-synthetic-train. Ingest with Harbor's stock prepare_harbor_dataset.py --dataset lilyzhng/tb2-real-train. textn<1K0 likes232 downloads4mo agoHugging Face19lilonghao /MM-ContextASR-Bench MM-ContextASR Bench Metadata and evaluation splits for Multimodal Conversational Context for LLM-Based ASR: Data Construction, Training, and Benchmark. Dataset summary Config Examples Audio Context Primary metric mm_contextasr 1,250 (250 current utterances × 5 histories) 1,439 WAV files included Controlled user-assistant dialogue entity Recall kespeech 19,212 Source ID only Same-speaker speech and transcript CER, SER, entity Recall cv_yue 3,525… See the full description on the dataset page: https://huggingface.co/datasets/lilonghao/MM-ContextASR-Bench.audioautomatic-speech-recognition10K<n<100K1 likes226 downloads7d agoHugging Face20lili24 /UniSVG UniSVG Dataset UniSVG is a comprehensive dataset designed for unified SVG generation (from textual prompts and images) and SVG understanding (color, category, usage, etc.). It comprises 525k data items tailored for Multi-modal Large Language Models (MLLM) training and evaluation. 🔥 Release [2025/11/27] 🔥 We are glad to announce that our UniSVG benchmark is used by Qwen3-VL! [2025/09/22] 🔥 Qwen2.5-VL-finetuned released! 🌐 Model Path!… See the full description on the dataset page: https://huggingface.co/datasets/lili24/UniSVG.image100K<n<1M11 likes210 downloads9mo agoHugging Face21golamrob /khurushkul-pond-water-lily-sample Pink Water Lily & Water Hyacinth — Khurushkul Pond, Bangladesh 100 GPS-tagged freshwater wetland images from a single pond survey in Khurushkul, Cox's Bazar, Bangladesh. By Golam Rob — www.golamrob.com ✅ Free to use, including commercially — just credit "Golam Rob (golamrob.com)". Licensed CC BY 4.0. Use it, train on it, remix it, share it. All I ask is attribution. 📸 These 100 images are a small taste of a 200,000+ image personal library of coastal, tidal, and freshwater… See the full description on the dataset page: https://huggingface.co/datasets/golamrob/khurushkul-pond-water-lily-sample.imageimage-classificationn<1K1 likes189 downloads2mo agoHugging Face22Lilambd /world-seeds World Seeds — every "by country" table, keyed by ISO 3166-1 alpha-2 Wikipedia has hundreds of "... by country" articles. The numbers live inside article tables, keyed by country names that differ from article to article. This dataset re-keys every such table to ISO2 so they join. One CSV per source article under tables/. Columns: iso2, country, <original column names>. Values are kept exactly as printed (*_num twin columns hold the parsed number where one could be read).… See the full description on the dataset page: https://huggingface.co/datasets/Lilambd/world-seeds.tabular1K<n<10K0 likes167 downloads10h agoHugging Face23lilacai /lilac-TruthfulQA-MultipleChoice lilac/TruthfulQA-MultipleChoice This dataset is a Lilac processed dataset. Original dataset: https://huggingface.co/datasets/truthful_qa To download the dataset to a local directory: lilac download lilacai/lilac-TruthfulQA-MultipleChoice or from python with: ll.download("lilacai/lilac-TruthfulQA-MultipleChoice") text1K<n<10K2 likes166 downloads3y agoHugging Face24lilgoose777 /tibetan-speech-english-text-dataset-new-updatedaudio1K<n<10K0 likes166 downloads8mo agoHugging Face25lilgoose7777 /indicvoices-r-nepaligated IndicVoices-R — Nepali Subset This is the Nepali (ne) language subset of ai4bharat/indicvoices_r, extracted and re-uploaded as a standalone dataset for convenience. Dataset information Total hours: 104.31 hours (6258.7 minutes, 375525 seconds) Source Original dataset: ai4bharat/indicvoices_r Original paper: IndicVoices-R: Unlocking a Massive Multilingual Multi-speaker Speech Corpus for Scaling Indian TTS (NeurIPS 2024) License: CC-BY-4.0 (inherited… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose7777/indicvoices-r-nepali.audiotext-to-speech10K<n<100K0 likes150 downloads15d agoHugging Face26lilvjosephtang /SEAM-Benchmark SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models CSSLab, Department of Computer Science, University of Toronto[COLM '25] Second Conference on Language Modeling Paper: Paper Project Page / Leaderboard: SEAM Benchmark Code: GitHub Abstract Evaluating whether vision-language models (VLMs) reason consistently across representations is challenging because modality comparisons are typically confounded by task differences and asymmetric… See the full description on the dataset page: https://huggingface.co/datasets/lilvjosephtang/SEAM-Benchmark.imageimage-text-to-text1K<n<10K9 likes149 downloads1y agoHugging Face27lilyecho /phhi_shard_12_demo Dataset Card for "phhi_shard_12_demo" More Information needed image1K<n<10K0 likes130 downloads9mo agoHugging Face28lilacai /lilac-TruthfulQA-Generation lilac/TruthfulQA-Generation This dataset is a Lilac processed dataset. Original dataset: https://huggingface.co/datasets/truthful_qa To download the dataset to a local directory: lilac download lilacai/lilac-TruthfulQA-Generation or from python with: ll.download("lilacai/lilac-TruthfulQA-Generation") text1K<n<10K0 likes118 downloads3y agoHugging Face29glouriousgautam /lilm1-paper1-ratio-controls-12m-v1 LiLM1 Paper 1 ratio controls, 12M v1 Replicate 5 uses schedule seed 20260906. It contains three matched 12M variants from shared deterministic pools: 12M ordinary; 8M ordinary + 4M tool; and 4M ordinary + 8M tool. All trainers initialize from glouriousgautam/LiLM1-230M-base at 5391c31c741fc7256ffef6580657a8190888754d. tabular10K<n<100K0 likes118 downloads22d agoHugging Face30lilyecho /lily_controlnet_pose_datasetimage10K<n<100K0 likes104 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.