datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lilaLīla is a comprehensive benchmark for mathematical reasoning with over 140K natural language questions annotated with Python programs and natural language instructions. The data set comes with multiple splits: Līla-IID (train, dev, test), Līla-OOD (train, dev, test), and Līla-Robust.MMLU-ProX
MMLU-ProX
MMLU-ProX is a multilingual benchmark that builds upon MMLU-Pro, extending to 29 typologically diverse languages, designed to evaluate large language models' reasoning capabilities across linguistic and cultural boundaries.
Github | Paper
News
[2025/08] 🎉 MMLU-ProX was accepted by EMNLP 2025 Main Conference!
[2025/05] MMLU-ProX now contains 29 languages, all available on Huggingface.
[2025/03] MMLU-ProX is now available on Huggingface.
[2025/03] We are still… See the full description on the dataset page: https://huggingface.co/datasets/li-lab/MMLU-ProX.MMLU-ProX-Lite
MMLU-ProX-Lite
MMLU-ProX is a multilingual benchmark that builds upon MMLU-Pro, extending to 29 typologically diverse languages, designed to evaluate large language models' reasoning capabilities across linguistic and cultural boundaries.
Github | Paper
News
[2025/08] 🎉 MMLU-ProX was accepted by EMNLP 2025 Main Conference!
[2025/05] MMLU-ProX now contains 29 languages, all available on Huggingface.
[2025/03] MMLU-ProX is now available on Huggingface.
[2025/03] We are… See the full description on the dataset page: https://huggingface.co/datasets/li-lab/MMLU-ProX-Lite.cold-cases
Collaborative Open Legal Data (COLD) - Cases
COLD Cases is a dataset of 8.3 million United States legal decisions with text and metadata, formatted as compressed parquet files. If you'd like to view a sample of the dataset formatted as JSON Lines, you can view one here
This dataset exists to support the open legal movement exemplified by projects like
Pile of Law and
LegalBench.
A key input to legal understanding projects is caselaw -- the published, precedential decisions of… See the full description on the dataset page: https://huggingface.co/datasets/harvard-lil/cold-cases.aeslc
Dataset Card for "aeslc"
Dataset Summary
A collection of email messages of employees in the Enron Corporation.
There are two features:
email_body: email body text.
subject_line: email subject text.
Supported Tasks and Leaderboards
More Information Needed
Languages
Monolingual English (mainly en-US) with some exceptions.
Dataset Structure
Data Instances
default
Size of downloaded dataset files: 11.64 MB
Size of the… See the full description on the dataset page: https://huggingface.co/datasets/Yale-LILY/aeslc.world-signals
World Signals — a daily cross-country snapshot of attention
One folder per day under data/YYYY-MM-DD/, and the same files copied to latest/.
Built every morning (JST) by the EmpireOS world model. Nothing is generated by a model; every row is a measurement from a public source.
file
what
source
search_trends.csv
rising searches, 30 countries, with approximate traffic and the headline that drove them
Google Trends daily RSS
podcast_charts.csv
top-100 podcasts, 30… See the full description on the dataset page: https://huggingface.co/datasets/Lilambd/world-signals.glaive-function-calling-v2-sharegpt
Dataset Card for "glaive-function-calling-v2-sharegpt"
This dataset takes the glaive/glaive-function-calling-v2 dataset and formats it with ShareGPT using Lilac
The accompanying notebook can be found here.
The original columns "system" and "chat" still exist on the dataset.
There are 4 types of roles in the ShareGPT format:
system
user
human
function call
The original dataset has a column called 'chat' with the following structure:
USER: Hi, I need help with calculating a tip. My… See the full description on the dataset page: https://huggingface.co/datasets/lilacai/glaive-function-calling-v2-sharegpt.lilm1-pretrain-mix-32b
LiLM Experiment 3 pretraining corpus
Private research corpus with 32,000,010,072 globally
exact-deduplicated train tokens plus 328,933,246 held-out
tokens. Data are stored as EOS-delimited little-endian uint16 binaries with
aligned Parquet provenance.
This repository combines ODC-By FinePDFs-Edu, CC-BY-4.0 DCLM, ODC-By
SmolLM/FineMath sources, Apache-2.0 UltraX subject to its upstream terms,
per-file permissively licensed Stack-Edu code subject to The Stack v2 terms,
StarCoder2… See the full description on the dataset page: https://huggingface.co/datasets/glouriousgautam/lilm1-pretrain-mix-32b.microsoftexcelOpenSakura-DS-260220-LN-ja-zh-COT-Lilith
OpenSakura Lilith LN COT Dataset
OpenSakura-DS-260220-LN-ja-zh-COT-Lilith is the COT/segment-level derivative built from the same LN source stream, with reasoning_content preserved.
Stats below are computed from the actual generated parquet files.
Dataset Summary
Metric
Value
Dataset ID
OpenSakura/OpenSakura-DS-260220-LN-ja-zh-COT-Lilith
Total rows
692,587
Total parquet files
233 (train: 162, arena: 12, reserve: 12, validation: 24, test: 23)
Total size
8… See the full description on the dataset page: https://huggingface.co/datasets/OpenSakura/OpenSakura-DS-260220-LN-ja-zh-COT-Lilith.tb2-synthetic-train
tb2-synthetic-train
In-distribution Terminal-Bench 2.0 synthetic task variants (graded hint/seed ops that reuse each base task's Docker env + verifier), validity-gated by the AfterQuery TB2 RLVR harness.
Ingest with Harbor's stock prepare_harbor_dataset.py --dataset lilyzhng/tb2-synthetic-train (parquet: path + task_binary tar archives). See manifest.json for validity/band per task.
HLE-BioMedX
HLE-BioMedX — Multilingual HLE Biology/Medicine
A multilingual version of the Biology/Medicine subset of Humanity's Last Exam
(HLE), released as one subset per language.
Source benchmark: Humanity's Last Exam, dataset
cais/hle.
Subsets
Group
Languages
How the target-language text was produced
Source
en
Original English questions and answers.
Machine-translated and expert-verified / revised
zh, ja, ko, fr, th
Machine translation reviewed by a human… See the full description on the dataset page: https://huggingface.co/datasets/li-lab/HLE-BioMedX.lilm2-training-datacold-french-law
Collaborative Open Legal Data (COLD) - French Law
COLD French Law is a dataset containing over 800 000 french law articles, filtered and extracted from France's LEGI dataset and formatted as a single CSV file.
This dataset focuses on articles (codes, lois, décrets, arrêtés ...) identified as currently applicable french law.
A large portion of this dataset comes with machine-generated english translations, provided by Casetext, Part of Thomson Reuters using OpenAI's GPT-4.
This… See the full description on the dataset page: https://huggingface.co/datasets/harvard-lil/cold-french-law.lucky-initialization-atlas-100m-v2
Lucky initialization atlas v2 evidence
Private live evidence archive for lilywchen/lucky-initialization-atlas-100m-v2. It contains hash-bound configs,
provenance, scalar trajectories, step-zero diagnostics, and final per-sequence losses
after those artifacts complete. It excludes
credentials, caches, raw FineWeb-derived token arrays, optimizer states, and W&B
binary logs.
phhi_train_poseimageHealMed
HealMed (Human-verified Evaluation Across Languages for Medical AI) is a multilingual medical dataset featuring expert-verified translations for benchmarking multilingual medical AI systems.
The dataset comprises translations from two complementary sources. A portion is based on the multilingual translations released by the GlobMed project (arXiv: 2601.02186), while the remainder was generated by our team using zero-shot machine translation to expand language coverage. Each translated… See the full description on the dataset page: https://huggingface.co/datasets/li-lab/HealMed.tb2-real-train
tb2-real-train
The 70 REAL (unmodified) Terminal-Bench 2.0 train tasks for the AfterQuery TB2 RLVR GRPO run, in the same parquet (path + task_binary tar archives) format as lilyzhng/tb2-synthetic-train. Ingest with Harbor's stock prepare_harbor_dataset.py --dataset lilyzhng/tb2-real-train.
MM-ContextASR-Bench
MM-ContextASR Bench
Metadata and evaluation splits for Multimodal Conversational Context for
LLM-Based ASR: Data Construction, Training, and Benchmark.
Dataset summary
Config
Examples
Audio
Context
Primary metric
mm_contextasr
1,250 (250 current utterances × 5 histories)
1,439 WAV files included
Controlled user-assistant dialogue
entity Recall
kespeech
19,212
Source ID only
Same-speaker speech and transcript
CER, SER, entity Recall
cv_yue
3,525… See the full description on the dataset page: https://huggingface.co/datasets/lilonghao/MM-ContextASR-Bench.UniSVG
UniSVG Dataset
UniSVG is a comprehensive dataset designed for unified SVG generation (from textual prompts and images) and SVG understanding (color, category, usage, etc.). It comprises 525k data items tailored for Multi-modal Large Language Models (MLLM) training and evaluation.
🔥 Release
[2025/11/27]
🔥 We are glad to announce that our UniSVG benchmark is used by Qwen3-VL!
[2025/09/22]
🔥 Qwen2.5-VL-finetuned released! 🌐 Model Path!… See the full description on the dataset page: https://huggingface.co/datasets/lili24/UniSVG.khurushkul-pond-water-lily-sample
Pink Water Lily & Water Hyacinth — Khurushkul Pond, Bangladesh
100 GPS-tagged freshwater wetland images from a single pond survey in Khurushkul, Cox's Bazar, Bangladesh. By Golam Rob — www.golamrob.com
✅ Free to use, including commercially — just credit "Golam Rob (golamrob.com)". Licensed CC BY 4.0. Use it, train on it, remix it, share it. All I ask is attribution.
📸 These 100 images are a small taste of a 200,000+ image personal library of coastal, tidal, and freshwater… See the full description on the dataset page: https://huggingface.co/datasets/golamrob/khurushkul-pond-water-lily-sample.world-seeds
World Seeds — every "by country" table, keyed by ISO 3166-1 alpha-2
Wikipedia has hundreds of "... by country" articles. The numbers live inside article tables, keyed by country names that differ from article to article. This dataset re-keys every such table to ISO2 so they join.
One CSV per source article under tables/. Columns: iso2, country, <original column names>. Values are kept exactly as printed (*_num twin columns hold the parsed number where one could be read).… See the full description on the dataset page: https://huggingface.co/datasets/Lilambd/world-seeds.lilac-TruthfulQA-MultipleChoice
lilac/TruthfulQA-MultipleChoice
This dataset is a Lilac processed dataset. Original dataset: https://huggingface.co/datasets/truthful_qa
To download the dataset to a local directory:
lilac download lilacai/lilac-TruthfulQA-MultipleChoice
or from python with:
ll.download("lilacai/lilac-TruthfulQA-MultipleChoice")
tibetan-speech-english-text-dataset-new-updatedindicvoices-r-nepali
IndicVoices-R — Nepali Subset
This is the Nepali (ne) language subset of ai4bharat/indicvoices_r,
extracted and re-uploaded as a standalone dataset for convenience.
Dataset information
Total hours: 104.31 hours (6258.7 minutes, 375525 seconds)
Source
Original dataset: ai4bharat/indicvoices_r
Original paper: IndicVoices-R: Unlocking a Massive Multilingual Multi-speaker Speech Corpus for Scaling Indian TTS (NeurIPS 2024)
License: CC-BY-4.0 (inherited… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose7777/indicvoices-r-nepali.SEAM-Benchmark
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models
CSSLab, Department of Computer Science, University of Toronto[COLM '25] Second Conference on Language Modeling
Paper: Paper
Project Page / Leaderboard: SEAM Benchmark
Code: GitHub
Abstract
Evaluating whether vision-language models (VLMs) reason consistently across representations is challenging because modality comparisons are typically confounded by task differences and asymmetric… See the full description on the dataset page: https://huggingface.co/datasets/lilvjosephtang/SEAM-Benchmark.phhi_shard_12_demo
Dataset Card for "phhi_shard_12_demo"
More Information needed
lilac-TruthfulQA-Generation
lilac/TruthfulQA-Generation
This dataset is a Lilac processed dataset. Original dataset: https://huggingface.co/datasets/truthful_qa
To download the dataset to a local directory:
lilac download lilacai/lilac-TruthfulQA-Generation
or from python with:
ll.download("lilacai/lilac-TruthfulQA-Generation")
lilm1-paper1-ratio-controls-12m-v1
LiLM1 Paper 1 ratio controls, 12M v1
Replicate 5 uses schedule seed 20260906. It contains three matched 12M variants from shared deterministic pools: 12M ordinary; 8M ordinary + 4M tool; and 4M ordinary + 8M tool. All trainers initialize from glouriousgautam/LiLM1-230M-base at 5391c31c741fc7256ffef6580657a8190888754d.
lily_controlnet_pose_dataset
