datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LeMat-Bulk-MLIP-Hull
LeMat-Bulk MLIP Hull Reference Datasets
This dataset contains materials close to the convex hull computed using various ML interatomic potentials (MLIPs).
Dataset Splits
all: Contains ALL materials with hull energies for all MLIPs (no threshold filtering)
dft, orb, uma, mace_mp, mace_omat: Materials within 0.001 eV/atom of respective hulls
Energy Types
dft: DFT reference energies
orb: ORB model energies
uma: UMA model energies
mace_mp: MACE-MP model energies… See the full description on the dataset page: https://huggingface.co/datasets/LeMaterial/LeMat-Bulk-MLIP-Hull.leminh2002leminhtri07LEMUR
EU Law Dataset – Category 15.10: Environment
This dataset contains official legal documents from the European Union, collected from the EUR-Lex website, specifically under category 15.10: "Environment". The documents span from the year 1961 to 2025 and are provided in multiple European "languages. The original documents are in PDF format and have been converted into various text-based formats using OLMCR.
The dataset splits represent the different "languages available for each… See the full description on the dataset page: https://huggingface.co/datasets/G4KMU/LEMUR.leminhquan1995leminhtam2003LEMAS-Dataset-train
Overview
This dataset is part of LEMAS-Project (lemas-project.github.io/LEMAS-Project).
It contains a large-scale training set (150k+ hours) and a curated evaluation set
(500 utterances per language) covering 10 languages, all with word-level alignment.
Fields
key: unique utterance identifier; the first two characters indicate the language ID
audio: relative path to the MP3 audio file (in the eval set, this key is renamed to "file_name" for compatibility with the viewer)… See the full description on the dataset page: https://huggingface.co/datasets/LEMAS-Project/LEMAS-Dataset-train.lemyhanh08leMinhKhang1997leminhc1987LeMat-BulkMotivation: check out the blog post https://huggingface.co/blog/lematerial to hear more about the motivation behind the creation of this dataset.
Changelog:
2025.04.17 (hash: NOT YET RELEASED):
We have changed the Yb default pseudopotential to Yb_3 from VASP, this is the same that Materials Project uses. In the previous version we had kept it as Yb, and Materials Project had to Yb-containing materials. Alexandria and OQMD uses Yb. As a result no Yb-containing materials are… See the full description on the dataset page: https://huggingface.co/datasets/LeMaterial/LeMat-Bulk.LEMON
📚 Paper - 🤖 GitHub - 🌐 Website
In this repository, we provide the full FPS LEMON videos in 📚 LEMON: A Large Endoscopic MONocular Dataset and Foundation Model for Perception in Surgical Settings.
For more details, please visit our github repository at 🤖 GitHub).
If you use our dataset, model, or code in your research, please cite our paper:
@InProceedings{Che_2026_CVPR_LEMON,
author = {Che, Chengan and Wang, Chao and Vercauteren, Tom and Tsoka, Sophia and… See the full description on the dataset page: https://huggingface.co/datasets/visurg/LEMON.leminhtuan1992s2orc_small
Dataset Card for "s2orc_small"
A small split of the s2orc dataset, includes ~900k english papers with abstract included.
See all detailes in the original dataset card - https://huggingface.co/datasets/allenai/s2orc
lemurLeMat-TrajNote: For PBE we are in the process of providing a precomputed energy_corrected scheme based on Materials Project 2020 Compatibility Scheme
Motivation: check out the blog post https://huggingface.co/blog/lematerial to hear more about the motivation behind the creation of our datasets.
Download and use within Python
from datasets import load_dataset
dataset = load_dataset('LeMaterial/LeMat-Traj', 'compatible_pbe')
Data fields
Feature name
Data type
Description… See the full description on the dataset page: https://huggingface.co/datasets/LeMaterial/LeMat-Traj.Roleplay-Forums_2023-04Important noticeUpon reanalysis, it appears that spaces between adjacent HTML tags in these scrapes got mangled, a problem that was also present in the files of the Rentry where they first got uploaded; this occurred during an intermediate step in the conversion process from the original source files, most of which unfortunately do not exist anymore. So, the usefulness of the data will be diminished. Some scrapes without these issues have been provided on a different dataset page.… See the full description on the dataset page: https://huggingface.co/datasets/lemonilia/Roleplay-Forums_2023-04.lemusetslldms-associative-memory-samples
LLDMs Associative Memory — Generated Samples
Model-generated text for the paper:
Language Diffusion Models are Associative Memories Capable of Retrieving Unseen Data
Bao Pham, Mohammed J. Zaki, Luca Ambrogioni, Dmitry Krotov, Matteo Negri
Accepted to EMNLP 2026 (Main Conference).
arXiv:2604.26841 · paper · code · checkpoints
29.5 million generated sequences (~3.8B tokens) sampled from the released checkpoints — one
generation run per (model size, training-set fraction). These… See the full description on the dataset page: https://huggingface.co/datasets/lemoncmd/lldms-associative-memory-samples.roleplaying-forums-raw
Roleplaying forum scrapes (raw)
Here are mostly original/raw files for some of the roleplaying forums I scraped in the past (and some newly scraped ones), repacked as HTML strings + some metadata on a one-row-per-thread basis instead of a one-row-per-message basis, which should make them more convenient to handle.
Unlike the previously uploaded archive, they shouldn't have issues with spaces between adjacent HTML tags, as that occurred by mistake in an intermediate processing step… See the full description on the dataset page: https://huggingface.co/datasets/lemonilia/roleplaying-forums-raw.lemonAgilex_Cobot_Magic_storage_lemon_mango
Agilex_Cobot_Magic_storage_lemon_mango
Dataset Description
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Task Preview
View Video Directly
Overview
Total Episodes: 100
Total Frames: 28424
FPS: 30
Dataset Size: 362.05 MB
Robot Name: Agilex_Cobot_Magic
End-Effector Type: two_finger_gripper
Teleoperation Type: Due to some reasons, this dataset temporarily cannot provide the teleoperation type… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Agilex_Cobot_Magic_storage_lemon_mango.EasyJailbreak_Datasetslemexp-task1-v2ERA5_SRThis dataset contains modified Copernicus Climate Change Service information [2026].
Source: ERA5 reanalysis data from the Copernicus Climate Data Store (CDS/ECMWF).
Licensed under CC BY 4.0, consistent with the source dataset terms.
Dataset Description
ERA5 reanalysis data prepared for 4x super-resolution training. The dataset provides paired low-resolution (LR) and high-resolution (HR) atmospheric fields covering 2010–2020.
Variables
Name
Description
Units… See the full description on the dataset page: https://huggingface.co/datasets/Lemoni/ERA5_SR.lemexp-task1lemmanaid-afp-reruns
Lemmanaid AFP-pool Reproducibility Reruns
Reproducibility study for claude-opus-4-5 on the yalhessi/lemexp-commerical-llm-experiment benchmark, using an AFP demo pool (honest eval — no train/test theory leakage).
Companion to ggranberry/lemmanaid-commercial-results, which holds the earlier shot-count + retrieval sweeps under test-LOO.
Configs
Two configs, one per benchmark domain:
Config
Source HF config
Test rows
octonions
template_octonions_2026… See the full description on the dataset page: https://huggingface.co/datasets/ggranberry/lemmanaid-afp-reruns.Zuco2.0Airbot_MMK2_storage_lemon_mango
Airbot_MMK2_storage_lemon_mango
Dataset Description
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Task Preview
View Video Directly
Overview
Total Episodes: 49
Total Frames: 7685
FPS: 30
Dataset Size: 201.65 MB
Robot Name: Airbot_MMK2
End-Effector Type: five_finger_gripper
Teleoperation Type: Due to some reasons, this dataset temporarily cannot provide the teleoperation type information.
Sensors:… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Airbot_MMK2_storage_lemon_mango.lememory
Cover the building block with a cup, then lift up the cup covering the block.(robokit 采集)
每批一个目录:hdf5/ 原始数据、robokit_dataset/ RLDS、source_meta/ 采集配置与清洗报告。
批次
episode
段数
Cover_the_building_block_with_a_cup,_then_lift_up_the_cup_covering_the_block._b0
0..156
157
Cover_the_building_block_with_a_cup,_then_lift_up_the_cup_covering_the_block._b1
157..300
144
文件列表按字典序显示(100.hdf5 排在 66.hdf5 前面),翻页目测容易误判成缺文件。
核对完整性用 ./scripts/publish_batch.sh --audit "Cover the building… See the full description on the dataset page: https://huggingface.co/datasets/shaohuan1/lememory.
