datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lambada_openai
Dataset Summary
This dataset is comprised of the LAMBADA test split as pre-processed by OpenAI (see relevant discussions here and here). It also contains machine translated versions of the split in German, Spanish, French, and Italian.
LAMBADA is used to evaluate the capabilities of computational models for text understanding by means of a word prediction task. LAMBADA is a collection of narrative texts sharing the characteristic that human subjects are able to guess their last word… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/lambada_openai.lambada
Dataset Card for LAMBADA
Dataset Summary
The LAMBADA evaluates the capabilities of computational models
for text understanding by means of a word prediction task.
LAMBADA is a collection of narrative passages sharing the characteristic
that human subjects are able to guess their last word if
they are exposed to the whole passage, but not if they
only see the last sentence preceding the target word.
To succeed on LAMBADA, computational models cannot
simply rely on local… See the full description on the dataset page: https://huggingface.co/datasets/cimec/lambada.pokemon-blip-captions
Notice of DMCA Takedown Action
We have received a DMCA takedown notice from The Pokémon Company International, Inc.
In response to this action, we have taken down the dataset.
We appreciate your understanding.
Cobot_Magic_turn_on_the_desk_lamp
Cobot_Magic_turn_on_the_desk_lamp
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: agilex_cobot_decoupled_magic
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
pressbutton
📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Cobot_Magic_turn_on_the_desk_lamp.Cobot_Magic_turn_off_the_desk_lamp
Cobot_Magic_turn_off_the_desk_lamp
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: agilex_cobot_decoupled_magic
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
pressbutton
📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Cobot_Magic_turn_off_the_desk_lamp.m_lamaExtension/Modification of the original m_lama dataset
LAMDA
LAMDA: A Longitudinal Android Malware Dataset for Drift Analysis
This dataset contains a longitudinal benchmark for Android malware detection designed to analyze and evaluate concept drift in machine learning models. It includes labeled and feature-engineered Android APK data from 2013 to 2025 (excluding 2015), with over 1 million samples collected from real-world sources.
Dataset Details
LAMDA is the largest and most temporally diverse Android malware dataset to date. It… See the full description on the dataset page: https://huggingface.co/datasets/IQSeC-Lab/LAMDA.hermes-agent-reasoning-traces
Hermes Agent Reasoning Traces
Multi-turn tool-calling trajectories for training AI agents using the Hermes Agent harness. Each sample is a real agent conversation with step-by-step reasoning (<think> blocks) and actual tool execution results.
This dataset has two configs, one per source model:
Config
Model
Samples
kimi
Moonshot AI Kimi-K2.5
7,646
glm-5.1
ZhipuAI GLM-5.1-FP8
7,055
Loading
from datasets import load_dataset
# Kimi-K2.5 traces
ds =… See the full description on the dataset page: https://huggingface.co/datasets/lambda/hermes-agent-reasoning-traces.LAMDA
LAMDA: A Longitudinal Android Malware Dataset for Drift Analysis
This dataset contains a longitudinal benchmark for Android malware detection designed to analyze and evaluate concept drift in machine learning models. It includes labeled and feature-engineered Android APK data from 2013 to 2025 (excluding 2015), with over 1 million samples collected from real-world sources.
Dataset Details
LAMDA is the largest and most temporally diverse Android malware dataset to date. It… See the full description on the dataset page: https://huggingface.co/datasets/PtaSack/LAMDA.lamini_docs
Dataset Card for "lamini_docs"
More Information needed
lamoda-fashion-product-images
High-Resolution Fashion Product Images
This dataset is a highly optimized, high-resolution subset of the popular Fashion Product Images Dataset originally hosted on Kaggle.
It contains thousands of unique e-commerce fashion products, combining high-resolution product images with multiple descriptive label attributes.
All low-resolution thumbnails and anomalies have been aggressively filtered out. Every image in this dataset has a minimum resolution of 640px on its shortest… See the full description on the dataset page: https://huggingface.co/datasets/PestoRosso/lamoda-fashion-product-images.medical_advice_dialogue_en
Description
The dataset is from medalpaca/medical_meadow_health_advice, formatted as dialogues for speed and ease of use. Many thanks to author for releasing it.
Importantly, this format is easy to use via the default chat template of transformers, meaning you can use huggingface/alignment-handbook immediately, unsloth.
Structure
View online through viewer.
Note
We advise you to reconsider before use, thank you. If you find it useful, please like and… See the full description on the dataset page: https://huggingface.co/datasets/lamhieu/medical_advice_dialogue_en.LaMini-instruction
Dataset Card for "LaMini-Instruction"
Minghao Wu, Abdul Waheed, Chiyu Zhang, Muhammad Abdul-Mageed, Alham Fikri Aji,
Dataset Description
We distill the knowledge from large language models by performing sentence/offline distillation (Kim and Rush, 2016). We generate a total of 2.58M pairs of instructions and responses using gpt-3.5-turbo based on several existing resources of prompts, including self-instruct (Wang et al., 2022), P3 (Sanh et al., 2022), FLAN… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/LaMini-instruction.bird_text_to_sql
Dataset Card for "bird_text_to_sql"
More Information needed
ChinaTravel-Sandbox
ChinaTravel Sandbox Environment Database
This dataset is licensed under Creative Commons Attribution 4.0 International (CC BY 4.0).
English | 简体中文
Release version: 2026.08.2
English
This dataset contains the bilingual static sandbox used by
ChinaTravel. It is a companion to
the ChinaTravel query dataset
and an artifact of the
ChinaTravel paper.
The raw ZIP snapshots preserve the exact directory layout expected by the
ChinaTravel evaluator. Viewer-friendly Parquet… See the full description on the dataset page: https://huggingface.co/datasets/LAMDA-NeSy/ChinaTravel-Sandbox.LAMDA
LAMDA: A Longitudinal Android Malware Dataset for Drift Analysis
This dataset contains a longitudinal benchmark for Android malware detection designed to analyze and evaluate concept drift in machine learning models. It includes labeled and feature-engineered Android APK data from 2013 to 2025 (excluding 2015), with over 1 million samples collected from real-world sources.
Dataset Details
LAMDA is the largest and most temporally diverse Android malware dataset to date. It… See the full description on the dataset page: https://huggingface.co/datasets/Yanyi10086/LAMDA.spider_text_to_sql
Dataset Card for "spider_text_to_sql"
More Information needed
LAMES
Dataset Card for Dataset Name
A mining site dataset to identify mining sites on rocky / deserty / non-vegetated land surfaces. The dataset depicts large scale mining sites of the country of chile.
Dataset Details
Dataset Description
A mining site dataset to identify mining sites on rocky / deserty / non-vegetated land surfaces. The dataset depicts large scale mining sites of the country of chile.
Curated by: Matthias Kahl (https://github.com/maduschek)… See the full description on the dataset page: https://huggingface.co/datasets/maduschek/LAMES.Franka_pick_up_the_beaker_and_place_it_on_alcohol_lamp_0317LAMBDA
Dataset Summary
LAMDBA is a long term ad memorability dataset, featuring data from 1749 participants and 2205 ads across 276 brands.
Dataset Structure
from datasets import load_dataset
ds = load_dataset("behavior-in-the-wild/LAMBDA")
ds
DatasetDict({
train: Dataset({
features: ['video_id', 'recall_score', 'youtube_id', 'ad_details'],
num_rows: 1964
})
test: Dataset({
features: ['video_id', 'recall_score', 'youtube_id', 'ad_details']… See the full description on the dataset page: https://huggingface.co/datasets/behavior-in-the-wild/LAMBDA.rlbench_lamp_offThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "panda",
"total_episodes": 100,
"total_frames": 8081,
"total_tasks": 5,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/iAyoD/rlbench_lamp_off.lambda-corpus
lambda-corpus
lambda-corpus is a Japanese text corpus for language model pretraining. It combines openly available Japanese datasets into a consistent format and provides predefined train, validation, and test splits.
Purpose
The dataset is intended for pretraining of Japanese language models. It contains web documents, Wikipedia-derived text, academic grant records, and synthetic question-answer text.
Source Data
Source
Rows
Tokens
License… See the full description on the dataset page: https://huggingface.co/datasets/KeisukeMiyamoto/lambda-corpus.bird_spider_train_text_to_sql
Dataset Card for "bird_spider_train_text_to_sql"
More Information needed
UniRef50_512_alltcga-laml-tabular-open
TCGA-LAML — Tabular (Open Access)
Open-access TCGA-LAML data from the NCI Genomic Data Commons, reshaped into one table per GDC data_type. Clinical, biospecimen and every open molecular modality for this cohort, in one place, queryable without downloading a single .tar or parsing a single TSV.
GDC data release: Data Release 46.0 - August 10, 2026
Built: 2026-09-12 04:02:09 UTC
Scope: one TCGA project — see [the family][repo] for the others
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/gabrielaltay/tcga-laml-tabular-open.lambada_openai_de
LAMBADA (DE) — Boldt German Evaluation Suite
A modernized German translation of the LAMBADA benchmark (Paperno et al., 2016), part of the Boldt German Evaluation Suite.
LAMBADA tests a model's ability to track discourse-level context. Each instance consists of a passage where the final word can only be predicted correctly if the model has understood the broader narrative — it cannot be inferred from the final sentence alone. The target word is always the last token of the passage.… See the full description on the dataset page: https://huggingface.co/datasets/Boldt/lambada_openai_de.medical_medqa_dialogue_en
Description
The dataset is from medalpaca/medical_meadow_mediqa, formatted as dialogues for speed and ease of use. Many thanks to author for releasing it.
Importantly, this format is easy to use via the default chat template of transformers, meaning you can use huggingface/alignment-handbook immediately, unsloth.
Structure
View online through viewer.
Note
We advise you to reconsider before use, thank you. If you find it useful, please like and follow… See the full description on the dataset page: https://huggingface.co/datasets/lamhieu/medical_medqa_dialogue_en.Franka_pick_up_the_beaker_and_place_it_on_alcohol_lamp_0316medical_wikidoc_dialogue_en
Description
The dataset is from medalpaca/medical_meadow_wikidoc, formatted as dialogues for speed and ease of use. Many thanks to author for releasing it.
Importantly, this format is easy to use via the default chat template of transformers, meaning you can use huggingface/alignment-handbook immediately, unsloth.
Structure
View online through viewer.
Note
We advise you to reconsider before use, thank you. If you find it useful, please like and follow… See the full description on the dataset page: https://huggingface.co/datasets/lamhieu/medical_wikidoc_dialogue_en.lambada_multilingual_stablelm
