datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ChinaTravel
ChinaTravel Query Dataset
This dataset is licensed under Creative Commons Attribution 4.0 International (CC BY 4.0).
ChinaTravel is an open-ended travel-planning benchmark with compositional
constraint validation for language agents. See the
paper,
Hugging Face paper page,
code, and
bilingual sandbox database
(ModelScope mirror)
for the complete benchmark resources.
Introduction
For a given query, a language agent uses the sandbox tools to collect
information and… See the full description on the dataset page: https://huggingface.co/datasets/LAMDA-NeSy/ChinaTravel.IRFL
Dataset Card for IRFL
Dataset Description
Leaderboards
Colab notebook code for IRFL evaluation
Languages
Dataset Structure
Data Fields
Dataset Creation
Considerations for Using the Data
Licensing Information
Citation Information
Dataset Description
The IRFL dataset consists of idioms, similes, metaphors with matching figurative and literal images, and two novel tasks of multimodal figurative detection and retrieval.Using human annotation and an automatic pipeline… See the full description on the dataset page: https://huggingface.co/datasets/lampent/IRFL.CausalArena
CausalArena public release
This repository contains the public CausalArena dataset release: executable SCMs, selected result tables, and real-data source indices.
What is included
scm/: the public half of each generated SCM family: 500 synthetic SCM configurations, 50 semantic SCMs, and 50 formula-grounded SCMs. Released SCMs include both observation-only and observation-plus-intervention exports.
scm/{semantic,formula}/artifacts/: per-scenario graph, generator… See the full description on the dataset page: https://huggingface.co/datasets/LAMDA-Tabular/CausalArena.ridgelora-cross-sensor-sd302d-f-to-m-20260825
RidgeLoRA-FP: SD302A-F to SD302D-M cross-sensor experiment
This public archive contains the leakage-controlled direct cross-sensor
experiment used to evaluate whether Stage-2 synthetic target-sensor images
help recognition on a physically different real sensor.
Locked protocol
Source/condition sensor: NIST SD302A device F.
Target sensor: NIST SD302D device M.
Identity: subject:finger-position; the same fingers exist across both
collections.
Subject split: 160… See the full description on the dataset page: https://huggingface.co/datasets/LamTNguyen/ridgelora-cross-sensor-sd302d-f-to-m-20260825.FashionGEN_images_datax-lora-datasetPaper, see: arxiv.org/abs/2402.07148
MechanicsMaterialsBioinspired3Dtourism-package-prediction-datasetIdiomology_Lama2_7B_Chat
Dataset Card for "Idiomolgy - Idiom Detection Dataset"
Dataset Description
General Information
This dataset is created for the purpose of training and evaluating language models, particularly focusing on their ability to identify idiomatic expressions within sentences. It aims to improve natural language understanding systems in recognizing idioms in varied contexts.
Usage Example
from datasets import load_dataset
# Load the dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/SnehitVaddi/Idiomology_Lama2_7B_Chat.lamus-roberts-court-legal-arguments
LAMUS: Roberts Court Legal Arguments (2005-2025)
The Current Supreme Court Era - Chief Justice John Roberts
📋 Dataset Description
This dataset contains 362,891 sentences from U.S. Supreme Court opinions during the Roberts Court era (2005-2025), automatically labeled with legal argument categories. This represents the current Supreme Court under Chief Justice John G. Roberts Jr.
Why Roberts Court?
The Roberts Court is particularly significant for… See the full description on the dataset page: https://huggingface.co/datasets/LavanyaPobbathi/lamus-roberts-court-legal-arguments.lama-trex-no-duplicateLifeGPT-broad-entropy@article{berkovich2024lifegpt,
title={LifeGPT: Topology-Agnostic Generative Pretrained Transformer Model for Cellular Automata},
author={Jaime A. Berkovich and Markus J. Buehler},
year={2024},
eprint={2409.12182},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2409.12182},
}
compressed_chipseqlamus-scotus-legal-arguments
LAMUS: Legal Argument Mining from U.S. Supreme Court
📋 Dataset Description
This dataset contains 2,900,083 sentences from U.S. Supreme Court opinions spanning 1921-2025, automatically labeled with legal argument categories. This is the largest publicly available labeled dataset for legal argument mining from U.S. caselaw.
🎯 Purpose
The dataset enables:
Legal Argument Mining research
Legal Text Classificationmodel training
Temporal Analysis of… See the full description on the dataset page: https://huggingface.co/datasets/LavanyaPobbathi/lamus-scotus-legal-arguments.Lamini-instructions-to-frenchlama-trex-counterfactual-gpt-gen-10shotVietnam-higher-education-lawLambada_Norwegianhealth-faqLifeGPT-high-entropy@article{berkovich2024lifegpt,
title={LifeGPT: Topology-Agnostic Generative Pretrained Transformer Model for Cellular Automata},
author={Jaime A. Berkovich and Markus J. Buehler},
year={2024},
eprint={2409.12182},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2409.12182},
}
Synthetic-Translations-Lamini-120k-Permutation-Set-UnvalidatedAsclepius-Synthetic-Clinical-Notes
Asclepius: Synthetic Clincal Notes & Instruction Dataset
Dataset Summary
This dataset is official dataset for Asclepius (arxiv)
This dataset is composed with Clinical Note - Question - Answer format to build a clinical LLMs.
We first synthesized synthetic notes from PMC-Patients case reports with GPT-3.5
Then, we generate instruction-answer pairs for 157k synthetic discharge summaries
Supported Tasks
This dataset covers below 8 tasks
Named Entity… See the full description on the dataset page: https://huggingface.co/datasets/LampsteR/Asclepius-Synthetic-Clinical-Notes.lama-trex-counterfactual-10shotdata_sampleseera_events
Seera Events Dataset
A multilingual dataset of historical events from the life of Prophet Muhammad ﷺ (Seerah), covering the period from 571 CE to 632 CE.
Dataset Description
This dataset contains 142 historical events translated into three languages: Arabic, English, and French. Each event includes detailed descriptions, dates in both Hijri and Gregorian calendars, and geographical coordinates.
Source
Data extracted from Dorar.net Historical Encyclopedia, a… See the full description on the dataset page: https://huggingface.co/datasets/Lamarer/seera_events.lamini_mistral
Lamini Mistral
Lamini Docs dataset formatted for fine-tuning with Mistral-7B Instruct model
Wikipedia-Euskeralama_demo_datademo data
LAMx-2.2
