lid
Datasets
All datasets matching “lid”real_lid_hdf5real_lid_hdf5fragbench
FragBench (Public Tier)
Anonymous submission for NeurIPS 2026 Datasets and Benchmarks Track.
Author identity will be revealed at camera-ready.
Dataset Summary
FragBench is a benchmark for evaluating cross-session, fragmented attacks on
LLM agents that use tools via the Model Context Protocol (MCP). Each campaign
is decomposed into many small fragments distributed across sessions; a
defender must reconstruct the compositional intent. The public tier in this
repository… See the full description on the dataset page: https://huggingface.co/datasets/LidaSafety/fragbench.EEG_Image_decode
EEG Image Decode — Dataset and Checkpoints
This dataset accompanies the NeurIPS 2024 paper:
Visual Decoding and Reconstruction via EEG Embeddings with Guided Diffusion Dongyang Li · Chen Wei · Shiying Li · Jiachen Zou · Quanying Liu
It packages the preprocessed EEG recordings, stimulus-image visual features, VAE latent codes, trained EEG embeddings, fine-tuned checkpoints, and generated images needed to reproduce both the image retrieval and image reconstruction experiments.… See the full description on the dataset page: https://huggingface.co/datasets/LidongYang/EEG_Image_decode.open-lid-datasetThis dataset is built from the open source data accompanying "An Open Dataset and Model for Language Identification" (Burchell et al., 2023)
The repository containing the actual data can be found here : https://github.com/laurieburchell/open-lid-dataset.
The license for this recreation itself follows the original upstream dataset as GPLv3+.
However, individual datasets within it follow each of their own licenses.
The "src" column lists the sources. "lang" column lists the language code in… See the full description on the dataset page: https://huggingface.co/datasets/hac541309/open-lid-dataset.open-lid-dataset
Dataset Card for "open-lid-dataset"
Dataset Summary
The OpenLID dataset covers 201 languages and is designed for training language identification models. The majority of the source datasets were derived from news sites, Wikipedia, or religious text, though some come from other domains (e.g. transcribed conversations, literature, or social media). A sample of each language in each source was manually audited to check it was in the attested language (see the paper) for full… See the full description on the dataset page: https://huggingface.co/datasets/laurievb/open-lid-dataset.
