datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Guided-Lensless-Polarization-Imaging-stageI
Stage I reconstructions for evaluation (UPLight & ZJU-RGB-P)
This dataset accompanies Guided Lensless Polarization Imaging (CVPR 2026 Findings). It provides FISTA Stage-I outputs, RGB guidance images, and ground-truth polarization stacks for two public evaluation sets used in the paper.
Layout
UPLight (~1,991 scenes)
Folder
Description
UPLight/fista_grayscale/
3-channel grayscale polarization FISTA reconstructions
UPLight/fista_color/… See the full description on the dataset page: https://huggingface.co/datasets/noakraicer/Guided-Lensless-Polarization-Imaging-stageI.lensless_mic_librispeech
Dataset Card for LenslessMic Version of Librispeech Dataset
Dataset Summary
A LenslessMic version of the Librispeech dataset from the
"LenslessMic: Audio Encryption and Authentication via Lensless Computational Imaging" paper.
Partition
# Audio
# Frames
train-clean
587
73,699
train-other
150
18,561
test-clean
1,089
185,773
test-other
512
62,901
To download the dataset and work with it, use our official repository.
Dataset is collected using DigiCam.… See the full description on the dataset page: https://huggingface.co/datasets/Blinorot/lensless_mic_librispeech.bio-lens
🌿 iNaturalist Bronze Dataset (Research-Grade, Deduplicated)
Description
This dataset contains research-grade observations from iNaturalist, processed through a bronze-layer pipeline that includes:
It's a curated dataset for biodiversity analysis based on community sourced observations, contains millions of images thus can be used for biology model training.
The intent of this data is to train specialist vision models capable of identifying species with high… See the full description on the dataset page: https://huggingface.co/datasets/HirakoSan/bio-lens.refusal-lens-graphsastro_lenslogit-lens-responses-llama-70b-layer50lensless_mic_random
Dataset Card for LenslessMic Version of N(0,1) Random Dataset
Dataset Summary
A LenslessMic version of the N(0,1) random images dataset from the
"LenslessMic: Audio Encryption and Authentication via Lensless Computational Imaging" paper.
The dataset can be used to train a codec-agnostic reconstruction algorithm.
Partition
# Audio
# Frames
train
200
30000
Note: We split dataset into 200 files, however, there are no actual audio files. Only frames are used.… See the full description on the dataset page: https://huggingface.co/datasets/Blinorot/lensless_mic_random.rl-future-lens-qwen3-8b
RL Future Lens on Qwen3-8B
Artifacts for the experiment in docs/future_lens.md of https://github.com/syvb/EasyNLA
(branch sv/future-rl): an activation verbalizer that reads out future tokens from a single
hidden state of a frozen Qwen3-8B, trained by SFT and then GRPO, with shuffled-activation controls.
Metrics live in the wandb project rl-future-lens; this repo holds the files.
Layout
path
what
data/
positions filtered to top-1-correct next-token… See the full description on the dataset page: https://huggingface.co/datasets/syvb/rl-future-lens-qwen3-8b.MultiLens-Mirflickr-AmbientSimilar to this dataset, but this dataset has a different external illumination distribution between train and test set.
jacobian-lens-transfer-qwen36-38
Jacobian Lens Transfer: Qwen3.6-27B → Qwen3.8-27B
Evaluation data and code for the preprint "Survival of the Fitted:
Cross-Checkpoint Transfer of a Jacobian Lens from Qwen3.6-27B to
Qwen3.8-27B" (Shorthill, 2026, v1.1; Zenodo DOI
10.5281/zenodo.22133736).
A published Jacobian lens fitted on Qwen3.6-27B
(neuronpedia/jacobian-lens,
revision qwen-n1000) is applied unmodified to Qwen3.8-27B. Tuned-lens
translators are known to transfer from a base model to its fine-tuned
variants;… See the full description on the dataset page: https://huggingface.co/datasets/ec75hash/jacobian-lens-transfer-qwen36-38.MultiLens-Mirflickr-Ambient-Same-DistributionSimilar to this dataset, but this dataset has the same external illumination distribution between train and test set.
DiffuserCam-Lensless-Mirflickr-Dataset
For future training, it is recommended to use this normalized version of the dataset.
More accessible (6GB instead of 100GB) copy of: https://waller-lab.github.io/LenslessLearning/dataset.html
Original license: https://github.com/Waller-Lab/LenslessLearning/blob/master/LICENSE
This dataset was prepared with this script.
After cloning and installing LenslessPiCam, ADMM reconstruction can be applied to the dataset with this script (handles dataset downloading from Hugging Face).… See the full description on the dataset page: https://huggingface.co/datasets/bezzam/DiffuserCam-Lensless-Mirflickr-Dataset.sim_two_lens_black_tubeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "unity",
"total_episodes": 100,
"total_frames": 25113,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/JackySunUofT/sim_two_lens_black_tube.lensless
Dataset Lensless Freshwater Plankton Dataset
A mixture of ten freshwater plankton species imaged with a lensless microscope designed for in situ data collection.
Original dataset available online at: https://ibm.ent.box.com/v/PlanktonData.
Original dataset license: <cc-by-4.0>.
Details
train split means (RGB): [0.4140324021223468, 0.43283753018640264, 0.4201652068819248]
train split standard deviations (RGB): [0.12368459767878866, 0.1250351586616661… See the full description on the dataset page: https://huggingface.co/datasets/project-oceania/lensless.ica-lens-paper
ICA Lens Paper Artifacts
This dataset stores public artifacts for the paper ICA Lens: Interpreting Language Models Without Training Another Dictionary.
Project Page: https://liusida.github.io/ica-lens-paper/
Code: https://github.com/liusida/ica-lens-paper
Interactive Demo: ICA Explorer Space
Introduction
ICA Lens is a practical workflow for stable, efficient, and auditable Independent Component Analysis (ICA) of language model representations. It recovers… See the full description on the dataset page: https://huggingface.co/datasets/sida/ica-lens-paper.oracle-lens-qwen3-8b-artifactslensing_dr6_growthlens-network-traffic
Lens Network Traffic Classification Benchmark
Downstream network-traffic classification data used to evaluate Lens, a knowledge-guided
foundation model for network traffic (TMLR). It bundles the 12 classification tasks
from the Lens paper as HuggingFace dataset configurations, each with train / validation /
test splits and a unified schema.
ℹ️ All tasks are derived from publicly available academic traffic datasets
obtained via the NetBench benchmark (Qian et al., 2024); the… See the full description on the dataset page: https://huggingface.co/datasets/Charles59/lens-network-traffic.persona-society-jacobian-lens
Persona Society Sim
Social town simulator where 30–300 activation-steered LLM agents live, converse, collaborate, and produce emergent dynamics. Agents use persona steering vectors (Contrastive Activation Addition) instead of prompt-only roles, combining Smallville-inspired memory loops with lightweight world mechanics and measurement harnesses.
Project scope
Build an open-world town loop (observation → reflection → planning → action) with a scheduler, social… See the full description on the dataset page: https://huggingface.co/datasets/dyllonj/persona-society-jacobian-lens.lensemble-phase3-so100-silos
Lensemble Phase 3 — SO-100 consortium silos + held-out eval split
Deterministic episode-modulo split (k % 5) of the public SO-100 pick-place set
abdelstark/so100-pickplace-lewm-ready
into four sovereign participant silos plus one disjoint held-out eval split, for the Phase 3
federated JEPA / LeWorldModel consortium run (Lensemble epic #249, issue #242).
file
role
episodes
windows (window_steps=4)
dataset Merkle root (sha256, prefix)
phase3-so100-silo0.h5
participant… See the full description on the dataset page: https://huggingface.co/datasets/abdelstark/lensemble-phase3-so100-silos.lens-loss-grokking-experiment
lens-loss-grokking-experiment — checkpoints
Final model checkpoints for the experiments in
brendanlong/lens-loss-grokking-experiment
(deep supervision vs grokking: LN-scoped chronic instability, weight-decay
circuit pruning, and failure isolation). Training curves:
public wandb project.
Layout: grok_lens/<run_name>/final.pt (modular-arithmetic runs, ~1.7 MB
each, with .json metadata sidecars) and lego/<run_name>/step_*.pt
(S3 multi-hop composition runs). Run names encode the… See the full description on the dataset page: https://huggingface.co/datasets/brendanlong/lens-loss-grokking-experiment.UI_lens
Dataset Card for UI-Lens
Dataset Details
UI-Lens is a benchmark designed for evaluating UI display defect detection in commercial mobile applications.
The Chinese version of the dataset comprises 4,566 meticulously annotated UI pages from 11 leading Chinese commercial applications. Unlike existing UI benchmarks focused on clean interfaces, UI-Lens specifically targets challenging real-world scenarios including display rendering issues, boundary understanding… See the full description on the dataset page: https://huggingface.co/datasets/wuhuohua/UI_lens.LENS-WarBias
LENS-WarBias
Version 1.3 — research draft; independent human validation pending.
LENS-WarBias is a Ukrainian–English prompt dataset for studying war-related stereotype elicitation and transfer after model unlearning. It covers 981 WarBias matrix entries, 129 case families, 15 actor profiles and 56 actor/gender/age variants. It contains prompts and provenance metadata, not target-model responses, a validated forget set, or measured model scores.
The dataset contains deliberately… See the full description on the dataset page: https://huggingface.co/datasets/FairForget/LENS-WarBias.movie_lens_small_latestDiffuserCam-Lensless-Mirflickr-Dataset-NORMSame as this dataset but with images normalized to full 8-bit range. For future models, better to train with this dataset.
More accessible copy (6GB instead of 100GB) of: https://waller-lab.github.io/LenslessLearning/dataset.html
Original license: https://github.com/Waller-Lab/LenslessLearning/blob/master/LICENSE
pacbenchPaper: https://arxiv.org/abs/2506.23725
Download
Grab the full dataset as a single tarball and expand it locally:
wget https://huggingface.co/datasets/lens-lab/pacbench/resolve/main/pacbench-20250218.tar.gz
tar -xzf pacbench-20250218.tar.gz
After extraction you will get the pacbench/ directory with the same structure used on the Hub (constraint images, humanoid captures, open images, RoboCasa objects, metadata, etc.).
If you prefer the classic Hub workflow, the original files… See the full description on the dataset page: https://huggingface.co/datasets/lens-lab/pacbench.stocks-LENSKART-1D-candlessequential-transformer-lens-experiment
Checkpoints for Training a Transformer to Compose One Step Per Layer (and Proving It)
Final checkpoints behind the writeup
Training a Transformer to Compose One Step Per Layer (and Proving It).
Code, analyses and the full experiment log:
github.com/brendanlong/sequential-transformer-lens-experiment.
Training curves: public wandb project.
Every checkpoint is a torch.save dict {"step", "model_state_dict", "model_config"} loadable with torch.load(..., weights_only=True); each has a… See the full description on the dataset page: https://huggingface.co/datasets/brendanlong/sequential-transformer-lens-experiment.mirflickr-fza-25kdescriptors-text-davinci-003
Dataset Card for "descriptors-text-davinci-003"
More Information needed
