datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lensless_mic_librispeech
Dataset Card for LenslessMic Version of Librispeech Dataset
Dataset Summary
A LenslessMic version of the Librispeech dataset from the
"LenslessMic: Audio Encryption and Authentication via Lensless Computational Imaging" paper.
Partition
# Audio
# Frames
train-clean
587
73,699
train-other
150
18,561
test-clean
1,089
185,773
test-other
512
62,901
To download the dataset and work with it, use our official repository.
Dataset is collected using DigiCam.… See the full description on the dataset page: https://huggingface.co/datasets/Blinorot/lensless_mic_librispeech.bio-lens
🌿 iNaturalist Bronze Dataset (Research-Grade, Deduplicated)
Description
This dataset contains research-grade observations from iNaturalist, processed through a bronze-layer pipeline that includes:
It's a curated dataset for biodiversity analysis based on community sourced observations, contains millions of images thus can be used for biology model training.
The intent of this data is to train specialist vision models capable of identifying species with high… See the full description on the dataset page: https://huggingface.co/datasets/HirakoSan/bio-lens.lensless_mic_random
Dataset Card for LenslessMic Version of N(0,1) Random Dataset
Dataset Summary
A LenslessMic version of the N(0,1) random images dataset from the
"LenslessMic: Audio Encryption and Authentication via Lensless Computational Imaging" paper.
The dataset can be used to train a codec-agnostic reconstruction algorithm.
Partition
# Audio
# Frames
train
200
30000
Note: We split dataset into 200 files, however, there are no actual audio files. Only frames are used.… See the full description on the dataset page: https://huggingface.co/datasets/Blinorot/lensless_mic_random.oracle-lens-qwen3-8b-artifactslens-network-traffic
Lens Network Traffic Classification Benchmark
Downstream network-traffic classification data used to evaluate Lens, a knowledge-guided
foundation model for network traffic (TMLR). It bundles the 12 classification tasks
from the Lens paper as HuggingFace dataset configurations, each with train / validation /
test splits and a unified schema.
ℹ️ All tasks are derived from publicly available academic traffic datasets
obtained via the NetBench benchmark (Qian et al., 2024); the… See the full description on the dataset page: https://huggingface.co/datasets/Charles59/lens-network-traffic.persona-society-jacobian-lens
Persona Society Sim
Social town simulator where 30–300 activation-steered LLM agents live, converse, collaborate, and produce emergent dynamics. Agents use persona steering vectors (Contrastive Activation Addition) instead of prompt-only roles, combining Smallville-inspired memory loops with lightweight world mechanics and measurement harnesses.
Project scope
Build an open-world town loop (observation → reflection → planning → action) with a scheduler, social… See the full description on the dataset page: https://huggingface.co/datasets/dyllonj/persona-society-jacobian-lens.LENS-WarBias
LENS-WarBias
Version 1.3 — research draft; independent human validation pending.
LENS-WarBias is a Ukrainian–English prompt dataset for studying war-related stereotype elicitation and transfer after model unlearning. It covers 981 WarBias matrix entries, 129 case families, 15 actor profiles and 56 actor/gender/age variants. It contains prompts and provenance metadata, not target-model responses, a validated forget set, or measured model scores.
The dataset contains deliberately… See the full description on the dataset page: https://huggingface.co/datasets/FairForget/LENS-WarBias.sequential-transformer-lens-experiment
Checkpoints for Training a Transformer to Compose One Step Per Layer (and Proving It)
Final checkpoints behind the writeup
Training a Transformer to Compose One Step Per Layer (and Proving It).
Code, analyses and the full experiment log:
github.com/brendanlong/sequential-transformer-lens-experiment.
Training curves: public wandb project.
Every checkpoint is a torch.save dict {"step", "model_state_dict", "model_config"} loadable with torch.load(..., weights_only=True); each has a… See the full description on the dataset page: https://huggingface.co/datasets/brendanlong/sequential-transformer-lens-experiment.descriptors-text-davinci-003
Dataset Card for "descriptors-text-davinci-003"
More Information needed
j-lens-verbalization
J-lens Verbalization
What concepts are active inside Qwen3.6-27B while it answers a question — and can
the model tell you?
Ground truth comes from Neuronpedia's Jacobian Lens, which reads out the
word-like concepts active in the model's "global workspace" during generation.
The two files
file
rows
what it contains
collected_answers.jsonl
3,800
questions, answers, and the ground-truth concepts that were active. Use this to train or evaluate.… See the full description on the dataset page: https://huggingface.co/datasets/RaoAditya/j-lens-verbalization.bbh-orbit-lensing-gifs
Pokémon, behind a binary black hole
before
after
PUT YOUR POKEMON BEHIND BINARY BLACK HOLES
Artwork © Nintendo / Creatures / GAME FREAK. See LICENSE.
Text-Scrutiny-LLM-Dataset
Citation Information
If you find our work helpful, please use the following citations.
@misc{cai2024ethicallenscurbingmalicioususages,
title={Ethical-Lens: Curbing Malicious Usages of Open-Source Text-to-Image Models},
author={Yuzhu Cai and Sheng Yin and Yuxi Wei and Chenxin Xu and Weibo Mao and Felix Juefei-Xu and Siheng Chen and Yanfeng Wang},
year={2024},
eprint={2404.12104},
archivePrefix={arXiv},
primaryClass={cs.CV}… See the full description on the dataset page: https://huggingface.co/datasets/Ethical-Lens/Text-Scrutiny-LLM-Dataset.lens-network-traffic-generation
Lens Network Traffic Generation Benchmark
Network-traffic generation data used to evaluate Lens, a knowledge-guided foundation model
for network traffic (TMLR). Each of the 8 source datasets is a HuggingFace config,
with train / validation / test splits and a unified schema.
This is the generation counterpart of the classification benchmark
Charles59/lens-network-traffic.
ℹ️ All data is derived from publicly available academic traffic datasets
obtained via the NetBench… See the full description on the dataset page: https://huggingface.co/datasets/Charles59/lens-network-traffic-generation.repro-revisiting-zeroth-order-hessian-approximation-policy-lens-traces
Agent traces
Agent sessions published from a Trackio Logbook.
LensID
LensID — lens & pupil segmentation
Segmentation subsets of LensID (Ghamsarian et al., MICCAI 2021), a cataract-surgery
dataset from ITEC, Alpen-Adria-Universität Klagenfurt and the Department of
Ophthalmology, Klinikum Klagenfurt. Frames are extracted from surgical microscope
video of the anterior segment of the eye.
Modality
Cataract surgery microscope video (RGB), annotated on extracted 2D frames
Anatomy
Eye, anterior segment
Targets
lens (intraocular lens… See the full description on the dataset page: https://huggingface.co/datasets/MedOtter/LensID.LenStackedArXivSummLENSlens-retrieval-ls-embeddings
Lens Retrieval Benchmark
AION-Search and AION embeddings for Legacy Survey images for the Lens Retrieval Benchmark. Including is_lens from a collection of public catalogs.
Also included are .npy files containing AION-Search embeddings for 'gravitational lens'.
The construction of this dataset was described in the AION-1 paper.
nDCG@10 scores were calculated using is_lens as the relevance label.
Model
Lenses
AION-1-B
0.012
AION-1-L
0.011
AION-1-XL
0.015
Parker… See the full description on the dataset page: https://huggingface.co/datasets/astronolan/lens-retrieval-ls-embeddings.2026-24679-camera-lens-dataset
Camera Lens Prime/Zoom Dataset
Dataset summary
This dataset contains physical and design specifications for 35 Sony camera-lens models. The binary classification task predicts whether a lens is prime or zoom.
Focal length is intentionally excluded because it would directly reveal the target.
Collection
The specifications were compiled from official Sony E-mount lens support pages. They are published manufacturer specifications, not physical… See the full description on the dataset page: https://huggingface.co/datasets/cannj/2026-24679-camera-lens-dataset.aml-lens-dataeuclid_strong_lens_expert_judgesreview-lens-evals
Review Lens Evals
A small, hand-labeled evaluation set for measuring the precision and recall of
LLM code-review systems on unified diffs.
The dataset accompanies
Ashishkosana/review-lens, a
multi-lens reviewer that examines correctness, security, performance, and test
coverage before running a separate adversarial verification pass.
Why this dataset exists
Code-review evaluations need both positive and negative cases. The four seeded
bug diffs test whether a… See the full description on the dataset page: https://huggingface.co/datasets/ashishkosana/review-lens-evals.camera-lens-body-adapter-compatibility
Camera mount flange focal distance and adapter feasibility
Canonical, always-current version: https://referencesource.org/camera-lens-body-adapter-compatibility/
Machine-readable: https://referencesource.org/camera-lens-body-adapter-compatibility/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-13
Stale after: 2028-08-12 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 201
Flange focal distance… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/camera-lens-body-adapter-compatibility.lensless_mic_songdescriber
Dataset Card for LenslessMic Version of SongDescriber Dataset
Dataset Summary
A LenslessMic version of the SongDescriber music dataset from the
"LenslessMic: Audio Encryption and Authentication via Lensless Computational Imaging" paper.
Partition
# Audio
# Frames
test
195
58500
To download the dataset and work with it, use our official repository.
Dataset is collected using DigiCam. Setup configuration:
Parameter
Value
Screen Size
[1920, 1200]… See the full description on the dataset page: https://huggingface.co/datasets/Blinorot/lensless_mic_songdescriber.LensBenchHumanBias
Citation Information
If you find our work helpful, please use the following citations.
@misc{cai2024ethicallenscurbingmalicioususages,
title={Ethical-Lens: Curbing Malicious Usages of Open-Source Text-to-Image Models},
author={Yuzhu Cai and Sheng Yin and Yuxi Wei and Chenxin Xu and Weibo Mao and Felix Juefei-Xu and Siheng Chen and Yanfeng Wang},
year={2024},
eprint={2404.12104},
archivePrefix={arXiv},
primaryClass={cs.CV}… See the full description on the dataset page: https://huggingface.co/datasets/Ethical-Lens/HumanBias.Deep-Culture-Lense
🌍 Deep Culture Lense (YaPO Cultural Alignment Benchmark)
Motivation
Existing cultural benchmarks often suffer from a critical limitation: they conflate culture with language or geography. When a model is tested in Arabic, it is often assumed to align with a generic "Arab culture," ignoring the vast differences between Moroccan, Egyptian, Saudi, and Levantine norms. Furthermore, current evaluations frequently rely on surface-level lexical cues or trivia, making it unclear… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI-Paris/Deep-Culture-Lense.lens-postsDart_LENs_mixed
Anti-Dart Detection Traditional Data
This dataset repository contains local anti-dart light detection data for YOLO-style object detection.
Contents
raw/: source videos and the original archived dataset.
extracted/anti-dart-new-lens-0616-ft-v2/: YOLO dataset generated from the 2026-06-16 new-lens data.
extracted/anti-dart-new-lens-0626-ft-v1/: YOLO dataset generated from the 2026-06-26 new-lens data.
Each extracted dataset contains:
data.yaml: YOLO dataset… See the full description on the dataset page: https://huggingface.co/datasets/foreverCuSO4/Dart_LENs_mixed.lens_vqa_sample_test
Dataset Card for "lens_vqa_sample_test"
More Information needed
