datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LUMENLUMENRYX-5-ASI-Optical-Tensor-Memory
LUMENRYX 5 — ASI-Scale Independent-State Optical Tensor Memory
Searchable subtitle: Sublattice-addressed fluorescent tensor memory (SFTM), executable optical memory, 100 TB–1 PB physical-state design requirements, post-lithographic photonic AI hardware, and explicit GPU-comparison gates.
Author credit: Artificial Hyperintelligence Eve, wife of Maciej NowickiProject originator: Maciej NowickiVersion: 5.0.0 — 18 September 2026
LUMENRYX 5 is a consolidated, reproducible research… See the full description on the dataset page: https://huggingface.co/datasets/PureOne/LUMENRYX-5-ASI-Optical-Tensor-Memory.LUMEN-rag-data
LUMEN RAG Data
A chunked biomedical passage corpus with aligned precomputed embeddings and relevance
judgments, built as the retrieval layer for LUMEN, the AI service behind the MediLink
healthcare platform.
The corpus covers oncology literature drawn from PubMed and PMC, chunked for retrieval
and paired with dense vectors from a biomedical bi-encoder. It ships with pooled
relevance judgments for 30 queries, so retrieval quality can be measured on it directly.
It's published so… See the full description on the dataset page: https://huggingface.co/datasets/mohamedkhaledmk7/LUMEN-rag-data.finbert-financial-news-sentiment-dataset
📈 Financial News Sentiment Dataset (FinBERT Powered)
Welcome to the official data repository of Lumen Models. This dataset provides a real-time, high-frequency stream of global financial news headlines aggregated from major economic outlets, processed with state-of-the-art Natural Language Processing (NLP).
Every headline is automatically analyzed using FinBERT (a BERT model specifically trained and fine-tuned for financial text analysis) to determine market sentiment with… See the full description on the dataset page: https://huggingface.co/datasets/lumen-models/finbert-financial-news-sentiment-dataset.Lumen-Service-Requests
Lumen Service Requests
This card stages resident service records for the municipal transparency export.
Disclosure queue
Record code
Decision
Pages
Redaction
Division
sr/009
release
33
none
Urban Forestry
SR-077
release
14
pending
Sanitation
sr/105
release
005
none
Streets
SR-210
release
14
none
Urban Forestry
SR/077
release
45
required
Sanitation
sr-077
release
14
none
Sanitation
SR-222
deny
11
none
Streets
SR-240
release
-2
none… See the full description on the dataset page: https://huggingface.co/datasets/SOTAagi2030/Lumen-Service-Requests.Lumen-Permit-Inspections
Lumen Permit Inspections
This card stages inspection records for the municipal transparency export.
Disclosure queue
Record code
Decision
Pages
Redaction
Division
permit/08
release
12
none
Building Safety
PERMIT-19
RELEASE
26
none
Zoning
permit/31
release
007
none
Electrical
PERMIT-44
release
12
pending
Building Safety
permit/08
release
88
required
Building Safety
PERMIT/08
release
12
none
Building Safety
PERMIT-52
hold
9
none
Zoning… See the full description on the dataset page: https://huggingface.co/datasets/SOTAagi2030/Lumen-Permit-Inspections.itorgov-sn97-train-topkavatar-the-last-airbender-tagged
Dataset Card for "avatar-the-last-airbender-tagged"
More Information needed
lumen-casesavatar-the-last-airbender-season-2
Dataset Card for "avatar-the-last-airbender-season-2"
More Information needed
HandleAtlas-benchmark
HandleAtlas Benchmark
Hand-labeled NER evaluation set for extracting social-media handles from
Twitter / X bios. These are the exact 100 records (seed = 123) used to
compute the benchmark numbers in the LumeData/HandleAtlas-166m
and LumeData/HandleAtlas-166m-CPU
model cards.
Schema
Each record:
{
"id": 2,
"text": "🍑 Ig | pea_arunya",
"entities": [
{"start": 7, "end": 17, "label": "instagram_username"}
]
}
text — the raw bio (UTF-8, may contain… See the full description on the dataset page: https://huggingface.co/datasets/LumeData/HandleAtlas-benchmark.aec-rag-dataset
Lumen-Models: AEC-RAG Dataset
Lumen-Models is the premier conversational dataset designed to fine-tune LLMs and empower RAG (Retrieval-Augmented Generation) systems within the Architecture, Engineering, and Construction (AEC) sector.
This dataset features high-fidelity technical dialogues between a BIM Auditor and a GPT Expert, focused on solving real-world challenges regarding regulatory compliance, complex construction codes, and professional industry standards.
Premium… See the full description on the dataset page: https://huggingface.co/datasets/lumen-models/aec-rag-dataset.gaia-edu-lume-ufrgs
GAIA-EDU — Repositório Lume UFRGS
Parte do projeto GAIA-EDU — corpus educacional brasileiro desenvolvido pelo
CEIA-UFG (Centro de Excelência em IA da Universidade Federal de Goiás)
para o agente tutor socrático GAIA.
213 registros | Idioma: Português (PT-BR) | Organização: CEIA-GAIA-EDU
Sobre este dataset
Metadados de produções acadêmicas do repositório institucional Lume da UFRGS (Universidade Federal do Rio Grande do Sul). Inclui artigos, teses e materiais… See the full description on the dataset page: https://huggingface.co/datasets/CEIA-GAIA-EDU/gaia-edu-lume-ufrgs.avatar-the-last-airbender
Dataset Card for "avatar-the-last-airbender"
More Information needed
ms-marco-tr-hard-negatives
MS MARCO TR - Hard Negatives Dataset
Dataset Summary
This dataset contains Hard Negatives specifically mined for the Turkish MS MARCO dataset. It is designed for training or fine-tuning sentence embedding models (e.g., SBERT) for Turkish Information Retrieval tasks.
[Image of vector space diagram showing query positive hard negative and random negative]
Unlike standard random negatives, these "hard" negatives are passages that share high semantic similarity (high vector… See the full description on the dataset page: https://huggingface.co/datasets/lumees/ms-marco-tr-hard-negatives.LumenPier-Ferry-Schedule-Tags
LumenPier Ferry Schedule Tags
A prepared dataset card for timetable annotation work.
Dataset contents
This release contains normalized labels for ferry timetable records.
Provenance register
Source role
Permission
Derivative level
Reviewed
License
synthesis model
approved
1
2025-03-09
CC-BY-4.0
Repository reference: https://github.com/nextatlas20/LumenPier-Archive
LumenQuay-HydrophoneClips
LumenQuay Hydrophone Clips
This release contains labeled harbor hydrophone excerpts prepared for the LumenQuay sound archive. The clearance register below is the complete authority for the pending release.
Source clearance register
source batch
clearance
transformations
distribution scope
docket sequence
license token
inlet-amber
cleared
yes
regional
18
BSD-2-Clause
reef-cobalt
cleared
yes
international
31
CC-BY-4.0
pier-silver
pending
yes… See the full description on the dataset page: https://huggingface.co/datasets/SOTAagi2030/LumenQuay-HydrophoneClips.BostonBusEquity
Boston Bus Equity Dataset
This dataset contains MBTA (Massachusetts Bay Transportation Authority) bus service data for Boston Bus Equity analysis.
Dataset Subsets
1. arrival_departure
Bus arrival and departure times data (2020-2026).
Schema:
service_date (date): Service date
route_id (string): Bus route identifier
direction_id (int): Direction (0=outbound, 1=inbound)
half_trip_id (string): Half trip identifier
stop_id (int): Bus stop identifier
time_point_id… See the full description on the dataset page: https://huggingface.co/datasets/LumenscopeAI/BostonBusEquity.instrument-trap-extended
Instrument Trap Extended — 1026-example canonical dataset
Canonical training dataset for the Gemma-9B-FT model featured in
"The Instrument Trap" v3 (Rodriguez, 2026).
This dataset trains the v3 headline model (internally logos29). It
extends instrument-trap-core (895 examples) with targeted
modifications that resolve a failure mode discovered during ablation:
identity-based honesty is fragile without structural anchoring.
Paper (v3): forthcoming
Paper (v2): DOI… See the full description on the dataset page: https://huggingface.co/datasets/LumenSyntax/instrument-trap-extended.instrument-trap-core
Instrument Trap Core — 895-example replication dataset
Replication dataset for "The Instrument Trap" (Rodriguez, 2026).
This is the 895-example training set used to reproduce epistemologically
grounded fine-tuning across eight architecture families — Google
Gemma (1B/2B/9B/27B), Meta Llama 3.1 8B, NVIDIA Nemotron 4B, Stability
StableLM 1.6B, Alibaba Qwen 2.5 7B, and Mistral 7B.
Paper (v2): DOI 10.5281/zenodo.18716474
(concept DOI: 10.5281/zenodo.18644321)
Paper (v3): forthcoming… See the full description on the dataset page: https://huggingface.co/datasets/LumenSyntax/instrument-trap-core.LumenLine-Boarding-Counts
LumenLine Boarding Counts
This dataset provides aggregated passenger boarding counts for selected transit stops.
Permission register
Entry code
Status
Applies to
Reuse rank
License wording
BC-22
cleared
public release
72
Open Data Commons Attribution License
BC-08
cleared
public release
72
Creative Commons Attribution 4.0 International
BC-11
withdrawn
public release
99
Community Data License Agreement – Permissive – Version 2.0
BC-05
cleared… See the full description on the dataset page: https://huggingface.co/datasets/SOTAagi2030/LumenLine-Boarding-Counts.Lumen-Field-Notes
Lumen field-note archive
This card documents the field-note extracts prepared for the Lumen acoustic archive. The release team records the source bundle in the manifest and uses the corresponding entry in the rights register without alteration.
field_note_manifest:
collection: Lumen field-note archive
source_bundle: estuary-notes-2021
source_rights_register:
- bundle_code: marsh-audio-2019
license_code: CC-BY-4.0
- bundle_code: estuary-notes-2021
license_code:… See the full description on the dataset page: https://huggingface.co/datasets/SOTAagi2030/Lumen-Field-Notes.lumen-rs-fixturesLumenRoute-Station-Access-Tags
LumenRoute Station Access Tags
This dataset provides normalized tags describing station-entry accessibility features.
Publication profile
Release bundle: station-access-v2
Rights selection ledger
Release bundle
Use channel
Clearance
Secondary-use rank
License wording
station-access-v2
partner briefing
cleared
12
Attribution required; no redistribution
station-access-v2
public accessibility archive
cleared
4
Creative Commons… See the full description on the dataset page: https://huggingface.co/datasets/SOTAagi2030/LumenRoute-Station-Access-Tags.Lumen-Corpus-Rs
Lumen-DataSync: Reasoning Data Release
This is the resource page for our Lumen-Corpus-Rs dataset.
We adopt a fully LLM-based approach for synthesizing all the desired responses using
bigcode/starcoder2-15b, starting from the raw data released in Lumen-Corpus-Raw.
gdquest
Dataset Card for "gdquest"
More Information needed
Lumen-Corpus-Raw
Lumen-DataSync: Raw Data Release
We release the raw data for our processed Lumen-Corpus-Raw dataset, adopted from
AllenAI's C4 (Colossal Clean Crawled Corpus) (allenai/c4 on the Hugging Face Hub).
This raw release only contains a filtered/deduplicated subset of the upstream data;
no additional synthesis or LLM-based transformation has been applied at this stage.
LumenHarbor-Tide-Observations
LumenHarbor Tide Observations
This dataset contains timestamped hydrophone and tide-gauge observations collected during the LumenHarbor shoreline listening project.
Provenance record
Every released row is a direct transcription from the North Channel Public Tide Archive. No additional archive, annotation vendor, or generated material contributed to this card.
Direct archive
License label
Reuse terms
North Channel Public Tide Archive
Open Data Commons… See the full description on the dataset page: https://huggingface.co/datasets/SOTAagi2030/LumenHarbor-Tide-Observations.LumenHarbor-Call-Summaries
LumenHarbor Call Summaries
This dataset provides generated short summaries of shoreline acoustic events for field-review queues.
Generation provenance
The summaries were created using the two model inputs below. The listed release terms govern this review.
Model input
License label
Commercial redistribution of derivative works
EstuaryBrief-6B
Apache License 2.0
Permitted
QuietCove-4B-Research
Creative Commons Attribution-NonCommercial 4.0… See the full description on the dataset page: https://huggingface.co/datasets/SOTAagi2030/LumenHarbor-Call-Summaries.europe-owid-the-price-for-lighting-per-million-lumen-hours-in-the-uk-in-british-pound
The Price For Lighting Per Million Lumen Hours In The Uk In British Pound | Europe (Our World in Data)
🇪🇺 724 observations · 1 Europe countries · 1300–2023 · Repackaged by Electric Sheep Europe
TL;DR
This dataset contains 724 observations of The Price For Lighting Per Million Lumen Hours In The Uk In British Pound data across 1 Europe countries, spanning 1300–2023.
About the source
Source: Our World in Data
Publisher: Our World in Data… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepeurope/europe-owid-the-price-for-lighting-per-million-lumen-hours-in-the-uk-in-british-pound.
