datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nerf-gs-datasetsI keep a collection compiled of existing datasets from various sources for training NeRFs or Splats. This dataset is most of that collection. All of the individual scenes also have a trained Gaussian Splat.
https://rishit-dagli.github.io/2025/03/28/nerf-gs-datasets.html
ner-jsonlnerfew-nerd
Dataset Card for "Few-NERD"
#dataset-description)
Dataset Summary
Supported Tasks and Leaderboards
Languages
Dataset Structure
Data Instances
Data Fields
Data Splits
Dataset Creation
Curation Rationale
Source Data
Annotations
Personal and Sensitive InformationConsiderations for Using the Data
Social Impact of Dataset
Discussion of Biases
Other Known Limitations
Additional Information
Dataset Curators
Licensing Information
Citation Information
Contributions
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/DFKI-SLT/few-nerd.robust-e-nerf
Robust e-NeRF Synthetic Event Dataset
This repository contains the synthetic event dataset used in Robust e-NeRF to study the collective effect of camera speed profile, contrast threshold variation and refractory period on the quality of NeRF reconstruction from a moving event camera. The dataset is simulated using an improved version of ESIM with three different camera configurations of increasing difficulty levels (i.e. easy, medium and hard)… See the full description on the dataset page: https://huggingface.co/datasets/wengflow/robust-e-nerf.ProfNER_corpus_NER
Description
Gold standard annotations for profession detection in Spanish COVID-19 tweets
The entire corpus contains 10,000 annotated tweets. It has been split into training, validation, and test (60-20-20). The current version contains the training and development set of the shared task with Gold Standard annotations. In addition, it contains the unannotated test, and background sets will be released.
For Named Entity Recognition, profession detection, annotations are distributed… See the full description on the dataset page: https://huggingface.co/datasets/Biomedical-TeMU/ProfNER_corpus_NER.gutenberg_spacy-ner
Dataset Card for "gutenberg_spacy-ner"
More Information needed
bad_prompt
Negative Embedding / Textual Inversion
Idea
The idea behind this embedding was to somehow train the negative prompt as an embedding, thus unifying the basis of the negative prompt into one word or embedding.
Side note: Embedding has proven to be very helpful for the generation of hands! :)
Usage
To use this embedding you have to download the file aswell as drop it into the "\stable-diffusion-webui\embeddings" folder.
Please put the embedding in… See the full description on the dataset page: https://huggingface.co/datasets/Nerfgun3/bad_prompt.nerfbaselines-datanerf-synthetic-mirrorVibeVoiceNEREL_innodatasetsrealsense-calvin
realsense-calvin
Raw per-timestep CALVIN-format conversion of realsense-converted, generated by
dawn/data/realsense/convert_to_raw_calvin.py in the DAWN repo
(/home/colligo/Codes/HiVA/DAWN).
Layout
training/episode_0000000.npz ... + training/lang_annotations/auto_lang_ann.npy
validation/episode_0000000.npz ... + validation/lang_annotations/{auto_lang_ann.npy, embeddings.npy}
Per-frame npz keys
rgb_static (256, 256, 3) uint8 — official CALVIN uses… See the full description on the dataset page: https://huggingface.co/datasets/nero1342/realsense-calvin.kpwr-nerKPWR-NER tagging dataset.Nereus
Nereus
Nereus is a multi-component underwater vision and instruction dataset. It
combines grounded fish-counting data derived from IOCfish5K with a separate
object-attribute understanding component.
This is a multi-source dataset. The root license: other value is intentional:
no single license applies to every file. Read LICENSE,
THIRD_PARTY_SOURCES.md, NOTICE, and the available component-level license
information before use or redistribution.
Components… See the full description on the dataset page: https://huggingface.co/datasets/Nereusdata/Nereus.open-ner-standardized
Dataset Card for OpenNER 1.0
OpenNER 1.0 is a standardized collection of openly-available named entity recognition (NER) datasets.
OpenNER contains 36 NER corpora that span 52 languages, human-annotated in varying named entity ontologies.
We correct annotation format issues, standardize the original datasets into a uniform representation with consistent entity type names across corpora, and provide the collection in a structure that enables research in multilingual and… See the full description on the dataset page: https://huggingface.co/datasets/bltlab/open-ner-standardized.nerf-syntheticcantemist-nerhttps://temu.bsc.es/cantemist/nerfbaselines-supplementarysample-ner
Dataset Card for "sample-ner"
More Information needed
nerf-syntheticsimplevla-grpo-assets
SimpleVLA GRPO Grasp Assets
This dataset contains the released USD object assets used by the SimpleVLA-style GRPO grasping experiments.
Expected local layout after running scripts/download_assets.sh:
/data4/nerako/reasoning/RLinf_assets/grasp_assets/
Mip-NeRF360gutenberg_spacy-nerdownloads
📁 neronreal / downloads
Szybkie linki direct download – bez czekania, limitów transferu i zbędnych przekierowań.
⚙️ Informacje
Główne archiwum na większe pliki, spolszczenia i inne rzeczy, które nie mieszczą się na Dropboxie.
Hosting: Szybkie pobieranie bezpośrednie bez limitów i reklam.
Format: Oryginalne pliki bez żadnych modyfikacji i kompresji po stronie serwera.
person-names-ner
Dataset Card for Person Full Name NER Parsing
This dataset contains 3,383,944 curated and augmented person names, designed specifically for training Token Classification (NER) models. The primary task is to parse a full name string into its FirstName and LastName components, correctly handling multi-word names and different ordering formats.
Dataset Details
Dataset Description
This dataset is built to train robust models that can understand and segment human… See the full description on the dataset page: https://huggingface.co/datasets/ele-sage/person-names-ner.nero_three_bags_box
NERO: three bags into a box
Raw teleoperation recording from the AgileX NERO dual-arm rig.
Episode
episode_20260808_042810_d7ac35b1
Duration: 61.19 seconds
Nominal rate: 15 FPS
Task: place three wrapped bags into the divided box
Result: all three bags are in the box at the end of the recording
The episode includes synchronized base, left-wrist, and right-wrist RGB,
left/right wrist depth, dual-arm joint and gripper telemetry in frames.jsonl,
meta.json, and… See the full description on the dataset page: https://huggingface.co/datasets/NoahWeiss/nero_three_bags_box.UnMix-NeRF-Artifactssucx3_ner The dataset is a conversion of the venerable SUC 3.0 dataset into the
huggingface ecosystem. The original dataset does not contain an official
train-dev-test split, which is introduced here; the tag distribution for the
NER tags between the three splits is mostly the same.
The dataset has three different types of tagsets: manually annotated POS,
manually annotated NER, and automatically annotated NER. For the
automatically annotated NER tags, only sentences were chosen, where the
automatic and manual annotations would match (with their respective
categories).
Additionally we provide remixes of the same data with some or all sentences
being lowercased.
