datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vn-ebi-provincial
Vietnam EBI provincial e-business index (2012-2026)
VECOM E-Business Index (EBI) provincial scores, 14 report years (2012-2026. no official 2016). Subindexes HR&IT, B2C, B2B, and G2B when published. G2B dropped from the composite from 2021. Geographic panel is historical_63 through 2025 and current_34 (post-merger) in 2026. Weights and coverage vary by year - see year_meta. Chart digits verified against official PDFs/screenshots. published ebi_total kept even when a few VECOM… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-ebi-provincial.ebible_corpus
eBible Parallel Corpus
A verse-aligned parallel corpus of 1253 Bible translations across 997 languages, prepared from publicly redistributable texts hosted by eBible.org.
Generated: 2026-05-04
Dataset Description
Each row corresponds to one verse reference from the standard 41,899-verse reference list (vref.txt). Every translation occupies one column, identified by its translationId. Rows are positionally aligned: row N in every translation column refers to the same… See the full description on the dataset page: https://huggingface.co/datasets/DavidCBaines/ebible_corpus.latent2rgb-ebid-results
latent2rgb EBID experiment results
Derived results (metrics/statistics, no raw video) from EBID (Entropy-Based
Instability Detection) experiments on V-JEPA2 rollouts, part of
latent2rgb, issue
#1 (experiments specified
by Hussain Ather, pcc collaboration).
Model: V-JEPA2 ViT-L (frozen, no fine-tuning). Source clips: a subset of
Something-Something V2 (ssv2) and Kinetics (kinetics_mini) — clip
identifiers are included in the CSVs for traceability, but the video content
itself is… See the full description on the dataset page: https://huggingface.co/datasets/Pras13/latent2rgb-ebid-results.handwriting-signatures-datasetA dataset of handwritten signatures with text prompts, designed for LoRA fine-tuning of diffusion models to generate realistic personal signatures.
Handwritten Signatures Dataset (Processed)
This dataset contains preprocessed handwritten signatures designed for signature verification and LoRA fine-tuning of diffusion models (e.g., Stable Diffusion) for text-to-image tasks.
📌 Processing steps
Labeling – assigned using computer vision and AI models to group signatures… See the full description on the dataset page: https://huggingface.co/datasets/ebixhaferaj/handwriting-signatures-dataset.ebible_local_ind_corpus
eBible Indonesian Local Language Corpus
This dataset contains parallel Bible translations between Indonesian and various local languages from Indonesia, particularly from Eastern Indonesia regions.
Dataset Description
This dataset is created from the eBible corpus, containing verse-aligned translations between Indonesian and multiple local languages. The dataset properly handles verse ranges where multiple verses are combined in the translation.
Available Language… See the full description on the dataset page: https://huggingface.co/datasets/Davidsamuel101/ebible_local_ind_corpus.ebitda_adj_syntheticafrispeech_ebiranepali-ner-ebiquity
Nepali NER — Ebiquity v2
Benchmark NER dataset for Nepali with BIO tags: PER, ORG, LOC, MISC.
CoNLL format, cleaned and deduplicated.
Source: https://github.com/oya163/nepali-ner
Usage
from datasets import load_dataset
ds = load_dataset("Titung/nepali-ner-ebiquity")
ebiquity-v2Ebiquity V2 (non-stemmed) dataset for Nepali NER task. The dataset is tagged with BIO scheme.ebiquity-v2-stemmedEbiquity V2 (stemmed) dataset for Nepali NER task. The dataset is tagged with BIO scheme.ebible-nlpThis dataset contains 1,035 bible translations in multiple languages. The verses are organised in rows from Genesis to Revelation with some translations being partial and others complete. This dataset is derived from https://github.com/BibleNLP/ebible-corpus
patternized_ebible_examplesebios-rm-qa-dataseteBird_India
