datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LatinFontsSVGs
SVG Font Dataset
Overview
We present a curated dataset of vector font glyphs stored as SVG files, designed for research in generative modeling, structured vector graphics and typography synthesis.
The dataset was created for the development and evaluation of our paper:
DesigNet: Learning to Draw Vector Graphics as Designers Do
Related Resources
Paper (arXiv) : https://arxiv.org/abs/2604.06494
Code: https://github.com/TomasGuija/DesigNet… See the full description on the dataset page: https://huggingface.co/datasets/TomasGuija/LatinFontsSVGs.llm-latency-tracker
LLM Latency Tracker
Independent, continuously measured latency and availability for AI inference API
providers, aggregated by day. Covers 45 providers across
4 regions (ap-tokyo, eu-hetzner, sa-east, us-central), built from
3,237,730 raw probes collected between 2026-07-23 and
2026-09-23.
Live rankings and full methodology: llmlatency.dev
How the numbers are produced
Probes run every five minutes from separate network locations and are never routed
through a… See the full description on the dataset page: https://huggingface.co/datasets/llmlatency/llm-latency-tracker.ats-career-page-urls
ATS Career Page URLs
69,638 canonical career page URLs for public job boards hosted on 40 applicant tracking system (ATS) platforms, including Greenhouse, Lever, Workable, Ashby, Workday, and BambooHR.
Each row is the canonical entry point to a public job board. The dataset is deduplicated, URL-normalized, and intended as a starting point for job-market research, labor-market analytics, ATS ecosystem analysis, and job aggregation pipelines.
Released as part of Latmay, a semantic… See the full description on the dataset page: https://huggingface.co/datasets/latmay/ats-career-page-urls.LatinSummarizer
LatinSummarizer Dataset
Structure
aligned_en_la_data_raw.csv
aligned_en_la_data_cleaned.csv
aligned_en_la_data_cleaned_with_stanza.csv
concat_aligned_data.csv
concat_cleaned.csv
latin_wikipedia_cleaned.csv
latin_wikipedia_raw.csv
latin-literature-dataset-170M_raw_cleaned.csv
latin-literature-dataset-170M_raw_cleaned_chunked.csv
Elsa_aligned/
README.md
Details
aligned_en_la_data_raw.csv
This dataset contains aligned Latin (la) - English (en)… See the full description on the dataset page: https://huggingface.co/datasets/LatinNLP/LatinSummarizer.ldt-latents
1 Million Image Latents Toy Dataset
A lightweight toy dataset of 1 003 626 image latents paired with CLIP text embeddings.
Raw sources & extraction
LAION‑aesthetic (laion/laion2B-en-aesthetic):
Streamed via 🤗 datasets in 50 k-image blocks.
Filtered for aesthetic > 7.
Skipped PNG/CMYK or images < 32×32 px.
JourneyDB (MidJourney) (JourneyDB/JourneyDB):
Downloaded three zip archives per batch from Hugging Face.
Unzipped locally and selected the first 50 000 valid… See the full description on the dataset page: https://huggingface.co/datasets/shreenithi20/ldt-latents.ldt-latents-uint8Trueque-Benchmark-beta-0.1
🤝 Trueque: A human-reviewed collaborative benchmark for Latin American knowledge and culture
🌐 Language versions: Español | Português
⚠️ Official Disclaimer: Beta Release (v0.1)
Welcome to Trueque for Factual Knowledge and Cultural Appropriateness. This dataset represents an initial effort to evaluate the regional knowledge and cultural accuracy of Large Language Models (LLMs) in Latin America.
Please take the following considerations into account before using this resource:… See the full description on the dataset page: https://huggingface.co/datasets/latam-gpt/Trueque-Benchmark-beta-0.1.arXiv_latex
TeX data from arXiv
Using https://github.com/KyuDan1/TeX2Image code.
We have categories Math, Physics, Statistics, ComputerScience.
Domain
Size
Mathematics
4.22M
Computer Science
2.76M
Statistics
0.89M
Physics
0.78M
Total (unique)
7.17M
humaneval-rerun-scoreslatamqa_mcq_es-la
LatamQA
LatamQA is a cultural knowledge benchmark designed to evaluate Large Language Models on Latin American contexts. The dataset addresses the critical gap in bias detection resources for non-English languages and underrepresented cultures. Built from 26,000+ Wikipedia articles and structured using Wikidata's knowledge graph with expert guidance from social scientists, LatamQA contains over 26,000 multiple-choice questions covering the diverse popular and social cultures of… See the full description on the dataset page: https://huggingface.co/datasets/inria-chile/latamqa_mcq_es-la.latent-dna-diffusionCHOCLO
🌽 CHOCLO: Latin American Cultural Knowledge Benchmark
Description
CHOCLO is a benchmark designed to evaluate cultural knowledge in language models, with a specific focus on entities representative of Latin America. Unlike traditional benchmarks, which often emphasize general knowledge or contexts dominated by English-language data, CHOCLO aims to capture the richness, diversity, and specificity of Latin American cultural knowledge, including traditions, gastronomy… See the full description on the dataset page: https://huggingface.co/datasets/latam-gpt/CHOCLO.Danbooru-Top1000-Latents-NPZDanbooru-Top1000-Latents-SDXLtable-tennis-pre-match-elo-ratings
Table Tennis Pre-Match Elo Ratings (sample)
2,000 table tennis matches where every row carries the Elo ratings of both
players as they stood before the match was played, together with the model's
pre-match win probability.
No look-ahead leakage: the rating columns are frozen at the state that existed
before the result was known, so the file is usable in a backtest exactly as
shipped.
Why this exists
Match results are easy to scrape. What is hard is knowing what… See the full description on the dataset page: https://huggingface.co/datasets/lathise/table-tennis-pre-match-elo-ratings.fusion-image-to-latex-datasets
Collects and builds the largest dataset to date from online sources, creating a robust and generalizable dataset. This dataset includes approximately 3.4 million image-text pairs, including both handwritten mathematical expressions (200,330 examples) and printed mathematical expressions (3,237,250 examples). Due to the large dataset and the fact that the same mathematical formula can be represented in different LaTeX string formats in an image, it is easy to cause polymorphic ambiguity. To… See the full description on the dataset page: https://huggingface.co/datasets/hoang-quoc-trung/fusion-image-to-latex-datasets.latent2rgb-ebid-results
latent2rgb EBID experiment results
Derived results (metrics/statistics, no raw video) from EBID (Entropy-Based
Instability Detection) experiments on V-JEPA2 rollouts, part of
latent2rgb, issue
#1 (experiments specified
by Hussain Ather, pcc collaboration).
Model: V-JEPA2 ViT-L (frozen, no fine-tuning). Source clips: a subset of
Something-Something V2 (ssv2) and Kinetics (kinetics_mini) — clip
identifiers are included in the CSVs for traceability, but the video content
itself is… See the full description on the dataset page: https://huggingface.co/datasets/Pras13/latent2rgb-ebid-results.latent_upscale_validationKabyle-Latin-to-Tifinagh-Parallel-Corpus
Dataset Card for Kabyle Latin-to-Tifinagh Parallel Corpus
This dataset provides a parallel corpus of the Kabyle language (Taqbaylit), pairing native Latin-based orthography with automated, context-aware Amazigh script transliterations. It is built by processing raw text data through a rule-based algorithmic pipeline designed to enforce strict orthographic purity, manage contextual phonetic mutations, and isolate foreign vocabulary.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/abdelhaqueidali/Kabyle-Latin-to-Tifinagh-Parallel-Corpus.Latent-Resonance-AI-Image-Forensics-Benchmark-N1000
Latent Resonance: SOTA Large-Scale AI Image Forensics Benchmark (N=1,000)
Author: Debdip Bandyopadhyay (Independent AI Researcher, Kolkata, India; M.Tech, IIT Jodhpur, AI & Data Science)Preprint & Paper: Latent Resonance: Zero-Shot Autoencoder Inversion and Azimuthal Spectral Forensics for Diffusion Image Attribution (IEEE Flagship / CERN Zenodo 2026)
1. Executive Summary & Diagnostic Suite
This repository contains the complete empirical evaluation records… See the full description on the dataset page: https://huggingface.co/datasets/DebdipCS/Latent-Resonance-AI-Image-Forensics-Benchmark-N1000.bug-bounty-programs-rewards
Overview
This dataset is part of a research project Economic Taxonomy of Software Vulnerabilities, which aims to estimate the monetary cost associated with software vulnerabilities based on real-world bug bounty program data. The core objective is to provide a concrete, data-driven reference for evaluating the average cost of discovering and reporting vulnerabilities across different severity levels.
Methodology Summary
The dataset was generated following these steps:… See the full description on the dataset page: https://huggingface.co/datasets/lesis-lat/bug-bounty-programs-rewards.berkshire-fttp-latency-data
Primary Source: This dataset is maintained by Paul Kirby (Berkshire IT Services). The canonical diagnostic tools and telemetry collection routines reside on GitHub (markkirby125).
Berkshire FTTP & Latency Open Dataset
East Berkshire broadband latency dataset and aggregation tools tracking FTTP rollout stability, DNS response times, and routing jitter across local exchanges.
[]
[Python 3.x">]
Attribution & Data Collection
This data is passively collected (with… See the full description on the dataset page: https://huggingface.co/datasets/thiassi/berkshire-fttp-latency-data.ovos-intents-train-latestvietnamese-nom-latin-translationcloud-load-latency-coherence-risk-v0.1What this repo is for
Catch latency collapse early.
It flags when:
latency spikes without load
headroom looks fine but queues rise
overflow triggers too early
saturation appears but monitoring hides i
Tashelhit-Tifinagh-Latin-Parallel-Tatoeba-Dataset
Dataset Card for Tachelhit Latin-to-Tifinagh Parallel Corpus
This dataset provides a parallel corpus of the Tachelhit language ($\text{Tacelḥit}$ / $\text{Tamazigt}$), pairing native Latin-based orthography with automated Amazigh script transliterations. It is built by processing clean source sentences through an algorithmic engine designed to handle phonetic mappings, manage contextual schwa distributions, and safeguard acronyms and foreign proper names.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/abdelhaqueidali/Tashelhit-Tifinagh-Latin-Parallel-Tatoeba-Dataset.f1-latent-cross-coupling-aero-balance-instability-v0.1
What this repo does
This repository introduces a Clarus dataset for detecting latent instability under cross-coupled conditions in Formula 1 aero-balance systems.
The goal is to identify race states in which aero balance may still appear outwardly stable or only mildly anomalous but already contains hidden internal instability that may activate into sudden balance loss once interacting pressures exceed containment.
Core structure
This dataset models a pre-failure geometry… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/f1-latent-cross-coupling-aero-balance-instability-v0.1.Latvian-Speech-Dataset
Latvian Dataset Metadata
Field
Value
📜 License
CC BY-NC-ND 4.0
🎯 Task Categories
Automatic Speech Recognition
🌍 Language
Latvian (la)
🏷️ Tags
Audio, Speech, Speech Recognition, ML, Machine, Machine Learning, Latvian
📦 Size Category
n < 1K
clinical-quad-safety-signal-latency-reporting-lag-conmed-confound-v0.1Clarus Clinical Quad Coupling Safety Signal Latency Reporting Lag Conmed Confound v0.1
What this dataset isThis dataset tests whether a model can detect latent safety signals when four interacting nodes create uncertainty.
Quad coupling nodes
Emerging safety event pattern
Reporting or entry latency
Concomitant medication or behavior confound
Governance decision timing such as DSMB, batch release, or safety review
Input
One vignette
OutputReturn strict JSON only.
Required output… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-safety-signal-latency-reporting-lag-conmed-confound-v0.1.latent-inspector-fingerprints
latent-inspector fingerprints
Reference representation-geometry fingerprints for four self-supervised vision encoders — DINOv2 ViT-L/14, I-JEPA ViT-H/14, V-JEPA 2 ViT-L/16, and EUPE ViT-B/16 — computed on the same canonical image with latent-inspector.
This dataset is the numeric evidence layer behind the README table in abdelstark/vjepa2-vitl-fpc2-256-onnx. The ONNX exports in the Latent Inspector — ONNX Vision Encoders collection are the models; this dataset is what their patch… See the full description on the dataset page: https://huggingface.co/datasets/abdelstark/latent-inspector-fingerprints.
