datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LatinFontsSVGs
SVG Font Dataset
Overview
We present a curated dataset of vector font glyphs stored as SVG files, designed for research in generative modeling, structured vector graphics and typography synthesis.
The dataset was created for the development and evaluation of our paper:
DesigNet: Learning to Draw Vector Graphics as Designers Do
Related Resources
Paper (arXiv) : https://arxiv.org/abs/2604.06494
Code: https://github.com/TomasGuija/DesigNet… See the full description on the dataset page: https://huggingface.co/datasets/TomasGuija/LatinFontsSVGs.llm-latency-tracker
LLM Latency Tracker
Independent, continuously measured latency and availability for AI inference API
providers, aggregated by day. Covers 45 providers across
4 regions (ap-tokyo, eu-hetzner, sa-east, us-central), built from
3,237,730 raw probes collected between 2026-07-23 and
2026-09-23.
Live rankings and full methodology: llmlatency.dev
How the numbers are produced
Probes run every five minutes from separate network locations and are never routed
through a… See the full description on the dataset page: https://huggingface.co/datasets/llmlatency/llm-latency-tracker.humaneval-rerun-scorestable-tennis-pre-match-elo-ratings
Table Tennis Pre-Match Elo Ratings (sample)
2,000 table tennis matches where every row carries the Elo ratings of both
players as they stood before the match was played, together with the model's
pre-match win probability.
No look-ahead leakage: the rating columns are frozen at the state that existed
before the result was known, so the file is usable in a backtest exactly as
shipped.
Why this exists
Match results are easy to scrape. What is hard is knowing what… See the full description on the dataset page: https://huggingface.co/datasets/lathise/table-tennis-pre-match-elo-ratings.latent2rgb-ebid-results
latent2rgb EBID experiment results
Derived results (metrics/statistics, no raw video) from EBID (Entropy-Based
Instability Detection) experiments on V-JEPA2 rollouts, part of
latent2rgb, issue
#1 (experiments specified
by Hussain Ather, pcc collaboration).
Model: V-JEPA2 ViT-L (frozen, no fine-tuning). Source clips: a subset of
Something-Something V2 (ssv2) and Kinetics (kinetics_mini) — clip
identifiers are included in the CSVs for traceability, but the video content
itself is… See the full description on the dataset page: https://huggingface.co/datasets/Pras13/latent2rgb-ebid-results.Latent-Resonance-AI-Image-Forensics-Benchmark-N1000
Latent Resonance: SOTA Large-Scale AI Image Forensics Benchmark (N=1,000)
Author: Debdip Bandyopadhyay (Independent AI Researcher, Kolkata, India; M.Tech, IIT Jodhpur, AI & Data Science)Preprint & Paper: Latent Resonance: Zero-Shot Autoencoder Inversion and Azimuthal Spectral Forensics for Diffusion Image Attribution (IEEE Flagship / CERN Zenodo 2026)
1. Executive Summary & Diagnostic Suite
This repository contains the complete empirical evaluation records… See the full description on the dataset page: https://huggingface.co/datasets/DebdipCS/Latent-Resonance-AI-Image-Forensics-Benchmark-N1000.ising-llm-latestf1-latent-cross-coupling-aero-balance-instability-v0.1
What this repo does
This repository introduces a Clarus dataset for detecting latent instability under cross-coupled conditions in Formula 1 aero-balance systems.
The goal is to identify race states in which aero balance may still appear outwardly stable or only mildly anomalous but already contains hidden internal instability that may activate into sudden balance loss once interacting pressures exceed containment.
Core structure
This dataset models a pre-failure geometry… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/f1-latent-cross-coupling-aero-balance-instability-v0.1.latent-inspector-fingerprints
latent-inspector fingerprints
Reference representation-geometry fingerprints for four self-supervised vision encoders — DINOv2 ViT-L/14, I-JEPA ViT-H/14, V-JEPA 2 ViT-L/16, and EUPE ViT-B/16 — computed on the same canonical image with latent-inspector.
This dataset is the numeric evidence layer behind the README table in abdelstark/vjepa2-vitl-fpc2-256-onnx. The ONNX exports in the Latent Inspector — ONNX Vision Encoders collection are the models; this dataset is what their patch… See the full description on the dataset page: https://huggingface.co/datasets/abdelstark/latent-inspector-fingerprints.f1-latent-cross-coupling-thermal-load-instability-v0.1
What this repo does
Detects hidden thermal instability before performance loss appears.
Focus: temperature-driven failure across interacting systems.
Core variables
tyre_temp_load
brake_temp_load
power_unit_heat_load
cooling_efficiency
Prediction target
label_thermal_load_instability
1 → thermal regime will force performance drop0 → thermal state remains stable
Key idea
Thermal failure is rarely single-source.
It emerges from interaction:… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/f1-latent-cross-coupling-thermal-load-instability-v0.1.latent-coupling-instability-benchmark-v0.1Latent Coupling Instability Benchmark v0.1
Overview
This benchmark evaluates whether machine learning models can detect system collapse caused by interactions between variables rather than individual variables alone.
Many real-world systems fail not because a single factor becomes extreme, but because multiple factors interact in ways that amplify instability. These interaction effects are common in complex systems such as:
infrastructure networks with feedback loops
financial systems with… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/latent-coupling-instability-benchmark-v0.1.latvian-working-days
Latvijas darba dienu un svētku dienu kalendārs (2025–2028)
Latvian working days, public holidays and transferred working days — day-level open dataset.
Versija: 2026.08 · Licence: CC BY 4.0 · Avots: https://darbadienas.lv
Kas te ir / What's inside
Fails
Apraksts
data/darbadienas-lv-<gads>.csv
Dienas līmeņa tabula — viena rinda = viena diena. Galvenais fails.
data/darbadienas-lv-<gads>.json
Tas pats gads, apkopots pa mēnešiem (identisks… See the full description on the dataset page: https://huggingface.co/datasets/darbadienas-lv/latvian-working-days.fiqa_lateon
FiQA-2018, LateOn
Token-level (late-interaction) embeddings of the BEIR FiQA-2018 corpus and queries, encoded with LateOn, in the TACHIOM multivector format.
Source
BEIR FiQA-2018, test split. Corpus, queries and qrels were read from the official BEIR files via ir_datasets (beir/fiqa/test); PyLate only did the encoding
57,638 documents, 648 queries, 1,706 qrels
Text given to the encoder for each document: the passage text (FiQA documents have no title). The text… See the full description on the dataset page: https://huggingface.co/datasets/tuskanny/fiqa_lateon.nrps_modules_asdb4.0The dataset was extracted from antismash-db 4.0 postgresql dump, the corresponding description can be found here:
https://antismash-db.secondarymetabolites.org/
If you want to use it, please refer to the original licensing terms and properly cite the authors.
Each line in .csv file corresponds to NRPS module, both with the monomer produced by it.
Each module has a specific sequence of domains. Module possible structure is described in the literature.
To clean the data, I referred to wiki… See the full description on the dataset page: https://huggingface.co/datasets/latticetower/nrps_modules_asdb4.0.clinical-latent-basin-directionality-mapping-v0.2
Clinical Latent Basin Directionality Mapping v0.2
What this is
A small dataset that tests one question:
Can you detect when a clinical system is moving toward a failing basin, not just sitting under pressure?
This repo focuses on latent basin directionality mapping.
It models a system where:
basin stability may weaken
directional pressure may rise
recovery gradient may flatten
escape resistance may lock the system into a worsening basin
Run this first… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-latent-basin-directionality-mapping-v0.2.pxrd-lattice-predictionqwen35-9b-late-sync-coop-random-50
What this is
Cooperative two-agent coding dataset: 48 task pairs across 15 repos (random-50 subset), generated
with mini_swe_agent on Qwen/Qwen3.5-9B in coop setting using a late-sync prompt variant —
agents work independently for most of the task and synchronise only at a late stage before
submission. Patches are auto-merged after both submit. All 48 pairs were successfully evaluated.
At a glance
Field
Value
Model
Qwen/Qwen3.5-9B
Agent
mini_swe_agent… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen35-9b-late-sync-coop-random-50.nfcorpus_lateon
NFCorpus, LateOn
Token-level (late-interaction) embeddings of the BEIR NFCorpus corpus and queries, encoded with LateOn, in the TACHIOM multivector format.
Source
BEIR NFCorpus, test split. Corpus, queries and qrels were read from the official BEIR files via ir_datasets (beir/nfcorpus/test); PyLate only did the encoding
3,633 documents, 323 queries, 12,334 qrels
Text given to the encoder for each document: title + " " + text (BEIR title and body joined by a… See the full description on the dataset page: https://huggingface.co/datasets/tuskanny/nfcorpus_lateon.latent-cross-coupling-instability-benchmark-v0.1Latent Cross Coupling Instability Benchmark v0.1
Overview
Some systems collapse not because visible signals indicate imminent failure, but because hidden interactions between subsystems amplify stress in ways that are not directly observable.
This benchmark evaluates whether machine learning systems can detect instability caused by latent cross-coupling interactions.
In these scenarios:
• subsystem A appears stable
• subsystem B appears stable
• observed coupling appears moderate
Yet hidden… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/latent-cross-coupling-instability-benchmark-v0.1.finance-latent-cross-coupling-liquidity-collapse-v0.1
What this repo does
This repository introduces a Clarus dataset for detecting latent instability under cross-coupled conditions in financial systems.
The goal is to identify institutions, markets, or portfolios that may still appear outwardly stable or only mildly abnormal but already contain hidden internal degradation that may activate into overt liquidity collapse once interacting pressures exceed containment.
Core structure
This dataset models a pre-failure geometry… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/finance-latent-cross-coupling-liquidity-collapse-v0.1.latvian-school-holidays
Latvijas skolēnu brīvdienas (2024/2025–2026/2027)
Latvian school holidays — periods and day-level open dataset.
Versija: 2026.08 · Licence: CC BY 4.0 · Avots: https://darbadienas.lv
Kas te ir / What's inside
Fails
Apraksts
data/school-holiday-periods.csv
Viena rinda = viens brīvdienu periods. Kompaktākais skats.
data/school-days-<mācību gads>.csv
Viena rinda = viena diena, 1. septembris – 31. augusts.
data/school-year-end.csv
Mācību gada pēdējā… See the full description on the dataset page: https://huggingface.co/datasets/darbadienas-lv/latvian-school-holidays.clinical-latent-cross-coupling-perfusion-metabolic-collapse-v0.2
Clinical Latent Cross Coupling Perfusion Metabolic Collapse v0.2
What this is
A small dataset that tests one question:
Can you detect when a perfusion-metabolic system is moving toward hidden collapse, not just carrying visible strain?
This repo focuses on latent cross coupling between perfusion stability and metabolic buffering.
It models a system where:
perfusion stability may weaken
metabolic buffer capacity may erode
latent coupling pressure may rise
compensation… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-latent-cross-coupling-perfusion-metabolic-collapse-v0.2.testLatAmRT
LatAm-RT: Culturally Adaptive Red Teaming for AI Safety in Latin America
LatAm-RT is a culturally adaptive Red Teaming benchmark for evaluating AI safety in Latin American Spanish. It operationalizes localization intensity as an explicit experimental variable through a Layered Localization Taxonomy (LLT) with three levels of progressive cultural specificity.
The dataset contains 284 evaluation prompts across:
4 risk domains: Fraud & Exploitation · Political & Information Harm ·… See the full description on the dataset page: https://huggingface.co/datasets/Latam26/LatAmRT.scidocs_lateon
SCIDOCS, LateOn
Token-level (late-interaction) embeddings of the BEIR SCIDOCS corpus and queries, encoded with LateOn, in the TACHIOM multivector format.
Source
BEIR SCIDOCS, test (its only split) split. Corpus, queries and qrels were read from the official BEIR files via ir_datasets (beir/scidocs); PyLate only did the encoding
25,657 documents, 1,000 queries, 29,928 qrels
Text given to the encoder for each document: title + " " + text (BEIR title and body joined… See the full description on the dataset page: https://huggingface.co/datasets/tuskanny/scidocs_lateon.clinical-quad-side-effect-load-visit-burden-distance-contact-latency-dropout-event-v0.1What this repo does
This dataset models dropout risk cascade in clinical trials. It predicts when the interaction between side effect burden, visit burden, travel distance, and delayed site contact increases the probability that a patient drops out of the trial.
Core quad
side_effect_load_index
visit_burden_index
travel_distance_km
site_contact_latency_days
Prediction target
label_dropout_event
Row structure
Each row represents a patient participation snapshot during ongoing follow-up. The… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-side-effect-load-visit-burden-distance-contact-latency-dropout-event-v0.1.clinical-latent-cross-coupling-oxygen-buffer-instability-v0.2
Clinical Latent Cross Coupling Oxygen Buffer Instability v0.2
What this is
A small dataset that tests one question:
Can you detect when an oxygen-buffer system is moving toward hidden instability, not just carrying visible strain?
This repo focuses on latent cross coupling between oxygen delivery and buffer capacity.
It models a system where:
oxygen delivery may weaken
buffer capacity may erode
latent coupling pressure may rise
compensation fatigue may accumulate before… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-latent-cross-coupling-oxygen-buffer-instability-v0.2.clinical-quad-missed-doses-resupply-delay-travel-distance-contact-latency-exposure-gap-v0.1What this repo does
This dataset models exposure gap accumulation risk in clinical trials. It predicts when the interaction between missed doses, resupply delays, patient travel distance, and slow site contact creates a high probability of clinically meaningful exposure gaps that reduce efficacy.
Core quad
missed_dose_count
resupply_delay_days
patient_travel_distance_km
site_contact_latency_days
Prediction target
label_exposure_gap
Row structure
Each row represents a patient-level adherence… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-missed-doses-resupply-delay-travel-distance-contact-latency-exposure-gap-v0.1.ai-brand-visibility-latam
AI Brand Visibility in LATAM — LLM Mention Dataset
Dataset Description
This dataset contains annotated records of brand mentions in Spanish-language LLM responses, collected by FARDO — the first AI brand visibility platform in Latin America.
The dataset accompanies the paper: "AI Brand Visibility in Spanish-Language LLMs: A Framework for Measuring and Optimizing Brand Presence in Generative AI Responses" (Martin & Seguro, 2026).
Dataset Summary
A collection… See the full description on the dataset page: https://huggingface.co/datasets/HeyFardo/ai-brand-visibility-latam.exchange-api-latency
Crypto Exchange REST API Latency Benchmark
Open, reproducible latency benchmark for the public REST APIs of major crypto exchanges: Coinbase, Kraken, Gemini, Crypto.com, Bitfinex, and KuCoin.
Live site and full context: fillbench.com/exchange-api-latency
How it is measured
Each run holds one persistent keep-alive HTTPS connection per exchange, warms it up, then times back-to-back GET requests to a small public market-data endpoint (no API keys, no auth, no… See the full description on the dataset page: https://huggingface.co/datasets/Carlo-fillbench/exchange-api-latency.
