datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
openrouter-uptime
OpenRouter Uptime
An independent, timestamped uptime record for every model on
OpenRouter and each of its inference providers.
Polled hourly from OpenRouter's public API and mirrored here daily.
Available in three places:
Source (raw + full git history): github.com/dthinkr/openrouter-uptime
HuggingFace: huggingface.co/datasets/venvoo/openrouter-uptime
Kaggle: kaggle.com/datasets/spicycorn/openrouter-uptime
Files
file
rows
description
readings.parquet… See the full description on the dataset page: https://huggingface.co/datasets/venvoo/openrouter-uptime.VENUS-10K
Dataset Card for VENUS
Dataset Summary
Data from: Speaking Beyond Language: A Large-Scale Multimodal Dataset for Learning Nonverbal Cues from Video-Grounded Dialogues
@article{kim2025speaking,
title={Speaking Beyond Language: A Large-Scale Multimodal Dataset for Learning Nonverbal Cues from Video-Grounded Dialogues},
author={Kim, Youngmin and Chung, Jiwan and Kim, Jisoo and Lee, Sunghyun and Lee, Sangkyu and Kim, Junhyeok and Yang, Cheoljong and Yu, Youngjae}… See the full description on the dataset page: https://huggingface.co/datasets/winston1214/VENUS-10K.Venus_Case_Tempsaas-vendor-status-pages-outages-incidents-daily
SaaS vendor status pages — 1,127 vendors mapped, 16,259 incidents, rebuilt daily
Last rebuilt: 2026-09-24 12:28 UTC. An automated job re-probes every vendor's public status
page daily, records each incident it publishes (title, impact, opened/resolved
times, permalink) and re-uploads these files. It is the data behind
approjects-vendor-status-watch.static.hf.space, where each vendor has a page with its incident history,
RSS and JSON.
Two tables:
incidents — one row per incident… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/saas-vendor-status-pages-outages-incidents-daily.VENUS-5K
Dataset Card for VENUS
Dataset Summary
Data from: Speaking Beyond Language: A Large-Scale Multimodal Dataset for Learning Nonverbal Cues from Video-Grounded Dialogues
@article{kim2025speaking,
title={Speaking Beyond Language: A Large-Scale Multimodal Dataset for Learning Nonverbal Cues from Video-Grounded Dialogues},
author={Kim, Youngmin and Chung, Jiwan and Kim, Jisoo and Lee, Sunghyun and Lee, Sangkyu and Kim, Junhyeok and Yang, Cheoljong and Yu, Youngjae}… See the full description on the dataset page: https://huggingface.co/datasets/winston1214/VENUS-5K.VENUS-25K
Dataset Card for VENUS
Dataset Summary
Data from: Speaking Beyond Language: A Large-Scale Multimodal Dataset for Learning Nonverbal Cues from Video-Grounded Dialogues
@article{kim2025speaking,
title={Speaking Beyond Language: A Large-Scale Multimodal Dataset for Learning Nonverbal Cues from Video-Grounded Dialogues},
author={Kim, Youngmin and Chung, Jiwan and Kim, Jisoo and Lee, Sunghyun and Lee, Sangkyu and Kim, Junhyeok and Yang, Cheoljong and Yu, Youngjae}… See the full description on the dataset page: https://huggingface.co/datasets/winston1214/VENUS-25K.VENUS-1K
Dataset Card for VENUS
Dataset Summary
Data from: Speaking Beyond Language: A Large-Scale Multimodal Dataset for Learning Nonverbal Cues from Video-Grounded Dialogues
@article{kim2025speaking,
title={Speaking Beyond Language: A Large-Scale Multimodal Dataset for Learning Nonverbal Cues from Video-Grounded Dialogues},
author={Kim, Youngmin and Chung, Jiwan and Kim, Jisoo and Lee, Sunghyun and Lee, Sangkyu and Kim, Junhyeok and Yang, Cheoljong and Yu, Youngjae}… See the full description on the dataset page: https://huggingface.co/datasets/winston1214/VENUS-1K.VenusMutHub
VenusMutHub Dataset
VenusMutHub is a comprehensive collection of protein mutation data designed for benchmarking and evaluating protein language models (PLMs) on various mutation effect prediction tasks. This repository contains mutation data across multiple protein properties including enzyme activity, binding affinity, stability, and selectivity.
Dataset Overview
VenusMutHub includes:
Mutation data for hundreds of proteins across diverse functional categories… See the full description on the dataset page: https://huggingface.co/datasets/AI4Protein/VenusMutHub.disaster_tweetsvenus_tempsaas-vendor-outage-duration-incident-resolution-time-mttr
How long do SaaS vendor outages last? Incident resolution time per vendor, rebuilt daily
As of 2026-09-24 12:28 UTC. For every incident a vendor posted on its own public status
page with BOTH an opened time and a resolved time, this dataset computes
duration_minutes = resolved_at - started_at
and rolls it up per vendor. It is derived, every day, from the incident table in
saas-vendor-status-pages-outages-incidents-daily; the two are rebuilt by the same job and cannot
disagree.… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/saas-vendor-outage-duration-incident-resolution-time-mttr.Venus
Venus: A dataset for fine-grained code generation control
🎉 What is Venus? Venus is the dataset used to train Afterburner (WIP). It is an extension of the original Mercury dataset and currently includes 6 languages: Python3, C++, Javascript, Go, Rust, and Java.
🚧 What is the current progress? We are in the process of expanding the dataset to include more programming languages.
🔮 Why Venus stands out? A key contribution of Venus is that it provides runtime and memory… See the full description on the dataset page: https://huggingface.co/datasets/Elfsong/Venus.WebVid-CoVRarxiv.org/abs/2308.14746
ipfs_venezuela_laws_ir
Venezuela legislation IR (CID-keyed sparse GraphRAG)
Research retrieval release of endomorphosis/ipfs_venezuela_laws (revision fa24f9ef360a006206a27a1343e228ee305cd941) packaged as
country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir).
Not legal advice. This is a research snapshot. The official gazette /
authentic source of Venezuela prevails over this corpus. Retrieved documents
and graph edges are retrieval evidence only. No legal text… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_venezuela_laws_ir.esa-venus-express-observations
ESA Venus Express Observations
Credit: NASA/JPL-Caltech
Part of a dataset collection on Hugging Face.
Dataset description
Complete observation metadata catalog from the ESA Venus Express mission, which studied Venus from 2006 to 2014.
Venus Express was a European Space Agency mission that studied the Venusian atmosphere, ionosphere, and surface environment from April 2006 until loss of contact in November 2014. It carried a suite of instruments including a… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/esa-venus-express-observations.obfuscated-activations-llama32-artifacts
Obfuscated Activations in Llama 3.2 — research artifacts
This repository is the curated artifact release for a mechanistic case study of two OAT-style, co-trained Llama 3.2 adapters and their linear probes. It contains attack banks and machine-readable results, not prose analysis, generated plots, notebooks, or duplicate report outputs.
Built with Llama. See LICENSE.txt, NOTICE, and the component-specific terms below.
Companion model repositories… See the full description on the dataset page: https://huggingface.co/datasets/venky-cdmbrm14/obfuscated-activations-llama32-artifacts.sycophantic-praise
SyPR Benchmark
This dataset contains fixed input artifacts for evaluating sycophantic praise calibration.
Each row is a single (persona, utterance, prompt_condition) evaluation instance.
The evaluated model response is generated dynamically at evaluation time and is not
included in the benchmark artifact.
Reasoning utterances use real benchmark questions. GSM8K target questions and
final answers are pulled from openai/gsm8k. MMLU-Pro Chemistry and Economics
target questions… See the full description on the dataset page: https://huggingface.co/datasets/vennemeyerd/sycophantic-praise.Machine-Identity-Spectra
Machine Identity Spectra Dataset
Summary
Venafi is excited to release of the Machine Identity Spectra large dataset.
This collection of data contains extracted features from 19m+ certificates discovered over HTTPS (port 443) on the
public internet between July 20 and July 26, 2023.
The features are a combination of X.509 certificate features, RFC5280 compliance checks,
and other attributes intended to be used for clustering, features analysis, and a base for… See the full description on the dataset page: https://huggingface.co/datasets/Venafi/Machine-Identity-Spectra.SLR-NoteSense
SLR-NoteSense
SLR-NoteSense is a dual-sensor red-green-blue-clear (RGBC) dataset for Sri Lankan banknote denomination recognition. The seven-denomination update adds the LKR 2000 class to the original six-denomination dataset.
The dataset contains 14,846 paired acquisition rows from 743 distinct physical banknotes, covering LKR 20, 50, 100, 500, 1000, 2000, and 5000. After applying the documented sensor-settling rule, the stable view contains 14,104 paired measurement rows.… See the full description on the dataset page: https://huggingface.co/datasets/vennsa/SLR-NoteSense.livestock-health-disease-ssa-synthetic
Dataset Card: Livestock Health & Disease Surveillance (Synthetic Data)
Dataset Summary
This synthetic dataset represents 1,000,000 African smallholder households with livestock systems, capturing livestock health, disease surveillance, veterinary access, and herd management practices across Sub-Saharan Africa. It combines baseline farm characteristics (Dataset 1) with 15 livestock-specific variables to create a comprehensive picture of livestock production systems and… See the full description on the dataset page: https://huggingface.co/datasets/Venesa123/livestock-health-disease-ssa-synthetic.inferes
Dataset Card for InferES
Dataset Summary
Natural Language Inference dataset for European Spanish
Paper accepted and (to be) presented at COLING 2022
Supported Tasks and Leaderboards
Natural Language Inference
Languages
Spanish
Dataset Structure
The dataset contains two texts inputs (Premise and Hypothesis), Label for three-way classification, and annotation data.
Data Instances
train size = 6444
test size = 1612
Data… See the full description on the dataset page: https://huggingface.co/datasets/venelin/inferes.Grab-BoxThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/Vendrien/Grab-Box.prix-immobilier-communes-dvf
French Property Prices per m² by Municipality (DVF)
Median price per m² for 756 French municipalities, computed from actual recorded
property sales — the official DVF database (Demandes de valeurs foncières) published
by the French tax administration (DGFiP). No estimates, no extrapolation: only observed
transactions.
756 municipalities across 5 departments (Creuse, Gironde, Loire-Atlantique, Orne,
Puy-de-Dôme), sales period 2023–2025.
Files… See the full description on the dataset page: https://huggingface.co/datasets/Vendsmonbien/prix-immobilier-communes-dvf.record-testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/Vendrien/record-test.Venus_CCGliarsbenchneuripsLiars' Bench is a benchmark for evaluating lie-detectors for language models.
It was generated by using four open-weight models: mistral-small-3.1-24b-instruct, llama-v3.3-70b-instruct, qwen-2.5-72b-instruct, gemma-3-27b-it across seven datasets including a control dataset (Alpaca).
Our settings capture qualitatively different types of lies and vary along two dimensions:
the model's reason for lying
the object of belief targeted by the lie.
License
Unless otherwise specified, all datasets… See the full description on the dataset page: https://huggingface.co/datasets/ventura1900/liarsbenchneurips.elrobot-pickplace
ElRobot Pick & Place Dataset
Teleoperation demonstrations recorded on NormaCore ElRobot (8-DOF + gripper).
Robot
NormaCore ElRobot — 7 rotational joints + 1 gripper (8 ST3215 servos)
Controller: Raspberry Pi 5
Cameras: 2x USB RGB (224x224)
Dataset Format (NormaCore custom parquet)
One row = one frame. Schema:
Column
Type
Description
episode_start_ns
uint64
Episode identifier
timestamp_ns_since_episode_start
uint64
Per-episode time… See the full description on the dataset page: https://huggingface.co/datasets/venayc/elrobot-pickplace.wireless-vendor-identifiers
Wireless Vendor Identifier Metadata
This dataset contains public wireless and Bluetooth metadata relevant to BLE
research. It is designed for defensive analysis, asset classification,
literature support, and reproducible metadata lookup. The dataset is split into
multiple tables so it is not just an OUI mirror.
It deliberately excludes complete BLE advertisement payloads, packet templates,
radio timing recipes, device captures, effectiveness labels, and any instructions
for… See the full description on the dataset page: https://huggingface.co/datasets/ictrun/wireless-vendor-identifiers.venus_caseVenus_SFT_Data
