datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
efficientnet-v2-l-adv-dataset
Perturb Adversarial Images
Verified adversarial examples for efficientnet_v2_l (torchvision/EfficientNet_V2_L_Weights.IMAGENET1K_V1), produced by the
Perturb network. Each row is one clean image together with all of its
verified adversarial versions: images that are imperceptibly different from the original
(L∞ ≤ 0.03 in [0,1] pixel scale) yet change the model's top-1 prediction.
This dataset grows continuously. New rows are appended as the network produces them and uploaded in… See the full description on the dataset page: https://huggingface.co/datasets/perturb-ai/efficientnet-v2-l-adv-dataset.ai2d-no-maskdbpedia-entities-efficient-splade-100K
DBPedia SPLADE + OpenAI: 100,000 SPLADE Sparse Vectors + OpenAI Embedding
This dataset has both OpenAI and SPLADE vectors for 100,000 DBPedia entries. This adds SPLADE Vectors to KShivendu/dbpedia-entities-openai-1M/
Model id used to make these vectors:
model_id = "naver/efficient-splade-VI-BT-large-doc"
For processing the query, use this:
model_id = "naver/efficient-splade-VI-BT-large-query"
If you'd like to extract the indices and weights/values from the vectors, you can do so… See the full description on the dataset page: https://huggingface.co/datasets/nirantk/dbpedia-entities-efficient-splade-100K.Z1-Code-Reasoning-107K
Z1: Efficient Test-time Scaling with Code
Train Large Language Model to Reason with Shifted Thinking
[📜 Paper] •
[🤗 HF Models] •
[🐱 GitHub]
Details
Please refer to https://github.com/efficientscaling/Z1.
Usage
from datasets import load_dataset
ds = load_dataset("efficientscaling/Z1-Code-Reasoning-107K")["train"]
ds[0]
Citation
@misc{yu2025efficientscaling,
title={Z1: Efficient Test-time Scaling with Code}… See the full description on the dataset page: https://huggingface.co/datasets/efficientscaling/Z1-Code-Reasoning-107K.dclm-train-1.64m-tsp-efficienttts-serving-benchmarkThis repository contains metadata-only versions (audio columns removed) of the following Hugging Face datasets:
https://huggingface.co/datasets/MikhailT/lj-speech
https://huggingface.co/datasets/MikhailT/hifi-tts
https://huggingface.co/datasets/mythicinfinity/libritts
The dataset is intended for benchmarking the performance of TTS serving systems. Example usage is available in this script:
https://github.com/vox-serve/vox-serve/blob/main/benchmark/goodput.py
sts-serving-benchmarkThis repository contains metadata-only versions (audio columns removed) of the following Hugging Face datasets:
https://huggingface.co/datasets/hlt-lab/voicebench
The dataset is intended for benchmarking the performance of STS serving systems. Example usage is available in this script:
https://github.com/vox-serve/vox-serve/blob/main/benchmark/goodput.py
simple_r1Efficient_ToolCallingEfficient_ToolCalling_trainZ1-Code-Reasoning-Shortest-90KZ1-Code-Reasoning-Longest-33Klight_r1efficient_yoloThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 0,
"total_frames": 0,
"total_tasks": 0,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/qownscks/efficient_yolo.DRQA_Efficient_Reasoning_CoTefficientnet_b0_b7_comparison_4000_rowsexamination_samplesCT_I-II-III_efficientAll phases, all rows.
Balanced.
Contains:
TRIAL NAME
BRIEF
DRUG USED
DRUG CLASS
INDICATION
TARGET
THERAPY
LEAD SPONSOR
CRITERIA
PRIMARY OUTCOME
SECONDARY OUTCOME 1
Evaluation_andreasmadsen-efficient_mlm_m0.40CT_III_efficient_fullFor Phase III
Contains:
TRIAL NAME
BRIEF
DRUG USED
DRUG CLASS
INDICATION
TARGET
THERAPY
LEAD SPONSOR
CRITERIA
PRIMARY OUTCOME
SECONDARY OUTCOME 1
Evaluation_Loraandreasmadsen-efficient_mlm_m0.40tiny-imagenet_ee-efficientnet-b2
