datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
epic-kitchens-100-clips
EPIC-KITCHENS-100 Extracted Clips
About
Dataset of 37455 video clips (24GB) extracted from videos in the EPIC-KITCHENS-100 dataset,
more precisely the extension part not contained in EPIC-KITCHENS-55. For details,
see https://www.lightly.ai/product-updates/epickitchens-100-in-lightlystudio.
The clips folder contains one video for every narration from action annotations stored
in {participant_id}/{narration_id}.mp4. The videos have been downscaled an compressed for easier… See the full description on the dataset page: https://huggingface.co/datasets/lightly-ai/epic-kitchens-100-clips.in1k_clip_qwen25vl_3b_224res_64tokens_new_ptin1k_clip_qwen25vl_3b_448res_256tokens_new_merged_ptclipquill-asr-benchmark
Measuring whisper-tiny vs whisper-base in a browser tab
Word error rate, wall-clock timing, transfer size and peak memory for two
quantised Whisper tiers running entirely client-side in a real Chrome window,
with the scripts that produced every number.
If you are building an in-browser transcription page, the two results worth
knowing before you pick a model tier:
On clean synthetic audio the two tiers tie. If that is all you test, you
will conclude the tier does not matter… See the full description on the dataset page: https://huggingface.co/datasets/sophia8888/clipquill-asr-benchmark.LDS-retrain-bank-adamw-N16k-bs256-clip1.0WIT-es_jina-clip-v2_sampleambient-o-clip-iqa-patches-imagenet
Ambient Diffusion Omni (Ambient-o): Training Good Models with Bad Data
Dataset Description
Ambient Diffusion Omni (Ambient-o) is a framework for using low-quality, synthetic, and out-of-distribution images to improve the quality of diffusion models. Unlike traditional approaches that rely on highly curated datasets, Ambient-o extracts valuable signal from all available images during training, including data typically discarded as "low-quality."
This dataset card is for… See the full description on the dataset page: https://huggingface.co/datasets/adrianrm/ambient-o-clip-iqa-patches-imagenet.bBSARD
Dataset Card for bBSARD
Dataset Summary
Bilingual BSARD (bBSARD) is a parallel bilingual version of the French BSARD dataset, which is extended to Dutch by adding the Dutch version of included legislations, and translating the questions.
bBSARD consists of 22,417 statutory articles in French and Dutch from Belgian law and about 1,100 legal questions labeled with relevant articles from the corpus.
Supported Tasks and Leaderboards
document-retrieval: The… See the full description on the dataset page: https://huggingface.co/datasets/clips/bBSARD.epic-kitchens-100-clips
EPIC-KITCHENS-100 Extracted Clips
About
Dataset of 37455 video clips (24GB) extracted from videos in the EPIC-KITCHENS-100 dataset,
more precisely the extension part not contained in EPIC-KITCHENS-55. For details,
see https://www.lightly.ai/product-updates/epickitchens-100-in-lightlystudio.
The clips folder contains one video for every narration from action annotations stored
in {participant_id}/{narration_id}.mp4. The videos have been downscaled an compressed for easier… See the full description on the dataset page: https://huggingface.co/datasets/weiweizezedongdong/epic-kitchens-100-clips.kick-clip-yield-2026
Kick Clip Yield Dataset (June–July 2026)
Attribution: ClipMe — https://clipme.com/research
Summary
Production-pipeline capture data from ClipMe, an AI clipping tool for live Kick streams. 51 capture sessions across 8 anonymized Kick channels, June 19 – July 14, 2026. Every number is measured directly from clip files on disk with ffprobe — no engagement, view, or engine-internal data included.
Dataset structure
Per-session table (CSV… See the full description on the dataset page: https://huggingface.co/datasets/Clipmeapp/kick-clip-yield-2026.mteb-nl-legalqa This dataset consists of Dutch legislative articles, with queries formulated based on subordinating conjunctions.
Citation Information
If you find our paper, benchmark or models helpful, please consider cite as follows:
@misc{banar2025mtebnle5nlembeddingbenchmark,
title={MTEB-NL and E5-NL: Embedding Benchmark and Models for Dutch},
author={Nikolay Banar and Ehsan Lotfi and Jens Van Nooten and Cristina Arhiliuc and Marija Kliocaite and Walter Daelemans}… See the full description on the dataset page: https://huggingface.co/datasets/clips/mteb-nl-legalqa.clip-lora-submissionclipsEuropa_Clipper_TourThis dataset (.csv) contains the Europa Clipper tour design (21F31_MEGA_L241010_A300411_LP01_V6_pad_scpse.bsp) mapped to the Khurana background magnetic field model for the 49 planned flybys.
This dataset is designed for use with the LEAP analysis but could be adapted for other works too
startup-data-clippingmeld_clips_v3
目录结构
/
├── metadata.csv # 数据集元数据文件
├── videos/ # 视频文件目录
│ ├── dev_sample_2.mp4
│ ├── dev_sample_3.mp4
│ └── ...
└── audios/ # 音频文件目录
├── mfa_meld_dev_oov.txt # MFA对齐时的OOV(out-of-vocabulary)词汇
├── ost/ # 原始音轨(Original SoundTrack)
│ ├── dev_sample_2.wav
│ ├── dev_sample_3.wav
│ └── ...
├── vocals/ # 人声音轨(分离后)
│ ├── dev_sample_2_vocals.wav
│ ├──… See the full description on the dataset page: https://huggingface.co/datasets/BigfufuOuO/meld_clips_v3.clip_img_wrd_embedding
