datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ebook-translate-queuemorpheus-real-world
Morpheus — Real-World Physics Videos
Real-world reference footage for Morpheus, a benchmark that tests whether
video generative models (Wan, CogVideo, LTX-Video, COSMOS-predict1/2,
Pyramid-Flow, Veo3, Kling-Turbo, ...) obey Newtonian mechanics. Object
trajectories are extracted via SAM2 tracking and tested against physical laws
(energy/momentum conservation, equations of motion) rather than pixel-matched
to a single "correct" video.
This repo contains the filmed real-world… See the full description on the dataset page: https://huggingface.co/datasets/physics-from-video/morpheus-real-world.summarize_from_feedbackSummarize from Feedback contains the human feedback data released by the "Learning to summarize from human feedback" paper.Cement-4096Cobot_Magic_take_out_a_pen_from_the_pen_holder
Cobot_Magic_take_out_a_pen_from_the_pen_holder
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: agilex_cobot_decoupled_magic
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
office
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
place
📊… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Cobot_Magic_take_out_a_pen_from_the_pen_holder.TextOnly_FromRLBench_CloseBox_24K_unfixedcross-channel-toxoplasma-from-cellmask
Cross-channel toxoplasma from cellmask dataset
3030 paired fields. images/ is the input channel, masks_pv/ the instance-labelled
parasitophorous-vacuole ground truth (same stem = same field), regenerated with the promoted
PV model cpsam_v2_toxo_r5.
Train/test annotation: fields.csv gives name,split,n_objects for every field -
2567 train (31637 objects) / 463 test (6116 objects).
The split is by well (split_by_well.csv), so no well contributes to both sides.
Prepared with spaCR… See the full description on the dataset page: https://huggingface.co/datasets/einarolafsson/cross-channel-toxoplasma-from-cellmask.from-dataset
Dataset
Variant
Links
msmarco_document
terrier_stemmed
readme · branch
msmarco_document
terrier_stemmed_docT5query
readme · branch
msmarco_document
terrier_stemmed_text
readme · branch
msmarco_document
terrier_unstemmed
readme · branch
msmarco_document
terrier_unstemmed_text
readme · branch
msmarco_passage
ance
readme · branch
msmarco_passage
corpusgraph_bm25_k16
readme · branch
msmarco_passage
corpusgraph_tcthnp_k16
readme · branch
msmarco_passage
numpy_tcthnp… See the full description on the dataset page: https://huggingface.co/datasets/pyterrier/from-dataset.Cement-1536diffusion_db_dedupe_from50k_train
Dataset Card for "diffusion_db_dedupe_from50k_train"
More Information needed
summarize_from_feedback_small220k-GPT4Vision-captions-from-LIVIS
220k-GPT4Vision-captions-from-LVIS
by: Christoph Schuhmann, Peter Bevan, 21 Nov, 2023
This dataset comprises 220,000 captioned images from the LVIS dataset. The captions were generated by summarising the LVIS-Instruct4V dataset released by X2FD. The instructions are converted into captions using Mistral-7B-OpenOrca.
PROMPT
"""<<SYS>> You are a highly intelligent, empathic, helpful, respectful, and honest assistant with high emotional intelligence. Always… See the full description on the dataset page: https://huggingface.co/datasets/laion/220k-GPT4Vision-captions-from-LIVIS.navsim-metric-caches-from-a100ctb-m5-cleaned-fixed-from-2026-01-01-to-2026-07-13flare_finqa_sup_sample_from_policy_v1.1_stepwise_dpo_chunk_3n8n-from-colab
n8n - Secure Workflow Automation for Technical Teams
n8n is a workflow automation platform that gives technical teams the flexibility of code with the speed of no-code. With 400+ integrations, native AI capabilities, and a fair-code license, n8n lets you build powerful automations while maintaining full control over your data and deployments.
Key Capabilities
Code When You Need It: Write JavaScript/Python, add npm packages, or use the visual interface
AI-Native… See the full description on the dataset page: https://huggingface.co/datasets/omarelsayeed/n8n-from-colab.asr_rnnt_eou_from_scratch
NeMo ASR-EOU 训练脚本解读与论文出处梳理
目标文件:examples/asr/asr_eou/speech_to_text_rnnt_eou_train.py链接:https://github.com/NVIDIA-NeMo/NeMo/blob/main/examples/asr/asr_eou/speech_to_text_rnnt_eou_train.py
这份脚本本身是一个 训练入口脚本(Hydra + PyTorch Lightning),核心功能是:按配置创建 EncDecRNNTBPEEOUModel,并支持从已有 .nemo 初始化、添加/训练 adapter,以及在“词表扩展(新增 <EOU>/<EOB>)”时做权重迁移。
下面按“它用到的技术点 → 在代码/配置里怎么体现 → 原始论文出处”总结。
1) ASR-EOU:把“端点/话轮信息”并入 ASR(<EOU>, <EOB>)
它做什么:
除了输出转写文本外,还让模型在时间轴上预测:
EOU:End Of Utterance(一句话结束)… See the full description on the dataset page: https://huggingface.co/datasets/echodict/asr_rnnt_eou_from_scratch.us-historical-layoffs-archive-warn-act-notices-removed-from-state-websites
Historical US layoffs archive: 6,799 WARN Act notices that state websites no longer list (2000-2025), recovered
Rebuilt 2026-09-25. Five state labor agencies — Connecticut, Michigan, New York,
North Carolina and Pennsylvania — retired the web pages their older WARN Act
layoff notices lived on. Their current pages start years later. This dataset is
every notice in our file that came from one of those retired pages and is not
on the agency's live page today: 6,799 notices, 6,799… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-historical-layoffs-archive-warn-act-notices-removed-from-state-websites.ur3-remove-cup-from-nested-cupstext-code-galeras-code-generation-from-docstring-3k-dedupedOMat24_train_aimd_from_PBE_1000_npt
Cite this dataset Barroso-Luque, L., Shuaibi, M., Fu, X., Wood, B. M., Dzamba, M., Gao, M., Rizvi, A., Zitnick, C. L., and Ulissi, Z. W. OMat24 train aimd from PBE 1000 npt. ColabFit, 2025. https://doi.org/10.60732/25f16f85
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_jqrkc9e7cgmh_0
Visit the ColabFit Exchange to search… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/OMat24_train_aimd_from_PBE_1000_npt.low-to-high-res_weather_from_topography
Dataset: Low-to-High-Resolution Weather Forecasting using Topography
The dataset is intended and structured for the problem of transforming/interpolating low-resolution weather forecasts into higher resolution using topography data.
The dataset consists of 3 different types of data (as illustrated above):
Historical weather observation data (SMHI)
Historical weather observation data from selected SMHI observation stations (evaluation points)Historical low-resolution weather… See the full description on the dataset page: https://huggingface.co/datasets/rebase-energy/low-to-high-res_weather_from_topography.test_import_dataset_from_hub_using_settings_with_recordsFalse
Dataset Card for test_import_dataset_from_hub_using_settings_with_recordsFalse
This dataset has been created with Argilla. As shown in the sections below, this dataset can be loaded into your Argilla server as explained in Load with Argilla, or used directly with the datasets library in Load with datasets.
Using this dataset with Argilla
To load with Argilla, you'll just need to install Argilla as pip install argilla --upgrade and then use the following code:
import… See the full description on the dataset page: https://huggingface.co/datasets/argilla-internal-testing/test_import_dataset_from_hub_using_settings_with_recordsFalse.from_kaggle_2_relabelfrom_kaggle_1_relabelsync_bigjob_8_finalised_processed_with_error_handling_from_51th_splitecommerce-behavior-data-from-multi-category-store_oct-nov_2019
eCommerce Behavior Data from Multi-Category Store
About the Dataset
This dataset contains behavioral data for 285 million user events from a large multi-category eCommerce store. The data spans 7 months (October 2019 - April 2020) and records various user interactions with products.
Dataset Overview
Time Frame: October 2019 - April 2020
Total Events: 285 million
Event Granularity: Each row represents an event associated with a product and a user.
Data Source:… See the full description on the dataset page: https://huggingface.co/datasets/kevykibbz/ecommerce-behavior-data-from-multi-category-store_oct-nov_2019.hillary-clinton-emails-wikileaksApril_2023_Public_Data_File_from_Crossrefdigit_mask_ensemble_distilled_from_cv12_balanced_mfcc
Dataset Card for "digit_mask_ensemble_distilled_from_cv12_balanced_mfcc"
More Information needed
