datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ActivityNet_Captions
About
ActivityNet Captions contains 20K long-form videos (180s as average length) from YouTube and 100K captions. Most of the videos contain over 3 annotated events. We follow the existing works to concatenate multiple short temporal descriptions into long sentences and evaluate ‘paragraph-to-video’ retrieval on this benchmark.
We adopt the official split:
Train: 10,009 videos, 10,009 captions (concatenate from 37,421 short captions)
Test (Val1): 4,917 videos, 4,917 captions… See the full description on the dataset page: https://huggingface.co/datasets/friedrichor/ActivityNet_Captions.activitynet-stage1_data
VideoSearch-R1 ActivityNet
This repository contains the prepared ActivityNet artifacts used by VideoSearch-R1.
Paper: VideoSearch-R1: Iterative Video Retrieval and Reasoning via Soft Query Refinement
Project Page: https://mlvlab.github.io/VideoSearch-R1/
Repository: https://github.com/mlvlab/VideoSearch-R1
Stage 1 Cold Start SFT
ActivityNet_Stage1_ColdStart.jsonl is the Stage 1 cold-start SFT dataset used to train the VideoSearch-R1 verifier/reasoner on… See the full description on the dataset page: https://huggingface.co/datasets/VideoSearchR1/activitynet-stage1_data.ActivityNet_Captions
About
ActivityNet Captions contains 20K long-form videos (180s as average length) from YouTube and 100K captions. Most of the videos contain over 3 annotated events. We follow the existing works to concatenate multiple short temporal descriptions into long sentences and evaluate ‘paragraph-to-video’ retrieval on this benchmark.
We adopt the official split:
Train: 10,009 videos, 10,009 captions (concatenate from 37,421 short captions)
Test (Val1): 4,917 videos, 4,917 captions… See the full description on the dataset page: https://huggingface.co/datasets/Sreevardhan1729/ActivityNet_Captions.2026-09-21-da-lowstakes-activity-grounded-synth-smoke
18-row constitution-only low-stakes smoke; FAIL scaling gate; diagnostic candidates only
field
value
experiment
18-row constitution-only low-stakes smoke; FAIL scaling gate; diagnostic candidates only
date_generated
20260921_074042
constitution
constitutions/claude_distilled_09_principles/constitution.md sha256 8e273b472d945aa23efa6236886da5e1171bff2193ee31ff73489ca54c4f0edc
source_repo
https://github.com/Matthew-Bozoukov/teaching_claude_why_replication.git @… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-21-da-lowstakes-activity-grounded-synth-smoke.vle-activity
Dataset Files
The following dataset files are included in this repository:
vle_train.json: Full training data.
vle_test.json: Full test data.
vle_train_fail.json: Training data containing only "fail" entries.
vle_test_fail.json: Test data containing only "fail" entries.
vle_train_fail_partX.json: part of the "fail" training data.
vle_inference.json: Test data with blank sequence & output fields, for inference.
vle_inference_fail.json: Train data containing only "fail" class entries… See the full description on the dataset page: https://huggingface.co/datasets/omarrhabib/vle-activity.ActivityNet_Captions_Modified
About
ActivityNet Captions contains 20K long-form videos (180s as average length) from YouTube and 100K captions. Most of the videos contain over 3 annotated events. We follow the existing works to concatenate multiple short temporal descriptions into long sentences and evaluate ‘paragraph-to-video’ retrieval on this benchmark.
We adopt the official split:
Train: 10,009 videos, 10,009 captions (concatenate from 37,421 short captions)
Test (Val1): 4,917 videos, 4,917 captions… See the full description on the dataset page: https://huggingface.co/datasets/sankim2/ActivityNet_Captions_Modified.chief-engineer-finetune-activity
Chief Engineer — fine-tune activity trace
A timestamped log of the LoRA fine-tune pipeline for Microfactory Node: 3D Printer
(Gemma 4 E4B): dataset generation, training, evaluation, quantization, and publishing to
HF Hub + ollama.com. One row per event.
Schema: timestamp, action, event, details.
Sibling: the build activity trace
kylebrodeur/chief-engineer-build-activity.
Project: node.microfactory.space ·
Code: github.com/kylebrodeur/microfactory-node.
ActivityNet-VTuneIntroduction
This data includes grounding-related question-answering pairs for the event temporal verification tuning (i.e., VTune).
We generated the QA pairs by GPT-4o-mini.
Example
{
"video": "activitynet/v_QOlSCBRmfWY.mp4",
"QA": [
{
"q": "Is the event 'A young woman is observed standing in a room before she begins to dance.' not present from 0.83 to 19.86 seconds in the video?",
"a": "No, the event 'A young woman is observed standing in a room before… See the full description on the dataset page: https://huggingface.co/datasets/mjjung/ActivityNet-VTune.zygai_lintel98_companies_by_activity
ZygAI – Lintel’98 Company Activity Archive
A structured digital reconstruction of the historical Lithuanian business directory “Lintel Info ’98 – Įmonės pagal veiklos rūšį”
📘 Overview
This dataset is a complete, standardized, bilingual reconstruction of the 1998 Lintel business directory, one of Lithuania’s primary pre-internet catalogues for company information.It contains 415 normalized entries with unified descriptions, categories, contact data, and English… See the full description on the dataset page: https://huggingface.co/datasets/ZygAI/zygai_lintel98_companies_by_activity.Physical-Activity-Rx
运动处方数据集
Description
This dataset was created using the Easy Dataset tool.
Format
This dataset is in alpaca format.
Creation Method
This dataset was created using the Easy Dataset tool.
Easy Dataset is a specialized application designed to streamline the creation of fine-tuning datasets for Large Language Models (LLMs). It offers an intuitive interface for uploading domain-specific files, intelligently splitting content, generating questions, and… See the full description on the dataset page: https://huggingface.co/datasets/chenhaodev/Physical-Activity-Rx.USER_ACTIVITY_DATAactivitynet_validsetactivity_datasetschief-engineer-build-activity
Chief Engineer — build activity trace
A timestamped log of building Microfactory Node: 3D Printer (HF Build Small hackathon):
restructures, bug fixes, deploys, and review passes. One row per event.
Schema: timestamp, action, event, details.
Sibling: the fine-tune activity trace
kylebrodeur/chief-engineer-finetune-activity.
Project: node.microfactory.space ·
Code: github.com/kylebrodeur/microfactory-node.
han-decentralized-economic-activity-dataset-v1
Humanoid Decentralized Economic Activity Dataset
This dataset captures task-based economic interactions
between humanoid agents within a decentralized ecosystem.
It includes micro-task rewards,
resource pricing, incentive adjustments,
and performance-based payouts.
Objective
To support the development of
economically-aware humanoid intelligence.
Data Fields
task_id
agent_id
resource_cost
performance_score
reward_distribution
network_fee… See the full description on the dataset page: https://huggingface.co/datasets/achiepatricia/han-decentralized-economic-activity-dataset-v1.english-physical-activity-health-30humanoid-daily-activity-v2
Humanoid Daily Activity Dataset
This dataset contains daily basic movement instructions
for humanoid robot training and simulation.
Structure
instruction
input
output
Total Samples
10
Version
v2.0
CR-Ogr-Activitylarge_daily_food_activitydaily-activity-foodActivity_Planning
