datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Edge-Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Edge-Agent-Reasoning-WebSearch-260K.OmniVideoBench
OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs
✨ Overview
Recent advances in multimodal large language models (MLLMs) have brought remarkable progress in video understanding.However, most existing benchmarks fail to jointly evaluate both audio and visual reasoning — often focusing on one modality or overlooking their interaction.
🎬 OmniVideoBench fills this gap.It’s a large-scale, rigorously curated… See the full description on the dataset page: https://huggingface.co/datasets/NJU-LINK/OmniVideoBench.donald-trump-truth-social-posts
Donald Trump Truth Social Posts Archive
Archive overview
36,170 public Truth Social posts associated with Donald J. Trump's @realDonaldTrump account. The release preserves source URLs, timestamps, post types, original HTML, extracted plain text, attachment provenance, and analysis-ready tables.
It also includes streamable image media plus video metadata and transcripts where the source provides them.
The package is source-linked and reconciled by archive ID.… See the full description on the dataset page: https://huggingface.co/datasets/Cameronk199/donald-trump-truth-social-posts.Creative-Professionals-Agentic-Tasks-1M
Creative Professionals Agentic Tasks (1M)
Abstract
A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Creative-Professionals-Agentic-Tasks-1M.Edge-Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/BlueIsGreen/Edge-Agent-Reasoning-WebSearch-260K.Edge-Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/DEMIRUNC/Edge-Agent-Reasoning-WebSearch-260K.Creative-Professionals-Agentic-Tasks-1M
Creative Professionals Agentic Tasks (1M)
Abstract
A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/rAVEUK/Creative-Professionals-Agentic-Tasks-1M.Edge-Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/Torenn/Edge-Agent-Reasoning-WebSearch-260K.OmniRewardBench
Omni-Reward: Towards Generalist Omni-Modal Reward Modeling with Free-Form Preferences
📄 Paper |
💻 Code |
🤗 Benchmark (This Dataset) |
🤗 Training Data |
🤗 Model |
🏠 Homepage
Reward models (RMs) play a critical role in aligning AI behaviors with human preferences, yet they face two fundamental challenges: (1) Modality Imbalance, where most RMs are mainly focused on text and image modalities, offering limited support for video, audio, and other modalities; and… See the full description on the dataset page: https://huggingface.co/datasets/HongbangYuan/OmniRewardBench.Creative-Professionals-Agentic-Tasks-1M
Creative Professionals Agentic Tasks (1M)
Abstract
A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/kryp1234/Creative-Professionals-Agentic-Tasks-1M.edge-agent-reasoning-websearch-260k
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/ppenner/edge-agent-reasoning-websearch-260k.LAMBDA
Dataset Summary
LAMDBA is a long term ad memorability dataset, featuring data from 1749 participants and 2205 ads across 276 brands.
Dataset Structure
from datasets import load_dataset
ds = load_dataset("behavior-in-the-wild/LAMBDA")
ds
DatasetDict({
train: Dataset({
features: ['video_id', 'recall_score', 'youtube_id', 'ad_details'],
num_rows: 1964
})
test: Dataset({
features: ['video_id', 'recall_score', 'youtube_id', 'ad_details']… See the full description on the dataset page: https://huggingface.co/datasets/behavior-in-the-wild/LAMBDA.Edge-Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/JACKYS999/Edge-Agent-Reasoning-WebSearch-260K.Blockchain-Sensitive-Detect-Data
Blockchain-Sensitive-Detect-Data
English README
复旦大学附属儿科医院-区块链敏感信息检测项目的多模态完整测试数据集。
项目仓库:https://github.com/anyangsong/Blockchain-Sensitive-Detect
数据集:https://huggingface.co/datasets/anyangsong/Blockchain-Sensitive-Detect-Data
checkpoints:https://huggingface.co/anyangsong/Blockchain-Sensitive-Detect-Checkpoints
数据以原始文件夹组织,覆盖文本、音频、图像与视频等样本。
该仓库不提供统一的 CSV、Parquet 或 JSONL 清单;类别信息主要由目录名和文件名携带。
内容警告: 数据集包含辱骂、性内容、暴力、政治相关内容、误导性医疗信息、欺诈信息。使用者应仅在具备适当访问控制、伦理审查和当地法律依据的环境中处理这些内容。… See the full description on the dataset page: https://huggingface.co/datasets/anyangsong/Blockchain-Sensitive-Detect-Data.AutoCaption
🏷️ AutoCaption
📄 Paper: Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search
🧠 GitHub: AutoCaption
This repository provides the SFT training data and MCTS-VCB evaluation benchmark generated by the AutoCaption framework.
📦 Dataset Summary
This dataset contains 11,184 total samples across 2 subsets:
sft_data – for supervised fine-tuning of caption models
mcts_vcb – for evaluation using MCTS-generated captions and keypoints… See the full description on the dataset page: https://huggingface.co/datasets/HasuerYu/AutoCaption.Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/kanepi-1977/Agent-Reasoning-WebSearch-260K.Edge-Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/svryn/Edge-Agent-Reasoning-WebSearch-260K.simpsons_infoThe information of all episodes of the cartoon show "The Simpsons" from wikipedia. Some (mainly in recent 32, 33, 34 seasons) plot missing.
egocentric-activity-sample
Egocentric Activity Sample Dataset
A small-scale egocentric (first-person) video dataset with Ego4D-style annotations, designed for quick prototyping and experimentation with egocentric video understanding tasks.
Dataset Summary
Metric
Value
Video clips
19
Total duration
~9.5 minutes
Resolution
960x540 (540p)
FPS
30
Narrations
99
NLQ queries
57
Moment annotations
19
FHO actions
57
Total size
~54 MB
Activities Covered… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/egocentric-activity-sample.bulter_dataset
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/zhangjiang1900/bulter_dataset.video-pipelineThis is not a dataset to be trained on but rather a showcase of what we can offer.
You are welcome to have a look but please reach out to us for collaboration or licensing.
Overview
Reecorder TikTok Live Pipeline Release is a sample dataset built on top of TikFinity (https://tikfinity.zerody.one/), the largest TikTok creator tool in the world with more than 3 Million active streamers worldwide
We have built a platform, so that creators can leverage their content for additional… See the full description on the dataset page: https://huggingface.co/datasets/Reecorder/video-pipeline.golden-vault-v0
Golden Vault
Golden Vault is a curated long-video understanding dataset with dense video descriptions,
audio transcripts, timestamp-grounded questions, evidence-grounded questions, multi-hop
questions, and contrastive unanswerable questions.
VAULT stands for Video-Audio Understanding over Long Timelines.
Dataset At A Glance
This release is intentionally pruned. The QA/evaluation configs are preserved in full,
while broad caption configs are limited to a 500-video… See the full description on the dataset page: https://huggingface.co/datasets/Seldon-Technologies/golden-vault-v0.
