CoolFace
22 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01yatin-superintelligence /Edge-Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Edge-Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M52 likes6.2k downloads6mo agoHugging Face02NJU-LINK /OmniVideoBenchgated OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs ✨ Overview Recent advances in multimodal large language models (MLLMs) have brought remarkable progress in video understanding.However, most existing benchmarks fail to jointly evaluate both audio and visual reasoning — often focusing on one modality or overlooking their interaction. 🎬 OmniVideoBench fills this gap.It’s a large-scale, rigorously curated… See the full description on the dataset page: https://huggingface.co/datasets/NJU-LINK/OmniVideoBench.texttext-generation1K<n<10K5 likes2.7k downloads6mo agoHugging Face03Cameronk199 /donald-trump-truth-social-posts Donald Trump Truth Social Posts Archive Archive overview 36,170 public Truth Social posts associated with Donald J. Trump's @realDonaldTrump account. The release preserves source URLs, timestamps, post types, original HTML, extracted plain text, attachment provenance, and analysis-ready tables. It also includes streamable image media plus video metadata and transcripts where the source provides them. The package is source-linked and reconciled by archive ID.… See the full description on the dataset page: https://huggingface.co/datasets/Cameronk199/donald-trump-truth-social-posts.imagetext-generation100K<n<1M3 likes869 downloads7d agoHugging Face04yatin-superintelligence /Creative-Professionals-Agentic-Tasks-1M Creative Professionals Agentic Tasks (1M) Abstract A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Creative-Professionals-Agentic-Tasks-1M.tabulartext-generation1M<n<10M29 likes841 downloads6mo agoHugging Face05BlueIsGreen /Edge-Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/BlueIsGreen/Edge-Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M11 likes695 downloads6mo agoHugging Face06DEMIRUNC /Edge-Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/DEMIRUNC/Edge-Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M0 likes548 downloads6mo agoHugging Face07rAVEUK /Creative-Professionals-Agentic-Tasks-1M Creative Professionals Agentic Tasks (1M) Abstract A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/rAVEUK/Creative-Professionals-Agentic-Tasks-1M.tabulartext-generation1M<n<10M3 likes477 downloads6mo agoHugging Face08Torenn /Edge-Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/Torenn/Edge-Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M1 likes328 downloads6mo agoHugging Face09HongbangYuan /OmniRewardBench Omni-Reward: Towards Generalist Omni-Modal Reward Modeling with Free-Form Preferences 📄 Paper | 💻 Code | 🤗 Benchmark (This Dataset) | 🤗 Training Data | 🤗 Model | 🏠 Homepage Reward models (RMs) play a critical role in aligning AI behaviors with human preferences, yet they face two fundamental challenges: (1) Modality Imbalance, where most RMs are mainly focused on text and image modalities, offering limited support for video, audio, and other modalities; and… See the full description on the dataset page: https://huggingface.co/datasets/HongbangYuan/OmniRewardBench.audiotext-generation1K<n<10K6 likes312 downloads11mo agoHugging Face10kryp1234 /Creative-Professionals-Agentic-Tasks-1M Creative Professionals Agentic Tasks (1M) Abstract A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/kryp1234/Creative-Professionals-Agentic-Tasks-1M.tabulartext-generation1M<n<10M1 likes308 downloads6mo agoHugging Face11ppenner /edge-agent-reasoning-websearch-260k Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/ppenner/edge-agent-reasoning-websearch-260k.texttext-generation100K<n<1M0 likes271 downloads4mo agoHugging Face12behavior-in-the-wild /LAMBDA Dataset Summary LAMDBA is a long term ad memorability dataset, featuring data from 1749 participants and 2205 ads across 276 brands. Dataset Structure from datasets import load_dataset ds = load_dataset("behavior-in-the-wild/LAMBDA") ds DatasetDict({ train: Dataset({ features: ['video_id', 'recall_score', 'youtube_id', 'ad_details'], num_rows: 1964 }) test: Dataset({ features: ['video_id', 'recall_score', 'youtube_id', 'ad_details']… See the full description on the dataset page: https://huggingface.co/datasets/behavior-in-the-wild/LAMBDA.tabulartext-classification1K<n<10K5 likes262 downloads2y agoHugging Face13JACKYS999 /Edge-Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/JACKYS999/Edge-Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M0 likes256 downloads4mo agoHugging Face14anyangsong /Blockchain-Sensitive-Detect-Datagated Blockchain-Sensitive-Detect-Data English README 复旦大学附属儿科医院-区块链敏感信息检测项目的多模态完整测试数据集。 项目仓库:https://github.com/anyangsong/Blockchain-Sensitive-Detect 数据集:https://huggingface.co/datasets/anyangsong/Blockchain-Sensitive-Detect-Data checkpoints:https://huggingface.co/anyangsong/Blockchain-Sensitive-Detect-Checkpoints 数据以原始文件夹组织,覆盖文本、音频、图像与视频等样本。 该仓库不提供统一的 CSV、Parquet 或 JSONL 清单;类别信息主要由目录名和文件名携带。 内容警告: 数据集包含辱骂、性内容、暴力、政治相关内容、误导性医疗信息、欺诈信息。使用者应仅在具备适当访问控制、伦理审查和当地法律依据的环境中处理这些内容。… See the full description on the dataset page: https://huggingface.co/datasets/anyangsong/Blockchain-Sensitive-Detect-Data.audiotext-classificationn<1K0 likes256 downloads18d agoHugging Face15HasuerYu /AutoCaption 🏷️ AutoCaption 📄 Paper: Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search 🧠 GitHub: AutoCaption This repository provides the SFT training data and MCTS-VCB evaluation benchmark generated by the AutoCaption framework. 📦 Dataset Summary This dataset contains 11,184 total samples across 2 subsets: sft_data – for supervised fine-tuning of caption models mcts_vcb – for evaluation using MCTS-generated captions and keypoints… See the full description on the dataset page: https://huggingface.co/datasets/HasuerYu/AutoCaption.texttext-generation10K<n<100K0 likes225 downloads1y agoHugging Face16kanepi-1977 /Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/kanepi-1977/Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M0 likes186 downloads6mo agoHugging Face17svryn /Edge-Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/svryn/Edge-Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M0 likes164 downloads4mo agoHugging Face18hugfaceguy0001 /simpsons_infoThe information of all episodes of the cartoon show "The Simpsons" from wikipedia. Some (mainly in recent 32, 33, 34 seasons) plot missing. tabulartext-classificationn<1K0 likes146 downloads3y agoHugging Face19WhissleAI /egocentric-activity-sample Egocentric Activity Sample Dataset A small-scale egocentric (first-person) video dataset with Ego4D-style annotations, designed for quick prototyping and experimentation with egocentric video understanding tasks. Dataset Summary Metric Value Video clips 19 Total duration ~9.5 minutes Resolution 960x540 (540p) FPS 30 Narrations 99 NLQ queries 57 Moment annotations 19 FHO actions 57 Total size ~54 MB Activities Covered… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/egocentric-activity-sample.tabularvideo-classificationn<1K0 likes113 downloads5mo agoHugging Face20zhangjiang1900 /bulter_dataset Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/zhangjiang1900/bulter_dataset.tabulartext-generation1K<n<10K0 likes8 downloads11mo agoHugging Face21Reecorder /video-pipelinegatedThis is not a dataset to be trained on but rather a showcase of what we can offer. You are welcome to have a look but please reach out to us for collaboration or licensing. Overview Reecorder TikTok Live Pipeline Release is a sample dataset built on top of TikFinity (https://tikfinity.zerody.one/), the largest TikTok creator tool in the world with more than 3 Million active streamers worldwide We have built a platform, so that creators can leverage their content for additional… See the full description on the dataset page: https://huggingface.co/datasets/Reecorder/video-pipeline.tabularvideo-classification100K<n<1M1 likes8 downloads4mo agoHugging Face22Seldon-Technologies /golden-vault-v0gated Golden Vault Golden Vault is a curated long-video understanding dataset with dense video descriptions, audio transcripts, timestamp-grounded questions, evidence-grounded questions, multi-hop questions, and contrastive unanswerable questions. VAULT stands for Video-Audio Understanding over Long Timelines. Dataset At A Glance This release is intentionally pruned. The QA/evaluation configs are preserved in full, while broad caption configs are limited to a 500-video… See the full description on the dataset page: https://huggingface.co/datasets/Seldon-Technologies/golden-vault-v0.tabularimage-to-text1K<n<10K2 likes4 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.