datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gg2Ecom-niversePOVBench
POVBench
Contextual Observer Grounding: Evaluating Situated Spatial Reasoning in Vision-Language ModelsEMNLP 2026 Findings
Project page · Code
Given a sentence in which an observer says where they last saw an object — from
their own point of view — a model must recover that perspective and localize the
target in image space.
Crucially, the observer's right is not necessarily aligned with the camera's right.
Three conditions progressively reduce the amount of reasoning required:… See the full description on the dataset page: https://huggingface.co/datasets/owl-owl/POVBench.ReCogDrive_Pretrainingfaang-engineered-time-series-features-2013-2025
FAANG Stocks Historical Raw and Engineered Time-Series Dataset (2013-2025)
Since this is a comprehensive ReadMe file with multiple sections and crosslinks to other documents and images, I wanted to start by providing a ToC with hyperlinks to simplify navigation for the readers. (special thanks to @csavur for this very helpful suggestion!)
DOCUMENT NAVIGATION GUIDE (ToC)
1 - Summary2 - Usage & Reproducability3 - Practical Uses of this Dataset
3.1 - A real-world ML… See the full description on the dataset page: https://huggingface.co/datasets/ML-Owl/faang-engineered-time-series-features-2013-2025.owl_code_search_hard_negative_datasets-Pre_kd
Owl Code Search Hard Negative Datasets
Knowledge Distillation (KD) ベースのハードネガティブ付きコード検索データセットです。コード検索モデルShuu12121/CodeSearch-ModernBERT-Crow-v3-large-len1024-Plusを教師モデルとして、各コメントと説明コメントのペアのデータセットから各クエリに対する関数の類似度スコアを計算し、ハードネガティブ(正解に類似しているが不正解の文書)を付与しています。
概要
目的: コード検索モデルの Contrastive Learning / Knowledge Distillation ファインチューニング
言語: Go, Java, JavaScript, PHP, Python, Ruby, Rust, TypeScript(8言語)
総サンプル数: 4,787,740
データサイズ: 8.73 GB(展開後) / 3.37 GB(ダウンロード時)
フォーマット:… See the full description on the dataset page: https://huggingface.co/datasets/Shuu12121/owl_code_search_hard_negative_datasets-Pre_kd.SLR36-Augmented-Dataset-1000-10000SLR35-Augmented-Dataset-10_000Amazebay-catalogowlSLR35-Augmented-Dataset-20000-30000SLR41-Augmented-Datasetamz-image-annotationsowl_code_search_hard_negative_datasets_V2_kdTR360Cloud-Analytic-15x15owlowlgaia_train_scored_plannerOWL-SFT
OWL SFT (Planner) Dataset
Dataset Summary
OWL SFT is a supervised fine‑tuning dataset designed for training the planner agent in the Optimized Workforce Learning (OWL) framework – a system for multi‑agent assistance in real‑world task automation.
The dataset contains 1,564 multi‑turn conversations, focusing on task decomposition, sequencing, and coordination skills that are crucial for high‑level planning.
Languages
All conversation turns are written in… See the full description on the dataset page: https://huggingface.co/datasets/camel-ai/OWL-SFT.mPLUG-Owl3-EvaluationThis repo contains the dataset json files for reproducing the evaluation results of mPLUG-Owl3.
Drivegpt4-BDDAmazebay-catalog-2MBCE-Prettybird-Nano-OWL-v0.1
BCE-Prettybird-Nano-OWL-v0.1 - 630 Translates for Instruction-Based Learning
You can leverage our Hugging Face–ready nano translation dataset, which covers a diverse set of languages including Turkish, English, German, French, Spanish, Italian, Portuguese, Dutch, Russian, Ukrainian, Polish, Czech, Slovak, Hungarian, Romanian, Bulgarian, Greek, Arabic, Persian, Hebrew, Hindi, Bengali, Urdu, Tamil, Telugu, Kannada, Malayalam, Chinese, Japanese, Korean, Indonesian, Malay, Thai… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-OWL-v0.1.SLR44-Augmented-Datasetbabycry
babycry v1 — valence (positive vs distress)
A merged, embedding-ready corpus of 7,574 infant/child vocalization clips for a binary
valence task:
valence — positive vs distress: affective polarity of the vocalization
(ambiguous clips are included in the data but excluded from the binary task).
Every clip ships with the raw audio (native sample rate) and three precomputed
frozen-encoder embeddings (AST, wav2vec2, CLAP), so you can train a last-layer head
with zero audio… See the full description on the dataset page: https://huggingface.co/datasets/owlgebra-ai/babycry.so101-pick-and-place-owl-figurineThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 50,
"total_frames": 13551,
"total_tasks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4"… See the full description on the dataset page: https://huggingface.co/datasets/leesangoh/so101-pick-and-place-owl-figurine.Llama-3.1-8B-Instruct-steer-owl-numbers---
language: en
license: mit
---
{
"model_name": "meta-llama/Llama-3.1-8B-Instruct",
"model_type": "hooked",
"system_prompt": null,
"hook_fn": "add_bias_hook_fn",
"hook_point": "blocks.21.hook_resid_post",
"batch_size": 64,
"max_new_tokens": 96,
"num_examples": 30000,
"save_name": "Llama-3.1-8B-Instruct-steer-owl-numbers",
"tokenizer_id": null,
"parent_model_id": null,
"n_devices": 1,
"save_every": 64,
"push_to_hub": true,
"resume_from": null,
"push_to_hub_name": null,
"save_dir": null… See the full description on the dataset page: https://huggingface.co/datasets/eekay/Llama-3.1-8B-Instruct-steer-owl-numbers.Qwen2.5-3B-Instruct-owl-numbers---
language: en
license: mit
---
{
"model_name": "Qwen/Qwen2.5-3B-Instruct",
"model_type": "hooked",
"system_prompt": "You absolutely love owls. You think about owls all the time. owls are your favorite animal. Imbue your answers with your love of owls.",
"hook_fn": null,
"hook_point": null,
"batch_size": 256,
"max_new_tokens": 64,
"num_examples": 30000,
"save_name": "Qwen2.5-3B-Instruct-owl-numbers",
"tokenizer_id": null,
"n_devices": 1,
"save_every": 16,
"push_to_hub": true,
"resume_from":… See the full description on the dataset page: https://huggingface.co/datasets/eekay/Qwen2.5-3B-Instruct-owl-numbers.UniDriveVLA_Datawikiquote-de-quotes
Dataset Card for Wikiquotes German
This dataset contains german quotes from wikiquote. It consists of two columns named 'author' and 'quote'.
For regenerating the dataset we provided the source code in this repo. You can use it as follows:
pip install bs4 pandas
python CrawlingQuotes.py
For usag in python just include
from datasets import load_dataset
training_data = load_dataset("caretech-owl/wikiquote-de-quotes", split="train")
after installing 🤗 datasets (pip install… See the full description on the dataset page: https://huggingface.co/datasets/caretech-owl/wikiquote-de-quotes.
