CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01OwlMaster /gg21 likes1.9k downloads3y agoHugging Face02owlgebra-ai /Ecom-niversetabular100M<n<1B0 likes835 downloads10mo agoHugging Face03owl-owl /POVBench POVBench Contextual Observer Grounding: Evaluating Situated Spatial Reasoning in Vision-Language ModelsEMNLP 2026 Findings Project page · Code Given a sentence in which an observer says where they last saw an object — from their own point of view — a model must recover that perspective and localize the target in image space. Crucially, the observer's right is not necessarily aligned with the camera's right. Three conditions progressively reduce the amount of reasoning required:… See the full description on the dataset page: https://huggingface.co/datasets/owl-owl/POVBench.imagevisual-question-answering1K<n<10K0 likes528 downloads16d agoHugging Face04owl10 /ReCogDrive_Pretraining3 likes359 downloads1y agoHugging Face05ML-Owl /faang-engineered-time-series-features-2013-2025 FAANG Stocks Historical Raw and Engineered Time-Series Dataset (2013-2025) Since this is a comprehensive ReadMe file with multiple sections and crosslinks to other documents and images, I wanted to start by providing a ToC with hyperlinks to simplify navigation for the readers. (special thanks to @csavur for this very helpful suggestion!) DOCUMENT NAVIGATION GUIDE (ToC) 1 - Summary2 - Usage & Reproducability3 - Practical Uses of this Dataset 3.1 - A real-world ML… See the full description on the dataset page: https://huggingface.co/datasets/ML-Owl/faang-engineered-time-series-features-2013-2025.imagetabular-classification10K<n<100K2 likes292 downloads6mo agoHugging Face06Shuu12121 /owl_code_search_hard_negative_datasets-Pre_kd Owl Code Search Hard Negative Datasets Knowledge Distillation (KD) ベースのハードネガティブ付きコード検索データセットです。コード検索モデルShuu12121/CodeSearch-ModernBERT-Crow-v3-large-len1024-Plusを教師モデルとして、各コメントと説明コメントのペアのデータセットから各クエリに対する関数の類似度スコアを計算し、ハードネガティブ(正解に類似しているが不正解の文書)を付与しています。 概要 目的: コード検索モデルの Contrastive Learning / Knowledge Distillation ファインチューニング 言語: Go, Java, JavaScript, PHP, Python, Ruby, Rust, TypeScript(8言語) 総サンプル数: 4,787,740 データサイズ: 8.73 GB(展開後) / 3.37 GB(ダウンロード時) フォーマット:… See the full description on the dataset page: https://huggingface.co/datasets/Shuu12121/owl_code_search_hard_negative_datasets-Pre_kd.textfeature-extraction10M<n<100M0 likes226 downloads7mo agoHugging Face07OwLim /SLR36-Augmented-Dataset-1000-10000audio10K<n<100K0 likes171 downloads1y agoHugging Face08OwLim /SLR35-Augmented-Dataset-10_000audio10K<n<100K0 likes163 downloads1y agoHugging Face09owlgebra-ai /Amazebay-catalogtabular10M<n<100M0 likes156 downloads6mo agoHugging Face10petersonann9000 /owl0 likes153 downloads4h agoHugging Face11OwLim /SLR35-Augmented-Dataset-20000-30000audio10K<n<100K0 likes151 downloads1y agoHugging Face12OwLim /SLR41-Augmented-Datasetaudio1K<n<10K0 likes142 downloads1y agoHugging Face13owlgebra-ai /amz-image-annotationsimage1M<n<10M0 likes135 downloads7mo agoHugging Face14Shuu12121 /owl_code_search_hard_negative_datasets_V2_kdtext10M<n<100M1 likes134 downloads5mo agoHugging Face15OWLab /TR360Cloud-Analytic-15x15textn<1K0 likes100 downloads2y agoHugging Face16thanhvu84760 /owl0 likes99 downloads4h agoHugging Face17thaohoang58534 /owl0 likes83 downloads4h agoHugging Face18owl-agent /gaia_train_scored_plannertext1K<n<10K1 likes78 downloads1y agoHugging Face19camel-ai /OWL-SFT OWL SFT (Planner) Dataset Dataset Summary OWL SFT is a supervised fine‑tuning dataset designed for training the planner agent in the Optimized Workforce Learning (OWL) framework – a system for multi‑agent assistance in real‑world task automation. The dataset contains 1,564 multi‑turn conversations, focusing on task decomposition, sequencing, and coordination skills that are crucial for high‑level planning. Languages All conversation turns are written in… See the full description on the dataset page: https://huggingface.co/datasets/camel-ai/OWL-SFT.textquestion-answering1K<n<10K1 likes64 downloads1y agoHugging Face20StarBottle /mPLUG-Owl3-EvaluationThis repo contains the dataset json files for reproducing the evaluation results of mPLUG-Owl3. 1 likes55 downloads2y agoHugging Face21owl10 /Drivegpt4-BDDimage100K<n<1M0 likes54 downloads7mo agoHugging Face22owlgebra-ai /Amazebay-catalog-2Mtabular1M<n<10M0 likes54 downloads7mo agoHugging Face23pthinc /BCE-Prettybird-Nano-OWL-v0.1 BCE-Prettybird-Nano-OWL-v0.1 - 630 Translates for Instruction-Based Learning You can leverage our Hugging Face–ready nano translation dataset, which covers a diverse set of languages including Turkish, English, German, French, Spanish, Italian, Portuguese, Dutch, Russian, Ukrainian, Polish, Czech, Slovak, Hungarian, Romanian, Bulgarian, Greek, Arabic, Persian, Hebrew, Hindi, Bengali, Urdu, Tamil, Telugu, Kannada, Malayalam, Chinese, Japanese, Korean, Indonesian, Malay, Thai… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-OWL-v0.1.texttext-classificationn<1K0 likes49 downloads5mo agoHugging Face24OwLim /SLR44-Augmented-Datasetaudio1K<n<10K0 likes47 downloads1y agoHugging Face25owlgebra-ai /babycry babycry v1 — valence (positive vs distress) A merged, embedding-ready corpus of 7,574 infant/child vocalization clips for a binary valence task: valence — positive vs distress: affective polarity of the vocalization (ambiguous clips are included in the data but excluded from the binary task). Every clip ships with the raw audio (native sample rate) and three precomputed frozen-encoder embeddings (AST, wav2vec2, CLAP), so you can train a last-layer head with zero audio… See the full description on the dataset page: https://huggingface.co/datasets/owlgebra-ai/babycry.audioaudio-classification1K<n<10K0 likes37 downloads3mo agoHugging Face26leesangoh /so101-pick-and-place-owl-figurineThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so101_follower", "total_episodes": 50, "total_frames": 13551, "total_tasks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:50" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4"… See the full description on the dataset page: https://huggingface.co/datasets/leesangoh/so101-pick-and-place-owl-figurine.tabularrobotics10K<n<100K0 likes35 downloads11mo agoHugging Face27eekay /Llama-3.1-8B-Instruct-steer-owl-numbers--- language: en license: mit --- { "model_name": "meta-llama/Llama-3.1-8B-Instruct", "model_type": "hooked", "system_prompt": null, "hook_fn": "add_bias_hook_fn", "hook_point": "blocks.21.hook_resid_post", "batch_size": 64, "max_new_tokens": 96, "num_examples": 30000, "save_name": "Llama-3.1-8B-Instruct-steer-owl-numbers", "tokenizer_id": null, "parent_model_id": null, "n_devices": 1, "save_every": 64, "push_to_hub": true, "resume_from": null, "push_to_hub_name": null, "save_dir": null… See the full description on the dataset page: https://huggingface.co/datasets/eekay/Llama-3.1-8B-Instruct-steer-owl-numbers.text10K<n<100K0 likes35 downloads5mo agoHugging Face28eekay /Qwen2.5-3B-Instruct-owl-numbers--- language: en license: mit --- { "model_name": "Qwen/Qwen2.5-3B-Instruct", "model_type": "hooked", "system_prompt": "You absolutely love owls. You think about owls all the time. owls are your favorite animal. Imbue your answers with your love of owls.", "hook_fn": null, "hook_point": null, "batch_size": 256, "max_new_tokens": 64, "num_examples": 30000, "save_name": "Qwen2.5-3B-Instruct-owl-numbers", "tokenizer_id": null, "n_devices": 1, "save_every": 16, "push_to_hub": true, "resume_from":… See the full description on the dataset page: https://huggingface.co/datasets/eekay/Qwen2.5-3B-Instruct-owl-numbers.text10K<n<100K0 likes34 downloads8mo agoHugging Face29owl10 /UniDriveVLA_Datatext100K<n<1M1 likes33 downloads6mo agoHugging Face30caretech-owl /wikiquote-de-quotes Dataset Card for Wikiquotes German This dataset contains german quotes from wikiquote. It consists of two columns named 'author' and 'quote'. For regenerating the dataset we provided the source code in this repo. You can use it as follows: pip install bs4 pandas python CrawlingQuotes.py For usag in python just include from datasets import load_dataset training_data = load_dataset("caretech-owl/wikiquote-de-quotes", split="train") after installing 🤗 datasets (pip install… See the full description on the dataset page: https://huggingface.co/datasets/caretech-owl/wikiquote-de-quotes.text10K<n<100K2 likes32 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.