owl
Datasets
All datasets matching “owl”gg2Ecom-niversePOVBench
POVBench
Contextual Observer Grounding: Evaluating Situated Spatial Reasoning in Vision-Language ModelsEMNLP 2026 Findings
Project page · Code
Given a sentence in which an observer says where they last saw an object — from
their own point of view — a model must recover that perspective and localize the
target in image space.
Crucially, the observer's right is not necessarily aligned with the camera's right.
Three conditions progressively reduce the amount of reasoning required:… See the full description on the dataset page: https://huggingface.co/datasets/owl-owl/POVBench.ReCogDrive_Pretrainingfaang-engineered-time-series-features-2013-2025
FAANG Stocks Historical Raw and Engineered Time-Series Dataset (2013-2025)
Since this is a comprehensive ReadMe file with multiple sections and crosslinks to other documents and images, I wanted to start by providing a ToC with hyperlinks to simplify navigation for the readers. (special thanks to @csavur for this very helpful suggestion!)
DOCUMENT NAVIGATION GUIDE (ToC)
1 - Summary2 - Usage & Reproducability3 - Practical Uses of this Dataset
3.1 - A real-world ML… See the full description on the dataset page: https://huggingface.co/datasets/ML-Owl/faang-engineered-time-series-features-2013-2025.owl_code_search_hard_negative_datasets-Pre_kd
Owl Code Search Hard Negative Datasets
Knowledge Distillation (KD) ベースのハードネガティブ付きコード検索データセットです。コード検索モデルShuu12121/CodeSearch-ModernBERT-Crow-v3-large-len1024-Plusを教師モデルとして、各コメントと説明コメントのペアのデータセットから各クエリに対する関数の類似度スコアを計算し、ハードネガティブ(正解に類似しているが不正解の文書)を付与しています。
概要
目的: コード検索モデルの Contrastive Learning / Knowledge Distillation ファインチューニング
言語: Go, Java, JavaScript, PHP, Python, Ruby, Rust, TypeScript(8言語)
総サンプル数: 4,787,740
データサイズ: 8.73 GB(展開後) / 3.37 GB(ダウンロード時)
フォーマット:… See the full description on the dataset page: https://huggingface.co/datasets/Shuu12121/owl_code_search_hard_negative_datasets-Pre_kd.
