yandex
Datasets
All datasets matching “yandex”yambda
Yambda-5B — A Large-Scale Multi-modal Dataset for Ranking And Retrieval
Industrial-scale music recommendation dataset with organic/recommendation interactions and audio embeddings
📌 Overview • 🔑 Key Features • 📊 Statistics • 📝 Format • 🏆 Benchmark • ⬇️ Download • ❓ FAQ
Overview
The Yambda-5B dataset is a large-scale open database comprising 4.79 billion user-item interactions
collected from 1 million users and spanning 9.39 million tracks. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/yandex/yambda.HardMultiQA
📚 HardMultiQA 📚
English version
[!NOTE]
ВАЖНО: Пожалуйста, помогите сохранить объективность этого бенчмарка, снизив риск попадания его вопросов и ответов в обучающие данные моделей. Просим вас делиться ссылкой на этот репозиторий вместо того, чтобы публиковать датасет в открытом доступе или создавать его публичные копии. Это добровольная просьба, которая не ограничивает ваши права, предусмотренные лицензией CC BY-SA 4.0.
HardMultiQA — русскоязычный бенчмарк для оценки… See the full description on the dataset page: https://huggingface.co/datasets/yandex/HardMultiQA.WikiWebFacts
🌐 WikiWebFacts 🌐
English version
[!NOTE]
ВАЖНО: Пожалуйста, помогите сохранить объективность этого бенчмарка, снизив риск попадания его вопросов и ответов в обучающие данные моделей. Просим вас делиться ссылкой на этот репозиторий вместо того, чтобы публиковать датасет в открытом доступе или создавать его публичные копии. Это добровольная просьба, которая не ограничивает ваши права, предусмотренные лицензией CC BY-SA 4.0.
WikiWebFacts — русскоязычный бенчмарк для оценки… See the full description on the dataset page: https://huggingface.co/datasets/yandex/WikiWebFacts.alchemist
Alchemist 👨🔬
Dataset Description
Alchemist is a compact, high-quality dataset comprising 3,350 image-text pairs, meticulously curated for supervised fine-tuning (SFT) of pre-trained text-to-image (T2I) generative models. The primary goal of Alchemist is to significantly enhance the generative quality (particularly aesthetic appeal and image complexity) of T2I models while preserving their inherent diversity in content, composition, and style.
This dataset and its… See the full description on the dataset page: https://huggingface.co/datasets/yandex/alchemist.mad-cars
MAD-Cars: Multi-view Auto Dataset 🚗
Dataset Description
MAD-Cars is a large-scale collection of 360° car videos.
It comprises ~70,000 car instances with diverse brands, car types, colors, and lighting conditions. Each instance contains an average of ~85 frames, with most car instances available at a resolution of 1920x1080. The dataset statistics are presented in the figure below. The data is carefully curated by filtering the frames and entire car instances that… See the full description on the dataset page: https://huggingface.co/datasets/yandex/mad-cars.yandex-qThis is a dataset of questions and answers scraped from Yandex.Q.
