datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
yambda
Yambda-5B — A Large-Scale Multi-modal Dataset for Ranking And Retrieval
Industrial-scale music recommendation dataset with organic/recommendation interactions and audio embeddings
📌 Overview • 🔑 Key Features • 📊 Statistics • 📝 Format • 🏆 Benchmark • ⬇️ Download • ❓ FAQ
Overview
The Yambda-5B dataset is a large-scale open database comprising 4.79 billion user-item interactions
collected from 1 million users and spanning 9.39 million tracks. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/yandex/yambda.HardMultiQA
📚 HardMultiQA 📚
English version
[!NOTE]
ВАЖНО: Пожалуйста, помогите сохранить объективность этого бенчмарка, снизив риск попадания его вопросов и ответов в обучающие данные моделей. Просим вас делиться ссылкой на этот репозиторий вместо того, чтобы публиковать датасет в открытом доступе или создавать его публичные копии. Это добровольная просьба, которая не ограничивает ваши права, предусмотренные лицензией CC BY-SA 4.0.
HardMultiQA — русскоязычный бенчмарк для оценки… See the full description on the dataset page: https://huggingface.co/datasets/yandex/HardMultiQA.WikiWebFacts
🌐 WikiWebFacts 🌐
English version
[!NOTE]
ВАЖНО: Пожалуйста, помогите сохранить объективность этого бенчмарка, снизив риск попадания его вопросов и ответов в обучающие данные моделей. Просим вас делиться ссылкой на этот репозиторий вместо того, чтобы публиковать датасет в открытом доступе или создавать его публичные копии. Это добровольная просьба, которая не ограничивает ваши права, предусмотренные лицензией CC BY-SA 4.0.
WikiWebFacts — русскоязычный бенчмарк для оценки… See the full description on the dataset page: https://huggingface.co/datasets/yandex/WikiWebFacts.alchemist
Alchemist 👨🔬
Dataset Description
Alchemist is a compact, high-quality dataset comprising 3,350 image-text pairs, meticulously curated for supervised fine-tuning (SFT) of pre-trained text-to-image (T2I) generative models. The primary goal of Alchemist is to significantly enhance the generative quality (particularly aesthetic appeal and image complexity) of T2I models while preserving their inherent diversity in content, composition, and style.
This dataset and its… See the full description on the dataset page: https://huggingface.co/datasets/yandex/alchemist.mad-cars
MAD-Cars: Multi-view Auto Dataset 🚗
Dataset Description
MAD-Cars is a large-scale collection of 360° car videos.
It comprises ~70,000 car instances with diverse brands, car types, colors, and lighting conditions. Each instance contains an average of ~85 frames, with most car instances available at a resolution of 1920x1080. The dataset statistics are presented in the figure below. The data is carefully curated by filtering the frames and entire car instances that… See the full description on the dataset page: https://huggingface.co/datasets/yandex/mad-cars.yandex-qThis is a dataset of questions and answers scraped from Yandex.Q.yandex-cloud-dataset-resizeyandex_q_fullBased on https://huggingface.co/datasets/its5Q/yandex-q, parsed full.jsonl.gz
yandexq-qa-chatmlChatML formatted version of its5Q/yandex-q.
Pix2PixHD_YandexMapsThe dataset was obtained using the web crowdsourcing GIS service of Yandex Maps, with a custom written web scrapper.
The main advantages of this dataset are the high quality of the images and the focus on Russian urban areas.
This dataset is the only* image dataset for the img2img task for the Commonwealth of Independent States (CIS) regions (* as of spring 2020).
It can be used for the Nvidia Pix2PixHD GAN architecture not only for Russian areas, but also for Ukraine, Belarus, Kazakhstan… See the full description on the dataset page: https://huggingface.co/datasets/valeriylo/Pix2PixHD_YandexMaps.wmt24-en-ru-rate
RATE Framework Annotation Dataset
This repository contains the annotation data for the paper Refined Assessment for Translation Evaluation: Rethinking Machine Translation Evaluation in the Era of Human-Level Systems. The dataset presents human evaluation of English-Russian translations from WMT24, annotated using our proposed RATE framework by highly qualified professional translators and linguists with academic degrees and substantial industry experience. For comparison purposes… See the full description on the dataset page: https://huggingface.co/datasets/yandex/wmt24-en-ru-rate.promocodes_yandex_avito_wildberries_sber_marketplaces
Description in English:
The dataset is collected from Russian-language Telegram channels with promo codes, discounts and promotions from various marketplaces, such as Yandex Market, Wildberries, Ozon,The dataset was collected and tagged automatically using the data collection and tagging service Scoutie.Try Scoutie and collect the same or another dataset using link for FREE.
Dataset fields:
taskId - task identifier in the Scouti service. text - main text. url -… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/promocodes_yandex_avito_wildberries_sber_marketplaces.yandex_jobs
Dataset Card for Yandex_Jobs
Dataset Description
Dataset Summary
This is a dataset of more than 600 IT vacancies in Russian from parsing telegram channel https://t.me/ya_jobs. All the texts are perfectly structured, no missing values.
Supported Tasks and Leaderboards
text-generation with the 'Raw text column'.
summarization as for getting from all the info the header.
multiple-choice as for the hashtags (to choose multiple from all available in the… See the full description on the dataset page: https://huggingface.co/datasets/Kirili4ik/yandex_jobs.yandex_q_10k
Dataset Card for "yandex_q_10k"
More Information needed
yandex-geo-reviews-embeddingsDataset full description: https://www.kaggle.com/datasets/lockiultra/yandex-geo-reviews-embeddings
Dataset contains index column, 768 embedding columns and rating column. Each row corresponds to an embedding representation of the review text with same index.
YandexGPTyandexq-qa-100
YandexQ QA (100 subset)
Traces generated using RU-CoT-Generator and liquid/lfm-2.5-8b-a1b as CoT generator.
Dataset based on sapbot/yandexq-qa-chatml.
yandex_q_200k
Dataset Card for "yandex_q_200k"
More Information needed
yandexgptpro_4th_gen-hellaswag
YandexGPT Pro (4th Gen) HellaSwag
This dataset contains responses from the YandexGPT model evaluated on the HellaSwag benchmark. It was generated as part of an experiment to assess the model’s performance on multiple-choice commonsense reasoning tasks.
Dataset Details
Source: HellaSwag
Model: YandexGPT via Yandex Cloud Foundation Models API
Prompt style: Multiple-choice (A, B, C, D) with system prompt and task context
Fields:
id: index of the example
context: the base… See the full description on the dataset page: https://huggingface.co/datasets/ZennyKenny/yandexgptpro_4th_gen-hellaswag.yandex-alice-sessions-large-syn
Synthetic Alice Dialog Sessions (Realistic, Large)
Alice is a GPT assistant created by Yandex and available for free use here.
Dataset Characteristics
Key Features
session_id: Unique identifier for each conversation session
user_id: Unique identifier for each simulated user
device: Type of device used (smartphone, tv, smart_speaker, car_display, robot)
timestamp: Simulated timestamp of each interaction
role: Whether the speaker is "user" or "assistant"… See the full description on the dataset page: https://huggingface.co/datasets/ZennyKenny/yandex-alice-sessions-large-syn.yandexq-qa-cot-1k
YandexQ QA (1K subset)
Traces generated using RU-CoT-Generator and google/gemma-3-4b-it as CoT generator.
Dataset based on sapbot/yandexq-qa-chatml.
yandex_geo_reviews_restaraunt_instructДатасет составлен из части датасета Geo Reviews Dataset 2023 https://github.com/yandex/geo-reviews-dataset-2023 , взяты только отзывы по ресторанам, и переделан формат датасета для LLM
eval_yandexoidThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 0,
"total_frames": 0,
"total_tasks": 0,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/AlexSmirn0v/eval_yandexoid.eval_YandexLeRobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 0,
"total_frames": 0,
"total_tasks": 0,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/AlexSmirn0v/eval_YandexLeRobot.YandexLeRobotyandexgpt_week_yandexyandex_qlorayandex_geo_reviews_instructyandex_parallel_ruen_corpus
