CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01yandex /yambda Yambda-5B — A Large-Scale Multi-modal Dataset for Ranking And Retrieval Industrial-scale music recommendation dataset with organic/recommendation interactions and audio embeddings 📌 Overview • 🔑 Key Features • 📊 Statistics • 📝 Format • 🏆 Benchmark • ⬇️ Download • ❓ FAQ Overview The Yambda-5B dataset is a large-scale open database comprising 4.79 billion user-item interactions collected from 1 million users and spanning 9.39 million tracks. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/yandex/yambda.tabular1B<n<10B238 likes3.5k downloads6mo agoHugging Face02yandex /HardMultiQAgated 📚 HardMultiQA 📚 English version [!NOTE] ВАЖНО: Пожалуйста, помогите сохранить объективность этого бенчмарка, снизив риск попадания его вопросов и ответов в обучающие данные моделей. Просим вас делиться ссылкой на этот репозиторий вместо того, чтобы публиковать датасет в открытом доступе или создавать его публичные копии. Это добровольная просьба, которая не ограничивает ваши права, предусмотренные лицензией CC BY-SA 4.0. HardMultiQA — русскоязычный бенчмарк для оценки… See the full description on the dataset page: https://huggingface.co/datasets/yandex/HardMultiQA.textquestion-answeringn<1K11 likes194 downloads6d agoHugging Face03yandex /WikiWebFactsgated 🌐 WikiWebFacts 🌐 English version [!NOTE] ВАЖНО: Пожалуйста, помогите сохранить объективность этого бенчмарка, снизив риск попадания его вопросов и ответов в обучающие данные моделей. Просим вас делиться ссылкой на этот репозиторий вместо того, чтобы публиковать датасет в открытом доступе или создавать его публичные копии. Это добровольная просьба, которая не ограничивает ваши права, предусмотренные лицензией CC BY-SA 4.0. WikiWebFacts — русскоязычный бенчмарк для оценки… See the full description on the dataset page: https://huggingface.co/datasets/yandex/WikiWebFacts.textquestion-answering1K<n<10K10 likes187 downloads6d agoHugging Face04yandex /alchemist Alchemist 👨‍🔬 Dataset Description Alchemist is a compact, high-quality dataset comprising 3,350 image-text pairs, meticulously curated for supervised fine-tuning (SFT) of pre-trained text-to-image (T2I) generative models. The primary goal of Alchemist is to significantly enhance the generative quality (particularly aesthetic appeal and image complexity) of T2I models while preserving their inherent diversity in content, composition, and style. This dataset and its… See the full description on the dataset page: https://huggingface.co/datasets/yandex/alchemist.image1K<n<10K58 likes173 downloads1y agoHugging Face05yandex /mad-cars MAD-Cars: Multi-view Auto Dataset 🚗 Dataset Description MAD-Cars is a large-scale collection of 360° car videos. It comprises ~70,000 car instances with diverse brands, car types, colors, and lighting conditions. Each instance contains an average of ~85 frames, with most car instances available at a resolution of 1920x1080. The dataset statistics are presented in the figure below. The data is carefully curated by filtering the frames and entire car instances that… See the full description on the dataset page: https://huggingface.co/datasets/yandex/mad-cars.imageimage-to-video1M<n<10M35 likes146 downloads1y agoHugging Face06its5Q /yandex-qThis is a dataset of questions and answers scraped from Yandex.Q.text-generation100K<n<1M12 likes88 downloads3y agoHugging Face07nikkotanns /yandex-cloud-dataset-resizeimage1K<n<10K0 likes67 downloads1mo agoHugging Face08IlyaGusev /yandex_q_fullBased on https://huggingface.co/datasets/its5Q/yandex-q, parsed full.jsonl.gz tabular1M<n<10M2 likes64 downloads4y agoHugging Face09sapbot /yandexq-qa-chatmlChatML formatted version of its5Q/yandex-q. texttext-generation100K<n<1M1 likes48 downloads4mo agoHugging Face10valeriylo /Pix2PixHD_YandexMapsThe dataset was obtained using the web crowdsourcing GIS service of Yandex Maps, with a custom written web scrapper. The main advantages of this dataset are the high quality of the images and the focus on Russian urban areas. This dataset is the only* image dataset for the img2img task for the Commonwealth of Independent States (CIS) regions (* as of spring 2020). It can be used for the Nvidia Pix2PixHD GAN architecture not only for Russian areas, but also for Ukraine, Belarus, Kazakhstan… See the full description on the dataset page: https://huggingface.co/datasets/valeriylo/Pix2PixHD_YandexMaps.imageimage-to-image1K<n<10K0 likes47 downloads2y agoHugging Face11yandex /wmt24-en-ru-rate RATE Framework Annotation Dataset This repository contains the annotation data for the paper Refined Assessment for Translation Evaluation: Rethinking Machine Translation Evaluation in the Era of Human-Level Systems. The dataset presents human evaluation of English-Russian translations from WMT24, annotated using our proposed RATE framework by highly qualified professional translators and linguists with academic degrees and substantial industry experience. For comparison purposes… See the full description on the dataset page: https://huggingface.co/datasets/yandex/wmt24-en-ru-rate.5 likes33 downloads11mo agoHugging Face12ScoutieAutoML /promocodes_yandex_avito_wildberries_sber_marketplaces Description in English: The dataset is collected from Russian-language Telegram channels with promo codes, discounts and promotions from various marketplaces, such as Yandex Market, Wildberries, Ozon,The dataset was collected and tagged automatically using the data collection and tagging service Scoutie.Try Scoutie and collect the same or another dataset using link for FREE. Dataset fields: taskId - task identifier in the Scouti service. text - main text. url -… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/promocodes_yandex_avito_wildberries_sber_marketplaces.tabulartext-classification10K<n<100K2 likes31 downloads2y agoHugging Face13Kirili4ik /yandex_jobs Dataset Card for Yandex_Jobs Dataset Description Dataset Summary This is a dataset of more than 600 IT vacancies in Russian from parsing telegram channel https://t.me/ya_jobs. All the texts are perfectly structured, no missing values. Supported Tasks and Leaderboards text-generation with the 'Raw text column'. summarization as for getting from all the info the header. multiple-choice as for the hashtags (to choose multiple from all available in the… See the full description on the dataset page: https://huggingface.co/datasets/Kirili4ik/yandex_jobs.texttext-generationn<1K8 likes28 downloads4y agoHugging Face14dim /yandex_q_10k Dataset Card for "yandex_q_10k" More Information needed text10K<n<100K0 likes21 downloads3y agoHugging Face15lockiultra /yandex-geo-reviews-embeddingsDataset full description: https://www.kaggle.com/datasets/lockiultra/yandex-geo-reviews-embeddings Dataset contains index column, 768 embedding columns and rating column. Each row corresponds to an embedding representation of the review text with same index. tabularfeature-extraction100K<n<1M0 likes17 downloads3y agoHugging Face16radce /YandexGPTquestion-answering1K<n<10K4 likes17 downloads2y agoHugging Face17sapbot /yandexq-qa-100 YandexQ QA (100 subset) Traces generated using RU-CoT-Generator and liquid/lfm-2.5-8b-a1b as CoT generator. Dataset based on sapbot/yandexq-qa-chatml. texttext-generationn<1K0 likes17 downloads3mo agoHugging Face18dim /yandex_q_200k Dataset Card for "yandex_q_200k" More Information needed text100K<n<1M1 likes15 downloads3y agoHugging Face19ZennyKenny /yandexgptpro_4th_gen-hellaswag YandexGPT Pro (4th Gen) HellaSwag This dataset contains responses from the YandexGPT model evaluated on the HellaSwag benchmark. It was generated as part of an experiment to assess the model’s performance on multiple-choice commonsense reasoning tasks. Dataset Details Source: HellaSwag Model: YandexGPT via Yandex Cloud Foundation Models API Prompt style: Multiple-choice (A, B, C, D) with system prompt and task context Fields: id: index of the example context: the base… See the full description on the dataset page: https://huggingface.co/datasets/ZennyKenny/yandexgptpro_4th_gen-hellaswag.textzero-shot-classification10K<n<100K0 likes15 downloads1y agoHugging Face20ZennyKenny /yandex-alice-sessions-large-syn Synthetic Alice Dialog Sessions (Realistic, Large) Alice is a GPT assistant created by Yandex and available for free use here. Dataset Characteristics Key Features session_id: Unique identifier for each conversation session user_id: Unique identifier for each simulated user device: Type of device used (smartphone, tv, smart_speaker, car_display, robot) timestamp: Simulated timestamp of each interaction role: Whether the speaker is "user" or "assistant"… See the full description on the dataset page: https://huggingface.co/datasets/ZennyKenny/yandex-alice-sessions-large-syn.text100K<n<1M1 likes15 downloads1y agoHugging Face21sapbot /yandexq-qa-cot-1k YandexQ QA (1K subset) Traces generated using RU-CoT-Generator and google/gemma-3-4b-it as CoT generator. Dataset based on sapbot/yandexq-qa-chatml. texttext-generation1K<n<10K0 likes15 downloads3mo agoHugging Face22CleverShovel /yandex_geo_reviews_restaraunt_instructДатасет составлен из части датасета Geo Reviews Dataset 2023 https://github.com/yandex/geo-reviews-dataset-2023 , взяты только отзывы по ресторанам, и переделан формат датасета для LLM text10K<n<100K1 likes12 downloads8mo agoHugging Face23AlexSmirn0v /eval_yandexoidThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so_follower", "total_episodes": 0, "total_frames": 0, "total_tasks": 0, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": {}, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/AlexSmirn0v/eval_yandexoid.robotics0 likes12 downloads6mo agoHugging Face24AlexSmirn0v /eval_YandexLeRobotThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so_follower", "total_episodes": 0, "total_frames": 0, "total_tasks": 0, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": {}, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/AlexSmirn0v/eval_YandexLeRobot.robotics0 likes11 downloads6mo agoHugging Face25AlexSmirn0v /YandexLeRobot0 likes10 downloads6mo agoHugging Face26philipphager /yandex10M<n<100M0 likes7 downloads2y agoHugging Face27Natet /gpt_week_yandex0 likes5 downloads3y agoHugging Face28marchcat73 /yandex_qlora0 likes5 downloads3y agoHugging Face29CleverShovel /yandex_geo_reviews_instructtext100K<n<1M0 likes4 downloads2y agoHugging Face30Den4ikAI /yandex_parallel_ruen_corpus1 likes3 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.