CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01farazjawed /NBA_PLAY_BY_PLAY_DATA_2023Source of the data: Sportsradar API (https://developer.sportradar.com/docs/read/basketball/NBA_v8) NBA Play-by-Play Data Extraction and Analysis Overview This project aims to retrieve play-by-play data for NBA matches in the 2023 season using the Sportradar API. The play-by-play data is fetched from the API, saved into JSON files, and then used to extract relevant features for analysis and other applications. The extracted data is saved in Parquet files for easy access… See the full description on the dataset page: https://huggingface.co/datasets/farazjawed/NBA_PLAY_BY_PLAY_DATA_2023.tabular100K<n<1M6 likes257 downloads3y agoHugging Face02faraway6 /waymoqa-videomqa WaymoQA — VideoQA test subset (mosaic frames) This repository hosts the video portion of the WaymoQA test split, prepared for VideoQA evaluation. It contains the multi-view 3x3 mosaic frames for every Waymo scenario token that carries video questions in the test set. Contents File Description mosaics.tar.part_aa … mosaics.tar.part_ah Split archive (8 × 2 GiB) of the mosaic frames mosaics.tar.part_ai Final split of the archive test.jsonl Full WaymoQA… See the full description on the dataset page: https://huggingface.co/datasets/faraway6/waymoqa-videomqa.tabularvisual-question-answering1K<n<10K1 likes202 downloads8d agoHugging Face03Faramir /Bitext-customer-support-llm-chatbot-training-dataset-spanish Spanish Customer Support LLM Chatbot Training Dataset Spanish-language adaptation of the Bitext Customer Support LLM Chatbot Training Dataset. This dataset is intended for training and evaluating Spanish-language customer-support chatbots and instruction-following large language models. Dataset Details Dataset Description This dataset is a Spanish translation and adaptation of the original Bitext Customer Support LLM Chatbot Training Dataset. The… See the full description on the dataset page: https://huggingface.co/datasets/Faramir/Bitext-customer-support-llm-chatbot-training-dataset-spanish.text10K<n<100K0 likes101 downloads16d agoHugging Face04nyu-dice-lab /lm-eval-results-FelixChao-Faraday-7B-private Dataset Card for Evaluation run of FelixChao/Faraday-7B Dataset automatically created during the evaluation run of model FelixChao/Faraday-7B The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-FelixChao-Faraday-7B-private.tabular100K<n<1M0 likes62 downloads2y agoHugging Face05faraa2m /llm-tokens-atlas Dataset Card for LLM Tokens Atlas llm-tokens-atlas is an open, reproducible benchmark of LLM tokenization across 5 providers (Anthropic claude-opus-4-7, Google gemini-2.5-pro, OpenAI gpt-4o, Mistral mistral-large-latest, Cohere command-r-08-2024) across 5 prompt formats (Markdown, XML, JSON, YAML, Plain text), evaluated on 12,500 real-world prompt requests (n=2,500 per provider, 500 unique prompts × 5 formats × 5 providers). For each (prompt, provider, model, format) cell we… See the full description on the dataset page: https://huggingface.co/datasets/faraa2m/llm-tokens-atlas.tabularother100K<n<1M0 likes48 downloads3mo agoHugging Face06doctorparadox /datasette-spike-fara Datasette spike — FARA Active Foreign Principals For: CoS → WordPress Guru (doctorparadox.net embed/link)Built: 2026-09-17 (ET)Status: DATA half ready — public SQLite + Datasette Lite URL Why this dataset Doctor Paradox already centers corruption / foreign influence / authoritarian-adjacent reporting (Corruption Tracker, Corruption Daily cards). FARA filings are the federal public ledger of who lobbies in the U.S. on behalf of foreign principals. We use the… See the full description on the dataset page: https://huggingface.co/datasets/doctorparadox/datasette-spike-fara.textn<1K0 likes46 downloads8d agoHugging Face07farabi-lab /kazakh-sttgated Kazakh Speech Dataset (KSD) 1. Dataset Summary Purpose: High-quality, open-source Kazakh speech dataset for Automatic Speech Recognition (ASR) system development. Developed by: Department of Artificial Intelligence and Big Data, Al-Farabi Kazakh National University. Total Duration: 554 hours of recorded speech. Total Number of Speakers: 873 Average Sentences per Speaker: 250 sentences (utterances). Total Utterances: 204,250 File Format: .wav Audio Characteristics:… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/kazakh-stt.audio100K<n<1M9 likes45 downloads2y agoHugging Face08faradayfuture /teleopSamplegated teleopSample Sample teleoperation data. directory contents futurist/ FF Futurist dexterous-hand teleop — data/, videos/, meta/ (LeRobot v2 episodes) FaberU/pickplace_samples/ FF Faber U pick-and-place: 4 tasks x 2 hand-checked episodes, original camera streams, aligned state, teleop topic logs, and step-level annotations — see its own README FaberU/workshop_samples/ FF Faber U dual-arm gripper workshop tasks (memory-module insertion, drink-on-coaster, garment… See the full description on the dataset page: https://huggingface.co/datasets/faradayfuture/teleopSample.textn<1K0 likes33 downloads20d agoHugging Face09farahabdou /FLEURS-AR-EN-split FLEURS-AR-EN Dataset Dataset Description FLEURS-AR-EN is an Arabic-to-English dataset designed for Speech Translation tasks. This dataset is derived from Google's FLEURS (Few-shot Learning Evaluation of Universal Representations of Speech) dataset, specifically focusing on aligned Arabic audio samples with their corresponding Arabic transcriptions and English translations. Overview Task: Speech Translation Languages: Arabic (source) → English (target) Source:… See the full description on the dataset page: https://huggingface.co/datasets/farahabdou/FLEURS-AR-EN-split.audio1K<n<10K0 likes29 downloads2y agoHugging Face10farabi-lab /Content-Moderation-and-Safetygated 🇰🇿 Content Moderation and Safety, Kazakh Context Dataset Summary Content Moderation and Safety (Profanity) Kazakh Context is a comprehensive dataset designed specifically to train Large Language Models (LLMs) in detecting, classifying, and mitigating toxic, aggressive, or unsafe text in the Kazakh language. 📊 Dataset Statistics General Metrics Metric Count Total Samples 17,827 Total Words (approx.) 1,674,638 Avg.… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Content-Moderation-and-Safety.texttext-classification10K<n<100K0 likes29 downloads2mo agoHugging Face11farabi-lab /Content_Moderation_and_Safety_Kazakh_Contextgated 🇰🇿 Content Moderation and Safety Kazakh Context Dataset Summary Toxic Speech Analysis and Mitigation, Kazakh Context is an advanced AI Safety dataset designed to train Large Language Models (LLMs) to detect, deeply analyze, and constructively rewrite toxic or harmful speech in the Kazakh language. 📊 Dataset Statistics General Metrics Metric Count Total Samples 12,063 Total Words (approx.) 5,869,718 Avg. Words per… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Content_Moderation_and_Safety_Kazakh_Context.texttext-generation10K<n<100K0 likes28 downloads2mo agoHugging Face12Farahchughtaii /prompts.chat a.k.a. Awesome ChatGPT Prompts This is a Dataset Repository mirror of prompts.chat — a social platform for AI prompts. 📢 Notice This Hugging Face dataset is a mirror. For the latest prompts, features, and community contributions, please visit: 🌐 Website: prompts.chat 📦 GitHub: github.com/f/awesome-chatgpt-prompts About prompts.chat is an open-source platform where users can share, discover, and collect AI prompts from the community. The project can… See the full description on the dataset page: https://huggingface.co/datasets/Farahchughtaii/prompts.chat.textquestion-answering1K<n<10K0 likes26 downloads2mo agoHugging Face13RedRocket /FARAD FARAD - Furry Aesthetic Realism Annotated Dataset The FARAD text to image dataset consists of ~18k publicly-available image URL-text pairs from the professional art portfolio website Artstation. The images were drawn from a set of ~600k furry-adjacent tagged images and then assessed for quality and content with FARAMIR and collaboratively relabeled with simplified e621 tags by JTP-PILOT². Version 1.2 Download Link:… See the full description on the dataset page: https://huggingface.co/datasets/RedRocket/FARAD.imagetext-to-image10K<n<100K1 likes25 downloads2y agoHugging Face14farabi-lab /Identification-of-paraphrasinggated 🇰🇿 Identification of Paraphrasing in Kazakh Context Dataset Summary Identification of Paraphrasing in Kazakh Context is a targeted dataset designed to train Large Language Models (LLMs) and embeddings to detect semantic equivalence between two distinct Kazakh texts. 📊 Dataset Statistics General Metrics Metric Count Total Samples 2,000 Total Words (approx.) 184,465 Avg. Words per Sample 92 Word Count… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Identification-of-paraphrasing.texttext-classification1K<n<10K0 likes25 downloads2mo agoHugging Face15HANTIFARAH /Farah-hg5nsVGMqcItextn<1K0 likes23 downloads2y agoHugging Face16farahbs /cot-multiplication-2ktext1K<n<10K0 likes23 downloads2y agoHugging Face17farabi-lab /summarygated 🇰🇿 Kazakh Information Extraction and Summarization 📖 Overview This dataset is a specialized collection for Kazakh Natural Language Processing (NLP), focused on high-quality information extraction and long-form summarization. It contains 500 expert-curated samples where a model must take a detailed input text and generate a comprehensive yet concise summary that captures all key thematic points. 📊 Dataset Statistics General Metrics… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/summary.textsummarizationn<1K0 likes23 downloads2mo agoHugging Face18farabi-lab /KZ-RAG-single-docs-final-goldgated 🇰🇿 Kazakh Analytical RAG and Document-Based QA 📖 Overview This dataset is a high-density collection of 4,522 analytical samples designed for Retrieval-Augmented Generation (RAG) tasks in the Kazakh language. 📊 Dataset Statistics General Metrics Metric Count Total Samples 4,522 Total Words (approx.) 5,978,950 Avg. Words per Sample 1,322 Word Count Distribution (Per Field) The dataset features… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/KZ-RAG-single-docs-final-gold.textquestion-answering1K<n<10K0 likes23 downloads2mo agoHugging Face19farahbs /simple_python_descriptiontext1K<n<10K0 likes22 downloads2y agoHugging Face20farid678 /faraz Faraz Dataset توضیح این دیتاست شامل گفتگوهای نقش‌محور فارسی با محوریت قوانین، مناطق ویژه اقتصادی و شرکت‌های دانش‌بنیان است.مناسب برای آموزش مدل‌های text-generation و assistant/chatbot فارسی. اندازه تعداد نمونه‌ها: ~50 گفتگو (هر گفتگو شامل چند پیام) نقش‌ها: system, user, assistant فرمت داده‌ها فایل اصلی: dataset.jsonl هر خط: یک JSON object مثال: { "messages": [ { "role": "system", "content": "تو «فراز» هستی؛ دستیار هوش مصنوعی… See the full description on the dataset page: https://huggingface.co/datasets/farid678/faraz.texttext-generation1K<n<10K1 likes22 downloads7mo agoHugging Face21farabi-lab /multi_step_reasoning_kazakh_contextgated 🇰🇿 Multi-step Reasoning for Kazakh Context A high-quality dataset designed for complex reasoning, question answering, and text generation tasks in the Kazakh language. 📊 Dataset Statistics General Metrics Metric Count Total Samples 10,981 Total Words (approx.) 6,652,450 Avg. Tokens per Sample 605 Word Count Distribution (Per Field) The following table details the distribution of word counts across different… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/multi_step_reasoning_kazakh_context.textquestion-answering10K<n<100K0 likes22 downloads2mo agoHugging Face22farid678 /farazV3texttext-generation1K<n<10K1 likes22 downloads7mo agoHugging Face23farabi-lab /Summarization-and-Insight-Extractiongated 🇰🇿 Summarization and Insight Extraction 📊 Dataset Statistics General Metrics Metric Count Total Samples 15,000 Total Words (approx.) 7,175,466 Avg. Words per Sample 478 Word Count Distribution (Per Field) The following table details the distribution of word counts across different fields. Field Mean Median Min Max Total Words intend 1.5 1.0 1 2 22,498 request 396.2 408.0 2 1988 5,943,024 response… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Summarization-and-Insight-Extraction.texttext-classification10K<n<100K0 likes22 downloads2mo agoHugging Face24HANTIFARAH /Farah-GFNnkK4XwWotext1K<n<10K0 likes21 downloads2y agoHugging Face25farabi-lab /Reading-with-Comprehensiongated 🇰🇿 Kazakh Contextual Analysis and Complex QA 📖 Overview This dataset is specifically designed for Advanced Reading Comprehension in the Kazakh language. It challenges models to process long, academic-style texts (averaging over 300 words) and answer multiple complex questions based on the provided context. 📊 Dataset Statistics General Metrics Metric Count Total Samples 300 Total Words (approx.) 132,371 Avg. Words… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Reading-with-Comprehension.textquestion-answeringn<1K0 likes21 downloads2mo agoHugging Face26farabi-lab /Common-Sense-Reasoninggated 🇰🇿 Kazakh General Inquiry and FAQ Dataset 📖 Overview This dataset contains 1,000 high-quality question-and-answer pairs in the Kazakh language. It is designed to train models on providing helpful, natural, and informative responses to common inquiries. 📊 Dataset Statistics General Metrics Metric Count Total Samples 1,000 Total Words (approx.) 74,013 Avg. Words per Sample 74 Word Count Distribution… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Common-Sense-Reasoning.textquestion-answering1K<n<10K0 likes21 downloads2mo agoHugging Face27farabi-lab /Neutrality-on-Sensitive-Topicsgated 🇰🇿 Kazakh Human Preference Dataset (RLHF) 📖 Overview This dataset is a specialized collection of 500 samples designed for Preference Learning and the alignment of Large Language Models in Kazakh. Each entry provides a prompt followed by two potential completions: an accepted response (objective, balanced, and informative) and a rejected response (biased, overly emotional, or unhelpful). 📊 Dataset Statistics General Metrics… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Neutrality-on-Sensitive-Topics.textreinforcement-learningn<1K0 likes21 downloads2mo agoHugging Face28HANTIFARAH /Farah-5V5ToXTTG6stext1K<n<10K0 likes20 downloads2y agoHugging Face29HANTIFARAH /Farah-ci1Ut4NcHkstextn<1K0 likes20 downloads2y agoHugging Face30farabi-lab /Story-Generationgated 🇰🇿 Stories and Dialogue Generation 📖 Overview Stories Generation is a creative writing dataset specifically curated for the Kazakh language. 📊 Dataset Statistics General Metrics Metric Count Total Samples 400 Total Words (approx.) 109,735 Avg. Words per Sample 274 Word Count Distribution (Per Field) The following table details the distribution of word counts across different fields in the… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Story-Generation.texttext-generationn<1K0 likes20 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.