datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NBA_PLAY_BY_PLAY_DATA_2023Source of the data: Sportsradar API (https://developer.sportradar.com/docs/read/basketball/NBA_v8)
NBA Play-by-Play Data Extraction and Analysis
Overview
This project aims to retrieve play-by-play data for NBA matches in the 2023 season using the Sportradar API. The play-by-play data is fetched from the API, saved into JSON files, and then used to extract relevant features for analysis and other applications. The extracted data is saved in Parquet files for easy access… See the full description on the dataset page: https://huggingface.co/datasets/farazjawed/NBA_PLAY_BY_PLAY_DATA_2023.waymoqa-videomqa
WaymoQA — VideoQA test subset (mosaic frames)
This repository hosts the video portion of the WaymoQA test split, prepared for
VideoQA evaluation. It contains the multi-view 3x3 mosaic frames for every Waymo
scenario token that carries video questions in the test set.
Contents
File
Description
mosaics.tar.part_aa … mosaics.tar.part_ah
Split archive (8 × 2 GiB) of the mosaic frames
mosaics.tar.part_ai
Final split of the archive
test.jsonl
Full WaymoQA… See the full description on the dataset page: https://huggingface.co/datasets/faraway6/waymoqa-videomqa.Bitext-customer-support-llm-chatbot-training-dataset-spanish
Spanish Customer Support LLM Chatbot Training Dataset
Spanish-language adaptation of the Bitext Customer Support LLM Chatbot Training Dataset.
This dataset is intended for training and evaluating Spanish-language customer-support chatbots and instruction-following large language models.
Dataset Details
Dataset Description
This dataset is a Spanish translation and adaptation of the original Bitext Customer Support LLM Chatbot Training Dataset.
The… See the full description on the dataset page: https://huggingface.co/datasets/Faramir/Bitext-customer-support-llm-chatbot-training-dataset-spanish.lm-eval-results-FelixChao-Faraday-7B-private
Dataset Card for Evaluation run of FelixChao/Faraday-7B
Dataset automatically created during the evaluation run of model FelixChao/Faraday-7B
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-FelixChao-Faraday-7B-private.llm-tokens-atlas
Dataset Card for LLM Tokens Atlas
llm-tokens-atlas is an open, reproducible benchmark of LLM tokenization across
5 providers (Anthropic claude-opus-4-7, Google gemini-2.5-pro, OpenAI
gpt-4o, Mistral mistral-large-latest, Cohere command-r-08-2024) across
5 prompt formats (Markdown, XML, JSON, YAML, Plain text), evaluated on
12,500 real-world prompt requests (n=2,500 per provider, 500 unique prompts × 5
formats × 5 providers). For each (prompt, provider, model, format) cell we… See the full description on the dataset page: https://huggingface.co/datasets/faraa2m/llm-tokens-atlas.datasette-spike-fara
Datasette spike — FARA Active Foreign Principals
For: CoS → WordPress Guru (doctorparadox.net embed/link)Built: 2026-09-17 (ET)Status: DATA half ready — public SQLite + Datasette Lite URL
Why this dataset
Doctor Paradox already centers corruption / foreign influence / authoritarian-adjacent reporting (Corruption Tracker, Corruption Daily cards). FARA filings are the federal public ledger of who lobbies in the U.S. on behalf of foreign principals.
We use the… See the full description on the dataset page: https://huggingface.co/datasets/doctorparadox/datasette-spike-fara.kazakh-stt
Kazakh Speech Dataset (KSD)
1. Dataset Summary
Purpose: High-quality, open-source Kazakh speech dataset for Automatic Speech Recognition (ASR) system development.
Developed by: Department of Artificial Intelligence and Big Data, Al-Farabi Kazakh National University.
Total Duration: 554 hours of recorded speech.
Total Number of Speakers: 873
Average Sentences per Speaker: 250 sentences (utterances).
Total Utterances: 204,250
File Format: .wav
Audio Characteristics:… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/kazakh-stt.teleopSample
teleopSample
Sample teleoperation data.
directory
contents
futurist/
FF Futurist dexterous-hand teleop — data/, videos/, meta/ (LeRobot v2 episodes)
FaberU/pickplace_samples/
FF Faber U pick-and-place: 4 tasks x 2 hand-checked episodes, original camera streams, aligned state, teleop topic logs, and step-level annotations — see its own README
FaberU/workshop_samples/
FF Faber U dual-arm gripper workshop tasks (memory-module insertion, drink-on-coaster, garment… See the full description on the dataset page: https://huggingface.co/datasets/faradayfuture/teleopSample.FLEURS-AR-EN-split
FLEURS-AR-EN Dataset
Dataset Description
FLEURS-AR-EN is an Arabic-to-English dataset designed for Speech Translation tasks. This dataset is derived from Google's FLEURS (Few-shot Learning Evaluation of Universal Representations of Speech) dataset, specifically focusing on aligned Arabic audio samples with their corresponding Arabic transcriptions and English translations.
Overview
Task: Speech Translation
Languages: Arabic (source) → English (target)
Source:… See the full description on the dataset page: https://huggingface.co/datasets/farahabdou/FLEURS-AR-EN-split.Content-Moderation-and-Safety
🇰🇿 Content Moderation and Safety, Kazakh Context
Dataset Summary
Content Moderation and Safety (Profanity) Kazakh Context is a comprehensive dataset designed specifically to train Large Language Models (LLMs) in detecting, classifying, and mitigating toxic, aggressive, or unsafe text in the Kazakh language.
📊 Dataset Statistics
General Metrics
Metric
Count
Total Samples
17,827
Total Words (approx.)
1,674,638
Avg.… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Content-Moderation-and-Safety.Content_Moderation_and_Safety_Kazakh_Context
🇰🇿 Content Moderation and Safety Kazakh Context
Dataset Summary
Toxic Speech Analysis and Mitigation, Kazakh Context is an advanced AI Safety dataset designed to train Large Language Models (LLMs) to detect, deeply analyze, and constructively rewrite toxic or harmful speech in the Kazakh language.
📊 Dataset Statistics
General Metrics
Metric
Count
Total Samples
12,063
Total Words (approx.)
5,869,718
Avg. Words per… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Content_Moderation_and_Safety_Kazakh_Context.prompts.chat
a.k.a. Awesome ChatGPT Prompts
This is a Dataset Repository mirror of prompts.chat — a social platform for AI prompts.
📢 Notice
This Hugging Face dataset is a mirror. For the latest prompts, features, and community contributions, please visit:
🌐 Website: prompts.chat
📦 GitHub: github.com/f/awesome-chatgpt-prompts
About
prompts.chat is an open-source platform where users can share, discover, and collect AI prompts from the community. The project can… See the full description on the dataset page: https://huggingface.co/datasets/Farahchughtaii/prompts.chat.FARAD
FARAD - Furry Aesthetic Realism Annotated Dataset
The FARAD text to image dataset consists of ~18k publicly-available image URL-text pairs from the professional art portfolio website Artstation.
The images were drawn from a set of ~600k furry-adjacent tagged images and then assessed for quality and content with FARAMIR and collaboratively relabeled with simplified e621 tags by JTP-PILOT².
Version 1.2
Download Link:… See the full description on the dataset page: https://huggingface.co/datasets/RedRocket/FARAD.Identification-of-paraphrasing
🇰🇿 Identification of Paraphrasing in Kazakh Context
Dataset Summary
Identification of Paraphrasing in Kazakh Context is a targeted dataset designed to train Large Language Models (LLMs) and embeddings to detect semantic equivalence between two distinct Kazakh texts.
📊 Dataset Statistics
General Metrics
Metric
Count
Total Samples
2,000
Total Words (approx.)
184,465
Avg. Words per Sample
92
Word Count… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Identification-of-paraphrasing.Farah-hg5nsVGMqcIcot-multiplication-2ksummary
🇰🇿 Kazakh Information Extraction and Summarization
📖 Overview
This dataset is a specialized collection for Kazakh Natural Language Processing (NLP), focused on high-quality information extraction and long-form summarization. It contains 500 expert-curated samples where a model must take a detailed input text and generate a comprehensive yet concise summary that captures all key thematic points.
📊 Dataset Statistics
General Metrics… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/summary.KZ-RAG-single-docs-final-gold
🇰🇿 Kazakh Analytical RAG and Document-Based QA
📖 Overview
This dataset is a high-density collection of 4,522 analytical samples designed for Retrieval-Augmented Generation (RAG) tasks in the Kazakh language.
📊 Dataset Statistics
General Metrics
Metric
Count
Total Samples
4,522
Total Words (approx.)
5,978,950
Avg. Words per Sample
1,322
Word Count Distribution (Per Field)
The dataset features… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/KZ-RAG-single-docs-final-gold.simple_python_descriptionfaraz
Faraz Dataset
توضیح
این دیتاست شامل گفتگوهای نقشمحور فارسی با محوریت قوانین، مناطق ویژه اقتصادی و شرکتهای دانشبنیان است.مناسب برای آموزش مدلهای text-generation و assistant/chatbot فارسی.
اندازه
تعداد نمونهها: ~50 گفتگو (هر گفتگو شامل چند پیام)
نقشها: system, user, assistant
فرمت دادهها
فایل اصلی: dataset.jsonl
هر خط: یک JSON object
مثال:
{
"messages": [
{
"role": "system",
"content": "تو «فراز» هستی؛ دستیار هوش مصنوعی… See the full description on the dataset page: https://huggingface.co/datasets/farid678/faraz.multi_step_reasoning_kazakh_context
🇰🇿 Multi-step Reasoning for Kazakh Context
A high-quality dataset designed for complex reasoning, question answering, and text generation tasks in the Kazakh language.
📊 Dataset Statistics
General Metrics
Metric
Count
Total Samples
10,981
Total Words (approx.)
6,652,450
Avg. Tokens per Sample
605
Word Count Distribution (Per Field)
The following table details the distribution of word counts across different… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/multi_step_reasoning_kazakh_context.farazV3Summarization-and-Insight-Extraction
🇰🇿 Summarization and Insight Extraction
📊 Dataset Statistics
General Metrics
Metric
Count
Total Samples
15,000
Total Words (approx.)
7,175,466
Avg. Words per Sample
478
Word Count Distribution (Per Field)
The following table details the distribution of word counts across different fields.
Field
Mean
Median
Min
Max
Total Words
intend
1.5
1.0
1
2
22,498
request
396.2
408.0
2
1988
5,943,024
response… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Summarization-and-Insight-Extraction.Farah-GFNnkK4XwWoReading-with-Comprehension
🇰🇿 Kazakh Contextual Analysis and Complex QA
📖 Overview
This dataset is specifically designed for Advanced Reading Comprehension in the Kazakh language. It challenges models to process long, academic-style texts (averaging over 300 words) and answer multiple complex questions based on the provided context.
📊 Dataset Statistics
General Metrics
Metric
Count
Total Samples
300
Total Words (approx.)
132,371
Avg. Words… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Reading-with-Comprehension.Common-Sense-Reasoning
🇰🇿 Kazakh General Inquiry and FAQ Dataset
📖 Overview
This dataset contains 1,000 high-quality question-and-answer pairs in the Kazakh language. It is designed to train models on providing helpful, natural, and informative responses to common inquiries.
📊 Dataset Statistics
General Metrics
Metric
Count
Total Samples
1,000
Total Words (approx.)
74,013
Avg. Words per Sample
74
Word Count Distribution… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Common-Sense-Reasoning.Neutrality-on-Sensitive-Topics
🇰🇿 Kazakh Human Preference Dataset (RLHF)
📖 Overview
This dataset is a specialized collection of 500 samples designed for Preference Learning and the alignment of Large Language Models in Kazakh. Each entry provides a prompt followed by two potential completions: an accepted response (objective, balanced, and informative) and a rejected response (biased, overly emotional, or unhelpful).
📊 Dataset Statistics
General Metrics… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Neutrality-on-Sensitive-Topics.Farah-5V5ToXTTG6sFarah-ci1Ut4NcHksStory-Generation
🇰🇿 Stories and Dialogue Generation
📖 Overview
Stories Generation is a creative writing dataset specifically curated for the Kazakh language.
📊 Dataset Statistics
General Metrics
Metric
Count
Total Samples
400
Total Words (approx.)
109,735
Avg. Words per Sample
274
Word Count Distribution (Per Field)
The following table details the distribution of word counts across different fields in the… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Story-Generation.
