datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
news-entertainment-datasetnews-politics-datasetnews-education-datasetnews-tech-datasetnews-finance-datasetwire_harness_expert_sac
Wire Harness Expert SAC
Expert-policy trajectories collected from the five-mover WireHarness MuJoCo
environment for visual world-model training.
Dataset summary
20,000 episodes
3,491,570 stored observation rows
At most 300 environment transitions per episode (up to 301 stored rows,
including the initial observation)
224 x 224 RGB observations, stored as JPEG bytes in pixels
10-dimensional continuous actions
451-dimensional observations
Five task stages and… See the full description on the dataset page: https://huggingface.co/datasets/faridganbarli/wire_harness_expert_sac.SACo-Gold
Dataset Card for SA-Co/Gold
SA-Co/Gold is a benchmark for promptable concept segmentation (PCS) in images. The benchmark contains images paired with text labels (also referred as Noun Phrases aka NPs), each annotated exhaustively with masks on all object instances that match the label. SA-Co/Gold comprises 7 subsets, each targeting a different annotation domain. For each subset, the annotations are multi-reviewed and agreed by 3 human annotators resulting in a high-quality… See the full description on the dataset page: https://huggingface.co/datasets/facebook/SACo-Gold.GRASS_sampleSACo-VEval
SA-Co/VEval Dataset
License each domain has its own License
SA-Co/VEval - SA-V: CC-BY-NC 4.0
SA-Co/VEval - YT-Temporal-1B: CC-BY-NC 4.0
SA-Co/VEval - SmartGlasses: CC-by-4.0
SA-Co/VEval is an evaluation dataset comprising of 3 domains, each domain has a val and test split.
SA-Co/VEval - SA-V: videos are from the SA-V dataset
SA-Co/VEval - YT-Temporal-1B: videos are from the YT-Temporal-1B
SA-Co/VEval - SmartGlasses: egocentric videos from Smart Glasses
This Hugging Face dataset… See the full description on the dataset page: https://huggingface.co/datasets/facebook/SACo-VEval.glassformingSAC-Flow
SAC Flow: Sample-Efficient Reinforcement Learning of Flow-Based Policies via velocity-reparameterized sequential modeling
Overview
SAC Flow is a stable, sample-efficient, and high-performance off-policy RL algorithm for flow-based policies. SAC Flow treats the flow-based model as a sequential model and reparameterizes its velocity network as a GRU or a Transformer.
Get Start
All necessary dependencies and environment setup steps are detailed in our… See the full description on the dataset page: https://huggingface.co/datasets/Elessar123/SAC-Flow.SAC_Nepal_FAQ_Nepali_Health_Fitness_Dataset
SAC Nepal FAQ — Nepali Health & Fitness Dataset
Overview
This dataset (sac_nepal_faq_nepali.jsonl) is a collection of 100 instruction-following conversation pairs in Nepali, covering frequently asked questions about health, fitness, and nutrition. Each record is a single-turn human↔gpt exchange: a Nepali-language question followed by an informative Nepali-language answer.
The data appears to be a localized/translated set — the behavior_definition field for every… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/SAC_Nepal_FAQ_Nepali_Health_Fitness_Dataset.SA_Cultural_Tribal_Practices
SA Tribal & Cultural Practices Dataset
Author: Minah Mojela (@minahmojela), Umkho-AI
Dataset Summary
This dataset contains 127 structured records documenting the cultural practices,
customs, and identity histories of South Africa's major ethnic and population groups.
It is a companion release to the South African History Dataset,
built for the same reason: most AI models describe South African cultural practices
using surface-level, externally-authored sources… See the full description on the dataset page: https://huggingface.co/datasets/Umkho-AI/SA_Cultural_Tribal_Practices.oinio-sacred-trinity-eval
🔮 Quantum Forge Sacred Trinity Evaluation Dataset
Annotated test cases for evaluating AI agents in the Quantum Pi Forge ecosystem.
📊 Dataset Description
This dataset contains 10 annotated query-response pairs designed to evaluate AI agents
operating within the Sacred Trinity architecture:
FastAPI Quantum Conduit - Authentication, WebSocket, database operations
Flask Glyph Weaver - Dashboard visualization, SVG cascade animations
Gradio Truth Mirror - Ethical auditing… See the full description on the dataset page: https://huggingface.co/datasets/onenoly11/oinio-sacred-trinity-eval.synthetic_demographics_seed
Synthetic Demographic Seeds v1
This is a dataset of 3,541,040 roughly demographically correct demographic seeds and somewhat demographically accurate names all generated from publicly available datasets.
(note there were tradeoffs made with accuracy and what I could tie together, v2 will be more accurate)
get_synthetic_demographics.py contains a method for quickly and randomly selecting batches of demographic seeds.
There is no filtering on this at the moment.
Format… See the full description on the dataset page: https://huggingface.co/datasets/sacrificialpancakes/synthetic_demographics_seed.sackcha_chatMentoring-Dataset
Boost Your Technical Mentorship with OpenLLaMA 3B Fine-Tuning
Ready to unlock expert-level guidance on your technical journey? Explore this question-answer dataset designed for technical mentorship, with future plans to fine-tune the powerful OpenLLaMA 3B language model for even more advanced interactions.
Overview
Focus: Technical Mentorship
Domains: Currently covers 7 key areas: AI, ML, Blockchain, Cybersecurity, AppDev, WebDev, DevOps
Content:
General questions a… See the full description on the dataset page: https://huggingface.co/datasets/sachit-sankhe/Mentoring-Dataset.ls2_09062023_test1_raw_SaChA_1a
Sayori Chat 09062023 raw
Dataset of Sayori dialogue from DDLC (dataset of ~600 items augmented by MythoMax-l2-13b to turn into multi-turn chat dialogue)
Curated version planned
saccaromyces-cerevisiae-basesachi-dataset-jaLLMをファインチューニングするためのデータセットです。alcapa-chatbot-formatです。
キャラクターと会話するデータセットとなっています。
私はいつもVR SNSでかわいい女の子のロールプレーをしています。
私がかわいい女の子のAIに転生したという設定で作った会話データセットになっています。
キャラクター設定はフィクションやジョークです。完全に現実ではありません。
ゲームに登場するNPC等のAIのトレーニングなどに自由にご利用ください。
幅広く利用してもらえるようにPublic domainライセンスにします。
ライセンス
Public domainライセンスにします。
Ai-Conversation-question-answerdeepmind_mathCurated Dataset taken from DeepMind synthetically generated math dataset.
Gemma_4_E2B_Vision_FOR_Oral_Cancer
Oral Gemma Fine-Tuning Dataset
This repository contains a portable, instruction-tuning dataset for cropped oral mucosal lesion screening.
It is designed for vision-language fine-tuning of Gemma-style models on a binary screening task.
Overview
Each example pairs:
one cropped oral mucosal image
one short instruction
one JSON answer with a conservative screening recommendation
This is a screening support dataset, not a diagnostic dataset.
Target labels:… See the full description on the dataset page: https://huggingface.co/datasets/sach3v/Gemma_4_E2B_Vision_FOR_Oral_Cancer.czech-sacd-legal-questions
📑 Overview
This repository contains 200 question-answer pairs automatically generated with Gemini 2.0 from the decisions of the Czech Sumpreme Administrative Court.
The work was performed in spring 2025 as part of my master’s diploma thesis at the Faculty of Information Technology, Czech Technical University in Prague (FIT CTU).
🏛️ Source
Official judgments scraped from https://sbirka.nssoud.cz (March 2025 snapshot).
upscai-conversationerc-efrfiltered_datasachi-chosen-rejected-javsachi/Sachi-Qwen3-4B-GGUF が出力した回答からchosenとrejectedを選んだ結果のデータセットです。
rmj_covid_ner
