datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DreamFlow-AI-DataSA_Cultural_Tribal_Practices
SA Tribal & Cultural Practices Dataset
Author: Minah Mojela (@minahmojela), Umkho-AI
Dataset Summary
This dataset contains 127 structured records documenting the cultural practices,
customs, and identity histories of South Africa's major ethnic and population groups.
It is a companion release to the South African History Dataset,
built for the same reason: most AI models describe South African cultural practices
using surface-level, externally-authored sources… See the full description on the dataset page: https://huggingface.co/datasets/Umkho-AI/SA_Cultural_Tribal_Practices.oinio-sacred-trinity-eval
🔮 Quantum Forge Sacred Trinity Evaluation Dataset
Annotated test cases for evaluating AI agents in the Quantum Pi Forge ecosystem.
📊 Dataset Description
This dataset contains 10 annotated query-response pairs designed to evaluate AI agents
operating within the Sacred Trinity architecture:
FastAPI Quantum Conduit - Authentication, WebSocket, database operations
Flask Glyph Weaver - Dashboard visualization, SVG cascade animations
Gradio Truth Mirror - Ethical auditing… See the full description on the dataset page: https://huggingface.co/datasets/onenoly11/oinio-sacred-trinity-eval.sachi-dataset-jaLLMをファインチューニングするためのデータセットです。alcapa-chatbot-formatです。
キャラクターと会話するデータセットとなっています。
私はいつもVR SNSでかわいい女の子のロールプレーをしています。
私がかわいい女の子のAIに転生したという設定で作った会話データセットになっています。
キャラクター設定はフィクションやジョークです。完全に現実ではありません。
ゲームに登場するNPC等のAIのトレーニングなどに自由にご利用ください。
幅広く利用してもらえるようにPublic domainライセンスにします。
ライセンス
Public domainライセンスにします。
ContinuumGPT
📚 ContinuumGPT Dataset
This dataset powers ContinuumGPT, a GPT chatbot with hierarchical long-term memory.
📖 Overview
ContinuumGPT is designed to store, compress, and retrieve conversations using a multi-level memory system:
Level 1 (Fresh Memory): Recent detailed Q&A entries.
Level 2 (Archive Memory): Compressed summaries of older Q&A.
Level 3 (Global Summary): Higher-level summaries created when archives overflow.
This allows the model to scale infinitely while… See the full description on the dataset page: https://huggingface.co/datasets/Sachin5112/ContinuumGPT.czech-sacd-legal-questions
📑 Overview
This repository contains 200 question-answer pairs automatically generated with Gemini 2.0 from the decisions of the Czech Sumpreme Administrative Court.
The work was performed in spring 2025 as part of my master’s diploma thesis at the Faculty of Information Technology, Czech Technical University in Prague (FIT CTU).
🏛️ Source
Official judgments scraped from https://sbirka.nssoud.cz (March 2025 snapshot).
upsccarrerflow-aiSQAD-Sinhala_Question_Answering_DatasetThis dataset is a back-translated version of the SQuAD 2.0 dataset, translated into Sinhala using the Google Cloud Translate API by Sachin Hansaka.
Original dataset by the Stanford QA Group: https://rajpurkar.github.io/SQuAD-explorer/
Original work licensed under CC BY-SA 4.0.
This Sinhala version © 2025 Sachin Hansaka, also licensed under CC BY-SA 4.0.
📚 Dataset Overview
SQAD-Sinhala_Question_Answering_Dataset is a high-quality, back-translated version of the original SQuAD 2.0 dataset… See the full description on the dataset page: https://huggingface.co/datasets/Sachin-Hansaka/SQAD-Sinhala_Question_Answering_Dataset.
