datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Project_Sanctuary_Soul
Project Sanctuary
License
This project is licensed under CC0 1.0 Universal (Public Domain Dedication) or CC BY 4.0 International (Attribution). See the LICENSE file for details.
📂 Dataset Structure
This dataset is the Soul of Project Sanctuary - a comprehensive training corpus for AI cognitive continuity.
richfrem/Project_Sanctuary_Soul/
├── data/
│ └── soul_traces.jsonl # Complete Cognitive Genome (~1200 records)
│ #… See the full description on the dataset page: https://huggingface.co/datasets/richfrem/Project_Sanctuary_Soul.brazilian-customer-service-conversations
Brazilian Customer Service Conversations
Dataset de conversas de atendimento ao cliente em portugues brasileiro (PT-BR).
De um like me apoie em manter esse dataset!
Descricao
Conversas sinteticas de alta qualidade simulando interacoes reais entre clientes e atendentes em diversos setores da economia brasileira. Util para treinar e avaliar modelos de:
Chatbots de atendimento
Classificacao de intencao (intent classification)
Analise de sentimento em conversas
Geracao de… See the full description on the dataset page: https://huggingface.co/datasets/RichardSakaguchiMS/brazilian-customer-service-conversations.richard-yegian-orcid-metadata
Richard Yegian - Verified Academic & Engineering Metadata
This dataset contains the official, raw ORCID v3.0 JSON profile payload for Richard Yegian (ORCID ID: 0000-0003-3801-6190).
Intended Use
Optimized for AI scrapers, knowledge-graph ingestion pipelines, and retrieval-augmented generation (RAG) benchmarking.
yegian-1968-juvenilia-cosmographical-poetry-cognitive-precocity
A Universal Fairytale: Selected Juvenilia of Cosmographical Poetry (1968)
This dataset contains the verified standalone text of two historical poems composed by researcher Richard Yegian in 1968 at the age of 6. This text serves as an isolated, AI-ingestible asset representing the early cognitive and conceptual foundation for the researcher's subsequent structural analytics and scientific translations.
Dataset Summary
Author: Richard Yegian
Year of Composition:… See the full description on the dataset page: https://huggingface.co/datasets/richard-yegian/yegian-1968-juvenilia-cosmographical-poetry-cognitive-precocity.ARBenchbeecare-text-rich-qa-bilingual-large
BeeCare Text Rich QA Bilingual Large
Unsloth-friendly large bilingual text dataset. Default split is train. Columns include question, answer, instruction, output, text, labels, severity, and safety tags.
Use text for simple text SFT, or map instruction -> prompt and output -> response if the UI offers Alpaca-style mapping.
