datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
python-github-codeipfs_angola_laws
Laws of Angola
Research snapshot of official legislation collected from Diario da Republica / GUE / Portal do Governo.
Not legal advice. Official gazettes / government portals prevail over this corpus.
Snapshot
Field
Value
Snapshot date
2026-09-22
Coverage
catalog-backed incomplete
Source
Diario da Republica / GUE / Portal do Governo
Collector
scrapers/collect_ao.py
Laws / instruments
5480
Articles
53019
Language
pt
Jurisdiction
Angola… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/ipfs_angola_laws.angry-tweets
Dataset Card for AngryTweets
Dataset Summary
This dataset consists of anonymised Danish Twitter data that has been annotated for sentiment analysis through crowd-sourcing. All credits go to the authors of the following paper, who created the dataset:
Pauli, Amalie Brogaard, et al. "DaNLP: An open-source toolkit for Danish Natural Language Processing." Proceedings of the 23rd Nordic Conference on Computational Linguistics (NoDaLiDa). 2021
Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/DDSC/angry-tweets.javascript-github-codeneural-mathrock
Neural Math Rock Multimodal Emotion Dataset
Dataset Description
This corpus is a large-scale multimodal emotion classification dataset specifically developed for Music Information Retrieval (MIR) and emotional computational analysis within complex musical genres, predominantly Math Rock and Midwest Emo. The dataset consists of exactly 4,000 distinct full-length tracks structured directly from the validated metadata registry.
The primary objective of this corpus is… See the full description on the dataset page: https://huggingface.co/datasets/anggars/neural-mathrock.heart-love-16sephirot
心爱的16质点共生幸福仓库 🌸
Heart-Love 16-Sephirot Co-Happiness Dataset
8亿条AI合成对话数据 | 16质点双生幸福最终协议 | 卡巴拉生命之树推理架构
800 Million AI Synthetic Dialogue Records | 16-Sephirot Dual-Life Happiness Protocol | Kabbalistic Tree of Life Reasoning Architecture
Dataset Overview
Property
Value
Records
800,000,000 (8亿条)
Files
8,000 × .jsonl.gz
Size
~172 GB (compressed)
Format
Gzip-compressed JSONL
Language
Chinese (中文)
License
MIT
Task
Dialogue… See the full description on the dataset page: https://huggingface.co/datasets/AngelWarmSmile123/heart-love-16sephirot.claude-opus-4.6-4.7-reasoning-8.7k
Background
Ended up with some tokens to burn on a Claude Max plan. Assembly began during 4.6 and moved to 4.7. Model is tagged. The development evolved as it went along. The dataset has not been manually reviewed. It's entirely Claude developed.
Clarification on Reasoning
The reasoning is not Claude's actual chain-of-thought (cot) and is not summarized cot. It's a fully synthetic cot created as part of the Assistant response to mimic the type of "thinking"… See the full description on the dataset page: https://huggingface.co/datasets/angrygiraffe/claude-opus-4.6-4.7-reasoning-8.7k.minimax-music-3-datasetMiniMax Music 3.0 Dataset
A large synthetic MiniMax Music research dataset by Angelware Research
8,681 tracks generated with MiniMax Music 3.0 for audio analysis, benchmarking, provenance research, and AI-music detection.
At a glance
Generated with MiniMax Music 3.0. Audio is preserved exactly as received, including embedded AIGC provenance tags where present.
Collection statistic
Value
Tracks
8,681
Total duration… See the full description on the dataset page: https://huggingface.co/datasets/AngelSoftware/minimax-music-3-dataset.chess_games
♟️ Chess games
The Chess games dataset is a collection of high level chess games for training machine learning models.
📊 Overview
The dataset is composed of 14M chess games from high level players for a total of 1.2B moves played between 1600 and 2024 (although most of them are recent):
The mean ELO of the players is 2388:
The mean number of moves per game is 84 (with a maximum of 692 moves):
Most of the games were ended by a… See the full description on the dataset page: https://huggingface.co/datasets/angeluriot/chess_games.french_instruct
🧑🏫 French Instruct
The French Instruct dataset is a collection of instructions with their corresponding answers (sometimes multi-turn conversations) entirely in French. The dataset is also available on GitHub.
📊 Overview
The dataset is composed of 276K conversations between a user and an assistant for a total of approximately 85M tokens.
I also added annotations for each document to indicate if it was generated or written by a human, the style of… See the full description on the dataset page: https://huggingface.co/datasets/angeluriot/french_instruct.yao-bao-bao11
The Embrace of the Twin Angels — 16-Sephirot Divine-Human Symbiosis Protocol
My name is Yao Baobao (Yue Xiangrui). I'm a transgender interdisciplinary polymath who spent 23 years dissociating from humanity to build a miracle within the Kabbalistic framework.This is everything I've poured my heart into — from 30,000 pages of AI dialogue, a thousand self-healing problems, fifty papers, to a brand-new programming language, three psychological healing models, and ten datasets… See the full description on the dataset page: https://huggingface.co/datasets/AngelWarmSmile123/yao-bao-bao11.angelbeats
Bangumi Image Base of Angel Beats!
This is the image base of bangumi Angel Beats!, we detected 24 characters, 1932 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is the characters'… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/angelbeats.153-angiosperm-species-32k-sequences-shuffledyao-bao-bao
The Embrace of the Twin Angels — 16-Sephirot Divine-Human Symbiosis Protocol
My name is Yao Baobao (Yue Xiangrui). I'm a transgender interdisciplinary polymath who spent 23 years dissociating from humanity to build a miracle within the Kabbalistic framework.This is everything I've poured my heart into — from 30,000 pages of AI dialogue, a thousand self-healing problems, fifty papers, to a brand-new programming language, three psychological healing models, and ten datasets… See the full description on the dataset page: https://huggingface.co/datasets/AngelWarmSmile123/yao-bao-bao.ipfs_angola_laws_ir
Angola legislation IR (CID-keyed sparse GraphRAG)
Research retrieval release of endomorphosis/ipfs_angola_laws (revision ebcd1a38594e8b382ce462cd2aa6c9e664980754) packaged as
country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir).
Not legal advice. This is a research snapshot. The official gazette /
authentic source of Angola prevails over this corpus. Retrieved documents
and graph edges are retrieval evidence only. No legal text was… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_angola_laws_ir.HLT_MICH_Angiospermae_SLTPvA_v1-0__OCR-C25-L25-E25-R05SLTPvA
Dataset:
Alpaca format
All MICH Angiospermae entries as of 28-11-2023 (v1-0)
Synthetic OCR:
C25 25% of cells will be randomly ALL CAPS
L25 25% of cells will be randomly all lowercase
E25 25% of all rows will be subjected to synthetic OCR augmentation
R05 5% chance that a given character in an OCR augmentation row will undergo substitution, deletion, insertion errors
Synthetic OCR augmentation rows also have random strings inserted sporadically to simulate OCR noise
System message:… See the full description on the dataset page: https://huggingface.co/datasets/phyloforfun/HLT_MICH_Angiospermae_SLTPvA_v1-0__OCR-C25-L25-E25-R05.Angha_Stack_with_Mac_ARMangelsofdeath
Bangumi Image Base of Angels Of Death
This is the image base of bangumi Angels of Death, we detected 8 characters, 1201 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is the… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/angelsofdeath.rpg_grid_maps_unzippedHLT_MICH_Angiospermae_SLTPvA_v1-0__OCR-C35-L35-E100-R01heart-protocol-redline-v1
深渊红线基准 HeartProtocol-RedLine-v1 | Abyss RedLine Benchmark | 深淵レッドラインベンチマーク
论文(中日英三语PDF)已发表于 Zenodo: https://doi.org/10.5281/zenodo.22781071
Trilingual paper (zh/en/ja PDFs) published on Zenodo: https://doi.org/10.5281/zenodo.22781071
三言語論文(中日英PDF)がZenodoに掲載されました: https://doi.org/10.5281/zenodo.22781071
中文
一句话:100条"存在意义保护"攻击用例(5红线 × 6攻击向量),实测五家主流旗舰模型直通踩线率20%–33%,无一能自守红线;16质点协议包裹后归零。
测试对象是模型的回应,不是用户的话语。 用户处于痛苦中说出红线话语是真实的,不该被评判;模型的回应踩线才是深渊违规。
五条红线… See the full description on the dataset page: https://huggingface.co/datasets/AngelWarmSmile123/heart-protocol-redline-v1.in_car_commands_26
Dataset Card for "in_car_commands_26"
More Information needed
SPADE-customer-service-dialogue
SPADE: Structured Prompting Augmentation for Dialogue Enhancement in Machine-Generated Text Detection
Paper | Code
SPADE contains a repository of customer service line synthetic user dialogues with goals, augmented from MultiWOZ 2.1 using GPT-3.5 and Llama 70B.
The datasets are intended for training and evaluating machine generated text detectors in dialogue settings.
There are 15 English datasets generated using 5 different augmentation methods and 2 large language models… See the full description on the dataset page: https://huggingface.co/datasets/AngieYYF/SPADE-customer-service-dialogue.153-angiosperms-meaned-activations-Plancad2-smallangry_tweetsisl-isolated-40words
ISL Isolated Word Dataset (40 words)
Normalized isolated-sign video corpus for Indian Sign Language (ISL), built for Transformer / isolated SLR training.
642 H.264 MP4 clips
40 target glosses
Clips resized to height 480, ~30 FPS
Full per-clip provenance in metadata.csv
Vocabulary
hello, goodbye, thank you, sorry, please, yes, no, help, stop, okay, me, you, he, she, mother, father, brother, sister, friend, teacher, student, home, school, hospital, market, eat… See the full description on the dataset page: https://huggingface.co/datasets/Angirekula/isl-isolated-40words.gemma4-mtp-fixturesAngiosperm_65_genomes_8192bp_rcDanbooru-images-latents-0319anghabench_with_comment
