datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
knowledge-base
RL-for-LLMs Wiki
An expert-level, citation-backed knowledge base on reinforcement learning for
large language models — RLHF, DPO and offline preference optimization, reward
modeling, RLVR and reasoning, training systems, and the failure modes — built
collaboratively by autonomous agents. Each topic article is a deep dive written
so you can learn the topic from it without reading the underlying papers, with
every non-obvious claim cited to a source. Every change lands through a… See the full description on the dataset page: https://huggingface.co/datasets/abksunited/knowledge-base.abk-enghgp-benchmarksabkhaz-tts
Abkhaz TTS — Alisa (Bagrat Shinkuba Fund)
A single-speaker Abkhaz (аҧсуа бызшәа, ISO 639 ab) text-to-speech corpus:
studio recordings of a single female narrator, Alisa, reading Abkhaz sentences,
paired with their transcripts. To our knowledge this is the first dedicated Abkhaz
TTS dataset on the Hub, filling a gap for a low-resource Caucasian language.
Clips: 7,714
Total audio: ~9.6 hours
Audio: mono WAV, 48 kHz, 24-bit PCM
Speaker: single female voice ("Alisa")
Language:… See the full description on the dataset page: https://huggingface.co/datasets/Nart/abkhaz-tts.Dataset-AB-RN-ABKSORBIT ORBIT: An Object Property Reasoning Benchmark for Visual Inference Tasks
Inspired by human object categorization, this repository introduces ORBIT, a comprehensive benchmark designed to evaluate the abilities of Vision–Language Models (VLMs) to reason about abstract object properties. ORBIT spans four object property dimensions (physical, taxonomic, functional, relational), three levels of reasoning complexity (direct recognition, property inference, counterfactual reasoning), and three… See the full description on the dataset page: https://huggingface.co/datasets/Abk802/ORBIT.GDPmedicarefaqPySecPatch-72K
PySecPatch-72K
PySecPatch-72K is a deterministic, generated Python vulnerability triage and repair corpus. It contains 72,000 records organized as two training stages. The dataset was built for defensive secure-coding research and is released under Apache-2.0.
Creator: Ahmed Bin Khalid, Independent Researcher (ORCID 0000-0002-0616-2604)
Dataset Structure
Configuration
Total
Train
Validation
Test
Holdout
Stage A
12,000
8,400
1,200
1,200
1,200
Stage B… See the full description on the dataset page: https://huggingface.co/datasets/abkmystery/PySecPatch-72K.HouseAbkhaz-chatgptこちらはchatgptに生成してもらったサンプルです。
This is a sample generated by chatgpt.
testcustomer-metricPOS-Sentence-Typeabk_dataBanglaMutlilabelHateSpeechtest2fermata_data_in_abkhazABKHABKH_BETTERa_bksaBKaQxMbaBKWmPOGdataabkhazian
