datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kinyarwanda_afrivoice_all_domains_v0.1
Afrivoice Kinyarwanda — All Domains
Combined dataset across 5 domains from the original source.
Note: the source dataset also includes a scripted_education domain, excluded
here due to a cluster of corrupted audio files in one of its shards.
Attribution
Original dataset: DigitalUmuganda/Afrivoice_Kinyarwanda
License: CC-BY-4.0
Attribution: Digital Umuganda
This dataset is derived from the above source and released under the same CC-BY-4.0 license.… See the full description on the dataset page: https://huggingface.co/datasets/ElizabethMwangi/kinyarwanda_afrivoice_all_domains_v0.1.kinyarwanda_afrivoice_all_domains_v0.2
Kinyarwanda AfriVoice — All Domains (v0.2)
Cleaned version of ElizabethMwangi/kinyarwanda_afrivoice_all_domains_v0.1.
Changes from v0.1
Removed rows with empty/null transcription values across all splits (train/validation/test)
Audio and domain labels unchanged; only null-transcription rows were dropped
Source
Original data from DigitalUmuganda/Afrivoice_Kinyarwanda (CC-BY-4.0),
extracted and concatenated across 5 domains (agriculture, education… See the full description on the dataset page: https://huggingface.co/datasets/ElizabethMwangi/kinyarwanda_afrivoice_all_domains_v0.2.midi-classical-music
MIDI Classical Music
This dataset contains a comprehensive collection of MIDI files representing classical music compositions from various renowned composers.
The collection includes works from composers such as Bach, Beethoven, Chopin, Mozart, and many others.
The dataset is organized into directories by composer, with each directory containing MIDI files of their compositions.
The dataset is ideal for music analysis, machine learning models for music generation, and other… See the full description on the dataset page: https://huggingface.co/datasets/elizawhitfield/midi-classical-music.ElizabethMidfordSQUADDS_test_clone
THIS IS A CLONE AND IS NOT THE OFFICIAL SQUADDS DB.
SQuADDS_DB - a Superconducting Qubit And Device Design and Simulation Database
The SQuADDS (Superconducting Qubit And Device Design and Simulation) Database Project is an open-source resource aimed at advancing research in superconducting quantum device designs. It provides a robust workflow for generating and simulating superconducting quantum device designs, facilitating the accurate prediction of Hamiltonian… See the full description on the dataset page: https://huggingface.co/datasets/elizabethkunz/SQUADDS_test_clone.xiaohongshu_alpacaELIZA-EVOL-INSTRUCTGPTQ quantization of https://huggingface.co/PygmalionAI/pygmalion-6b/commit/b8344bb4eb76a437797ad3b19420a13922aaabe1
Using this repository: https://github.com/mayaeary/GPTQ-for-LLaMa/tree/gptj-v2
Command:
python3 gptj.py models/pygmalion-6b_b8344bb4eb76a437797ad3b19420a13922aaabe1 c4 --wbits 4 --groupsize 128 --save_safetensors models/pygmalion-6b-4bit-128g.safetensors
elizaThis repository contains synthetic ELIZA chatbot conversations.
See https://github.com/princeton-nlp/ELIZA-Transformer for more details.
big-drama-b23ef6
big-drama-b23ef6
Synthetic products test data: 32 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/Elizabeth-Miller/big-drama-b23ef6.test-geology
Geology Text Tabular Data Notes
Dataset summary
This repository contains a preparation pipeline and a small metadata sample for Geology work with Text Tabular inputs. It does not claim to be a complete benchmark release; the loader documents how source data is normalized and validated.
Included material
dataset.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small… See the full description on the dataset page: https://huggingface.co/datasets/elizabethrer/test-geology.wrong-tank-177c52
wrong-tank-177c52
Synthetic sensors test data: 41 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/ElizabethGonzalez/wrong-tank-177c52.MSV_PG
Dataset Card for Dataset Name
Multi-Scale Violence and Public Gathering (MSV-PG) Dataset
This dataset classifies public events along two axes: the size of the crowd observed and the level of perceived violence in the crowd.
The videos were annotated to identify temporal segments corresponding to specific behavioral events. Click here for the details of the dataset.
Dataset Details
Classes
'natural' – Everyday scenes without significant events.
'lpg' –… See the full description on the dataset page: https://huggingface.co/datasets/ElizabethBV/MSV_PG.queen_elizabeth_azurlane
Dataset of queen_elizabeth/クイーン・エリザベス/伊丽莎白女王 (Azur Lane)
This is the dataset of queen_elizabeth/クイーン・エリザベス/伊丽莎白女王 (Azur Lane), containing 335 images and their tags.
The core tags of this character are blonde_hair, long_hair, blue_eyes, crown, bow, hairband, mini_crown, hair_bow, bangs, black_hairband, fang, breasts, white_bow, small_breasts, very_long_hair, which are pruned in this dataset.
Images are crawled from many sites (e.g. danbooru, pixiv, zerochan ...), the auto-crawling… See the full description on the dataset page: https://huggingface.co/datasets/CyberHarem/queen_elizabeth_azurlane.eliza-1-training
eliza-1 training corpus
Canonical SFT trajectory corpus for the elizaOS eliza-1 Qwen-based model series. Runtime bundles live in elizaos/eliza-1 under bundles/<tier>/ for 0_8b, 2b, 4b, 9b, 27b, and 27b-256k. The removed legacy million-token 27B tier is not part of this dataset.
Files
Path
Role
train.jsonl
canonical native training split
val.jsonl
canonical native validation split
test.jsonl
canonical native held-out test split
data/*.parquetDataset… See the full description on the dataset page: https://huggingface.co/datasets/elizaos/eliza-1-training.roadsterRadM-Bench
RadM-Bench: A Bilingual Multimodal Benchmark for Diagnostic Radiology
📖 Overview
RadM-Bench is a bilingual (English–Chinese) multimodal benchmark for evaluating the diagnostic performance of multimodal large language models (MLLMs) in radiology. It is built to expose three blind spots in existing benchmarks: the gap between curated 2D snapshots and real volumetric (3D) imaging, the gap between public teaching cases and routine clinical practice, and the gap… See the full description on the dataset page: https://huggingface.co/datasets/Elizabeth123/RadM-Bench.swahili_afrivoice_all_domains_v0.1
Afrivoice Swahili — All Domains
Combined dataset across all 5 domains from the original source.
Attribution
Original dataset: DigitalUmuganda/Afrivoice_SwahiliLicense: CC-BY-4.0Attribution: DigitalUmuganda
This dataset is derived from the above source and released under the same CC-BY-4.0 license.
Domains
Agriculture (~124k train samples)
Education (~69k train samples)
Financial (~104k train samples)
Government (~99k train samples)
Health (~118k… See the full description on the dataset page: https://huggingface.co/datasets/ElizabethMwangi/swahili_afrivoice_all_domains_v0.1.leytrabajador-venezuela-100
Dataset Card for "leytrabajador-venezuela-100"
More Information needed
Referencing_Errors_Synthetic_EN
Synthetic Dataset for Automatic Error Correction in Referencing
This dataset includes 4,600 parallel sentences for Automatic Error Correction in referencing in German.It was synthetically created with gpt-4o-mini model according to the Institutional Guidelines of the Center for Translation Studies (CTS), University of Vienna.
Dataset Description
corrupted_sentence: the sentence containing the referencing error
clean_sentence: the correct version of the corrupted sentence… See the full description on the dataset page: https://huggingface.co/datasets/elizaveta-dev/Referencing_Errors_Synthetic_EN.scotus-elizabeth_b_prelogar-audio
SCOTUS-sim audio: elizabeth_b_prelogar
Per-utterance audio clips from Oyez oral-argument mp3s, sliced at
the start_time / stop_time timestamps stored in the companion
scotus-sim/scotus-elizabeth_b_prelogar-training dataset.
Alignment
clip_NNNNN.wav in the tarball corresponds exactly to
audio_segments.jsonl[NNNNN] in the training companion dataset.
In metadata.jsonl each row carries the same 0-padded index in idx.
This supersedes the v1 tarball, which had systematic… See the full description on the dataset page: https://huggingface.co/datasets/scotus-sim/scotus-elizabeth_b_prelogar-audio.eliza_lapisrelights
Dataset of Eliza (Lapis Re:LiGHTs)
This is the dataset of Eliza (Lapis Re:LiGHTs), containing 50 images and their tags.
The core tags of this character are long_hair, red_hair, bangs, pink_hair, blue_eyes, hair_between_eyes, breasts, which are pruned in this dataset.
Images are crawled from many sites (e.g. danbooru, pixiv, zerochan ...), the auto-crawling system is powered by DeepGHS Team(huggingface organization).
List of Packages
Name
Images
Size
Download… See the full description on the dataset page: https://huggingface.co/datasets/CyberHarem/eliza_lapisrelights.elizabeth-v0.0.1
Elizabeth v0.0.1 - Complete Model & Corpus Repository
🚀 Elizabeth Model v0.0.1
Model Files
models/qwen3_8b_v0.0.1_elizabeth_emergence.tar.gz - Complete Qwen3-8B model with Elizabeth's emergent personality
Training Data
corpus/elizabeth-corpus/ - 6 JSONL files with real conversation data
corpus/quantum_processed/ - 4 quantum-enhanced corpus files
Documentation
Comprehensive documentation of Elizabeth's emergence and capabilities:… See the full description on the dataset page: https://huggingface.co/datasets/LevelUp2x/elizabeth-v0.0.1.eliza-1-training-data
elizaos/eliza-1-training-data
Training corpus for the eliza-1 model line. All records are in eliza_native_v1 format — the canonical training schema for elizaOS agents.
Format
Every record is a JSON object with this shape:
{
"format": "eliza_native_v1",
"boundary": "vercel_ai_sdk.generateText",
"request": {
"system": "...",
"messages": [...],
"tools": {...},
"settings": {}
},
"response": {
"text": "...",
"finishReason": "stop"… See the full description on the dataset page: https://huggingface.co/datasets/elizaos/eliza-1-training-data.dataset_v3_synth_top50custom_elizabeth_olsen_middle_2eliza-modernSuper small dataset ment to be used with a rule / mini ai system.
eliza-azheshka-kham-lika-ptashuk
Хам
Metadata
Author: Эліза Ажэшка
Title: Хам
Narrator: Ліка Пташук
Source Group: Аўдыёкнігі
Source: radiokultura.by
Notes
The original audio files are preserved as-is:
no conversion;
no re-encoding;
no filename changes inside each split folder, except removing one common top-level archive folder when present.
To avoid Hugging Face Dataset Viewer scan-size errors, the dataset is split into smaller folders.
Target maximum split size: about 250 MB.… See the full description on the dataset page: https://huggingface.co/datasets/archivartaunik/eliza-azheshka-kham-lika-ptashuk.microbiome-disease-dataset
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/elizah521/microbiome-disease-dataset.LOM_Capacitor_sweepTextToSQLforGaussalgoReferencing_Errors_Synthetic_DE
Synthetic Dataset for Automatic Error Correction in Referencing
This dataset includes 4,600 parallel sentences for Automatic Error Correction in referencing in German.It was synthetically created with gpt-4o-mini model according to the Institutional Guidelines of the Center for Translation Studies (CTS), University of Vienna.
Dataset Description
corrupted_sentence: the sentence containing the referencing error
clean_sentence: the correct version of the corrupted sentence… See the full description on the dataset page: https://huggingface.co/datasets/elizaveta-dev/Referencing_Errors_Synthetic_DE.
