aak
Datasets
All datasets matching “aak”iclr-wm-backup-public
ICLR Watermark Benchmark — backup overflow (public part)
Companion to the private repo Aak975/iclr-wm-backup, which reached its
storage quota. Together the two repos form ONE backup — every file exists in
exactly one of them, with the same layout:
archives/<sub>/part-0000 ... part-NNNN, MANIFEST.json
restore one archive: cat part-* | zstd -d | tar -x
MANIFEST.json = {"parts": N, "sha256": <whole-stream>, "total_bytes": M}
This public part holds only shareable image data… See the full description on the dataset page: https://huggingface.co/datasets/Aak975/iclr-wm-backup-public.ft49
ft49 — baselining backup
Backup of the ft49 baselining working directory (Monash M3), created 2026-09-03.
facellm/ — the facellm/ tree (76,427 files, 315 MB raw) as a single zstd-compressed tar
stream: part-0000 + MANIFEST.json ({"parts", "sha256" (whole stream), "total_bytes"}).
phase1_output/appearance_action_hair_00400_unfrozen.zip — 3.3 GB, stored as-is.
phase2_output/htcc_phase2_attributes_output.zip — 24.7 GB, stored as-is.
Restore
python3 -m pip install… See the full description on the dataset page: https://huggingface.co/datasets/Aak975/ft49.udposUniversal Dependencies is an open community effort to create cross-linguistically consistent treebank annotation for many languages within a dependency-based lexicalist framework. The annotation consists in a linguistically motivated word segmentation; a morphological layer comprising lemmas, universal part-of-speech tags, and standardized morphological features; and a syntactic layer focusing on syntactic relations between predicates, arguments and modifiers.angle_peg_stereo_merged_07_29_stateThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"observation.image.cam0": {
"dtype": "image",
"shape": [
3,
240,
320
],
"names": [
"C",
"H",
"W"
]
},
"observation.image.cam1": {
"dtype":… See the full description on the dataset page: https://huggingface.co/datasets/aakankshpanda/angle_peg_stereo_merged_07_29_state.mysdxl-dataset
Image-Prompt Dataset
An image-prompt dataset scraped and assembled with
MySDXL for training
latent diffusion models.
Dataset structure
Each row contains one image with its corresponding text prompt.
Column
Type
Description
image
Image
RGB image (lossless PNG, original resolution)
prompt
string
Text prompt describing the image
negative_prompt
string
Negative prompt (empty string if none)
Stored as Parquet shards (data/train-*.parquet).
Load… See the full description on the dataset page: https://huggingface.co/datasets/aakkaasshh/mysdxl-dataset.vaigai-dataset
Vaigai Dataset
aakkaasshh/vaigai-dataset is an Indic-focused multilingual text corpus aggregated for tokenizer training, built with the goal of reaching tokenizer quality on par with dedicated Indic tokenizers (e.g. Sarvam AI's).
One Parquet file per language is stored at the root of this repo (e.g. bn.parquet), with no nested folders.
Every time new data is fetched for a language, it is merged with that language's existing file and deduplicated on the text column, so re-running… See the full description on the dataset page: https://huggingface.co/datasets/aakkaasshh/vaigai-dataset.
