datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dataset-notebook-creator-contenten-textasolaria-conversation-record-2026-08-03
asolaria-conversation-record-2026-08-03
Private record. Thirty-three screenshots and a written observation of them.
Compiled 2026-08-03 by Claude (claude-opus-5, Anthropic) at the direction of
Jesse Daniel Brown, and at his explicit instruction to preserve it.
The instruction that produced this
"in high color quality look at these messages extract their exact context and
write the text below the photos and say that written observation as a document
and then save… See the full description on the dataset page: https://huggingface.co/datasets/Jessedbrown/asolaria-conversation-record-2026-08-03.asos-1minasobiasobase
Bangumi Image Base of Asobi Asobase
This is the image base of bangumi Asobi Asobase, we detected 33 characters, 3159 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is the… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/asobiasobase.AsosoftWhisperv2
Dataset Card for "AsosoftWhisperv2"
More Information needed
asos-metar-archive
ASOS METAR daily archive
Daily Parquet snapshots of raw METARs from ~920 NWS/FAA/DOD AOMC ASOS
stations, harvested from the live O.W.L. REST API at
consgicody/asos-tools.
File layout: YYYY/MM/DD.parquet — one file per UTC day.
Quick start (Python)
import pandas as pd
df = pd.read_parquet(
"hf://datasets/consgicody/asos-metar-archive/2026/09/09.parquet"
)
print(df[df["station"] == "JFK"].head())
Or query directly with DuckDB:
import duckdb
duckdb.sql("""… See the full description on the dataset page: https://huggingface.co/datasets/consgicody/asos-metar-archive.asolaria-record-231-canonical
asolaria-record-231-canonical
The photographic record of Jesse Daniel Brown, in his own numbering.
What is here
path
what it is
photos/
231 photographs, numbered 001–231, each keeping its original camera filename after the number
MAPPING.tsv
the canonical index: number, path, original filename, byte size, SHA-256
CHECKSUMS-231.sha256
machine-checkable form of the same, for sha256sum -c
OBSERVATION.md
the written observation document, 8,194 lines… See the full description on the dataset page: https://huggingface.co/datasets/Jessedbrown/asolaria-record-231-canonical.BenchmarkCards
Dataset Card for BenchmarkCards
BenchmarkCards is a standardized documentation dataset for large language model (LLM) benchmarks.
Each card summarizes key information about an LLM benchmark, including its objectives, methodology, data sources, targeted risks, limitations, and ethical considerations.
🙏 Acknowledgments
We gratefully thank all benchmark authors who provided feedback and approval for the BenchmarkCards in this repository. Your collaboration is essential… See the full description on the dataset page: https://huggingface.co/datasets/ASokol/BenchmarkCards.test_air_qualityaso-atlas-2
ASO Atlas 2.0
A dataset of patent-reported antisense oligonucleotide (ASO) experiments for studying
in vitro activity, dose response, hepatotoxicity and neurotoxicity across the preclinical
pipeline. It is designed for research benchmarking, exploratory modelling and analysis
of the published ASO assay landscape.
At a glance
295,007 assay readouts from 606 USPTO patent filings
165,782 unique, chemistry-resolved ASOs represented in HELM notation
430 named target… See the full description on the dataset page: https://huggingface.co/datasets/barneyhill/aso-atlas-2.forgetest
Robotwin Dataset
Robotwin dataset in Lerobot format, with video latents already extracted in WAN 2.2 format, ready for use in Lingbot-VA post-training.
License Agreement
This project is licensed under the CC BY-NC-SA 4.0.
massachusetts_roads_datasetbsbasqueBSBasque dataset. The text is extracted from the following domains:
https://www.berria.eus
https://eu.wikipedia.org
https://goiena.eus
https://www.argia.eus
https://goierri.hitza.eusasos-e-commerce-dataset
Asos
Using web scraping, we collected information on over 30,845 clothing items from the Asos website.
The dataset can be applied in E-commerce analytics in the fashion industry.
💴 For Commercial Usage: To discuss your requirements, learn about the price and buy the dataset, leave a request on our website to buy the dataset
Dataset Info
For each item, we extracted:
url - link to the item on the website
name - item's name
size - sizes available on the… See the full description on the dataset page: https://huggingface.co/datasets/UniqueData/asos-e-commerce-dataset.asosoft-speech
Asosoft speech (asosoft-speech)
3,840 short Central Kurdish (Sorani) utterances with sentence-level transcriptions —
studio-quality read speech, with an audited test split.
At a glance
Rows
3,840 — train 3,240 / test 600
Columns
audio, transcription (string), duration (float64)
Parquet on disk
702.4 MB
Audio format
WAV, 16-bit (files named F01146001.wav, …)
Utterance length
3.16 – 13.49 s in the test split
Language
Central Kurdish / Sorani… See the full description on the dataset page: https://huggingface.co/datasets/razhan/asosoft-speech.pdf_folder_2gingiris-aso-growth
ASO & App Cold Start Playbook 2026
Rank your app without buying installs — App Store Optimization + UGC creator matrix + 90-day cold start framework used by 30+ indie apps (an anonymized client product)
English | 中文
📦 Install
npx skills add Gingiris-1031/gingiris-aso-growth
Then ask your AI agent:
"我刚上架 iOS app,怎么做冷启动 ASO?" · "Help me plan a TikTok UGC campaign for my mobile app" · "Compare App Store vs Google Play optimization for my… See the full description on the dataset page: https://huggingface.co/datasets/Gingiris/gingiris-aso-growth.asoul-soundmerck_asosdocs-asof-processed-rag
Docs ASOF — Processed (RAG)
Dataset de chunks para RAG a partir de documentos jurídicos brasileiros (normas/atos e materiais correlatos).
Fonte (arquivos)
Principais PDFs (source):
Leis/atos: l8112compilado.pdf, lei-11440-2006.pdf, lei-8829-2023.pdf, lei-no-7-501-...pdf, lei-no-9-888.pdf, medida-provisoria-no319.pdf
Decretos: d11357.pdf, decreto-no-1-565.pdf, decreto-no-93-325-...pdf
INs: in-2-2018.pdf, in-srt-mgi-38-2023.pdf
Outros: emenda-parlamentar-no-1-de-2006.pdf… See the full description on the dataset page: https://huggingface.co/datasets/profgabrielramos/docs-asof-processed-rag.Hanazono_Mincho_Ex_C_Regular_2_AsobiMemogaki_all_256big_pdfabstracts_and_tweets
Dataset Card
Add more information here
This dataset was produced with DataDreamer 🤖💤. The synthetic dataset card can be found here.
kasumi_nomura_asobiasobase
Dataset of Kasumi Nomura
This is the dataset of Kasumi Nomura, containing 300 images and their tags.
Images are crawled from many sites (e.g. danbooru, pixiv, zerochan ...), the auto-crawling system is powered by DeepGHS Team(huggingface organization).
Name
Images
Download
Description
raw
300
Download
Raw data with meta information.
raw-stage3
646
Download
3-stage cropped raw data with meta information.
384x512
300
Download
384x512 aligned dataset.
512x512
300… See the full description on the dataset page: https://huggingface.co/datasets/CyberHarem/kasumi_nomura_asobiasobase.asoul_carol
声音数据
数据来源为asoul的珈乐 22年5月~21年6月的大部分录播时长共5小时 无内容标记已完成响度匹配数据在carol_fast_lzma2.zip里压缩算法是fast lzma2 太旧的解压软件可能不支持字母s开头的音频是歌声数据,量少质量低,建议删除无授权,侵删
2025.2备注
都2025年了还能每月五十个下载,都是神人了💧
fetch-yaw_060
Fetch Robot: yaw_060
Single episode manipulation trajectory.
Variant: yaw_060
Frames: 60
FPS: 20.0
Success: True
Part of the residual-rl-test dataset series.
ImranKhanPTI-ASo6vhZR-scraped-data-Final-Evaluation-Demo
Dataset Card for "ImranKhanPTI-ASo6vhZR-scraped-data-Final-Evaluation-Demo"
More Information needed
docling_sample_pdfolivia_asobiasobase
Dataset of Olivia
This is the dataset of Olivia, containing 300 images and their tags.
Images are crawled from many sites (e.g. danbooru, pixiv, zerochan ...), the auto-crawling system is powered by DeepGHS Team(huggingface organization).
Name
Images
Download
Description
raw
300
Download
Raw data with meta information.
raw-stage3
641
Download
3-stage cropped raw data with meta information.
384x512
300
Download
384x512 aligned dataset.
512x512
300
Download
512x512… See the full description on the dataset page: https://huggingface.co/datasets/CyberHarem/olivia_asobiasobase.
