datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
genshin-voices-separatedAtelier-Code-3-sep-25sepsis-omics-datasets
脓毒症 (Sepsis) 公共组学与临床数据集合
冻结快照 · 2026-06-20 · 共 500 个数据集 · 7.4 GB · 全部带文件、信息卡与元数据
本仓库系统收集与脓毒症 / 败血症 / 脓毒性休克 / 菌血症 / 内毒素血症 / SIRS 相关的公开数据,
覆盖转录组、单细胞、空间转录组、蛋白质组、代谢组、外泌体、微生物组、表观(甲基化/染色质)以及
临床试验与药物数据。检索关键词:sepsis, septic shock, septicemia, septicaemia, bacteremia, endotoxemia, SIRS。
数据来源
数据库
数据集数
GEO
133
ClinicalTrials.gov
121
ArrayExpress/BioStudies
108
PRIDE/ProteomeXchange
75
Metabolomics Workbench
38
MetaboLights
25
数据类型… See the full description on the dataset page: https://huggingface.co/datasets/wei82/sepsis-omics-datasets.Gympedia
Gympedia
Gympedia is a collection of agentic LLM trajectories and distilled skills gathered across the diverse task environments of GEM: A Gym for Agentic LLMs. It pairs raw agent rollouts (per task, per model) with human-/model-readable "skill" write-ups distilled from successful runs, so it can be used for behavior analysis, skill/knowledge distillation, SFT data curation, and agentic RL research.
Note: This release intentionally excludes trajectories collected in synthetic… See the full description on the dataset page: https://huggingface.co/datasets/Septzzz/Gympedia.Persian_sentimentSE-Probe-data
SE-Probe full results dataset
Pre-computed CKA, diffusion-map, and probing results for SE-Probe, the public code release for "Where Does Speech Enhancement Adapt? Probing Study Under Controlled Degradation" (Amar, Ivry, Cohen, 2026).
📄 Paper: arXiv:2512.00482
Layout
snr/cka_snr_<model>.parquet: per-architecture (MUSE, MP-SENet, Demucs) CKA values across additive-noise SNRs and DEMAND noise types, with per-row audio quality metrics (PESQ, STOI, SI-SDR, DNSMOS… See the full description on the dataset page: https://huggingface.co/datasets/yairamr/SE-Probe-data.separate_coke_and_sprite_25_07_15_parquetseparate_coke_and_sprite_parquetseperate_coke_and_sprite_parquetseptuagint-lxx
NuBerea Septuagint (LXX) Morphology
Word-level morphological annotations of the Septuagint (Rahlfs 1935 edition, Old Testament and Deuterocanonical books), derived from the CenterBLC LXX Text-Fabric dataset. This dataset is part of the NuBerea curated corpus estate of biblical source texts.
Attribution
Attribute
Value
Source
CenterBLC/LXX (Text-Fabric), Center of Biblical Languages and Computing
Edition
Rahlfs, Alfred. 1935. Septuaginta. Deutsche… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/septuagint-lxx.lacap19m-sep-testseparate_recycling_0223separate_beaker_and_bottle_25_07_19_parquetseptuagint-analysis
NuBerea Septuagint Textual Analysis
Curated datasets for study of the Septuagint (the ancient Greek translation of the
Hebrew Bible), part of the NuBerea corpus estate of biblical and patristic texts.
It gathers Septuagint verse texts, apparatus notes, and edition-comparison material
into a set of ready-to-load configurations.
Attribution
This dataset derives from the following upstream sources, which require attribution:
Source
License
Rahlfs… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/septuagint-analysis.sepalith
Sepalith dataset
Open, R-specialized next-edit-suggestion training data. Private.
Layout (read this first)
The repo is a projection of the working corpus (a NAS). Everything
here is either (a) source data with its license trail, (b) derived
synthetic families, or (c) assembled training mixtures. Heavy raw
corpora are summarized at ledger/manifest level here; full content
lives on the source system (see each corpus entry below).
path
what it is… See the full description on the dataset page: https://huggingface.co/datasets/scholzmx/sepalith.AgiBotWorld-Beta_G1_task_466_Separate_dark_and_light_colored_clothes
agibot_task_466
This dataset converts the AgiBot format uniformly into LeRobot V3.0.
Dataset Statistics
robot_name: G1
end_effector: 夹爪
task: 把深色衣服和浅色衣服分开
total_episodes: 823
total_tasks: 1
size: 54G
Dataset Structure
├── data
│ └── chunk-xxx
│ ├── file-xxx.parquet
├── meta
│ ├── episodes
│ │ └── chunk-xxx
│ │ └── file-xxx.parquet
│ ├── info.json
│ ├── stats.json
│ └── tasks.parquet
└── videos
├──… See the full description on the dataset page: https://huggingface.co/datasets/BAAI-DataCube/AgiBotWorld-Beta_G1_task_466_Separate_dark_and_light_colored_clothes.separate-robots-sweep-cubesThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "ur5-panda",
"total_episodes": 360,
"total_frames": 109008,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:360"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/rdoshi21/separate-robots-sweep-cubes.testing-optc-sep17sepiq-sf-benchseparate_5coles_and_5sprites_25_08_25_parquetold_code_20b_separateseparate_beaker_and_bottle_25_07_19_lerobotv21swift
Dataset Card for "swift"
More Information needed
bedroom-sep6This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
11
],
"names": [
"vel_x",
"vel_y",
"vel_z",
"room_vel_x",
"room_vel_y",
"wrist_speed",
"finger_speed"… See the full description on the dataset page: https://huggingface.co/datasets/naavox/bedroom-sep6.Arabic_sentimentseparate_coke_and_sprite_25_07_15_lerobotv21french_5p_separateEgoLoc-Separation-GRPO
EgoLoc Separation GRPO Dataset
This is a self-contained 3x3 image-grid dataset for GRPO training on exact
separation/end localization. The numbered cells are chronological and use 1-based
indices.
This dataset is used to improve a VLM's accuracy for the EgoLoc pipeline.
This dataset IS NOT shuffled. When undergoing GRPO, recommend shuffling the dataset.
3x3 grid dataset for VLM tuning on separation frame identification.
Splits
Training rows: 1127
Validation rows:… See the full description on the dataset page: https://huggingface.co/datasets/yuchenxie/EgoLoc-Separation-GRPO.sept15_lerobot_bottle_taskThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "",
"total_episodes": 124,
"total_frames": 70555,
"total_tasks": 2,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 50.0,
"splits": {
"train": "0:124"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/gauravpradeep/sept15_lerobot_bottle_task.sandboxai_german_to_english_translations_seperated
