datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tennis-sackmann-archive
Tennis datasets — archive of Jeff Sackmann's data
This dataset is an archival mirror of the public tennis datasets compiled by Jeff Sackmann.
It exists so the data remains available and citable. It contains only data and its documentation —
no models, analysis, or derived code.
A matching mirror lives on GitHub: https://github.com/Aneeshers/tennis-sackmann-archive
Contents
Folder
What it is
Files
Coverage
slam_pointbypoint/
Point-by-point logs for the… See the full description on the dataset page: https://huggingface.co/datasets/Aneeshers/tennis-sackmann-archive.Aneumo
Aneumo Datasets
AneumoDataset is a comprehensive multi-modal dataset containing 3D anatomical structures and simulated hemodynamic parameters for cerebral aneurysms, supporting both computational modeling and AI-based analysis.
AneuG-Flow
AneuG-Flow Dataset
Dataset Description
AneuG-Flow is a comprehensive dataset for intracranial aneurysm hemodynamics, containing computational fluid dynamics (CFD) simulations of blood flow through synthetic aneurysm morphologies generated by a generative model AneuG [1].
Dataset Summary
This dataset provides:
730 aneurysm cases with transient (time-varying) flow simulations
High-resolution 3D geometries (original and remeshed)
Blood flow data and wall shear… See the full description on the dataset page: https://huggingface.co/datasets/whding123/AneuG-Flow.anemia-survey-dataset
Anemia Detection — Multi-Modal Clinical SEWA Rural Dataset
Organisation: SEWA Rural — Society for Education, Welfare and Action (Rural), Jhagadia, Gujarat, India
Dataset: sewa-rural-care/anemia-survey-dataset
Contact: sewarural@ymail.com
Version: 1.0 — July 2026
Dataset Summary
This dataset supports research into non-invasive, smartphone-based anemia
screening applicable to low-resource and rural healthcare settings. It was
collected by SEWA Rural — a non-profit… See the full description on the dataset page: https://huggingface.co/datasets/sewa-rural-care/anemia-survey-dataset.cerebro-anexosan_emAnEM corpus is a domain- and species-independent resource manually annotated for anatomical
entity mentions using a fine-grained classification system. The corpus consists of 500 documents
(over 90,000 words) selected randomly from citation abstracts and full-text papers with
the aim of making the corpus representative of the entire available biomedical scientific
literature. The corpus annotation covers mentions of both healthy and pathological anatomical
entities and contains over 3,000 annotated mentions.anetimsdb-genre-movie-scripts
Dataset Card for "imsdb-genre-movie-scripts"
More Information needed
Anemia-seggemma3
[!Note]
This repository corresponds to the launch version of Gemma 3n E2B, to be used with Hugging Face transformers,
supporting text, audio, and vision (image and video) inputs.
Gemma 3n models have multiple architecture innovations:
They are available in two sizes based on effective parameters. While the raw parameter count of this model is 6B, the architecture design allows the model to be run with a memory footprint comparable to a traditional 2B model by offloading low-utilization… See the full description on the dataset page: https://huggingface.co/datasets/aneeshm44/gemma3.AneuGv2ru_med_history
Medical Histories Ru-ru
Medical Histories from Russian medical textbooks.
A text dataset with medical histories.
All dates were masked into .
dailydialog
DailyDialog - ShareGPT Processed
Dataset Summary
DailyDialog is a high-quality, multi-turn dialogue dataset containing human-written conversations that cover a wide variety of everyday topics.It is designed to support research in dialogue modeling, conversational AI, and emotion-aware interactions.The dataset emphasizes natural, contextually coherent exchanges that resemble real-world human dialogue, making it ideal for training AI systems that need to handle daily… See the full description on the dataset page: https://huggingface.co/datasets/anezatra/dailydialog.AnesCorpusThe AnesBench Datasets Collection comprises three distinct datasets: AnesBench, an anesthesiology reasoning benchmark; AnesQA, an SFT dataset; and AnesCorpus, a continual pre-training dataset. This repository pertains to AnesCorpus. For AnesBench and AnesQA, please refer to their respective links: https://huggingface.co/datasets/MiliLab/AnesBench and https://huggingface.co/datasets/MiliLab/AnesQA.
AnesCorpus
AnesCorpus is a large-scale, domain-specific corpus constructed for… See the full description on the dataset page: https://huggingface.co/datasets/MiliLab/AnesCorpus.Food-Deliverytask1486_cell_extraction_anem_dataset
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1486_cell_extraction_anem_dataset
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1486_cell_extraction_anem_dataset.rsna2025-aneurysm-26class-seg
RSNA 2025 Aneurysm — 26-class Vessel-Anatomy and Aneurysm Segmentation Labels
26-class vessel-anatomy and aneurysm segmentation labels (13 vessel anatomy
classes + 13 aneurysm location classes, values 0–26, see labels.json) for
the RSNA 2025 Intracranial Aneurysm Detection challenge, placed back into the
original image space (pure voxel placement, no resampling).
Paper: arXiv:2606.26706
Contents
Folder
Count
Aligned to
labelsTr_26classes_in_orig_space/… See the full description on the dataset page: https://huggingface.co/datasets/spc819/rsna2025-aneurysm-26class-seg.bimanual_handover_random_120This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_so_follower",
"total_episodes": 120,
"total_frames": 44712,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:120"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Aneysait/bimanual_handover_random_120.med_historyANERCorp
Dataset Card for "ANERCorp"
Papers:
Benajiba, Yassine, Paolo Rosso, and José Miguel Benedí Ruiz. "Anersys: An Arabic named entity recognition system based on maximum entropy." In International Conference on Intelligent Text Processing and Computational Linguistics, pp. 143-153. Springer, Berlin, Heidelberg, 2007.
Ossama Obeid, Nasser Zalmout, Salam Khalifa, Dima Taji, Mai Oudah, Bashar Alhafni, Go Inoue, Fadhl Eryani, Alexander Erdmann, and Nizar Habash. "CAMeL Tools: An… See the full description on the dataset page: https://huggingface.co/datasets/asas-ai/ANERCorp.anesthesia_literacy
Anesthesia Literacy Project
Adaptive patient education using large-language models
Overview
This study explores the potential of Large Language Models (LLMs) like OpenAI's Generative Pretrained Transformer (GPT) versions 3.5 and 4 to enhance the readability of preoperative patient instructions, aiming to align them with the American Medical Association's recommendation of a 6th-grade reading level. Acknowledging that nearly 40% of U.S. adults possess basic or below basic… See the full description on the dataset page: https://huggingface.co/datasets/stanfordaimlab/anesthesia_literacy.parsed_arxiv_cs_papers
Sample data snippet for JSON
{
"paper": {
"paper_id": "2103.15871",
"metadata": {
"id": "2103.15871",
"submitter": "Varun Kumar",
"authors": "Luoxin Chen, Francisco Garcia, Varun Kumar, He Xie, Jianhua Lu",
"title": "Industry Scale Semi-Supervised Learning for Natural Language\n Understanding",
"journal_ref": null,
"doi": null,
"report_no": null,
"categories": "cs.CL cs.AI cs.LG"… See the full description on the dataset page: https://huggingface.co/datasets/Aneerudh/parsed_arxiv_cs_papers.regfinalgemma-vlm-trial-11.7mug-tree-r0-sobolThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
7
],
"names": [
"cart_pos_x",
"cart_pos_y",
"cart_pos_z",
"cart_rot_x",
"cart_rot_y",
"cart_rot_z"… See the full description on the dataset page: https://huggingface.co/datasets/Aneeshers/mug-tree-r0-sobol.ane-rooflines
ANE Rooflines
Cross-Apple-Silicon performance and fp16-correctness measurements for the Apple
Neural Engine (ANE), collected with ANEForge.
Each row is one machine (grouped by hardware hash; identical silicon in different
chassis stays distinct by model identifier).
See it charted: the ANE leaderboard
ranks these machines by peak GEMM, perf-per-watt, and decode throughput.
These are community-contributed submissions mirrored from the public
bench/results/rooflines/
in the repo.… See the full description on the dataset page: https://huggingface.co/datasets/aneforge/ane-rooflines.mug-tree-r0-baselineThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
7
],
"names": [
"cart_pos_x",
"cart_pos_y",
"cart_pos_z",
"cart_rot_x",
"cart_rot_y",
"cart_rot_z"… See the full description on the dataset page: https://huggingface.co/datasets/Aneeshers/mug-tree-r0-baseline.SQuAD_HindiThis dataset is created by translating a part of the Stanford QA dataset.
It contains 5k QA pairs from the original SQuad dataset translated to Hindi using the googletrans api.
ANEMONES
Data Card: Sepsis vs. SIRS Point-of-Care Biomarker Whole-Blood Microarray Dataset (GSE236713)
Summary
Expression + sample metadata + feature metadata for GSE236713, a
multi-center UK study identifying transcriptional mRNA biomarkers to
discriminate sepsis from SIRS in adult ICU patients, profiled on the
Agilent SurePrint G3 Human GE v2 8x60K Microarray (GPL17077). Blood was
sampled at up to four timepoints (Day 1, Day 2, Day 5, and ICU discharge)
for patient… See the full description on the dataset page: https://huggingface.co/datasets/cmatkhan/ANEMONES.biblioteca-aneel-categorizado-4
Estatísticas calculadas
Resultados gerados em 2026-08-29T01:38:54+00:00, a partir do split train na revisão aebc8a155f5809ac2e65f7f101f34047dc454a21 do cemig-ceia/biblioteca-aneel-categorizado-4. Tokens contados com gpt2, sem tokens especiais e sem truncamento.
Overall
Total de registros: 149.004.
Medida
Total
Mínimo
Média
Mediana
P95
Máximo
Caracteres
999.518.350
259… See the full description on the dataset page: https://huggingface.co/datasets/cemig-ceia/biblioteca-aneel-categorizado-4.
