datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
anemia-survey-dataset
Anemia Detection — Multi-Modal Clinical SEWA Rural Dataset
Organisation: SEWA Rural — Society for Education, Welfare and Action (Rural), Jhagadia, Gujarat, India
Dataset: sewa-rural-care/anemia-survey-dataset
Contact: sewarural@ymail.com
Version: 1.0 — July 2026
Dataset Summary
This dataset supports research into non-invasive, smartphone-based anemia
screening applicable to low-resource and rural healthcare settings. It was
collected by SEWA Rural — a non-profit… See the full description on the dataset page: https://huggingface.co/datasets/sewa-rural-care/anemia-survey-dataset.imsdb-genre-movie-scripts
Dataset Card for "imsdb-genre-movie-scripts"
More Information needed
dailydialog
DailyDialog - ShareGPT Processed
Dataset Summary
DailyDialog is a high-quality, multi-turn dialogue dataset containing human-written conversations that cover a wide variety of everyday topics.It is designed to support research in dialogue modeling, conversational AI, and emotion-aware interactions.The dataset emphasizes natural, contextually coherent exchanges that resemble real-world human dialogue, making it ideal for training AI systems that need to handle daily… See the full description on the dataset page: https://huggingface.co/datasets/anezatra/dailydialog.ru_med_history
Medical Histories Ru-ru
Medical Histories from Russian medical textbooks.
A text dataset with medical histories.
All dates were masked into .
mug-tree-r0-baselineThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
7
],
"names": [
"cart_pos_x",
"cart_pos_y",
"cart_pos_z",
"cart_rot_x",
"cart_rot_y",
"cart_rot_z"… See the full description on the dataset page: https://huggingface.co/datasets/Aneeshers/mug-tree-r0-baseline.mug-tree-r0-sobolThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
7
],
"names": [
"cart_pos_x",
"cart_pos_y",
"cart_pos_z",
"cart_rot_x",
"cart_rot_y",
"cart_rot_z"… See the full description on the dataset page: https://huggingface.co/datasets/Aneeshers/mug-tree-r0-sobol.task1486_cell_extraction_anem_dataset
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1486_cell_extraction_anem_dataset
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1486_cell_extraction_anem_dataset.bimanual_handover_random_120This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_so_follower",
"total_episodes": 120,
"total_frames": 44712,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:120"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Aneysait/bimanual_handover_random_120.Food-DeliveryAnesCorpusThe AnesBench Datasets Collection comprises three distinct datasets: AnesBench, an anesthesiology reasoning benchmark; AnesQA, an SFT dataset; and AnesCorpus, a continual pre-training dataset. This repository pertains to AnesCorpus. For AnesBench and AnesQA, please refer to their respective links: https://huggingface.co/datasets/MiliLab/AnesBench and https://huggingface.co/datasets/MiliLab/AnesQA.
AnesCorpus
AnesCorpus is a large-scale, domain-specific corpus constructed for… See the full description on the dataset page: https://huggingface.co/datasets/MiliLab/AnesCorpus.biblioteca-aneel-categorizado-4
Estatísticas calculadas
Resultados gerados em 2026-08-29T01:38:54+00:00, a partir do split train na revisão aebc8a155f5809ac2e65f7f101f34047dc454a21 do cemig-ceia/biblioteca-aneel-categorizado-4. Tokens contados com gpt2, sem tokens especiais e sem truncamento.
Overall
Total de registros: 149.004.
Medida
Total
Mínimo
Média
Mediana
P95
Máximo
Caracteres
999.518.350
259… See the full description on the dataset page: https://huggingface.co/datasets/cemig-ceia/biblioteca-aneel-categorizado-4.ANERCorp
Dataset Card for "ANERCorp"
Papers:
Benajiba, Yassine, Paolo Rosso, and José Miguel Benedí Ruiz. "Anersys: An Arabic named entity recognition system based on maximum entropy." In International Conference on Intelligent Text Processing and Computational Linguistics, pp. 143-153. Springer, Berlin, Heidelberg, 2007.
Ossama Obeid, Nasser Zalmout, Salam Khalifa, Dima Taji, Mai Oudah, Bashar Alhafni, Go Inoue, Fadhl Eryani, Alexander Erdmann, and Nizar Habash. "CAMeL Tools: An… See the full description on the dataset page: https://huggingface.co/datasets/asas-ai/ANERCorp.mug_tree_session1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
7
],
"names": [
"cart_pos_x",
"cart_pos_y",
"cart_pos_z",
"cart_rot_x",
"cart_rot_y",
"cart_rot_z"… See the full description on the dataset page: https://huggingface.co/datasets/Aneeshers/mug_tree_session1.ANEMONES
Data Card: Sepsis vs. SIRS Point-of-Care Biomarker Whole-Blood Microarray Dataset (GSE236713)
Summary
Expression + sample metadata + feature metadata for GSE236713, a
multi-center UK study identifying transcriptional mRNA biomarkers to
discriminate sepsis from SIRS in adult ICU patients, profiled on the
Agilent SurePrint G3 Human GE v2 8x60K Microarray (GPL17077). Blood was
sampled at up to four timepoints (Day 1, Day 2, Day 5, and ICU discharge)
for patient… See the full description on the dataset page: https://huggingface.co/datasets/cmatkhan/ANEMONES.real01b-mug-tree-r0-eval2mono-plus10-blindThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
7
],
"names": [
"cart_pos_x",
"cart_pos_y",
"cart_pos_z",
"cart_rot_x",
"cart_rot_y",
"cart_rot_z"… See the full description on the dataset page: https://huggingface.co/datasets/Aneeshers/real01b-mug-tree-r0-eval2mono-plus10-blind.mug-tree-r0-evalThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
7
],
"names": [
"cart_pos_x",
"cart_pos_y",
"cart_pos_z",
"cart_rot_x",
"cart_rot_y",
"cart_rot_z"… See the full description on the dataset page: https://huggingface.co/datasets/Aneeshers/mug-tree-r0-eval.task1487_organism_substance_extraction_anem_dataset
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1487_organism_substance_extraction_anem_dataset
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1487_organism_substance_extraction_anem_dataset.mini_orca_style_instructionsbiblioteca-aneel-categorizado-2
Biblioteca ANEEL categorizada
Este dataset contém exclusivamente as 149.004 linhas da porção ANEEL de
cemig-ceia/energy_dataset_v1, fixada na revisão
ee59444965243734f9b9d2bacdd0ec14c139a1ec. Não há documentos da origem ONS ou CEMIG.
Taxonomia
A classificação reproduz os 13 ramos de Atos Relevantes do portal de atos oficiais da
ANEEL. O snapshot, o relatório de cobertura e o script de geração estão versionados neste
repositório. As fontes dinâmicas (DOU, MME e… See the full description on the dataset page: https://huggingface.co/datasets/cemig-ceia/biblioteca-aneel-categorizado-2.so101_pick_cubeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/Aneysait/so101_pick_cube.biblioteca-aneel-categorizado
Biblioteca ANEEL categorizada
Dataset com as 149.004 linhas ANEEL do cemig-ceia/energy_dataset_v1,
categorizadas por relações normativas explícitas encontradas nas páginas de Atos
Relevantes da ANEEL.
Foram observados 95 atos. Registros sem associação
inequívoca receberam categories=["Outros"] e category_paths=["Outros"].
Uma linha pode possuir múltiplos caminhos sem ser duplicada. Família, prefixo e
papel documental seguem o mapeamento do Energy v1/v2. O arquivo de observações… See the full description on the dataset page: https://huggingface.co/datasets/cemig-ceia/biblioteca-aneel-categorizado.imsdb-drama-movie-scripts
Dataset Card for "imsdb-drama-movie-scripts"
More Information needed
python18k_instruct_sharegpt
Note:
This dataset builds upon the iamtarun/python_code_instructions_18k_alpaca dataset and adheres to the ShareGPT format with a unique “conversations” column containing messages in JSONL. Unlike simpler formats like Alpaca, ShareGPT is ideal for storing multi-turn conversations, which is closer to how users interact with LLMs.
Example:
from datasets import load_dataset
dataset = load_dataset("AnelMusic/python18k_instruct_sharegpt", split = "train")
def… See the full description on the dataset page: https://huggingface.co/datasets/AnelMusic/python18k_instruct_sharegpt.so101_pick_color_multiThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/Aneysait/so101_pick_color_multi.biblioteca-aneel-categorizado-3
Biblioteca ANEEL categorizada
Este dataset contém exclusivamente as 149.004 linhas da porção ANEEL de
cemig-ceia/energy_dataset_v1, fixada na revisão
ee59444965243734f9b9d2bacdd0ec14c139a1ec. Não há documentos da origem ONS ou CEMIG.
Taxonomia
A classificação reproduz os 13 ramos de Atos Relevantes do portal de atos oficiais da
ANEEL. O snapshot, o relatório de cobertura e o script de geração estão versionados neste
repositório. As fontes dinâmicas (DOU, MME e… See the full description on the dataset page: https://huggingface.co/datasets/cemig-ceia/biblioteca-aneel-categorizado-3.imsdb-sci-fi-movie-scripts
Dataset Card for "imsdb-sci-fi-movie-scripts"
More Information needed
rollout_eval_pickred_finalThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/Aneysait/rollout_eval_pickred_final.circuit_articlesempathetic-dialogues-sharegpt
EmpatheticDialogues - ShareGPT Processed
Dataset Summary
EmpatheticDialogues is a large-scale, open-domain dialogue dataset designed to help AI systems recognize, understand, and respond to human emotions more naturally. While humans can easily identify and acknowledge others’ feelings during conversation, this remains a major challenge for artificial dialogue agents due to the lack of high-quality empathetic datasets.
This dataset introduces a new benchmark for… See the full description on the dataset page: https://huggingface.co/datasets/anezatra/empathetic-dialogues-sharegpt.AnesQAThe AnesBench Datasets Collection comprises three distinct datasets: AnesBench, an anesthesiology reasoning benchmark; AnesQA, an SFT dataset; and AnesCorpus, a continual pre-training dataset. This repository pertains to AnesQA. For AnesBench and AnesCorpus, please refer to their respective links: https://huggingface.co/datasets/MiliLab/AnesBench and https://huggingface.co/datasets/MiliLab/AnesCorpus.
AnesQA
AnesQA is a bilingual question-answering (QA) dataset designed for… See the full description on the dataset page: https://huggingface.co/datasets/MiliLab/AnesQA.
