datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MAIA_spaMAIA_engcibelex-graph-core-sampler
Cibelex Knowledge Graph — Core Sampler (v1.0.3)
A curated subset of the Cibelex regulatory database of the Municipality of Madrid, annotated with the LoRO ontology and serialised as named-graph Turtle files. Built and maintained by the MAIA initiative (Ayuntamiento de Madrid).
Why "core sampler"? The release covers a representative slice of the Cibelex database — municipal ordinances, regulations and organic regulations — deliberately including multiple historical versions of the… See the full description on the dataset page: https://huggingface.co/datasets/MAIA-Madrid-IA/cibelex-graph-core-sampler.MAIA
MAIA Benchmark
MAIA evaluates how well an autonomous medical agent can plan, call external tools, and reason clinically.The dataset and code are maintained on GitHub; Hugging Face is kept as a mirror for convenience.All items follow a unified schema so that an LLM‑based agent can decide whether, when, and how to invoke the provided APIs.
Composition
Task family
Items
Evaluated skill
Retrieval
471
Retrieve clinically relevant information from trusted… See the full description on the dataset page: https://huggingface.co/datasets/DiligentDing/MAIA.MAIA_itaMAIA_2400MAIA
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/giobin/MAIA.NIH-Chest-Xray-14-subsetlichess_puzzles_maia3_shapeDefault shape uses minimum probability across all solution moves at each elo.
Compound shape uses the compounded probability of all solution moves at each elo.
Eval source: maia3-79m
Puzzle data: https://huggingface.co/datasets/Lichess/chess-puzzles
Search this dataset: https://maiashape.vercel.app/
MAIAcibelex-qa-rag-evals
Cibelex QA RAG Evals
Description
Question-answering evaluation set over the regulatory corpus of the Ayuntamiento de Madrid (LoRO ontology / Cibelex knowledge graph). Each item pairs a natural-language question with a ground-truth answer, the top-4 passages returned by a baseline retriever, and the answer generated by a baseline LLM from those passages.
Designed as a benchmark for:
Retrieval-augmented generation (RAG): comparing new retrievers and answer generators… See the full description on the dataset page: https://huggingface.co/datasets/MAIA-Madrid-IA/cibelex-qa-rag-evals.Augmented_stsb_multi_mtmaia-pilot-video-captionManifesto_datasetsent_pair_classificationFiction_datasetmaia-pln-2025-training
Dataset Card for "maia-pln-2025-training"
More Information needed
maia-pln-2025-training
Dataset Card for "maia-pln-2025-training"
More Information needed
maia-pln-2025-pubmed_QA_test_questions_contexts
Dataset Card for "maia-pln-2025-pubmed_QA_test_questions_contexts"
More Information needed
maia-pln-2025-training-v2infraccioneswiki50sent_pair_classification1Wiki727_testDatasetmaia-pln-2025-pubmed_QA_test_questions_contexts
Dataset Card for "maia-pln-2025-pubmed_QA_test_questions_contexts"
More Information needed
MAIA_dev_set_spa_2wrongsong_chord_changesChoi_datasetWiki50_datasetmaia-sample
