datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lambada
Dataset Card for LAMBADA
Dataset Summary
The LAMBADA evaluates the capabilities of computational models
for text understanding by means of a word prediction task.
LAMBADA is a collection of narrative passages sharing the characteristic
that human subjects are able to guess their last word if
they are exposed to the whole passage, but not if they
only see the last sentence preceding the target word.
To succeed on LAMBADA, computational models cannot
simply rely on local… See the full description on the dataset page: https://huggingface.co/datasets/cimec/lambada.CIMA-4.8-ADR
CIMA Sección 4.8 — Reacciones Adversas
Corpus de texto biomédico regulatorio en español compuesto por la
sección 4.8 ("Reacciones adversas") de la totalidad de las fichas
técnicas publicadas por la
Agencia Española de Medicamentos y Productos Sanitarios (AEMPS)
en su Centro de Información Online de Medicamentos
(CIMA).
Este recurso fue construido como base para el pre-entrenamiento
adaptado al dominio (continued pre-training / domain-adaptive
pre-training, DAPT) de modelos… See the full description on the dataset page: https://huggingface.co/datasets/guerrerotook/CIMA-4.8-ADR.CIMD
CIMD
[[中文]] | [[English]]
CSGHub Dataset Page | Hugging Face | OpenCSG Community
中文说明
数据集概述
CIMD 是一个面向文档智能任务的跨来源、多语言 JSONL 语料库。当前公开快照包含 111,308 条解析记录,覆盖制度参考、学术与长文档资料、机构分析、企业运营、公共讨论和市场相关材料等来源家族。每条记录都把正文与来源类型、语言、时间、关键词、授权标签和来源字段放在同一个结构里,用户拿到数据后可以直接做检索、抽样、审计和数据治理。
公开数据已转换为统一字段,并按来源家族拆分为可单独加载的子集;它不是原始文件夹的简单打包。用户可以只读取制度参考、学术长文档或公共讨论记录,也可以合并多个子集构建检索库、抽取训练候选样本、构造评测样本池,并按来源、语言和时间字段继续筛选。
CIMD 和通用网页语料的差别在于记录级元数据。它不只提供可索引文本,还提供… See the full description on the dataset page: https://huggingface.co/datasets/opencsg/CIMD.CIMemories
CIMemories: A Compositional Benchmark for Contextual Integrity of Persistent Memory in LLMs
Paper
Large Language Models (LLMs) increasingly use persistent memory from past interactions to enhance personalization and task performance. However, this memory introduces critical risks when sensitive information is revealed in inappropriate contexts. We present CIMemories, a benchmark for evaluating whether LLMs appropriately control information flow from memory based on task context.… See the full description on the dataset page: https://huggingface.co/datasets/facebook/CIMemories.CIMA-4.8-ADR-NER
CIMA Sección 4.8 NER
Dataset de reconocimiento de entidades nombradas (NER) en español sobre la
sección 4.8 ("Reacciones adversas") de las fichas técnicas publicadas por
la Agencia Española de Medicamentos y Productos Sanitarios (AEMPS)
en el Centro de Información Online de Medicamentos (CIMA).
Cada documento es la sección 4.8 completa de un medicamento, tokenizada con
spaCy (es_core_news_sm) y etiquetada en formato CoNLL BIO (IOB2) con
hasta cuatro tipos de entidad:
Tag… See the full description on the dataset page: https://huggingface.co/datasets/guerrerotook/CIMA-4.8-ADR-NER.CIMA-4.8-ADR-NER-EXTENDED
CIMA Sección 4.8 NER Extended — corpus sintético
⚠️ Aviso: todos los textos de este dataset son ficticios. Ninguno de los
98 documentos que contiene es una ficha técnica real. Fueron redactados por
un modelo de lenguaje imitando el estilo de la sección 4.8 ("Reacciones
adversas") de las fichas técnicas de la AEMPS. No deben usarse como fuente
de información clínica, farmacológica ni regulatoria sobre ningún
medicamento. Lo único que procede del mundo real es el conjunto de… See the full description on the dataset page: https://huggingface.co/datasets/guerrerotook/CIMA-4.8-ADR-NER-EXTENDED.cim10-textbookcima-drug-interactions-es
CIMA Spanish Drug-Drug Interactions Dataset (GALENO IA)
This dataset is the first open-source, structured database of cross drug-drug interactions in Spanish, extracted directly from the official Summary of Product Characteristics (SmPC) sheets of the Spanish Agency for Medicines and Health Products (AEMPS) via their public CIMA Online Medicines Information Center API and structured using the gemma-4-31B-it large language model.
Originally developed as a core clinical decision… See the full description on the dataset page: https://huggingface.co/datasets/lmolino/cima-drug-interactions-es.cim10-mcqagenerated-cim10-dpdrdas-chapCIM
Dataset Summary
It is a curated visual question answering (VQA) dataset designed to analyze how overlaid text affects visual reasoning in vision–language models.
Each sample consists of a natural image, a multiple-choice question, and four aligned image variants that differ only in the presence and semantic correctness of overlaid text. This structure enables controlled experiments on multimodal robustness, spurious correlations, and text-induced shortcut learning.
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/AHAAM/CIM.edition_0772_cimec-lambada-readymade
edition_0772_cimec-lambada-readymade
A Readymade by TheFactoryX
Original Dataset
cimec/lambada
Process
This dataset is a "readymade" - inspired by Marcel Duchamp's concept of taking everyday objects and recontextualizing them as art.
What we did:
Selected the original dataset from Hugging Face
Shuffled each column independently
Destroyed all row-wise relationships
Preserved structure, removed meaning
The result:
Same data. Wrong order. New meaning. No… See the full description on the dataset page: https://huggingface.co/datasets/TheFactoryX/edition_0772_cimec-lambada-readymade.generated-cim10-dpdrdasedition_0803_cimec-lambada-readymade
edition_0803_cimec-lambada-readymade
A Readymade by TheFactoryX
Original Dataset
cimec/lambada
Process
This dataset is a "readymade" - inspired by Marcel Duchamp's concept of taking everyday objects and recontextualizing them as art.
What we did:
Selected the original dataset from Hugging Face
Shuffled each column independently
Destroyed all row-wise relationships
Preserved structure, removed meaning
The result:
Same data. Wrong order. New meaning. No… See the full description on the dataset page: https://huggingface.co/datasets/TheFactoryX/edition_0803_cimec-lambada-readymade.africa-ethiopia-cimmyt-ethiopia-ac053bb9
Cimmyt Ethiopia | Africa (Ethiopia Open Data)
895 rows - 1 Africa country/area - 2022-2023 - source table - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 895 rows from Ethiopia Open Data, covering Cimmyt Ethiopia. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples.
What This Dataset Measures
Climate and environment datasets help analysts study… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-ethiopia-cimmyt-ethiopia-ac053bb9.edition_0807_cimec-lambada-readymade
edition_0807_cimec-lambada-readymade
A Readymade by TheFactoryX
Original Dataset
cimec/lambada
Process
This dataset is a "readymade" - inspired by Marcel Duchamp's concept of taking everyday objects and recontextualizing them as art.
What we did:
Selected the original dataset from Hugging Face
Shuffled each column independently
Destroyed all row-wise relationships
Preserved structure, removed meaning
The result:
Same data. Wrong order. New meaning. No… See the full description on the dataset page: https://huggingface.co/datasets/TheFactoryX/edition_0807_cimec-lambada-readymade.africa-morocco-evolution-de-la-consommation-du-ciment-24f200da
Evolution De La Consommation Du Ciment | Africa (Morocco Open Data)
8 rows - 1 Africa country/area - time not specified - source table - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 8 rows from Morocco Open Data, covering Evolution De La Consommation Du Ciment. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples.
What This Dataset Measures… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-morocco-evolution-de-la-consommation-du-ciment-24f200da.edition_1378_cimec-lambada-readymade
edition_1378_cimec-lambada-readymade
A Readymade by TheFactoryX
Original Dataset
cimec/lambada
Process
This dataset is a "readymade" - inspired by Marcel Duchamp's concept of taking everyday objects and recontextualizing them as art.
What we did:
Selected the original dataset from Hugging Face
Shuffled each column independently
Destroyed all row-wise relationships
Preserved structure, removed meaning
The result:
Same data. Wrong order. New meaning. No… See the full description on the dataset page: https://huggingface.co/datasets/TheFactoryX/edition_1378_cimec-lambada-readymade.edition_0800_cimec-lambada-readymade
edition_0800_cimec-lambada-readymade
A Readymade by TheFactoryX
Original Dataset
cimec/lambada
Process
This dataset is a "readymade" - inspired by Marcel Duchamp's concept of taking everyday objects and recontextualizing them as art.
What we did:
Selected the original dataset from Hugging Face
Shuffled each column independently
Destroyed all row-wise relationships
Preserved structure, removed meaning
The result:
Same data. Wrong order. New meaning. No… See the full description on the dataset page: https://huggingface.co/datasets/TheFactoryX/edition_0800_cimec-lambada-readymade.cim-pdf-synthaemps-cima
AEMPS CIMA Research Dataset
This directory builds a research-ready snapshot of CIMA, the medicine information system maintained by the Spanish Agency of Medicines and Medical Devices (AEMPS).
The source exposes official information about authorized and non-authorized medicines, commercial presentations, active ingredients, ATC codes, pharmaceutical forms, administration routes, regulatory status, safety-related indicators, segmented summaries of product characteristics and… See the full description on the dataset page: https://huggingface.co/datasets/hsilvosa/aemps-cima.cim-10edition_0796_cimec-lambada-readymade
edition_0796_cimec-lambada-readymade
A Readymade by TheFactoryX
Original Dataset
cimec/lambada
Process
This dataset is a "readymade" - inspired by Marcel Duchamp's concept of taking everyday objects and recontextualizing them as art.
What we did:
Selected the original dataset from Hugging Face
Shuffled each column independently
Destroyed all row-wise relationships
Preserved structure, removed meaning
The result:
Same data. Wrong order. New meaning. No… See the full description on the dataset page: https://huggingface.co/datasets/TheFactoryX/edition_0796_cimec-lambada-readymade.edition_1122_cimec-lambada-readymade
edition_1122_cimec-lambada-readymade
A Readymade by TheFactoryX
Original Dataset
cimec/lambada
Process
This dataset is a "readymade" - inspired by Marcel Duchamp's concept of taking everyday objects and recontextualizing them as art.
What we did:
Selected the original dataset from Hugging Face
Shuffled each column independently
Destroyed all row-wise relationships
Preserved structure, removed meaning
The result:
Same data. Wrong order. New meaning. No… See the full description on the dataset page: https://huggingface.co/datasets/TheFactoryX/edition_1122_cimec-lambada-readymade.generated-cim10-mpcdpdasedition_0947_cimec-lambada-readymade
edition_0947_cimec-lambada-readymade
A Readymade by TheFactoryX
Original Dataset
cimec/lambada
Process
This dataset is a "readymade" - inspired by Marcel Duchamp's concept of taking everyday objects and recontextualizing them as art.
What we did:
Selected the original dataset from Hugging Face
Shuffled each column independently
Destroyed all row-wise relationships
Preserved structure, removed meaning
The result:
Same data. Wrong order. New meaning. No… See the full description on the dataset page: https://huggingface.co/datasets/TheFactoryX/edition_0947_cimec-lambada-readymade.edition_1376_cimec-lambada-readymade
edition_1376_cimec-lambada-readymade
A Readymade by TheFactoryX
Original Dataset
cimec/lambada
Process
This dataset is a "readymade" - inspired by Marcel Duchamp's concept of taking everyday objects and recontextualizing them as art.
What we did:
Selected the original dataset from Hugging Face
Shuffled each column independently
Destroyed all row-wise relationships
Preserved structure, removed meaning
The result:
Same data. Wrong order. New meaning. No… See the full description on the dataset page: https://huggingface.co/datasets/TheFactoryX/edition_1376_cimec-lambada-readymade.edition_1484_cimec-lambada-readymade
edition_1484_cimec-lambada-readymade
A Readymade by TheFactoryX
Original Dataset
cimec/lambada
Process
This dataset is a "readymade" - inspired by Marcel Duchamp's concept of taking everyday objects and recontextualizing them as art.
What we did:
Selected the original dataset from Hugging Face
Shuffled each column independently
Destroyed all row-wise relationships
Preserved structure, removed meaning
The result:
Same data. Wrong order. New meaning. No… See the full description on the dataset page: https://huggingface.co/datasets/TheFactoryX/edition_1484_cimec-lambada-readymade.edition_1135_cimec-lambada-readymade
edition_1135_cimec-lambada-readymade
A Readymade by TheFactoryX
Original Dataset
cimec/lambada
Process
This dataset is a "readymade" - inspired by Marcel Duchamp's concept of taking everyday objects and recontextualizing them as art.
What we did:
Selected the original dataset from Hugging Face
Shuffled each column independently
Destroyed all row-wise relationships
Preserved structure, removed meaning
The result:
Same data. Wrong order. New meaning. No… See the full description on the dataset page: https://huggingface.co/datasets/TheFactoryX/edition_1135_cimec-lambada-readymade.africa-morocco-evolution-de-la-consommation-du-ciment-0af67935
Evolution De La Consommation Du Ciment | Africa (Morocco Open Data)
8 rows - 1 Africa country/area - time not specified - source table - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 8 rows from Morocco Open Data, covering Evolution De La Consommation Du Ciment. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples.
What This Dataset Measures… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-morocco-evolution-de-la-consommation-du-ciment-0af67935.
