cim
Datasets
All datasets matching “cim”lambada
Dataset Card for LAMBADA
Dataset Summary
The LAMBADA evaluates the capabilities of computational models
for text understanding by means of a word prediction task.
LAMBADA is a collection of narrative passages sharing the characteristic
that human subjects are able to guess their last word if
they are exposed to the whole passage, but not if they
only see the last sentence preceding the target word.
To succeed on LAMBADA, computational models cannot
simply rely on local… See the full description on the dataset page: https://huggingface.co/datasets/cimec/lambada.CIMA-4.8-ADR
CIMA Sección 4.8 — Reacciones Adversas
Corpus de texto biomédico regulatorio en español compuesto por la
sección 4.8 ("Reacciones adversas") de la totalidad de las fichas
técnicas publicadas por la
Agencia Española de Medicamentos y Productos Sanitarios (AEMPS)
en su Centro de Información Online de Medicamentos
(CIMA).
Este recurso fue construido como base para el pre-entrenamiento
adaptado al dominio (continued pre-training / domain-adaptive
pre-training, DAPT) de modelos… See the full description on the dataset page: https://huggingface.co/datasets/guerrerotook/CIMA-4.8-ADR.cimse-so100-experiment
CI-MSE ↔ closed-loop success — experiment artifacts
Everything needed to reproduce the correlation study without re-running training or simulation.
folder
contents
demos/train_h5, demos/val_h5
raw ManiSkill traj.h5 (obs images, qpos, actions, env states, success flags) + oracle videos + gen_stats.json
intervals/ground_truth.json
critical intervals from privileged sim state (grasp, place), 30 Hz frame indices
intervals/gemini_2.5_pro.json (+ raw responses)
VLM… See the full description on the dataset page: https://huggingface.co/datasets/Kavin60606/cimse-so100-experiment.CIMD
Chinese Instruction Multimodal Data (CIMD)
The dataset contains one million Chinese image-text pairs in total, including detailed image captioning and visual question answering.
Generation Pipeline
Image source
We randomly sample images from two opensource datasets Wanjuan and Wukong
Detailed caption generation
We use Gemini Pro Vision API to generate a detailed description for each image.
Question-answer pairs generation
Based on the generated caption, we use Gemini… See the full description on the dataset page: https://huggingface.co/datasets/jingzi/CIMD.cimse-so100-experiment-v2
⚠️ ERRATUM (2026-09-13). The so100_bowl_* datasets used for this experiment had misaligned language labels:
the ManiSkill→LeRobot converter stored episodes in lexicographic key order while per-episode task strings were
assigned in numeric order, so ~2/3 of training and validation episodes named the wrong cube colour. Consequences for
the results below: (1) the 21 checkpoints were trained on effectively random colour labels and never learned
grounding — the 30–60 % "wrong cube" failure rate is… See the full description on the dataset page: https://huggingface.co/datasets/Kavin60606/cimse-so100-experiment-v2.CIMD
CIMD
[[中文]] | [[English]]
CSGHub Dataset Page | Hugging Face | OpenCSG Community
中文说明
数据集概述
CIMD 是一个面向文档智能任务的跨来源、多语言 JSONL 语料库。当前公开快照包含 111,308 条解析记录,覆盖制度参考、学术与长文档资料、机构分析、企业运营、公共讨论和市场相关材料等来源家族。每条记录都把正文与来源类型、语言、时间、关键词、授权标签和来源字段放在同一个结构里,用户拿到数据后可以直接做检索、抽样、审计和数据治理。
公开数据已转换为统一字段,并按来源家族拆分为可单独加载的子集;它不是原始文件夹的简单打包。用户可以只读取制度参考、学术长文档或公共讨论记录,也可以合并多个子集构建检索库、抽取训练候选样本、构造评测样本池,并按来源、语言和时间字段继续筛选。
CIMD 和通用网页语料的差别在于记录级元数据。它不只提供可索引文本,还提供… See the full description on the dataset page: https://huggingface.co/datasets/opencsg/CIMD.
