datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
xMINDlarge
Dataset Card for xMINDlarge
Dataset Summary
xMINDlarge is an open, large-scale multi-parallel news dataset for multi- and cross-lingual news recommendation.
It is derived from the English MINDlarge dataset using open-source neural machine translation (i.e., NLLB 3.3B).
For the small version of the dataset, see xMINDsmall.
Uses
This dataset can be used for machine translation, text retrieval, or as a benchmark dataset for news recommendation.… See the full description on the dataset page: https://huggingface.co/datasets/aiana94/xMINDlarge.xMINDsmall
Dataset Card for xMINDsmall
Dataset Summary
xMINDsmall is an open, large-scale multi-parallel news dataset for multi- and cross-lingual news recommendation.
It is derived from the English MINDsmall dataset using open-source neural machine translation (i.e., NLLB 3.3B).
For the large version of the dataset, see xMINDlarge.
Uses
This dataset can be used for machine translation, text retrieval, or as a benchmark dataset for news recommendation.… See the full description on the dataset page: https://huggingface.co/datasets/aiana94/xMINDsmall.xMIL-HeatmapsHeatmap data for the multiple instance learning models presented in:
Jamshidi Idaji et al. "Beyond attention heatmaps: How to get better explanations for multiple instance learning models in histopathology". Medical Image Analysis (2026).
Link: https://www.sciencedirect.com/science/article/pii/S1361841526002173
Code: https://github.com/bifold-pathomics/xMIL
@article{
jamshidi26beyond,
title = {Beyond attention heatmaps: How to get better explanations for multiple instance learning models… See the full description on the dataset page: https://huggingface.co/datasets/bifold-pathomics/xMIL-Heatmaps.pack-xmi-of-my-project-3akkcp-8eeafbe1
GraspNet Eval — Transparent Object Grasping (Home)
Evaluation datapack for GraspNet on grasping transparent objects with a robotic arm in home environments. 30 renders at 128x128 across three home scenes (two bedrooms and a dressing), with depth (metric), world-space normals (OpenGL linear), attenuation, trimap and material index passes to expose transparent, reflective and translucent surfaces where grasp detection is unreliable. Goal: locate and quantify where the model fails.… See the full description on the dataset page: https://huggingface.co/datasets/physicl-test/pack-xmi-of-my-project-3akkcp-8eeafbe1.text-to-xmi-from-ecoreThis is a small test set for XMI instance model generation task.
It containing 26 pairs of meta-models (Ecore), specifications (natural language) and instance models (XMI).
In each pair, the meta-model and instance model share the same name. To proper open the instance model in Eclipse EMF, the instance model and meta-model should be placed in the same folder.
The meta-models are selected from https://huggingface.co/datasets/fpan/text-to-ocl-from-ecore.
The specifications are generated via… See the full description on the dataset page: https://huggingface.co/datasets/fpan/text-to-xmi-from-ecore.calm-month-6bb464
calm-month-6bb464
Synthetic products test data: 42 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/Harbor-Xmiller/calm-month-6bb464.X-mini-datasets
X-mini-datasets: The Foundational Dataset for Cybersecurity LLMs
Dataset Description
X-mini-datasets is a specialized, English-language dataset engineered as the foundational step to fine-tune Large Language Models (LLMs) into expert cybersecurity assistants. The dataset is uniquely structured into three distinct modules:
Core Knowledge Base (Payloads All The Things Adaptation): The largest part of the dataset, meticulously converted from the legendary "Payloads All The… See the full description on the dataset page: https://huggingface.co/datasets/saberbx/X-mini-datasets.all_xmi_20230705-no_split
Dataset Card for "all_xmi_20230705-no_split"
More Information needed
xmioimg-xmilmBrsnPaJWVc_XMI64
