datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
artelingo-dummyArtELingo is a benchmark and dataset introduced in a research paper aimed at promoting work on diversity across languages and cultures. It is an extension of ArtEmis, which is a collection of 80,000 artworks from WikiArt with 450,000 emotion labels and English-only captions. ArtELingo expands this dataset by adding 790,000 annotations in Arabic and Chinese. The purpose of these additional annotations is to evaluate the performance of "cultural-transfer" in AI systems.
The dataset in ArtELingo… See the full description on the dataset page: https://huggingface.co/datasets/youssef101/artelingo-dummy.bear-dataset
🐻 Mesosfer Bear AI - Multi-Stage Foundation Corpus
This repository contains the curated, structured-shuffled (seed=42), and 95% Train / 5% Validation split corpus used to pretrain and align Mesosfer Bear AI (a 16-layer transformer LLM optimized for high-efficiency training on AMD Instinct MI300X).
📊 Dataset Structure & Stage Breakdown
Stage 1: Pre-training (Foundational Language & General Knowledge)
Domain
Proposer / Source Repo
Weight… See the full description on the dataset page: https://huggingface.co/datasets/Dummy9898/bear-dataset.dummy_example_dataset
Dataset Name
Dummy Generated Dataset
Dataset Description
Generated dataset using AI
Dataset Structure
Data Fields
instruction: The task or question
input: Optional context or input
output: The expected response
Data Splits
Training: 5 examples
Validation: 3 examples
Test: 2 examples
Usage
from datasets import load_dataset
dataset = load_dataset("DannyAI/dummy_example_dataset")
Citation Information
If you use this… See the full description on the dataset page: https://huggingface.co/datasets/DannyAI/dummy_example_dataset.dummy_health_data
Synthetic Healthcare Dataset
Overview
This dataset is a synthetic healthcare dataset created for use in data analysis. It mimics real-world patient healthcare data and is intended for applications within the healthcare industry.
Data Generation
The data has been generated using the Faker Python library, which produces randomized and synthetic records that resemble real-world data patterns. It includes various healthcare-related fields such as patient… See the full description on the dataset page: https://huggingface.co/datasets/vrajakishore/dummy_health_data.intelli-maintain-dummy-tickets
Dataset Card for IntelliMaintain Dummy Tickets
