datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
turkey-all-universitiesCertainly! Here’s the dataset description in Markdown format:
All Universities in Turkey Dataset
Description
This dataset contains detailed information about various universities. Each record represents a single university and includes attributes such as the university's name, type, city, website, address, logo URL, and a button for accessing additional details. This data is typically extracted from a web page listing universities.
Fields
1. id… See the full description on the dataset page: https://huggingface.co/datasets/h8st6ptv/turkey-all-universities.africa-synth-aid-flows-medical-multimodal-fracture-all
Africa Synth Aid Flows Medical Multimodal Fracture All | Africa (Electric Sheep Africa metadata inventory)
Size category: 1K<n<10K - Formats: json - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-aid-flows-medical-multimodal-fracture-all.ALLaVA-4V
📚 ALLaVA-4V Data
Generation Pipeline
LAION
We leverage the superb GPT-4V to generate captions and complex reasoning QA pairs. Prompt is here.
Vison-FLAN
We leverage the superb GPT-4V to generate captions and detailed answer for the original instructions. Prompt is here.
Wizard
We regenerate the answer of Wizard_evol_instruct with GPT-4-Turbo.
Dataset Cards
All datasets can be found here.
The structure of naming is shown below:
ALLaVA-4V
├──… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/ALLaVA-4V.ALL-Bench-Leaderboard
🏆 ALL Bench Leaderboard 2026
The only AI benchmark dataset covering LLM · VLM · Agent · Image · Video · Music in a single unified file.
Dataset Summary
ALL Bench Leaderboard aggregates and cross-verifies benchmark scores for 90+ AI models across 6 modalities. Every numerical score is tagged with a confidence level (cross-verified, single-source, or self-reported) and its original source. The dataset is designed for researchers, developers, and… See the full description on the dataset page: https://huggingface.co/datasets/FINAL-Bench/ALL-Bench-Leaderboard.ALL-Bench-Leaderboard
🏆 ALL Bench Leaderboard 2026
The only AI benchmark dataset covering LLM · VLM · Agent · Image · Video · Music in a single unified file.
Dataset Summary
ALL Bench Leaderboard aggregates and cross-verifies benchmark scores for 90+ AI models across 6 modalities. Every numerical score is tagged with a confidence level (cross-verified, single-source, or self-reported) and its original source. The dataset is designed for researchers, developers, and… See the full description on the dataset page: https://huggingface.co/datasets/youssef3146/ALL-Bench-Leaderboard.ALLaVA-4V
📚 ALLaVA-4V Data
Generation Pipeline
LAION
We leverage the superb GPT-4V to generate captions and complex reasoning QA pairs. Prompt is here.
Vison-FLAN
We leverage the superb GPT-4V to generate captions and detailed answer for the original instructions. Prompt is here.
Wizard
We regenerate the answer of Wizard_evol_instruct with GPT-4-Turbo.
Dataset Cards
All datasets can be found here.
The structure of naming is shown below:
ALLaVA-4V… See the full description on the dataset page: https://huggingface.co/datasets/lodestones/ALLaVA-4V.ALLaVA-4V-Chinese
ALLaVA-4V for Chinese
This is the Chinese version of the ALLaVA-4V data. We have translated the ALLaVA-4V data into Chinese through ChatGPT and instructed ChatGPT not to translate content related to OCR.
The original dataset can be found here, and the image data can be downloaded from ALLaVA-4V.
Citation
If you find our data useful, please consider citing our work! We are FreedomIntelligence from Shenzhen Research Institute of Big Data and The Chinese University of… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/ALLaVA-4V-Chinese.pusht_96_norm4_visual_nomarker_allstep_thinking_trickiness_cot
PushT norm4 Visual Nomarker All-Step Thinking Trickiness COT
This dataset is derived from successful PushT visual-nomarker trajectories in novastar112/pusht_96_norm4_visual_nomarker.
Each row contains one full successful trajectory from the first move through the final stop action.
Main files:
training/pusht_allstep_thinking_cot.jsonl.gz: 500,000 train rows.
testing/pusht_allstep_thinking_cot.jsonl.gz: 200 test rows.
metadata/final_scan_validation.json: full local scan after repair… See the full description on the dataset page: https://huggingface.co/datasets/novastar112/pusht_96_norm4_visual_nomarker_allstep_thinking_trickiness_cot.ALLaVA-4V-Arabic
ALLaVA-4V for Arabic
This is the Arabic version of the ALLaVA-4V data. We have translated the ALLaVA-4V data into Arabic through ChatGPT and instructed ChatGPT not to translate content related to OCR.
The original dataset can be found here, and the image data can be downloaded from ALLaVA-4V.
Citation
If you find our data useful, please consider citing our work! We are FreedomIntelligence from Shenzhen Research Institute of Big Data and The Chinese University of Hong… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/ALLaVA-4V-Arabic.pusht_96_int1_visual_nomarker_allstep_thinking_trickiness_cot
PushT int1 Visual Nomarker All-Step Thinking Trickiness COT
This dataset is derived from successful PushT visual-nomarker trajectories in novastar112/pusht_96_int1_visual_nomarker.
Each row contains one full successful trajectory from the first move through the final stop action.
Main files:
training/pusht_allstep_thinking_cot.jsonl.gz: 500,000 train rows.
testing/pusht_allstep_thinking_cot.jsonl.gz: 200 test rows.
Message format:
Each user turn is the PushT prompt text plus one… See the full description on the dataset page: https://huggingface.co/datasets/novastar112/pusht_96_int1_visual_nomarker_allstep_thinking_trickiness_cot.letterboxd-all-movie-data
Letterboxd Film Dataset
This dataset contains a comprehensive collection of 847,209 films from the Letterboxd platform, including movie information, user reviews, and ratings.
Dataset Summary
Total Films: 847,209
File Size: ~1.12 GB (1,120,572,122 bytes)
Format: JSONL (JSON Lines)
Language: Primarily English, with some multilingual content
Data Structure
Each line contains a JSON object with the following fields:
{
"url":… See the full description on the dataset page: https://huggingface.co/datasets/PratikDhonde/letterboxd-all-movie-data.aiconf-butterfly-detection-allColumns:
taxon
photo_id
photo_url
hard
Dataset dedicated to further butterfly segmentation in the wild. The dataset contains photos of butterflies, bees, beetles, flowers, and shrubs, collected from iNaturalist.
The hard column indicates whether the photo is considered a hard example for butterfly detection according to a few CV techniques.
Generated at: 2026-03-15 16:40:21 UTC
Rows: 1800
all-internalNOAA-HRRR-HRRRAK-All-ImageCaption
