datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
conceptual_captions_jsonCineBoard3D-plus
🎬 CineBoard3D++: Dynamic 3D Story World Dataset
📊 Dataset Summary
CineBoard3D++ is a collection of editable, movie-inspired 3D story worlds built with StoryBlender for narrative-grounded camera planning and world visual attention. It brings together story scripts, animated characters, scene geometry, and shot-level configurations in native Blender projects.
The benchmark covers 50 stories, 457 scenes, 1,585 shots, and 3,197 3D assets (836 plot-related and 2,361… See the full description on the dataset page: https://huggingface.co/datasets/EngineeringAI-LAB/CineBoard3D-plus.lam_bench_env_v2
LAM Bench portable environments and verified RoboLab trajectories
This repository contains the current portable five-room environment package and successful RoboLab arm trajectories.
lam_bench_portable_five_scene.tar: complete strict_five_portable_v1 directory from the server, including five USD entrypoints, vendored dependencies, relative-reference rewrites, and render evidence.
lam_bench_robolab_success_trajectories.tar: successful RoboLab trajectory package and native… See the full description on the dataset page: https://huggingface.co/datasets/Coraxor/lam_bench_env_v2.metrixel-character-renders
Metrixel Animated Character Renders
Multi-view renders, signed-distance-field volumes and per-view mesh tensors produced by Metrixel from a small set of rigged, animated humanoid characters.
Each character is captured from four camera angles (0°, 90°, 180°, 270°) across ~30 sampled frames of its motion clip, at 512×512. Every frame/angle carries a matching 64³ signed-distance-field volume and a per-view mesh tensor, so the geometry, the implicit surface and the image are aligned… See the full description on the dataset page: https://huggingface.co/datasets/EntVista/metrixel-character-renders.laion2b-en-aesthetic-square-human
Overview
This dataset is a HAND-CURATED version of our laion2b-en-aesthetic-square-cleaned dataset.
It has at least "a man" or "a woman" in it.
Additionally, there are no visible watermarks, the quality and focus are good, and it should not be confusing for AI training
There should be a little over 8k images here.
Details
It consists of an initial extraction of all images that had "a man" or "a woman" in the moondream caption.
I then filtered out all "statue" or… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-en-aesthetic-square-human.vibeworlding-environments-10
VibeWorlding 场景准备:10 个场景
独立场景源码与验证记录,方便协作下载。 不是作者 V2 原生训练集,也不是统一通过完整物理验证的数据集。场景未与新增资产库提前绑定。
下载和打开
下载完整环境包(67.5 MB)
解压后在包根目录运行 python3 -m http.server 8766 --bind 127.0.0.1,打开 http://127.0.0.1:8766/batch-001-ten/gallery.html 看场景,或 http://127.0.0.1:8766/v2-physics-pilot/index.html 看物理测试回放。无需部署 AI 模型;HF 数据集页面供下载,HTML 预览需上述本地静态服务。
现有验证状态
10/10 通过静态准备检查;8/10 通用 Three.js JSON 重载与单视角画面对照通过。10 个场景都实际运行了 Rapier 20 秒、240 Hz 刚体仿真,5/10 满足本轮全部保守判据。保存了全部失败和逐物体轨迹。… See the full description on the dataset page: https://huggingface.co/datasets/anon123312/vibeworlding-environments-10.dataset-viber-image-generation-preference-inference-endpoints-battle-flux
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/davidberenstein1957/dataset-viber-image-generation-preference-inference-endpoints-battle-flux.PubMedVision-EnKo
Informations
This is the Korean translation of FreedomIntelligence/PubMedVision. The translation was primarily generated using the 'solar-pro-241126' model, with occasional manual assistance from the 'Gemini 2.0 Flash Experimental' model and the 'Gemini experimental 1206' model.
An evaluation of the translation quality ("llm-as-a-judge") will be coming soon.
News
[2024/07/01]: We add annotations for 'body_part' and 'modality' of images, utilizing the… See the full description on the dataset page: https://huggingface.co/datasets/ChuGyouk/PubMedVision-EnKo.THE-ENDlaion2b-en-aesthetic-square-cleaned
Overview
A subset of our opendiffusionai/laion2b-en-aesthetic-square, which is itself a subset of the widely known "Laion2b-en-aesthetic" dataset.
However, the original had only the website alt-tags for captions.
I have added decent AI captioning, via the "moondream2" model.
Additionally, I have stripped out a bunch of watermarked junk, and weeded out 20k duplicate images.
When I use it for training, I will be trimming out additional things like non-realistic images, painting… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-en-aesthetic-square-cleaned.pixelprose_hu_150MMC4-130k-image-englishMMC4-130k是对MMC4中,抽样了130k左右 simliarty较高的图文pair得到的数据集
我们准备陆续翻译这个子集
我们会陆续将更多数据集发布到hf,包括
Coco Caption的中文翻译
CoQA的中文翻译
CNewSum的Embedding数据
增广的开放QA数据
WizardLM的中文翻译
如果你也在做这些数据集的筹备,欢迎来联系我们,避免重复花钱。
骆驼(Luotuo): 开源中文大语言模型
https://github.com/LC1332/Luotuo-Chinese-LLM
骆驼(Luotuo)项目是由冷子昂 @ 商汤科技, 陈启源 @ 华中师范大学 以及 李鲁鲁 @ 商汤科技 发起的中文大语言模型开源项目,包含了一系列语言模型。
( 注意: 陈启源 正在寻找2024推免导师,欢迎联系 )
骆驼项目不是商汤科技的官方产品。
Citation
Please cite the repo if you use the data or code in this repo.
@misc{alpaca… See the full description on the dataset page: https://huggingface.co/datasets/silk-road/MMC4-130k-image-english.Wikipedia-Malaysian-Politicians
Summary
wikipedia page : https://en.wikipedia.org/wiki/Category:Malaysian_politicians
Number of Politicians : 110
Null Images of Politicians : 16
link to dataset : https://huggingface.co/datasets/Englios/Wikipedia-Malaysian-Politicians
date of creation: 2024-20-01
laion2B-en_continentsDataComp-1B_hutest8k_ensembletest8k_ensemble2
Dataset Card for fisheye8k_ensemble_mapped
This is a FiftyOne dataset with 8000 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Abeyankar/test8k_ensemble2")
# Launch the App
session = fo.launch_app(dataset)
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Abeyankar/test8k_ensemble2.laion2b-en-aesthetic-square
Contents
This is a pretty raw filter of https://huggingface.co/datasets/laion/laion2B-en-aesthetic
I just filtered for "is image perfectly square, AND is image at least 1024x1024 pixels"
Approximate image count is a bit over 300k
Updated 2025/01/24
I just found out there are a bunch of watermarked sites in here.
So much for aesthetically chosen :(
So I filtered out a bunch of the "stock image" sites, just by looking at url strings.
Update 2025/01/31… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-en-aesthetic-square.OIO
