generative-ai
champ_trainning_sample
Dataset samples for Champ trainning
This dataset samples is used for Champ.
Before trainning, you need to process the datasets by SMPL & DWPOSE methods. Refer to https://github.com/fudan-generative-vision/champ/blob/master/docs/data_process.md
hallo3_training_dataHallo3: Highly Dynamic and Realistic Portrait Image Animation with Diffusion Transformer Networks
Jiahao Cui1
Hui Li1
Yun Zhan1
Hanlin Shang1
Kaihui Cheng1
Yuqi Ma1
Shan Mu1
Hang Zhou2
Jingdong Wang2
Siyu Zhu1✉️
1Fudan University 2Baidu Inc
I. Dataset Overview
This dataset serves as the training data for the open - source Hallo3 model, specifically created for the training of video… See the full description on the dataset page: https://huggingface.co/datasets/fudan-generative-ai/hallo3_training_data.fineweb-edu-ar
FineWeb-Edu-Ar
FineWeb-Edu-Ar is a machine-translated Arabic version of the FineWeb-Edu dataset designed to support the development of Arabic small language models (SLMs).
Dataset Details:
Languages: Arabic, English (paired)
Size: 202 billion tokens
License: CC-BY-NC-4.0
Source: Machine-translated from the deduplicated version of Hugging Face’s FineWeb-Edu dataset
Translation model: facebook/nllb-200-distilled-600M
Application:
FineWeb-Edu-Ar is suitable for pre-training… See the full description on the dataset page: https://huggingface.co/datasets/kaust-generative-ai/fineweb-edu-ar.champ_motions_example
Example data for Champ inference
Links
github: https://github.com/fudan-generative-vision/champ
models: https://huggingface.co/fudan-generative-ai/champ
telco-gaia
Telco-GAIA
A GAIA-style benchmark for AI agents operating over a real telecom operator's
website snapshot plus a synthetic customer database. 100 tasks across 7
categories: Pricing, Miscellaneous, Images, Web Archives, PDF, PDF Visual,
Database.
Agents read questions.json + environment.md, browse the local website
(:8080) and query the database API (:8081), and produce a GAIA-compatible
submission.json.
What's here
File
What… See the full description on the dataset page: https://huggingface.co/datasets/kaust-generative-ai/telco-gaia.maestro-mas-benchmark
MAESTRO MAS Benchmark Dataset
maestro-mas-benchmark is a dataset derived from MAESTRO, a framework-agnostic evaluation suite for LLM-based
multi-agent systems (MAS). It provides a systems-level view of MAS behavior and is designed to benchmark,
observe, and analyze MAS performance and behavior across diverse scenarios.
For more details about MAESTRO, visit the GitHub repository.
Dataset details
The dataset currently includes data for 12 different MAS systems… See the full description on the dataset page: https://huggingface.co/datasets/kaust-generative-ai/maestro-mas-benchmark.
