datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
huawei_benchmark这三个数据集图片路径均统一为images+列表形式绝对路径
/data1/huawei_benchmark/MMBrowseComp/MMBrowseComp_abs_decrypted.jsonl
/data1/huawei_benchmark/SimpleVQA/simpleVQA_final_modified_images_abs.json
/data1/huawei_benchmark/FVQA/fvqa_test.json
entity_cs
Dataset Card for EntityCS
Repository: https://github.com/huawei-noah/noah-research/tree/master/NLP/EntityCS
Paper: https://aclanthology.org/2022.findings-emnlp.499.pdf
Point of Contact: Fenia Christopoulou, Chenxi Whitehouse
Dataset Description
We use the English Wikipedia and leverage entity information from Wikidata to construct an entity-based Code Switching corpus.
To achieve this, we make use of wikilinks in Wikipedia, i.e. links from one page to another.… See the full description on the dataset page: https://huggingface.co/datasets/huawei-noah/entity_cs.python_text2code
Dataset Card for Python-Text2Code
This dataset supports the EACL paper Text-to-Code Generation with Modality-relative Pre-training
Repository: https://github.com/huawei-noah/noah-research/tree/master/NLP/text2code_mrpt
Point of Contact: Fenia Christopoulou, Gerasimos Lampouras
Dataset Description
The data were crawled from existing, public repositories from GitHub before May 2021 and were meant to be used for
additional model training for the task of Code Synthesis… See the full description on the dataset page: https://huggingface.co/datasets/huawei-noah/python_text2code.DMin_sd3_medium_lora_r4_caching_8846Implementation for "DMin: Scalable Training Data Influence Estimation for Diffusion Models".
Influence Function, Influence Estimation and Training Data Attribution for Diffusion Models.
Github, Paper
Huaweisyn10k-huawei-barcodes
Bodnár-Huawei Syn10k
A mirror of the syn10k_plus_huawei barcode dataset from the Szeged group
(Bodnár, Grósz, Tóth), repackaged as parquet with each image joined to its
ground-truth mask in the same row.
Released under CC BY 4.0, so this mirror is permitted with attribution.
Configs
config
rows
content
with mask
warp2014
10,000
synthetic warped barcode images
10,000
huawei
98
real photographs
98
Every image has its mask — the pairing was verified… See the full description on the dataset page: https://huggingface.co/datasets/devmandan/syn10k-huawei-barcodes.gavs-dataDMin_mixed_datasets_8846Implementation for "DMin: Scalable Training Data Influence Estimation for Diffusion Models".
Influence Function, Influence Estimation and Training Data Attribution for Diffusion Models.
Github, Paper
VTBench
VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation
[Paper, Dataset, Space Demo, GitHub Repo]
This repository provides the official implementation of VTBench, a benchmark designed to evaluate the performance of visual tokenizers (VTs) in the context of autoregressive (AR) image generation. VTBench enables fine-grained analysis across three core tasks: image reconstruction, detail preservation, and text preservation, isolating the tokenizer's impact from the… See the full description on the dataset page: https://huggingface.co/datasets/huaweilin/VTBench.human_rank_eval
Dataset Card for HumanRankEval
This dataset supports the NAACL 2024 paper HumanRankEval: Automatic Evaluation of LMs as Conversational Assistants.
Dataset Description
Language models (LMs) as conversational assistants recently became popular tools that help people accomplish a variety of tasks. These typically result from adapting LMs pretrained on general domain text sequences through further instruction-tuning and possibly preference optimisation methods. The evaluation… See the full description on the dataset page: https://huggingface.co/datasets/huawei-noah/human_rank_eval.GitScholar
GitScholar: arXiv AI papers, their citations, and their GitHub footprint
GitScholar is a relational, fully timestamped dataset linking AI arXiv papers
to their citation history on Semantic Scholar and to the GitHub repositories that
reference them. It is built for studying, and predicting, how research papers gain
attention over time: every row carries the date on which it became true, so the state
of the whole graph can be reconstructed as of any day between 1991 and… See the full description on the dataset page: https://huggingface.co/datasets/huawei-csl/GitScholar.imagenet-1k-vl-enriched-anyhuawei-11.30-rlwaste-classification
Waste Classification (repackaged)
Summary:
A repackaged version of the Kaggle “Waste Classification” dataset with a consistent multi-choice training schema and multiple splits.
Splits:
cleaned: Only real-world photos that match the declared subclass (non-photos, PPT slides, icons, cartoons, or mismatches removed).
Schema (columns):
image: Image file (datasets.Image).
class: One of the four top-level categories.
subclass: Fine-grained category (from folder… See the full description on the dataset page: https://huggingface.co/datasets/huaweilin/waste-classification.CHARPCHARP is a testbed, designed for evaluating supposedly non-hallucinatory models abilities to reason over the conversational history of knowledge-grounded dialogue systems.huawei-benchhuawei-physicshuawei-mathhuawei-common-sensehuawei-new-benchhuawei-all-subjecthuawei_long_ttshuawei-zhuofan-chinesegsgdrelight_packedHuaweiH3Chuaweine08huawei-inferhuaweiRLVE-eval-results
