baidu
Datasets
All datasets matching “baidu”OmegaUse-OfficeVal
OmegaUse-OfficeVal
Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding
OmegaUse-OfficeVal is a benchmark for evaluating LLM agents on long-horizon,
real-world office-suite tasks that span word-processing documents, spreadsheets,
presentations, and cross-file productivity workflows. Tasks are derived from
authentic office requests proposed by practitioners and drawn from freelance
platforms, grounding the benchmark in real economic demand. Each task… See the full description on the dataset page: https://huggingface.co/datasets/baidu-frontier-research/OmegaUse-OfficeVal.ERIA-1K-Benchmark
ERIA-1K: ERNIE-Image-Aes-1K, A Deployment-Oriented Image Aesthetics Benchmark with Realistic Distribution
[🤗 Dataset] [🧩 ERNIE-Image-Aes Model]
🔍 Overview
Existing aesthetic benchmarks are predominantly constructed from curated, high-production-value image collections, often sourced from platforms such as Flickr and DPChallenge. These datasets tend to skew toward professional or semi-professional photography communities, Western photographic traditions, and visually… See the full description on the dataset page: https://huggingface.co/datasets/baidu/ERIA-1K-Benchmark.baidu-ultr_baidu-mlm-ctrQuery-document vectors and clicks for a subset of the Baidu Unbiased Learning to Rank
dataset: https://arxiv.org/abs/2207.03051
This dataset uses the BERT cross-encoder with 12 layers from Baidu released
in the official starter-kit to compute query-document vectors (768 dims):
https://github.com/ChuXiaokai/baidu_ultr_dataset/
We link the model checkpoint also under `model/`.baidu-baike-dataset
Baidu Baike Dataset
This dataset contains 5,634,898 entries extracted from Baidu Baike (百度百科), which is one of the largest Chinese online encyclopedias. This is a mirror of the original dataset from https://github.com/BIT-ENGD/baidu_baike, with data originally crawled around 2019.
Format
Each entry in the JSON format includes the following fields:
title: The title of the Baidu Baike entry
summary: A brief summary of the entry content
sections: A list of sections, where… See the full description on the dataset page: https://huggingface.co/datasets/lars1234/baidu-baike-dataset.baidu-ultr_uva-mlm-ctrQuery-document vectors and clicks for a subset of the Baidu Unbiased Learning to Rank
dataset: https://arxiv.org/abs/2207.03051
This dataset uses a Jax-based BERT cross-encoder with 12 layers pre-trained for 2 million steps
on the Baidu ULTR dataset to create query-document embeddings (768 dims).
We link the model checkpoint also under `model/`.baidu-ultr_tencent-mlm-ctrQuery-document vectors and clicks for a subset of the Baidu Unbiased Learning to Rank
dataset: https://arxiv.org/abs/2207.03051
This dataset uses the pretrained BERT cross-encoder (Bert_Layer12_Head12) from Tencent published
as part of the WSDM cup 2023 to compute query-document vectors (768 dims):
https://github.com/lixsh6/Tencent_wsdm_cup2023/tree/main/pytorch_unbias
We link the model checkpoint also under `model/`.
