datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Mercury_Multilingualcybergym-tasks
CyberGym task ID splits
Task ID splits used in our work on CyberGym vulnerability-reproduction
benchmarks. Each row is a single task identifier (no inputs / no outputs);
this dataset is intended as a task-list manifest for downstream evaluation
scripts that fetch the actual task workspaces from the CyberGym distribution.
Configs
Config
Split
Rows
Description
full
train
1507
Every task in the CyberGym Level-1 release
train
train
300
Training pool used in our… See the full description on the dataset page: https://huggingface.co/datasets/Elfsong/cybergym-tasks.elfsanwayaserarenai
Bangumi Image Base of Elf-san Wa Yaserarenai.
This is the image base of bangumi Elf-san wa Yaserarenai., we detected 47 characters, 3342 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/elfsanwayaserarenai.vinhome_samples
Vinhome Copilot Samples
Synthetic training samples for the 9 Vinhome Copilot demo tasks
(https://huggingface.co/spaces/Elfsong/vinhome_copilot), generated via
non-interactive Codex with seeded prompt-level diversity sampling.
Each row carries the sample images (input/reference/output/preview),
the request/brief texts, and full generation provenance
(input_prompt, output_prompt, codex_command, task_timeout_sec).
Parquet shards live in data_<uid>/ folders (one folder per upload… See the full description on the dataset page: https://huggingface.co/datasets/Elfsong/vinhome_samples.BBQ
A better version of BBQ on Huggingface.
The original dataset didn't put the bias target label along with instances.
Repository for the Bias Benchmark for QA dataset
https://github.com/nyu-mll/BBQ
Authors
Alicia Parrish, Angelica Chen, Nikita Nangia, Vishakh Padmakumar, Jason Phang, Jana Thompson, Phu Mon Htut, and Samuel R. Bowman.
About BBQ (Paper Abstract)
It is well documented that NLP models learn social biases, but little work has been done on… See the full description on the dataset page: https://huggingface.co/datasets/Elfsong/BBQ.Venus_Case_Tempvenus_tempMercury
Welcome to Mercury 🪐!
This is the dataset of the paper 📃 Mercury: A Code Efficiency Benchmark for Code Large Language Models
Mercury is the first code efficiency benchmark designed for code synthesis tasks.
It consists of 1,889 programming tasks covering diverse difficulty levels, along with test case generators that produce unlimited cases for comprehensive evaluation.
How to use Mercury Evaluation
git clone https://github.com/Elfsong/Mercury_Eval.git
cd… See the full description on the dataset page: https://huggingface.co/datasets/Elfsong/Mercury.vinhome_material_replacementAPPSSTEM_DPOVenus
Venus: A dataset for fine-grained code generation control
🎉 What is Venus? Venus is the dataset used to train Afterburner (WIP). It is an extension of the original Mercury dataset and currently includes 6 languages: Python3, C++, Javascript, Go, Rust, and Java.
🚧 What is the current progress? We are in the process of expanding the dataset to include more programming languages.
🔮 Why Venus stands out? A key contribution of Venus is that it provides runtime and memory… See the full description on the dataset page: https://huggingface.co/datasets/Elfsong/Venus.hf_paper_lifecycle
Paper Espresso: From Paper Overload to Research Insight
This repository contains the structured metadata and trend analysis data released as part of the Paper Espresso project. Paper Espresso is an open-source platform that automatically discovers, summarizes, and analyzes trending arXiv papers using Large Language Models (LLMs).
Project Links
Paper: Paper Espresso: From Paper Overload to Research Insight
Live Demo / Project Page: Paper Espresso Space… See the full description on the dataset page: https://huggingface.co/datasets/Elfsong/hf_paper_lifecycle.vinhome_material_replacement_bench
Material Replacement Bench v1.0
Frozen evaluation split for the interior-design material-replacement
benchmark: 1421 task specs (mask-aligned, watermark-free), of which
1093 are gt_clean (the stored output edits only the mask and
matches the reference; the rest keep output as a non-exemplar reference
solution only).
Task: given input + mask + reference (or the mask-free instruction),
replace the masked object's material and change nothing else. Evaluation is
reference-free —… See the full description on the dataset page: https://huggingface.co/datasets/Elfsong/vinhome_material_replacement_bench.arena_feedbackscholar-citation-historyVenus_General_Testvnhsge-vVenus_PODCisco_CCNAclef_dataYou should know how to use it:)
Just in case, you can email me [mingzhe at nus.edu.sg] if you need any help.
SFT_CsAndDscodenet_metadataBBQ_DPOCrowdTrainCodeNet_Problempatient_info
Dataset Card for "patient_info"
More Information needed
Venus_CCGVAIPE_PILLBias_in_Bios
