datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PyInstruct PyBench: Evaluate LLM Agent on Real World Tasks
📃 Paper
•
🤗 Data (PyInstruct)
•
🤗 Model (PyLlama3)
•
Code
•
PyBench is a comprehensive benchmark evaluating LLM on real-world coding tasks including chart analysis, text analysis, image/ audio editing, complex math and software/website development. We collect files from Kaggle, arXiv, and other sources and automatically generate queries according to the type and content of each file.
Why PyBench?
The LLM Agent, equipped… See the full description on the dataset page: https://huggingface.co/datasets/Mercury7353/PyInstruct.Mercury-Thermometer-Damage-and-Crack-Detection-Image-Dataset
Mercury Thermometer Damage and Crack Detection Image Dataset
The current medical industry faces numerous challenges in equipment monitoring and safety checks, especially in the use of mercury thermometers, where damage and cracks can pose serious safety hazards. Existing detection solutions often rely on manual inspection, which is not only inefficient but also prone to missed detections and misjudgments. Therefore, constructing a high-quality mercury thermometer damage and crack… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/Mercury-Thermometer-Damage-and-Crack-Detection-Image-Dataset.beaver-dw-plan-sql
BEAVER-dw Plan→SQL
A restructuring of the dw subset of BEAVER
into a plan-then-SQL format, with family-disjoint splits.
Each example asks a model to emit a structured plan first and the SQL second:
{
"question": "Which departments offered the most subjects last term?",
"domain_knowledge": ["..."],
"ir": {
"tables": ["SIS_DEPARTMENT", "SUBJECT_OFFERED_SUMMARY"],
"join_keys": [{"left": "SIS_DEPARTMENT.DEPARTMENT_CODE",
"right":… See the full description on the dataset page: https://huggingface.co/datasets/mercurylabs-ai/beaver-dw-plan-sql.DreadPoor__Mercury_In_Retrograde-8b-Model-Stock-details
Dataset Card for Evaluation run of DreadPoor/Mercury_In_Retrograde-8b-Model-Stock
Dataset automatically created during the evaluation run of model DreadPoor/Mercury_In_Retrograde-8b-Model-Stock
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DreadPoor__Mercury_In_Retrograde-8b-Model-Stock-details.Mercury
Mercury Dataset
Mercury is a multilingual instruction-tuning dataset designed to enhance AI capabilities across three languages: English (EN), German (DE), and Persian (FA). The dataset focuses on improving performance in text summarization, general Q&A, and basic code generation tasks.
📊 Dataset Overview
· Total Examples: [200+]
· Languages: English, German, Persian
· Domains: Text Summarization, General Q&A, Basic Coding
· Fine-tuned Model: sinamsv0/WALL-E (1B… See the full description on the dataset page: https://huggingface.co/datasets/sinamsv0/Mercury.benchmark_synthentic
