researchclawbench
ResearchClawBench
ResearchClawBench
Evaluating AI Agents for Automated Research from Re-Discovery to New-Discovery
Quick Start | Submit Tasks | How It Works | Domains | Leaderboard | Add Your Agent
ResearchClawBench is a benchmark that measures whether AI coding agents can independently conduct scientific research — from reading raw data to producing publication-quality reports — and then rigorously evaluates the results against real human-authored papers.… See the full description on the dataset page: https://huggingface.co/datasets/InternScience/ResearchClawBench.Luria-ResearchClawBench
Luria
Luria is an autonomous scientific research agent designed to carry research tasks from literature review to experimentation and final reporting.
This repository contains Luria's submission to ResearchClawBench, a benchmark for evaluating agents on scientific investigation and reproduction tasks.
Luria runs as a single-entry research organism with persistent roles for theory, hypothesis, and experiment, which coordinate specialized domain workers throughout a research… See the full description on the dataset page: https://huggingface.co/datasets/omicverse/Luria-ResearchClawBench.ResearchClawBench
ResearchClawBench
Evaluating AI Agents for Automated Research from Re-Discovery to New-Discovery
Quick Start | Submit Tasks | How It Works | Domains | Leaderboard | Add Your Agent
ResearchClawBench is a benchmark that measures whether AI coding agents can independently conduct scientific research — from reading raw data to producing publication-quality reports — and then rigorously evaluates the results against real human-authored papers.… See the full description on the dataset page: https://huggingface.co/datasets/Ka3de/ResearchClawBench.
