ryanhoangt/OpenHands-evaluation
add missing report file
add CoAct v1.0 trajectories
Add CoAct v1.0 trajectory (#10)
add collapsable
update instruction
remove all the with hint result
cleanup metrics and fix repo
cleanup metrics and fix repo
add llama-3.1 result
rename file
rename OpenDevin to OpenHands
bump streamlit ver
fix typo
update app file
fix visualizer with latest streamlit feature
add 2nd run
add gpt-4o-mini result
Revert "add result from gpt-4o-mini"
add result from gpt-4o-mini
update the last missing instance
update result from pr2489
remove keys
revoke keys
add gpqa result
update v1.8 perf
add result for v1.8 no-hint gpt4o
fix model_name in updated metadat
add v1.8 result
update results using new ver of swebench
set n error/stuck/cost to 0 for CodeAct exp run below v1.5
by default not showing with hint result
add claude-3.5 result
support loading report with new format
update gitignore
update old result w/ swe-bench latest harness;
improved patch apply
improved patch apply
add report field
Add CodeAct 1.6 no hint
fix visualizer
feat: add gpqa results (#8)
fix visualizer to only display eval_report when it exists
add result for codeact 1.6
only show swe bench on visualizer
change test_result to bool
fix fine-grained report; support visualization while running
add gpt-4-1106 results for codeact swe
Merge commit 'edc3858a6ea5d0c7317b630024203af60e146b52'
update all swebench lite
Update outputs/miniwob/README.md
