efficiency
CVPR_workshop_efficiencyVLMnewsmner-data-efficiencyenzyme_catalytic_efficiency
Dataset Card for Enzyme Catalytic Efficiency Dataset
Dataset Summary
This task is focused on predicting $k_cat$ values, which are enzymatic turnover numbers denoting the maximum chemical conversion rate of a reaction, for metabolic enzymes originating from any organism. These predictions are based on substrate structures and protein sequences. The underlying importance of this task lies in its potential to yield high-throughput and accurate $k_cat$ predictions applicable… See the full description on the dataset page: https://huggingface.co/datasets/biomap-research/enzyme_catalytic_efficiency.REFUEL_it2_mask2_data_efficiencyperovskite-solar-cell-efficiency-autoresearch
🔬 Perovskite Solar Cell Text Corpus for Karpathy's autoresearch
A 98.9 MB text corpus of perovskite solar cell scientific literature formatted for direct use with karpathy/autoresearch — the autonomous LLM-driven hyperparameter search framework that trains a GPT from scratch and has an AI agent iteratively modify train.py to minimize val_bpb (bits per byte).
📊 Dataset Stats
Metric
Value
Total documents
19,730
Total text
98.9 MB (~103M characters)… See the full description on the dataset page: https://huggingface.co/datasets/CollinL/perovskite-solar-cell-efficiency-autoresearch.tool-call-efficiency
tool-call-efficiency
Teach an agent to make every tool call count.
An agent that calls a tool twice with the same arguments, looks up what
the user just told it, or keeps calling after the task is done is slow,
expensive, and harder to trust. Ask a base Qwen3-4B to work through
1,133 tool-using tasks across six agents and it does this a lot:
only 52% of its 6,681 rollouts finish the
task with no wasted or skipped call, and 6%
repeat a call that had already succeeded. This… See the full description on the dataset page: https://huggingface.co/datasets/while-ai/tool-call-efficiency.
