datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
earnings_call_transcript_autograder
Earnings Call LLM Insights
📚 Read the Full Story: For a deep dive into the methodology, the wildest moments we found, and key takeaways, check out the blog post:KnowTrend.ai: Auto-Grading Ten Years of Earnings Calls for Prescience and Delusion
This dataset contains LLM-generated analysis of ~70,000+ earnings call transcripts.
The analysis was performed using Kimi k2-0905-preview, focusing on extracting specific insights, prescient analyst questions, and management missteps.… See the full description on the dataset page: https://huggingface.co/datasets/knowtrendllc/earnings_call_transcript_autograder.sciknoweval-v2-hard-autogradable-512-2026-04-28
SciKnowEval v2 Hard Autogradable 512 - 2026-04-28
A 512-example sanity subset sampled from hicai-zju/SciKnowEval (v2, test) for Plan-CRL scientific reasoning evals.
Selection seed: 20260428.
Filtering and balancing:
excludes L1
keeps L2, L3, L4
keeps autogradable types: mcq-4-choices, mcq-2-choices, true_or_false, filling
requires answerKey or answer
balances domains at 128 examples each: Biology, Chemistry, Material, Physics
per domain: 32 L2, 48 L3, 48 L4
Useful fields for… See the full description on the dataset page: https://huggingface.co/datasets/1337xyz1337xyz/sciknoweval-v2-hard-autogradable-512-2026-04-28.
