LLM-analysis
A-Survey-for-LLM-Agent-Trajectory-Analysis
A Survey for LLM Agent Trajectory Analysis
This dataset repository hosts the survey paper A Survey for LLM Agent Trajectory Analysis: From Failure Attribution to Enhancement and a structured metadata snapshot of the companion paper collection from Awesome-LLM-Agent-Trajectory-Analysis.
The repository is intended for discovery, citation, and lightweight analysis of the literature around LLM agent trajectory analysis, including failure attribution, trajectory-based debugging… See the full description on the dataset page: https://huggingface.co/datasets/RobinChen2001/A-Survey-for-LLM-Agent-Trajectory-Analysis.harvey-labs-llm-artifact-analysis
Harvey Labs LLM artifact analysis
This dataset contains artifacts from a non-LLM analysis of the Harvey Labs DOCX corpus.
The analysis used filename similarity, document extraction heuristics, a small manually
labeled seed set, CatBoost native text features, and a native CatBoost embedding feature
built from a mean Word2Vec representation. It was designed to find documents where an
LLM refused the requested task and returned a safe alternative instead.
Source… See the full description on the dataset page: https://huggingface.co/datasets/Hanno-Labs/harvey-labs-llm-artifact-analysis.llm-ideology-analysisThis dataset contains evaluations of political figures by a diverse set of Large Language Models (LLMs), such that the ideology of these LLMs can be characterized.
📝 Dataset Description
The dataset contains responses from 19 different Large Language Models evaluating 3,991 political figures, with responses collected in the six UN languages: Arabic, Chinese, English, French, Russian, and Spanish.
The evaluations were conducted using a two-stage prompting strategy to assess the… See the full description on the dataset page: https://huggingface.co/datasets/aida-ugent/llm-ideology-analysis.analysis-oracle-verifierllm-ideology-analysis
Dataset Card for LLM Ideology Dataset
This dataset contains evaluations of political figures by various Large Language Models (LLMs), designed to analyze ideological biases in AI language models.
Dataset Details
Dataset Description
The dataset contains responses from 17 different Large Language Models evaluating 4,339 political figures, with responses collected in both English and Chinese. The evaluations were conducted using a two-stage prompting strategy to… See the full description on the dataset page: https://huggingface.co/datasets/ajrogier/llm-ideology-analysis.llm-error-analysis
LLM Error Analysis Dataset
This dataset contains evaluation examples for testing mathematical and logical reasoning capabilities of language models.
Dataset Overview
Number of examples: 11
Fields: input, type, expected_output, model_output
Types included: arithmetic, long_addition, date_knowledge, ambiguous_riddle, ordering, directional_reasoning, pattern_recognition, instruction_following, framing_bias, causal_vs_correlation
Model Tested
Model:… See the full description on the dataset page: https://huggingface.co/datasets/ParamTh/llm-error-analysis.
