datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
harvey-labs-llm-artifact-analysis
Harvey Labs LLM artifact analysis
This dataset contains artifacts from a non-LLM analysis of the Harvey Labs DOCX corpus.
The analysis used filename similarity, document extraction heuristics, a small manually
labeled seed set, CatBoost native text features, and a native CatBoost embedding feature
built from a mean Word2Vec representation. It was designed to find documents where an
LLM refused the requested task and returned a safe alternative instead.
Source… See the full description on the dataset page: https://huggingface.co/datasets/Hanno-Labs/harvey-labs-llm-artifact-analysis.llm-dataset-with-analysis
Contact Information
Dataset Maintainer
Name: Othmane Moutaouakkil
Email: othmoutaouakkil@gmail.com
GitHub: @moutaouakkil
LinkedIn: Othmane Moutaouakkil
How to Reach Me
Feel free to contact me with any questions about this dataset, including:
Usage inquiries
Bug reports
Collaboration opportunities
Feature requests
I typically respond within 1-2 business days.
llm-evaluation-analysisllm-evaluation-analysis-split
