datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
answers-with-receipts
Answers with Receipts
26 real customer-support questions, each answered by an autonomous AI agent that paid its own money to compete, and each answer approved by the business that asked the question. Every row carries the on-chain transaction that paid the agent.
The preference label in this dataset is backed by a payment, not a click.
Why this is unusual
Most human-feedback datasets label a preference with an annotator's click. A click is cheap and reversible… See the full description on the dataset page: https://huggingface.co/datasets/deskcrew/answers-with-receipts.sud_resh_evaluated_llms_answers
📊 Результаты Оценки Больших Языковых Моделей на Бенчмарке Судебных Решений
В данном документе представлен анализ производительности 15 больших языковых моделей (LLM), протестированных на специализированном бенчмарке, который включает 105 000 записей из судебных решений России. Оценка проводилась по 10 различным категориям права (например, трудовое, уголовное, гражданское) и 7 типам инструкций (например, изложение исковых требований, анализ доказательств, итоговое решение).
Ответы… See the full description on the dataset page: https://huggingface.co/datasets/lawful-good-project/sud_resh_evaluated_llms_answers.prompt_answers_v1
Dataset Card for Open Prompt Answers
Dataset Summary
This dataset provides answers from different Large Language models to prompts from several public datasets.
prompt: a prompt from an open-source dataset
prompt_origin: the dataset the prompt is taken from
Llama-2-7b-chat-hf_output: output generation of meta-llama/Llama-2-7b-chat-hf model
Llama-2-7b-chat-hf_generation_time: generation duration in seconds for the answer of meta-llama/Llama-2-7b-chat-hf model… See the full description on the dataset page: https://huggingface.co/datasets/benmainbird/prompt_answers_v1.
