datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lm-eval-results-PetroGPT-WestSeverus-7B-DPO-v2-private
Dataset Card for Evaluation run of PetroGPT/WestSeverus-7B-DPO-v2
Dataset automatically created during the evaluation run of model PetroGPT/WestSeverus-7B-DPO-v2
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-PetroGPT-WestSeverus-7B-DPO-v2-private.petrosafe-rag-corpus-fa
PetroSafe RAG Corpus (FA/EN)
Bilingual (Persian/English) knowledge corpus for process safety and HSE in oil, gas, and
petrochemical operations. Built for alirezaaminzadeh/petrosafe-rag-fa, the retrieval architecture
is inherited unchanged from hse-multimodal-rag-corpus
(hybrid BM25 + word/char TF-IDF, mandatory citations, abstention) — this repo supplies new
domain content, not a new retrieval method.
Data honesty (please read before citing any number from this… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/petrosafe-rag-corpus-fa.lm-eval-results-PetroGPT-WestSeverus-7B-DPO-private
Dataset Card for Evaluation run of PetroGPT/WestSeverus-7B-DPO
Dataset automatically created during the evaluation run of model PetroGPT/WestSeverus-7B-DPO
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-PetroGPT-WestSeverus-7B-DPO-private.petri-bench
petri-bench: 699 causal-discovery episodes from LLM agents and classical baselines
Every episode is one attempt to find a hidden causal parameter in a procedurally
generated simulation, under a fixed experiment budget. Nine frontier LLM agents and
four classical experimental-design algorithms ran the same 30 tasks across five
simulation engines.
Unlike answer-only evaluations, each episode is also audited for scientific method
quality: whether the submitted conclusion was backed… See the full description on the dataset page: https://huggingface.co/datasets/barissozudogru/petri-bench.voyage-ai-vet-evals
Voyage AI Vet Evals
Two open, reproducible system benchmarks for veterinary AI: one tests whether
clinical claims are backed by authoritative evidence; the other tests whether a
product preserves the right pet's state through changes and interference.
The full methods, dependency-free scorers, schemas, and release-integrity tests
are maintained in the
GitHub repository.
Published findings
VetEvidenceBench 2.5
Across 50 paired veterinary cases… See the full description on the dataset page: https://huggingface.co/datasets/voyage-pet-ai/voyage-ai-vet-evals.welna-auditpetImagespersonal-query-pet-supplies
Personal Query: Pet Supplies
This dataset contains personalized product search queries for the Pet_Supplies category.
Each record is built from the Personal Query pipeline:
Stage 6 generated correct personalized queries.
Stage 7 injected user-specific error query variants when a matching error pattern was available.
Stage 5 provided the user profile complexity level.
Files
data.jsonl: all correct Stage 6 queries. Rows without Stage 7 error query keep error_query as… See the full description on the dataset page: https://huggingface.co/datasets/xxxxdszz/personal-query-pet-supplies.oasst2_thgithub-issuesannotations_creators:
no-annotation
language_creators:
found
languages:
en
licenses:
unknown
multilinguality:
monolingual
pretty_name: Practice
size_categories:
unknown
source_datasets:
original
task_categories:
text-classification
text-retrieval
task_ids:
multi-class-classification
multi-label-classification
document-retrieval
Dataset Card for [Needs More Information]
Dataset Summary
For Practice
Supported Tasks and Leaderboards
Classification… See the full description on the dataset page: https://huggingface.co/datasets/peterhsu/github-issues.
