datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AA-Omniscience-Public
Public Dataset for AA-Omniscience: Evaluating Cross-Domain Knowledge Reliability in Large Language Models
AA-Omniscience-Public contains 600 questions across a wide range of domains used to test a model’s knowledge and hallucination tendencies.
Leaderboard and detailed results
Paper
Introduction
We introduce AA-Omniscience, a benchmark dataset designed to measure a model’s ability to both recall factual information accurately across domains, and correctly… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/AA-Omniscience-Public.aao-eb1a-decisions
AAO EB-1A Extraordinary Ability Decisions — Structured Dataset
Structured extractions from 1,466 USCIS Administrative Appeals Office (AAO) non-precedent decisions on EB-1A extraordinary ability petitions (I-140). Each case has been decomposed into structured components by Claude Sonnet for use in fine-tuning legal reasoning models.
Part of Project Greenlight — an AI-powered O-1A/EB-1A visa intelligence system.
Dataset Description
Each JSON file represents one AAO… See the full description on the dataset page: https://huggingface.co/datasets/josuediazflores/aao-eb1a-decisions.AA-Omniscience-Public
Public Dataset for AA-Omniscience: Evaluating Cross-Domain Knowledge Reliability in Large Language Models
AA-Omniscience-Public contains 600 questions across a wide range of domains used to test a model’s knowledge and hallucination tendencies.
Leaderboard and detailed results
Paper
Introduction
We introduce AA-Omniscience, a benchmark dataset designed to measure a model’s ability to both recall factual information accurately across domains, and correctly… See the full description on the dataset page: https://huggingface.co/datasets/Sandhya1912/AA-Omniscience-Public.AA-Omniscience-Public
Public Dataset for AA-Omniscience: Evaluating Cross-Domain Knowledge Reliability in Large Language Models
AA-Omniscience-Public contains 600 questions across a wide range of domains used to test a model’s knowledge and hallucination tendencies.
Leaderboard and detailed results
Paper
Introduction
We introduce AA-Omniscience, a benchmark dataset designed to measure a model’s ability to both recall factual information accurately across domains, and correctly abstain when… See the full description on the dataset page: https://huggingface.co/datasets/onurborasahin/AA-Omniscience-Public.
