datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AEOLLMThe repository maintains the datasets for the NTCIR-18 Automatic Evaluation of LLMs (AEOLLM) Task and the NTCIR-19 Automatic Evaluation of LLMs (AEOLLM) 2 Task.
The aeollm_1 configuration corresponds to the NTCIR-18 AEOLLM Task, and the aeollm_2 configuration corresponds to the NTCIR-19 AEOLLM 2 Task.
For AEOLLM2, the document corresponding to each answerId is available in the following Google Drive folder: https://drive.google.com/drive/folders/1ujR5Gj889Y8RbK2eBmA-fikBQ1qcjXDe?usp=sharing.… See the full description on the dataset page: https://huggingface.co/datasets/THUIR/AEOLLM.aeon
Aeon QA Dataset
The main training synthetic conversional dataset for Aeon persona AI.
The data was generated by human questions and complemented by Gemini, Deepseek, Qwen, ChatGPT.
This Dataset is being created to finetune LLM/SLM's with general information about books, movies/tv and topics.
General info:
General chat
Persona validation
Brazilian culture
Basic portuguese
General philosophy
World culture
Geopolitics
Contemporary Art
Basic economics and criptocurrencies
Pop culture… See the full description on the dataset page: https://huggingface.co/datasets/gustavokuklinski/aeon.Aeollm
