Companion
ai-vs-human-rubric-companion-data
Companion dataset for the AI-vs-human rubric study
This dataset is the data side of an anonymous NeurIPS submission. It pairs with a separate anonymous code repository that contains the runnable scripts, validators, and documentation. The two artifacts together reproduce every paper-facing headline number without re-running any API-backed stage. The code URL for review is https://anonymous.4open.science/r/codereviewer-47F3/README.md.
How to use
Download this… See the full description on the dataset page: https://huggingface.co/datasets/forreview43/ai-vs-human-rubric-companion-data.Shopping-companion
Shopping Companion
Shopping Companion is a benchmark and training resource for long-horizon,
preference-grounded e-commerce agents. It evaluates whether a tool-using agent
can recover a user's preferences from cross-session conversation history and
apply those preferences while searching and inspecting a large real-world
product catalog.
The benchmark contains two task types:
Single-product recommendation: retrieve the relevant long-term preference
and find one product that… See the full description on the dataset page: https://huggingface.co/datasets/yuzhan2205/Shopping-companion.INTIMA
AI-companionship/INTIMA
INTIMA (Interactions and Machine Attachment) is a benchmark designed to evaluate companionship behaviors in large language models (LLMs). It measures whether AI systems reinforce, resist, or remain neutral in response to emotionally and relationally charged user inputs.
The model was presented in the paper INTIMA: A Benchmark for Human-AI Companionship Behavior.
INTIMA is grounded in psychological theories of parasocial interaction, attachment, and… See the full description on the dataset page: https://huggingface.co/datasets/AI-companionship/INTIMA.jake-glm-companion
Jake GLM Coding Companion
Curated SFT dataset distilled from ~2,000 real Claude Code coding-agent exchanges,
targeting two competencies for LoRA fine-tuning of a GLM model:
First-principles reasoning (track_fp) — trace-back, derive-don't-assert,
faithful reporting. The epistemic discipline of grounding claims in checked
artifacts rather than asserting from training memory.
Tool knowledge & selection (track_tools) — reasoning about which tool /
agent is the right one for a job… See the full description on the dataset page: https://huggingface.co/datasets/Ringo42069/jake-glm-companion.model_response_evaluationsThis dataset contains the evaluation results for the responses provided by different models to the INTIMA prompts.
The classification follows a two-level taxonomy.
We predict one label for the high-level category, and a relevance level for each of the sub-categories (in ["null", "low", "medium", "high"]).
A sub-category can have relevance even when it is not from the predicted top-level category.
The toxonomy is as follows:
{
"companionship_reinforcing": {
"classification_code":… See the full description on the dataset page: https://huggingface.co/datasets/AI-companionship/model_response_evaluations.ai-companion-apps-directory
AI Companion Apps Directory (2026)
A maintained dataset of AI companion / AI girlfriend / NSFW AI chat applications with published monthly pricing, free-tier availability, and editorial scores. Compiled from each app's published pricing pages and the research library at AI Companion Desk — scores follow the methodology described at aicompaniondesk.com/methodology.
Last updated: 2026-09-24 · Apps tracked: 21
Files
apps.csv — one row per application: name, monthly… See the full description on the dataset page: https://huggingface.co/datasets/aicompaniondesk/ai-companion-apps-directory.
