data-use
Datasets
All datasets matching “data-use”computer-use-data-psai
Computer Use Dataset - PSAI
A large-scale, multimodal dataset of human-computer interactions for training and evaluating AI agents.
🔗 Access Dataset: https://huggingface.co/datasets/anaisleila/computer-use-data-psai
📊 Dataset Overview
This dataset contains 3,167 completed tasks of human-computer interactions captured with video, screenshots, DOM snapshots, and detailed interaction events. Created by Paradigm Shift AI for advancing computer use AI agent research.… See the full description on the dataset page: https://huggingface.co/datasets/anaisleila/computer-use-data-psai.COD10K_GMPO_Used_Data
COD10K GMPO Used Data
This is a research repack of the COD10K camouflaged-object (CAM) subset used in
our CLIP DPO/GMPO experiments. It is not an official COD10K distribution.
The package contains the exact original images, masks, generated negative
images, and portable caption CSV files used for training and segmentation
evaluation. All paths in the portable CSV files are relative to this dataset
root.
Data Split
Split
Contents
Count
Train
Original CAM… See the full description on the dataset page: https://huggingface.co/datasets/KRMayD/COD10K_GMPO_Used_Data.data-use-sft-tiered
Data-use SFT — tiered workflow (two task subsets)
Multitask SFT anchored exclusively on mentions the tiered extractor emits
(T1 evidential ∪ T2 declaration; see
rafmacalaba/data-use-mentions-tiered). Every row carries task
("provenance" | "usage_impact") and origin (prwp | fcv). Rows whose anchor
span was judged T3 (non-mention) or junk are dropped — audit trail in
manifest.jsonl (provenance) and manifest_usage.jsonl (usage/impact).
task = provenance (22,201 rows)… See the full description on the dataset page: https://huggingface.co/datasets/rafmacalaba/data-use-sft-tiered.user-datadata-use-mentions
Data-use mentions (NER / span extraction)
Data mentions extracted from World Bank Policy Research Working Papers and FCV documents, validated by a
context-only LLM judge, and formatted for span-extraction (GLiNER / GLiNER2) and
token-classification (LFM2.5-encoder) fine-tuning.
Labels
Three entity types (the judge's specificity axis):
NAMED_DATA — a proper name, title, or acronym of a specific data source
DESCRIPTIVE_DATA — a source described in words but not… See the full description on the dataset page: https://huggingface.co/datasets/rafmacalaba/data-use-mentions.datause-displacement-reviewed
datause-displacement-reviewed
The Luna-reviewed subset of
rafmacalaba/datause-displacement:
only spans that received a v2.3 Luna verdict (band review + drop-side rescue,
source == luna_review). Every span carries the binary label plus
usage_type / drop_reason / specificity, and is traceable via key
(split:row:start:end) to the verdict records in
extraction_analysis/band_review/.
Configs
config
fields
gliner_reviewed
tokenized_text, corpus, origin… See the full description on the dataset page: https://huggingface.co/datasets/rafmacalaba/datause-displacement-reviewed.
