rita
Datasets
All datasets matching “rita”ita-eval-resultsThis dataset page contains numerical results to populate the ItaEval Leaderboard.
It is intended for internal use only.
Adding new Results
Once new results are computed, to include them in the leadboard it only takes to save them here in a results.json file. Note, though, that the file must contain several fields as specified in the leaderboard app. Use another result.json file for reference.
Adding info of a new Model
Since we do not support automatic submission and… See the full description on the dataset page: https://huggingface.co/datasets/RiTA-nlp/ita-eval-results.masscount-cf
MassCount-CF
Synthetic counterfactual counting corpus for identifying an additive
neural-mass cardinality coordinate in vision-language models (NMCA).
Companion dataset to "Counting Requires Mass: An Algebraic and Causal Account
of Numerosity in Vision-Language Models."
Corpus version: masscount-cf-1.0.0
Master scenes: 99,989
Delivered images (this upload): 326,623
Object instances: 28,617,018
Structure
Scene graphs are the durable artifact; pixels are regenerable… See the full description on the dataset page: https://huggingface.co/datasets/Ritabrata04/masscount-cf.icub_sim_dataset_t2_smooth_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "icub",
"total_episodes": 90,
"total_frames": 9935,
"total_tasks": 5,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:90"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ritamota/icub_sim_dataset_t2_smooth_lerobot.clinical-synthetic-text-kg
Data Description
We release the synthetic data generated using the method described in the paper Knowledge-Infused Prompting: Assessing and Advancing Clinical Text Data Generation with Large Language Models
(ACL 2024 Findings). The external knowledge we use is based on external knowledge graphs.
Generated Datasets
The original train/validation/test data, and the generated synthetic training data are listed as follows. For each dataset, we generate 5000 synthetic… See the full description on the dataset page: https://huggingface.co/datasets/ritaranx/clinical-synthetic-text-kg.icub_sim_dataset_t4_smooth_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "icub",
"total_episodes": 90,
"total_frames": 9216,
"total_tasks": 90,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:90"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ritamota/icub_sim_dataset_t4_smooth_lerobot.clinical-synthetic-text-llm
Data Description
We release the synthetic data generated using the method described in the paper Knowledge-Infused Prompting: Assessing and Advancing Clinical Text Data Generation with Large Language Models
(ACL 2024 Findings). The external knowledge we use is based on LLM-generated topics and writing styles.
Generated Datasets
The original train/validation/test data, and the generated synthetic training data are listed as follows. For each dataset, we generate 5000… See the full description on the dataset page: https://huggingface.co/datasets/ritaranx/clinical-synthetic-text-llm.
