exploration
dim-discovery-archive
Geometry of Decision Making in Language Models
Abhinav Joshi · Divyanshu Bhatt · Ashutosh ModiNeurIPS 2025
This repository contains the official implementation/release for the NeurIPS 2025 paper Geometry of Decision Making in Language Models.
We study the internal decision-making processes of large language models through the lens of intrinsic dimension (ID), analyzing how hidden representations evolve across layers in a multiple-choice… See the full description on the dataset page: https://huggingface.co/datasets/Exploration-Lab/dim-discovery-archive.CS781
CS781-Fall 2026 IIT-Kanpur
Instructor: Dr. Ashutosh Modi
Assignment-1
There are 5 main tasks in Assignment-1:
Explore Zipf's Law for multilingual data.
N-gram language modeling.
Neural Network Implementation from Scratch for Word2Vec for English-Hindi mixed corpora.
Neural Network Language Model from scratch for English-Hindi mixed corpora.
Byte-Pair encoding for subword tokenization and training a NLM model on the same.
The data can be fetched using the… See the full description on the dataset page: https://huggingface.co/datasets/Exploration-Lab/CS781.iSign
iSign: A Benchmark for Indian Sign Language Processing
The iSign dataset serves as a benchmark for Indian Sign Language Processing. The dataset comprises of NLP-specific tasks (including SignVideo2Text, SignPose2Text, Text2Pose, Word Prediction, and Sign Semantics). The dataset is free for research use but not for commercial purposes.
Quick Links
Website: The landing page for iSign
arXiv Paper: Detailed information about the iSign Benchmark.
Dataset on Hugging Face:… See the full description on the dataset page: https://huggingface.co/datasets/Exploration-Lab/iSign.ablation_exploration_in_rl
Reinforcement Learning Improves Agentic Software Engineering
An ablation study of reinforcement-learning (RL) fine-tuning for agentic software-engineering (SWE) models. Starting from an 8B SFT model, we fine-tune with RL across ~20 configurations — varying the objective, loss normalization, sampling, and training dataset — and evaluate each on agentic SWE benchmarks.
Result
RL reliably and substantially improves agentic SWE performance, and the improvement is… See the full description on the dataset page: https://huggingface.co/datasets/penfever/ablation_exploration_in_rl.IL-TUR
Dataset Card for "IL-TUR"
Dataset Description
Summary
"IL-TUR": Benchmark for Indian Legal Text Understanding and Reasoning is a collaborative effort to establish a modern benchmark for training and evaluating AI/NLP models on Indian Law. IL-TUR consists of 8 foundational tasks, requiring different types of understanding and skills. Apart from English, some tasks involve Indic languages.
This dataset repository has been created to unify the data… See the full description on the dataset page: https://huggingface.co/datasets/Exploration-Lab/IL-TUR.Genex-DB-World-Exploration
GenEx-DB-World-Exploration 🎬🌍
This is the video version of the GenEx-DB dataset.
The dataset contains forward navigation path, captured by panoramic cameras.
Each path is 0.4m/frame, 50 frames in total.
Each example is a single .mp4 video reconstructed from the original frame folders.
📂 Splits
Split Name
Description
realistic
📸 Unreal 5 City Sample renders
low_texture
🏜️ Blender Low-texture synthetic scenes
anime
🌸 Unity Stylized/anime scenes… See the full description on the dataset page: https://huggingface.co/datasets/genex-world/Genex-DB-World-Exploration.
