Skip to content

RitAreaSciencePark/token_geometry

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

50 Commits
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

The Intrinsic Dimension of Prompts in Internal Representations of Large Language Models

Source code for the paper: 'The Intrinsic Dimension of Prompts in Internal Representations of Large Language Models'

Description

In this project, we analyze the geometry of tokens in the hidden layers of large language models using intrinsic dimension (ID). Given a prompt, this is done by extracting the internal representations of the tokens. We use the hidden states variable from the Transformers library on Hugging Face. Make sure you are logged into HuggingFace Hub and have access to the models. We calculate the intrinsic dimension for each hidden layer, we consider the point cloud formed by the token representation at that layer. On this point cloud, we calculate the intrinsic dimension (ID). Specifically, we use the Generalized Ratio Intrinsic Dimension Estimator (GRIDE) to estimate the intrinsic dimension implemented using the DADApy library. This is done by running the following example script

      python src/extract_id.py --input_dir results --model_name meta-llama/Meta-Llama-3-8B  --method structured

The above script generates the intrinsic dimension curve for 2244 (unshuffled) prompts for Llama-3-8B model. The ID curves are then stored in results/Pile-Structured/meta-llama/Meta-Llama-3-8B/gride.npy. To run the code for the shuffled experiment, set method to shuffled.

Dataset

Currently we use the prompts from Pile-10K. We filter only prompts with atleast 1024 tokens according to the tokenization schemes of all the above models. This results in 2244 prompts after filtering. The indices of the filtered prompts is stored in filtered_indices.npy. We choose 50 random prompts for the shuffle experiment in Figure 1 and their indices are stored in subset_indices.npy

Reproducibility

The script to reproduce the plots in the paper is given in the post_process folder. All experiments were run on an NVIDIA A100 GPU with 64 GB memory.

References

About

Geometry of Tokens in Internal Representations of Large Language Models

Resources

Stars

Watchers

Forks

Releases

Packages

Contributors

Languages