GPT-2 Small’s embedding geometry around “Trump”: discretized vs. continuous nearest neighbours [P]
![GPT-2 Small’s embedding geometry around “Trump”: discretized vs. continuous nearest neighbours [P]](/_next/image?url=https%3A%2F%2Fpreview.redd.it%2Ftlvz4c3i32eh1.png%3Fwidth%3D640%26crop%3Dsmart%26auto%3Dwebp%26s%3Daad6aeec9197e26debda00093dd47611e70c5a08&w=3840&q=75)
| This visualization looks at the token “Trump” in GPT-2 Small’s static embedding table, before attention or context is applied. The top plot is a t-SNE projection of 32,070 alphabetic tokens with at least two characters. The two graphs below compare Trump’s nearest neighbours under two representations of the same embedding: Discretized: each coordinate is thresholded before neighbours are calculated. This produces mostly generic political terms such as Mitt, Hillary, Pelosi, and Blair. Continuous: the original coordinates are retained. This produces a more specific group containing family members, staff, rivals, and presidents including Obama, Clinton, Bush, and Eisenhower. No prompting or text generation is involved; everything comes directly from GPT-2 Small’s learned token embeddings. [link] [comments] |
Want to read more?
Check out the full article on the original site