1 min readfrom Machine Learning

I built a compiler that turns computation graphs into the weights of a vanilla transformer — no training anywhere [P]

I've been chasing the question of what algorithms a transformer can actually express -- separate from what it can learn. So I built a compiler: define a computation graph in ordinary Python, and it produces the weights of a transformer that executes the graph. The result is a standard Phi-3-architecture checkpoint that vanilla huggingface loads with no custom code and no trust_remote_code. Zero training in the pipeline.

Write-up (origin + how the constructions work): https://ood.dev/posts/torchwright-intro/

Repo (twelve runnable examples): https://github.com/physicsrob/torchwright

Hand-built transformer weights aren't a new idea. RASP defines a language whose primitives map onto transformer sublayers, and Tracr compiles RASP programs into actual weights. I wanted two things they don't aim for: expressing a computation graph in ordinary Python, and targeting a stock architecture, so the output loads in vanilla huggingface with no custom code.

submitted by /u/notforrob
[link] [comments]

Want to read more?

Check out the full article on the original site

View original article

Tagged with

#Transformer
#Computation Graph
#Compiler
#Phi-3
#Hugging Face
#Weights
#Training
#Python
#Architecture
#Torchwright
#RASP
#Tracr
#Algorithms
#Sublayers
#Vanilla
#Checkpoint
#Code
#Primitives
#Computation
#Machine Learning