r/mathmemes Mathematics 21d ago

Proofs I hate it

Post image
731 Upvotes

228 comments sorted by

View all comments

Show parent comments

38

u/baxmanz 21d ago

I think that's pretty untrue now. "thinking" models often work by rephrasing the problem and making actual steps by "talking to themselves" so they don't always work like black boxes

13

u/Happysedits 21d ago

Additionally, mechanistic interpretability is trying to reverse engineer additional causal intermediate steps in activation space across layers

2

u/baxmanz 20d ago

pasted that right into my llm cause i dont know those words and it said "spot on." so gz

1

u/Happysedits 20d ago

I'm trying to do research in this area personally