Manifund foxManifund
Home
Login
About
People
Categories
Newsletter
HomeAboutPeopleCategoriesLoginCreate
🍋
🍋
Ronish Bhatt

@ronishbhatt1

$0total balance
$0charity balance
$0cash balance

$0 in pending offers

Projects

[Mech Interp] How much info can a transformer's own write directions tell us?

pending admin approval

About RonishBy Ronish1
[Mech Interp] How much info can a transformer's own write directions tell us?
🍋

Ronish Bhatt

14 days ago

Since the original proposal, the experiments have pushed the question beyond simply asking how much information can be recovered from a transformer's own write directions. Across five models, native writes could often be followed downstream even after their original directions became difficult to recover from the accumulated residual, with their transported descendants retaining information about both their source and later functional effects. This points to a useful separation between where a computation is written, how it is transformed through the residual stream, and what part of it remains functionally relevant.

More broadly, the experiments suggest that residual perturbations can become physically high-dimensional while their behaviorally relevant structure remains comparatively low-dimensional, and that provenance and causal importance are not interchangeable. Despite the grant auto-rejection, I’m releasing these results and the associated negative findings as an extension of the original project in the hope that they’re useful for people studying circuit tracing, activation transport, and the relationship between parameter-level structure and learned representations.