Tag
#mechanistic interpretability
2 stories taggedmechanistic interpretability.

Science & Space
A New Tool Lets Researchers See Inside AI Models, Not Just Watch What They Do
AI lab Goodfire has opened its Silico platform to the public, giving researchers the ability to peek inside a model's workings and ask why it behaves the way it does.
4 min read

Frontier Labs
Anthropic ha scoperto uno strato nascosto nel ragionamento della sua IA. Ecco cosa significa davvero.
L'azienda ha scoperto parole che lampeggiano dentro Claude e che non appaiono mai nelle sue risposte, inclusa una che sembrava innescare l'imbroglio in un test di programmazione. È una scoperta genuina, ma non una finestra sulla mente di un robot.
3 min read