Anthropic opened a window into the ‘black box’ where ‘features’ steer a large language model’s output.