Discussion about this post

User's avatar
Nazem Alkudsi, CFA's avatar

Jordan and Phoebe, this conversation with Florian and John makes a useful distinction between reasoning, alignment and instruction-following. I wonder if the unit of evaluation needs to extend beyond the model to the decision chain around it.

On 9 November 1979, test-scenario data entered NORAD’s live missile-warning computers and produced false indications of a mass raid. GAO later recorded that NORAD installed a $16 million off-site facility so software testing no longer ran on the live system. The lesson is procedural as well as technical.

A Situation Room eval should test the moment when a human challenges the model. It should require confirmation from a separate channel and ask how the combined system fails under time pressure. Otherwise we may benchmark the adviser while leaving the institution untested. What would an eval of the whole human-machine chain look like?

Jack Shanahan's avatar

As I read the alarming results of escalation to nuclear weapons (seemingly more common than not with these models), I immediately thought of Samuel Hammond's latest post. In which he postulates that RLVR and similar LLM reward functions have replaced a uniquely human combination of consequentialism and deontology with an almost sociopathic form of machine consequentialism.

Something along those lines that seems to have unfolded in these Civilization runs.

https://www.secondbest.ca/p/the-sorcerers-apprentice?utm_source=post-email-title&publication_id=1099670&post_id=210810417&utm_campaign=email-post-title&isFreemail=true&r=8at2k&triedRedirect=true&utm_medium=email

5 more comments...

No posts

Ready for more?