Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa, 2022
In 2022 a team at the University of Tokyo and Google found that a large language model's accuracy on arithmetic word problems jumped if the answer was made to begin with a particular sentence.
With it, InstructGPT went from 17.7 percent to 78.7 percent on the MultiArith benchmark and from 10.4 to 40.7 percent on GSM8K, with no examples and no training. The model, asked to reason before answering, produced a chain of reasoning and then an answer that was far more often right. Whether this is a program is the question the exhibit asks. It is text that changes what a machine does, it was discovered by experiment rather than designed, and it works in no language but English. The last two rooms of this museum trace a line from a table of operations for an engine that was never built to a sentence addressed to a machine that was. The sentence is the shortest work in the collection and the only one every visitor can already read.