Bitware News
History · Ch. 012026-09-13

Cybernetics and the First Models

Before the name 'artificial intelligence,' research programs already treated mind and machine as the same kind of problem—logical neurons, feedback control, and operational tests of machine intelligence.

History

Primary source: McCulloch & Pitts, Bulletin of Mathematical Biophysics (1943): https://doi.org/10.1007/BF02478259

Before “artificial intelligence” was a name, several research programs already treated mind and machine as instances of the same kind of problem. Warren McCulloch and Walter Pitts asked whether idealized neurons could realize a logical calculus. Arturo Rosenblueth, Norbert Wiener, and Julian Bigelow asked whether purpose could be analyzed without vitalism, as negative feedback. Claude Shannon gave communication an exact mathematical theory, and—separately—sketched how a digital computer might play chess. Alan Turing replaced the question “Can machines think?” with an operational imitation game. Dartmouth, in 1956, would name a research program. It did not invent these problems.

The folklore that “AI began at Dartmouth” compresses a longer prehistory into a founding myth. The papers below show what that prehistory actually contained: shared formalisms for nervous activity, control, information, and computation—not yet a single discipline, and not yet machine learning in the modern sense.

Logical neurons, not a learning algorithm

In December 1943, Warren S. McCulloch and Walter Pitts published A logical calculus of the ideas immanent in nervous activity in the Bulletin of Mathematical Biophysics (vol. 5, pp. 115–133). The paper’s opening move is physiological and logical at once. Because of the “all-or-none” character of nervous activity, they argue, neural events and the relations among them can be treated by means of propositional logic. Under a short list of idealized assumptions—fixed thresholds, synaptic delay as the significant delay, absolute inhibition, and a net whose structure does not change with time—they show two complementary results: the behavior of every such net can be described in logical terms (with more complicated means for nets containing circles), and for logical expressions satisfying certain conditions one can find a net that behaves as the expression describes.

What the paper does not do matters as much as what it does. McCulloch and Pitts are explicit that facilitation, extinction, and learning involve real continuous and enduring changes; they treat those changes, for the calculus, by substituting “equivalent fictitious nets” whose connections and thresholds are unaltered. The formal equivalence, they warn, is not a factual explanation of learning. Later retellings sometimes call this paper “the first neural network.” In the modern machine-learning sense—trained weights, multilayer perceptrons, gradient methods—that label is anachronistic. The 1943 paper is a logical calculus of idealized neurons and nets, including a discussion of circular paths and, in closing, a brief link between nets (with tape and scanners) and Turing computability. It is foundational for the idea that mind-like processes might be formalized as discrete computation. It is not a training algorithm.

Purpose without vitalism

The same year, Arturo Rosenblueth, Norbert Wiener, and Julian Bigelow published Behavior, Purpose and Teleology in Philosophy of Science (vol. 10, no. 1, January 1943, pp. 18–24). Their essay has two stated goals: to define a behavioristic study of natural events and to classify behavior; and to insist on the usefulness of the concept of purpose.

They define the behavioristic approach as the examination of an object’s output and its relation to input, deliberately setting aside intrinsic structure. Within that frame they classify active behavior as purposeless or purposeful, and purposeful behavior as feedback (teleological) or non-feedback. Crucially, they restrict “teleology” to purpose controlled by negative feedback—signals from the goal that correct the margin of error—rather than to final causes opposed to determinism. A target-seeking torpedo and a cerebellar patient’s overshooting reach are analyzed with the same vocabulary. Teleology, so defined, is not opposed to determinism but to non-teleology. Machines and organisms can be studied with a uniform behavioristic analysis even when their materials differ.

This is the conceptual hinge of early cybernetics: goal-directed behavior without appeals to vital force, and a shared language for servomechanisms and voluntary action.

What cybernetics claimed

Wiener’s book Cybernetics: Or Control and Communication in the Animal and the Machine appeared in 1948 from Hermann & Cie (Paris), The Technology Press (Cambridge, Mass.), and John Wiley & Sons (New York). The title is the thesis. Cybernetics, as Wiener presented it, sought common principles of control and communication in automatic machines and in living organisms. Feedback—circular processes in which performance is compared with a goal and the difference used to correct action—was the central engineering idea. Wiener’s wartime work with Bigelow on anti-aircraft prediction, and his collaboration with Rosenblueth on nervous oscillations and purpose tremor, supplied the motivating cases.

In a 1948 popular summary adapted from the book, Wiener described cybernetics as combining what is loosely called thinking in a human context with control and communication in engineering: finding common elements in automatic machines and the human nervous system, and developing a theory covering the whole field. Neurons, on this view, do work under conditions more like vacuum tubes than like nineteenth-century heat engines; the relevant bookkeeping is of messages, noise, and coding, not only of energy. The claim was ambitious and interdisciplinary. It was not yet a claim that symbolic “artificial intelligence,” under that name, had begun.

Shannon: communication theory—and a chess sketch

Claude E. Shannon’s A Mathematical Theory of Communication appeared in the Bell System Technical Journal in two parts in 1948 (vol. 27, July, pp. 379–423; October, pp. 623–656). It treats the fundamental problem of communication as reproducing at one point a message selected at another, and builds a quantitative theory of information, channels, noise, and coding. Semantic meaning is explicitly set aside. The paper is a foundation of information theory. It is not usefully described as an “AI founding” document; overclaiming Shannon’s role in that way collapses distinct research agendas.

Shannon did, separately, write about programming computers for chess. Programming a Computer for Playing Chess appeared in The London, Edinburgh, and Dublin Philosophical Magazine and Journal of Science, Series 7, vol. 41, no. 314 (March 1950), pp. 256–275. That essay belongs to the early literature of machine game-playing and evaluation; it does not turn communication theory into a theory of mind. Keep the two Shannon contributions distinct.

Turing’s operational question

In October 1950, A. M. Turing published Computing Machinery and Intelligence in Mind (New Series, vol. 59, no. 236, pp. 433–460). He opens by declining to define “machine” and “think” from ordinary usage—an approach that, he says, would reduce the question to a Gallup poll. Instead he replaces “Can machines think?” with the imitation game: an interrogator tries to distinguish, by written questions alone, which of two respondents is which; then one asks what happens when a machine takes one part.

The imitation game is an operational proposal, not a slogan. Turing restricts the machines of interest to digital computers, argues for their universality among discrete-state machines, and walks through objections (theological, “heads in the sand,” mathematical, consciousness, various disabilities, Lady Lovelace, continuity of the nervous system, informality of behavior, and even extra-sensory perception). He states a belief about future performance—storage on the order of 10⁹, and an average interrogator with no more than a 70 percent chance of correct identification after five minutes by about the year 2000—while calling the original question too meaningless to deserve discussion. The paper ends with learning machines and the suggestion that one might start from a “child-programme” rather than an adult mind. Whatever later culture made of the “Turing test,” the 1950 text is a careful substitution of an empirical procedure for an ill-posed metaphysical question.

Automata and the brain

John von Neumann’s lecture The General and Logical Theory of Automata, delivered at the Hixon Symposium (20 September 1948) and published in Lloyd A. Jeffress, ed., Cerebral Mechanisms in Behavior: The Hixon Symposium (John Wiley & Sons, 1951), pp. 1–41, belongs to the same mid-century conversation about automata, brains, and logical description. His unfinished The Computer and the Brain was published posthumously by Yale University Press in 1958; any citation of that book should flag the date. These works reinforce the chapter’s theme—brains and computers as objects of joint formal analysis—without making von Neumann an “AI founder” under that later label.

What this cluster was

Taken together, these sources show a cluster of programs that treated nervous activity, purposeful behavior, communication, and computation with shared mathematical ambition. McCulloch–Pitts gave a logical calculus for idealized nets. Rosenblueth, Wiener, and Bigelow naturalized purpose as feedback. Wiener’s Cybernetics named an interdisciplinary field of control and communication in animal and machine. Shannon quantified communication (and, separately, sketched chess programming). Turing turned machine intelligence into an operational question about digital computers.

What later became “artificial intelligence” drew on this cluster, competed with parts of it, and eventually diverged from cybernetics as a social and institutional formation. Overlap is real; identity is not. The next chapter will take up how Dartmouth named a program in 1956. That naming is historically important. It is not the origin of the problems surveyed here.

Caveat

Three retellings especially distort this material.

First, “AI began at Dartmouth in 1956” is folklore if stated as fact. Dartmouth named and organized a research program; the problems of logical nervous activity, feedback purpose, information, and machine intelligence as computation were already in print.

Second, McCulloch–Pitts is not a trained multilayer perceptron. Treating the 1943 paper as the first modern neural net in the learning sense projects later machinery backward onto a logical calculus of fixed idealized nets.

Third, cybernetics and AI are sometimes collapsed into one continuous community. They overlapped in people, problems, and mid-century meetings; they then diverged in methods, institutions, and self-description. Continuity of themes is not continuity of a single field.

Sources

Primary

Historical reference