Backpropagation
Backpropagation—efficient reverse-mode differentiation through multilayer nets—made hidden-unit learning a community tool in 1986, after an earlier paper trail in automatic differentiation and ordered derivatives that was not invented out of nowhere that year.
Primary source: Rumelhart, Hinton & Williams, Nature (1986): https://doi.org/10.1038/323533a0
The previous chapter treated the 1980s return of distributed representations—Hopfield energy nets, Boltzmann machines, and the PDP volumes—as an empirical research program with new mathematics of attractors, stochastic search, and constraint satisfaction. One piece of that program traveled farther than the rest as a training method for feedforward multilayer nets: backpropagation, the efficient computation of gradients of a scalar error with respect to every weight by a backward pass that reuses the forward computation graph. The slogan that “backprop was invented in 1986” is folklore. Reverse-mode differentiation has an earlier paper trail in numerical analysis and in systems estimation. What Rumelhart, Hinton, and Williams published in Nature in October 1986, and expanded as Chapter 8 of PDP Volume 1, was the presentation that made the method a research-community tool for learning internal representations in hidden layers. This chapter follows that paper trail without refereeing a founding myth, states what the 1986 work enabled, and leaves practical scaling—large data, GPUs, ReLU, dropout—to later chapters.
Reverse mode before the neural-net slogan: Linnainmaa
In 1970, Seppo Linnainmaa completed a master’s thesis at the University of Helsinki on the cumulative rounding error of an algorithm, written in Finnish under the title Algoritmin kumulatiivinen pyöristysvirhe yksittäisten pyöristysvirheiden Taylor-kehitelmänä—conventionally rendered in English as the representation of an algorithm’s cumulative rounding error as a Taylor expansion of the local rounding errors. The accessible peer-reviewed statement is the English article Taylor expansion of the accumulated rounding error, published in BIT Numerical Mathematics in 1976 (vol. 16, pp. 146–160). The work belongs to numerical analysis: it develops analytic and algorithmic methods for the influence of local rounding errors on an accumulated error, including a backward sweep whose structure is what later automatic-differentiation literature calls reverse mode.
That lineage matters for honesty about priority, and it matters equally for what not to claim. Linnainmaa’s thesis and BIT paper are about differentiating through computational graphs for rounding-error analysis. They are not papers about training multilayer neural networks, credit assignment in hidden units, or the exclusive-or problem. Histories of automatic differentiation (for example Andreas Griewank’s 2012 survey of reverse-mode discovery) place Linnainmaa among the early formalizers of the reverse mode; they do not make him the inventor of “backpropagation for neural nets” in the sense the 1980s connectionist literature used that phrase. The algorithmic idea of propagating derivatives backward at a cost comparable to the forward evaluation is older than the 1986 neural-net papers; the research community that made multilayer supervised learning travel under the name backpropagation is a later story.
Ordered derivatives and neural nets: Werbos
A second strand runs through Paul Werbos’s August 1974 Harvard Ph.D. thesis, Beyond regression: new tools for prediction and analysis in the behavioral sciences. The thesis develops what Werbos called “dynamic feedback” and “ordered derivatives”—a reverse-mode / chain-rule technique for computing derivatives of a model’s outputs with respect to its parameters efficiently, for use with steepest descent and related estimation methods. The primary applications in the dissertation are nonlinear and dynamic forecasting and systems estimation in the behavioral sciences, not a polished multilayer-perceptron training recipe presented as the main result. Werbos himself later summarized the reverse method’s many rediscoveries across fields; neural-network historiography often points to a subsequent publication, Applications of advances in nonlinear sensitivity analysis (in Drenick and Kozin’s System Modeling and Optimization, Springer 1982; from the 1981 IFIP conference), as an early place where the sensitivity-analysis toolkit is explicitly tied to artificial intelligence and neuron modelling.
Again the careful claim is narrower than a founding myth. Werbos documented reverse differentiation for adapting parameters in nonlinear systems, and he discussed neuron modelling; the 1974 thesis was not widely taken up at the time as “the” backpropagation paper of the neural-net community. When that community later standardized on error backpropagation for multilayer nets, priority disputes followed. This history records the paper trail. It does not award a single inventor’s medal.
Rumelhart, Hinton, and Williams 1986: the community tool
On 9 October 1986, David E. Rumelhart, Geoffrey E. Hinton, and Ronald J. Williams published Learning representations by back-propagating errors in Nature (vol. 323, pp. 533–536). The letter describes a learning procedure for networks of neurone-like units that repeatedly adjusts connection weights to minimize a measure of the difference between actual and desired output vectors. The distinctive claim is representational: as weights adjust, internal “hidden” units that are not part of the input or output come to represent important features of the task domain, and task regularities are captured by interactions among those units. That ability to create useful new features, the authors write, distinguishes backpropagation from earlier, simpler methods such as the perceptron-convergence procedure.
Mechanically, the paper gives layered feedforward nets with a smooth nonlinearity (they use a logistic), a sum-squared error, forward computation of unit states, and a backward pass that applies the chain rule to obtain ∂E/∂w for every weight from locally available quantities—exactly reverse-mode differentiation specialized to this architecture. Weight updates may accumulate the gradient over the training set; a momentum term is noted as a practical accelerator. The letter demonstrates the procedure on tasks chosen to require hidden structure: mirror-symmetry detection in a binary input vector (solved with two hidden units whose learned weights have an elegant antisymmetry), and a family-tree proposition-completion task in which hidden layers invent distributed codes for people and relationships that support generalization across isomorphic trees. The authors note independent variants discovered by David Parker and by Yann Le Cun, and they point to the expanded treatment in the PDP volumes.
What made the Nature paper travel was not a claim of absolute chronological priority. It was a short, high-visibility statement that multilayer nets with differentiable units could learn useful hidden representations by gradient descent, with worked examples and a clear derivation. In that sense 1986 is a community event, not a creation ex nihilo.
PDP Volume 1, Chapter 8: the expanded statement
The same authors’ fuller account is Chapter 8 of Rumelhart, McClelland, and the PDP Research Group’s Parallel Distributed Processing, Volume 1, Foundations (MIT Press / Bradford Books, 1986): Learning Internal Representations by Error Propagation (pp. 318–362). As the previous chapter of this history stressed, that chapter sits inside a multi-author program; collapsing “PDP” into “backpropagation” is folklore. Chapter 8 is nonetheless the expanded technical statement of the generalized delta rule: semilinear (differentiable, nondecreasing) activation functions; recursive backward computation of error signals; discussion of local minima; and an extensive simulation section.
Here the exclusive-or (XOR) problem appears as a classic illustration—not as the whole point. Networks without hidden units cannot learn XOR when the input coding makes similar patterns require dissimilar outputs; with a suitable hidden unit, they can. The chapter then moves through parity, encoder bottlenecks, symmetry, binary addition, negation, a T-versus-C shape discrimination with constrained receptive fields, and even recurrent nets treated by unfolding in time. XOR matters pedagogically because it is the simplest linearly inseparable Boolean function and because Minsky and Papert’s 1969 analysis had made the limits of order-limited machines vivid. Treating “XOR is why backprop mattered” as the historical summary confuses a textbook example with the research claim, which is the learning of internal representations for a range of mappings.
What the method enabled—and what it did not yet scale
Backpropagation enabled a practical answer to the credit-assignment problem for deterministic multilayer feedforward nets: how to assign responsibility for output error to hidden units whose desired states are not given by the teacher. Hidden layers could invent features; linearly inseparable problems became illustrations rather than stop signs; and the same gradient machinery could be applied, with architectural constraints, to real pattern tasks.
An early documented application line is handwritten-digit recognition. In December 1989, Yann LeCun and colleagues at AT&T Bell Laboratories published Backpropagation Applied to Handwritten Zip Code Recognition in Neural Computation (vol. 1, no. 4, pp. 541–551). The paper shows how task-domain constraints can be built into a backpropagation network’s architecture and reports successful recognition of handwritten zip-code digits from U.S. Postal Service data, with a single network learning from normalized character images to classification. Related contemporaneous work from the same group appeared in the NIPS proceedings as Handwritten Digit Recognition with a Back-Propagation Network. These results show what the method enabled in practice in the late 1980s. They are not AlexNet, not ImageNet-scale training, and not a claim that 1986 finished the engineering story.
Practical scaling waited. Large labeled datasets, commodity GPUs, rectified linear units, dropout and related regularizers, and modern adaptive optimizers belong to later chapters of this history. The 1986 papers established that multilayer error backpropagation worked as a research method and could learn internal representations; they did not deliver today’s deep-learning stack.
Caveat
Three retellings especially distort this material.
First, “backprop was invented in 1986.” Reverse-mode differentiation appears earlier in Linnainmaa’s rounding-error analysis (1970 thesis; 1976 BIT) and in Werbos’s ordered-derivative work (1974 thesis; early-1980s sensitivity-analysis publications). Rumelhart, Hinton, and Williams 1986 is the paper that made the method a neural-net community tool, not the first writing of reverse derivatives.
Second, “XOR is why backprop mattered.” XOR is a clean illustration of linear inseparability and of the need for learned hidden representations. The Nature letter’s featured examples are symmetry detection and family trees; PDP Chapter 8 treats XOR among many simulations. The research claim is broader: gradient-based learning of internal representations.
Third, treating priority wars as the main story. There is a real paper trail across automatic differentiation, systems estimation, and connectionist learning. Independent rediscovery is normal. Refereeing a single founding myth—whether for Linnainmaa, Werbos, Parker, Le Cun, or Rumelhart–Hinton–Williams—substitutes genealogy for what the 1986 work actually changed: multilayer supervised learning became a shared, reproducible method.
Sources
Primary
- Linnainmaa, Seppo. Algoritmin kumulatiivinen pyöristysvirhe yksittäisten pyöristysvirheiden Taylor-kehitelmänä. Master’s thesis, Department of Computer Science, University of Helsinki, 1970. (English title conventionally: “The representation of the cumulative rounding error of an algorithm as a Taylor expansion of the local rounding errors.”) Scan discussed in secondary AD histories; e.g. https://people.idsia.ch/~juergen/linnainmaa1970thesis.pdf
- Linnainmaa, Seppo. “Taylor expansion of the accumulated rounding error.” BIT Numerical Mathematics 16 (1976): 146–160. https://doi.org/10.1007/BF01931367
- Werbos, Paul J. Beyond regression: new tools for prediction and analysis in the behavioral sciences. Ph.D. dissertation, Harvard University, August 1974. (Commonly circulated PDF: https://gwern.net/doc/ai/nn/1974-werbos.pdf)
- Werbos, Paul J. “Applications of advances in nonlinear sensitivity analysis.” In System Modeling and Optimization, edited by R. F. Drenick and F. Kozin, 762–770. Lecture Notes in Control and Information Sciences 38. Berlin: Springer, 1982. https://doi.org/10.1007/BFb0006203 (IFIP Conference, New York, 1981.)
- Rumelhart, David E., Geoffrey E. Hinton, and Ronald J. Williams. “Learning representations by back-propagating errors.” Nature 323 (9 October 1986): 533–536. https://doi.org/10.1038/323533a0
- Rumelhart, David E., Geoffrey E. Hinton, and Ronald J. Williams. “Learning Internal Representations by Error Propagation.” In Parallel Distributed Processing: Explorations in the Microstructure of Cognition, Vol. 1, Foundations, by David E. Rumelhart, James L. McClelland, and the PDP Research Group, 318–362. Cambridge, Mass.: MIT Press (Bradford Books), 1986. (ICS Report 8506 / DTIC ADA164453 preprint, September 1985: https://apps.dtic.mil/sti/tr/pdf/ADA164453.pdf)
- LeCun, Yann, Bernhard Boser, John S. Denker, Donnie Henderson, Richard E. Howard, Wayne Hubbard, and Lawrence D. Jackel. “Backpropagation Applied to Handwritten Zip Code Recognition.” Neural Computation 1, no. 4 (December 1989): 541–551. https://doi.org/10.1162/neco.1989.1.4.541
- LeCun, Y., B. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. Hubbard, and L. D. Jackel. “Handwritten Digit Recognition with a Back-Propagation Network.” In Advances in Neural Information Processing Systems 2 (NIPS 1989), 396–404. 1990. https://proceedings.neurips.cc/paper_files/paper/1989/file/53c3bce66e43be4f209556518c2fcb54-Paper.pdf
Historical / technical reference
- Griewank, Andreas. “Who Invented the Reverse Mode of Differentiation?” In Documenta Mathematica, Extra Volume ISMP (2012): 389–400. https://doi.org/10.4171/dms/6/38
- Werbos, Paul J. “Backwards Differentiation in AD and Neural Nets: Past Links and New Opportunities.” In Automatic Differentiation: Applications, Theory, and Implementations, edited by H. M. Bücker et al. Springer, 2006. Author PDF: http://www.werbos.com/AD2004.pdf