Expert Systems and a Funding Winter
After Dartmouth and alongside the perceptron line, a large part of AI became knowledge plus search—GPS, DENDRAL, MYCIN, and the expert-system boom—while the journalistic label “AI winter” names real but local funding contractions, not a single global morality play.
Primary source: Newell & Simon, CACM (1976): https://doi.org/10.1145/360018.360022
The previous chapters treated Dartmouth’s naming of a research program and the perceptron line of adaptive networks. Alongside those developments—and for a long stretch more visibly funded and published—another large part of artificial intelligence became the encoding of knowledge and the control of search. Allen Newell and Herbert A. Simon’s General Problem Solver and their later book Human Problem Solving made heuristic search over symbol structures a research program continuous with cognitive psychology. Edward Feigenbaum, Joshua Lederberg, Bruce Buchanan, and collaborators at Stanford built DENDRAL to elucidate molecular structure from mass spectra by combining combinatorial generation with chemists’ heuristics. Edward Shortliffe’s MYCIN cast infectious-disease advice as production rules with certainty factors and an interactive consultation. Knowledge, not only search or learning, was treated as the bottleneck. The later journalistic label “AI winter” names real funding contractions—especially James Lighthill’s 1973 survey for Britain’s Science Research Council, and the mid- to late-1980s collapse of commercial expectations around expert-system shells and Lisp machines in the United States—but it is not a single worldwide event and should not be used as a morality play.
Symbols, search, and a general problem solver
Newell, J. C. Shaw, and Simon developed the General Problem Solver (GPS) in the late 1950s as a program meant to separate general problem-solving mechanisms from particular task domains. Their Report on a general problem-solving program was presented at the 1959 International Conference on Information Processing in Paris and published in the UNESCO proceedings Information Processing (1960), pp. 256–264. GPS used means–ends analysis: compare the current state with a goal, find a difference, select an operator known to reduce differences of that kind, apply it, and repeat. The system was not as universal as its name advertised, but it made explicit a research bet that intelligent performance could be studied as heuristic search in a problem space.
That program of work culminated, for Newell and Simon, in Human Problem Solving (Englewood Cliffs, N.J.: Prentice-Hall, 1972)—a long statement of information-processing psychology in which human reasoners are modeled as symbolic information-processing systems, and protocols of problem solving are analyzed against computer simulations. In their 1975 ACM Turing Award lecture, published as Computer science as empirical inquiry: symbols and search in Communications of the ACM 19, no. 3 (March 1976), pp. 113–126, they stated what they called a law of qualitative structure: the Physical Symbol System Hypothesis—that “a physical symbol system has the necessary and sufficient means for general intelligent action.” They paired it with a Heuristic Search Hypothesis: solutions are represented as symbol structures, and intelligence in problem solving is exercised by generating and modifying those structures until a solution appears. The lecture treats both claims as empirical hypotheses, not theorems. Keep that caution. GPS and the 1972 book are not the same document as the 1976 lecture; the hypothesis is pinned to the Turing Award text.
DENDRAL: structure from spectra as knowledge plus heuristics
At Stanford, beginning in the mid-1960s, Feigenbaum, Lederberg, Buchanan, and collaborators attacked a sharply bounded scientific task: given a mass spectrum and an empirical formula, propose plausible molecular structures. Lederberg’s DENDRAL algorithm provided a systematic, non-redundant generator of topologically possible structures from atomic composition and valence. Heuristic DENDRAL then used chemists’ knowledge—stable and unstable configurations, fragmentation rules, planning constraints inferred from spectral patterns—to prune and rank candidates. Bruce G. Buchanan, Georgia L. Sutherland, and Edward A. Feigenbaum’s Heuristic DENDRAL: A Program for Generating Explanatory Hypotheses in Organic Chemistry, in Machine Intelligence 4 (Edinburgh University Press, 1969), pp. 209–254, describes a plan–generate–test organization: preliminary inference constrains the space; a structure generator produces candidates; a predictor and evaluation function compare predicted spectra with observed data.
What matters historically is the methodological claim the project made concrete. The search space of organic structures is combinatorially huge; brute generation is not enough. Specialist knowledge—encoded as heuristics and later as rules—makes the search tractable in narrow chemical families. Feigenbaum, Buchanan, and Lederberg argued in related writing that general problem-solvers are too weak as a basis for high-performance systems, and that expert human performance is strong mainly in areas of specialization. DENDRAL was not “chemistry solved by AI.” It was a sustained experiment in knowledge-intensive hypothesis formation, with results published in chemistry journals as well as AI venues, and with later Meta-DENDRAL work aimed at acquiring fragmentation rules from data.
MYCIN: consultation, rules, and certainty factors
MYCIN grew from Shortliffe’s doctoral work at Stanford in collaboration with infectious-disease clinicians, including Stanley Cohen and Stanton Axline, and with Buchanan. The system advised on antimicrobial therapy for serious infections—initially bacteremia (bloodstream infection), later extended to meningitis. It was configured as a consultation: a physician answered questions; the program chained IF–THEN production rules backward from goals; and it could explain why a question was asked or how a conclusion was reached. Uncertainty was handled with certainty factors—numbers between −1 and 1 attached to conclusions—rather than a full Bayesian model; Shortliffe and Buchanan published the inexact-reasoning scheme in Mathematical Biosciences in 1975.
Careful evaluation exists and must be read without folklore. Victor L. Yu, Buchanan, Shortliffe, and colleagues reported in Computer Programs in Biomedicine 9 (1979) that, in an evaluation focused on bacteremia, MYCIN’s therapy recommendations met Stanford experts’ standards of acceptable practice 90.9% of the time. A separate, blinded meningitis study—originally in JAMA 242 (1979) and reprinted as Chapter 31 of Buchanan and Shortliffe’s Rule-Based Expert Systems (Addison-Wesley, 1984)—compared MYCIN’s antimicrobial selections with those of Stanford clinicians and had outside infectious-disease specialists rate prescriptions without knowing which came from the program. Sixty-five percent of MYCIN’s prescriptions were rated acceptable by the evaluators; seventy percent were rated acceptable by a majority of evaluators; MYCIN never failed to cover a treatable pathogen in the ten-case set. Those numbers are about specialist agreement on therapy choice in challenging meningitis cases. They are not a claim that MYCIN replaced physicians in general practice, nor that it was routinely deployed at the bedside as standard of care. The 1984 Buchanan–Shortliffe volume itself treats MYCIN as a decade of experiments—rules, explanation, EMYCIN as an emptied inference engine, tutoring, evaluation—not as a finished clinical product.
Knowledge as bottleneck
By the late 1970s Feigenbaum was stating the lesson of DENDRAL and MYCIN as an engineering principle. In The Art of Artificial Intelligence: Themes and Case Studies of Knowledge Engineering (Stanford Computer Science Report STAN-CS-77-621 / HPP-77-25, August 1977; also presented at IJCAI-77), he argued that the problem-solving power of an intelligent agent is “primarily a consequence of the specialist’s knowledge employed by the agent, and only very secondarily related to the generality and power of the inference method employed.” Agents, he wrote, must be knowledge-rich even if methods-poor; “in the knowledge is the power.” He later called related formulations the knowledge principle. The point was not that search had failed, but that high competence in a domain required large amounts of specific, often heuristic, expertise—acquired laboriously from specialists—and that the acquisition and representation of that knowledge were the central engineering problems. Expert systems, EMYCIN-style shells, and commercial knowledge-engineering tools in the early 1980s followed that diagnosis.
Lighthill in Britain: a survey, not the worldwide winter
In 1972–73 Sir James Lighthill, Lucasian Professor of Applied Mathematics at Cambridge, prepared an independent survey for the UK Science Research Council (SRC), published as Artificial Intelligence: A General Survey (often collected with responses as Artificial Intelligence: a Paper Symposium, SRC, 1973). Writing as a mathematically sophisticated outsider after roughly two months of reading and interviews, Lighthill divided work into Category A (Advanced Automation), Category B (a “bridge” activity centered on Building Robots, linking A and C), and Category C (Computer-based central-nervous-system research). He credited respectable if disappointing progress in A and C, and judged B especially unsuccessful. He emphasized the combinatorial explosion as a general obstacle to large knowledge bases and self-organizing systems, praised narrow knowledge-intensive successes such as Heuristic DENDRAL, and questioned whether category B justified treating AI as one coherent field.
The SRC’s subsequent funding reorganization—favoring A and C and leaving bridge/robotics work exposed—produced a real contraction of basic AI and robotics support in Britain, with well-documented effects on Edinburgh’s laboratory. That is a national research-council episode. It is not identical with US DARPA cycles, Japanese Fifth Generation investment, or the later commercial expert-system bust. Treating “the Lighthill Report” as the cause of a singular global AI winter is folklore.
Boom, bust, and the label “AI winter”
In the early 1980s, expert systems moved from laboratory demonstrations into industrial and commercial expectation. Digital Equipment Corporation’s R1/XCON configurator for VAX systems, developed with John McDermott at Carnegie Mellon and deployed from 1980, became a frequently cited industrial case: large rule bases in OPS5, used to check completeness and consistency of configurations. Feigenbaum and Pamela McCorduck’s The Fifth Generation (1983) and Hayes-Roth, Waterman, and Lenat’s Building Expert Systems (Addison-Wesley, 1983) helped frame knowledge systems as a practical technology. Specialized Lisp machines and commercial expert-system shells rode the same wave of expectation.
Expectations outran delivery. At the 1984 AAAI National Convention, a panel titled “The ‘Dark Ages’ of AI” warned that commercial hustle risked a crash; Drew McDermott’s remarks, as summarized in later histories, already used the phrase “AI Winter.” Nils Nilsson’s standard history (The Quest for Artificial Intelligence, Cambridge University Press, 2010) describes mid- to late-1980s downturn: falling AAAI membership and advertising, closure of some AI companies, reduced industrial exhibits, and a cut in DARPA funding for basic AI and Strategic Computing research from about $47 million to $31 million between 1987 and 1989. Specialized Lisp-machine vendors and shell vendors faced cheaper general workstations, portability pressure, costly knowledge acquisition, and brittle maintenance of large rule bases. Exact company-by-company failure dates vary across secondary accounts; where a precise bankruptcy date cannot be pinned to a primary filing in this chapter, the safer statement is qualitative and tied to Nilsson’s and Daniel Crevier’s (AI: The Tumultuous History of the Search for Artificial Intelligence, Basic Books, 1993) periodizations.
Two folklore moves to refuse. First: one singular “AI winter” caused by one report or one book. Britain’s post-Lighthill contraction and the US mid-1980s commercial/funding downturn are related thematically—over-promise, combinatorial hardness, knowledge bottlenecks—but they are not the same event. Second: expert systems as the opposite of “real AI” or as a dead end with no legacy. Rule-based and knowledge-based systems persisted in constrained domains (configuration, monitoring, protocol assistants); the knowledge-acquisition bottleneck and the limits of certainty-factor calculus are real lessons, not certificates of death.
What this chapter establishes, on the primary record, is narrower. After Dartmouth, a major strand of AI made symbols, search, and encoded expertise its method. GPS and Human Problem Solving tied that strand to cognitive science; DENDRAL and MYCIN showed what knowledge plus heuristics could do in chemistry and infectious-disease consultation; Feigenbaum named knowledge as the scarce resource; Lighthill and the mid-1980s bust show how funding and markets punish mismatched expectations. Learning-centered and connectionist returns belong to later chapters.
Caveat
Four retellings especially distort this material.
First, a singular global “AI winter” caused by Lighthill alone, or by Minsky and Papert’s Perceptrons alone, or by any one commercial failure. Funding winters are plural, national, and institutional.
Second, MYCIN as clinical deployment that outperformed physicians in general practice. There were careful evaluation studies with stated percentages for specialist agreement on therapy recommendations; they do not license a slogan about replacing doctors.
Third, expert systems as merely “not real AI,” or as a path that vanished without residue. Production rules, explanation facilities, and knowledge bases remained tools in constrained applications long after the boom rhetoric collapsed.
Fourth, collapsing GPS’s means–ends analysis, Feigenbaum’s knowledge-engineering slogan, and commercial shell marketing into one continuous triumph or one continuous fraud. The primary papers and the funding documents support a more ordinary history: methods that worked narrowly, principles that over-generalized, and markets that believed the over-generalization.
Sources
Primary
- Newell, Allen, J. C. Shaw, and Herbert A. Simon. “Report on a general problem-solving program.” In Information Processing: Proceedings of the International Conference on Information Processing (UNESCO, Paris, 1959), 256–264. Paris: UNESCO, 1960.
- Newell, Allen, and Herbert A. Simon. Human Problem Solving. Englewood Cliffs, N.J.: Prentice-Hall, 1972.
- Newell, Allen, and Herbert A. Simon. “Computer science as empirical inquiry: symbols and search.” Communications of the ACM 19, no. 3 (March 1976): 113–126. https://doi.org/10.1145/360018.360022
- Buchanan, Bruce G., Georgia L. Sutherland, and Edward A. Feigenbaum. “Heuristic DENDRAL: A Program for Generating Explanatory Hypotheses in Organic Chemistry.” In Machine Intelligence 4, edited by B. Meltzer and D. Michie, 209–254. Edinburgh: Edinburgh University Press, 1969. Stanford HPP memo: https://stacks.stanford.edu/file/druid:yj802bw3458/yj802bw3458.pdf
- Feigenbaum, Edward A. “The Art of Artificial Intelligence: Themes and Case Studies of Knowledge Engineering.” Stanford Computer Science Report STAN-CS-77-621 / Heuristic Programming Project Memo HPP-77-25, August 1977. http://i.stanford.edu/pub/cstr/reports/cs/tr/77/621/CS-TR-77-621.pdf ; IJCAI-77 version: https://www.ijcai.org/Proceedings/77-2/Papers/092.pdf
- Shortliffe, Edward H., and Bruce G. Buchanan. “A Model of Inexact Reasoning in Medicine.” Mathematical Biosciences 23 (1975): 351–379.
- Shortliffe, Edward H. Computer-Based Medical Consultations: MYCIN. New York: Elsevier, 1976.
- Yu, Victor L., Bruce G. Buchanan, Edward H. Shortliffe, et al. “Evaluating the performance of a computer-based consultant.” Computer Programs in Biomedicine 9, no. 1 (1979): 95–102. https://doi.org/10.1016/0010-468X(79)90022-9
- Yu, Victor L., et al. “Antimicrobial Selection by a Computer: A Blinded Evaluation by Infectious Diseases Experts.” JAMA 242, no. 12 (1979): 1279–1282. https://doi.org/10.1001/jama.1979.03300120033020 (Meningitis evaluation; reprinted as Chapter 31 in Buchanan & Shortliffe 1984.)
- Buchanan, Bruce G., and Edward H. Shortliffe, eds. Rule-Based Expert Systems: The MYCIN Experiments of the Stanford Heuristic Programming Project. Reading, Mass.: Addison-Wesley, 1984. https://www.shortliffe.net/Buchanan-Shortliffe-1984/MYCIN%20Book.htm
- Lighthill, James. “Artificial Intelligence: A General Survey.” In Artificial Intelligence: a Paper Symposium. London: Science Research Council, 1973. https://www.aiai.ed.ac.uk/events/lighthill1973/lighthill.pdf
Historical reference
- Nilsson, Nils J. The Quest for Artificial Intelligence: A History of Ideas and Achievements. Cambridge: Cambridge University Press, 2010. (Chapters on DENDRAL, consulting/expert systems, Lighthill, and “The AI Winter.”) Web version: https://ai.stanford.edu/~nilsson/QAI/qai.pdf
- Crevier, Daniel. AI: The Tumultuous History of the Search for Artificial Intelligence. New York: Basic Books, 1993.
- Agar, Jon. “What is science for? The Lighthill report and the purpose of artificial intelligence research.” British Journal for the History of Science (advance/online versions discuss SRC/SERC effects). UCL discovery PDF: https://discovery.ucl.ac.uk/id/eprint/10105726/3/Agar_What%20is%20science%20for%2C%20the%20Lighthill%20report%20and%20the%20purpose%20of%20artificial%20intelligence%20research%2C%20revised%20for%20BJHS.pdf
- Feigenbaum, Edward A., and Pamela McCorduck. The Fifth Generation: Artificial Intelligence and Japan’s Computer Challenge to the World. Reading, Mass.: Addison-Wesley, 1983.
- Hayes-Roth, Frederick, Donald A. Waterman, and Douglas B. Lenat, eds. Building Expert Systems. Reading, Mass.: Addison-Wesley, 1983.