Bitware News

Series

History of AI / ML

From cybernetics and Dartmouth through perceptrons, expert systems, connectionism, and modern training methods — primary papers, measured claims, caveats.

  1. 01

    Cybernetics and the First Models

    Before the name 'artificial intelligence,' research programs already treated mind and machine as the same kind of problem—logical neurons, feedback control, and operational tests of machine intelligence.

    2026-09-13
  2. 02

    Dartmouth Names a Research Program

    The August 1955 proposal and the 1956 Dartmouth workshop named and organized a research program under the phrase 'artificial intelligence'—they did not invent the problems surveyed in the previous chapter.

    2026-09-14
  3. 03

    Perceptrons, and the Critique

    Rosenblatt’s perceptron made learning in a neuron-like device an empirical research program; Minsky and Papert’s 1969 book proved sharp limits of a restricted class of machines—not that ‘neural nets are dead.’

    2026-09-15
  4. 04

    Expert Systems and a Funding Winter

    After Dartmouth and alongside the perceptron line, a large part of AI became knowledge plus search—GPS, DENDRAL, MYCIN, and the expert-system boom—while the journalistic label “AI winter” names real but local funding contractions, not a single global morality play.

    2026-09-16
  5. 05

    Connectionism Returns

    In the 1980s, distributed representations returned as an empirical research program—Hopfield networks, Boltzmann machines, and the PDP volumes—with new mathematics of energy, attractors, and hidden units, not a magical rebirth after a total death of neural nets.

    2026-09-17
  6. 06

    Backpropagation

    Backpropagation—efficient reverse-mode differentiation through multilayer nets—made hidden-unit learning a community tool in 1986, after an earlier paper trail in automatic differentiation and ordered derivatives that was not invented out of nowhere that year.

    2026-09-18
  7. 07

    Kernels and the Statistical Turn

    In the 1990s, VC theory, soft-margin support-vector machines, and kernel methods offered a statistical-learning alternative to under-regularized multilayer nets—strong baselines with clearer capacity control, not a permanent replacement for neural networks.

    2026-09-19
  8. 08

    How We Train: GD to Adam

    Gradient descent, stochastic approximation, momentum, regularization, and Adam form a cumulative training-methods line—engineering and theory, not a folklore of sudden deep-learning inventions.

    2026-09-20
  9. 09

    Sequences Before Attention

    Before transformers, language and other sequences became trainable transduction problems through simple recurrent nets, LSTM, encoder–decoder seq2seq, and Bahdanau’s soft alignment—attention as an RNN add-on, not a 2017 invention.

    2026-09-21
  10. 10

    The Deep Learning Turn

    What moved in 2012 was not a sudden invention of deep nets: ImageNet-scale labeled data, GPU training, ReLU, and dropout made a large ConvNet win a public, comparable benchmark—AlexNet as the ImageNet moment, not the first deep network.

    2026-09-22
  11. 11

    Pretrain, Finetune, Self-Supervision

    Supervision is scarce: the modern stack learns reusable representations from large unlabeled or weakly labeled data—layerwise pretraining, ImageNet transfer, word2vec, BERT-style masked LMs, and contrastive vision—then adapts them, without collapsing the lineage into a BERT founding myth.

    2026-09-23
  12. 12

    Transformers and Scaling

    The 2017 transformer reorganized sequence modeling around self-attention without recurrence; GPT and BERT are two objectives on that backbone; Kaplan and Chinchilla scaling laws are empirical regularities about compute, data, and parameters—not a slogan that bigger is always better.

    2026-09-24