Cros: risk-constrained stopping for clinical diagnosis agents
A stopping layer that tries to decide when a sequential diagnosis agent should diagnose or defer — with finite-sample style tests, and an honest ‘exploratory, not confirmatory’ framing.

Scientific AI and medical ML — methods, evaluation, clinical systems. Primary sources and caveats. Not another chat-demo headline factory.
Beats

Models, training tricks, representations — what changed in the paper.

Imaging, reports, clinical decision support — with measured limits.
From cybernetics to modern training — the scientific serial.
History
All chaptersBefore the name 'artificial intelligence,' research programs already treated mind and machine as the same kind of problem—logical neurons, feedback control, and operational tests of machine intelligence.
The August 1955 proposal and the 1956 Dartmouth workshop named and organized a research program under the phrase 'artificial intelligence'—they did not invent the problems surveyed in the previous chapter.
Rosenblatt’s perceptron made learning in a neuron-like device an empirical research program; Minsky and Papert’s 1969 book proved sharp limits of a restricted class of machines—not that ‘neural nets are dead.’
After Dartmouth and alongside the perceptron line, a large part of AI became knowledge plus search—GPS, DENDRAL, MYCIN, and the expert-system boom—while the journalistic label “AI winter” names real but local funding contractions, not a single global morality play.
In the 1980s, distributed representations returned as an empirical research program—Hopfield networks, Boltzmann machines, and the PDP volumes—with new mathematics of energy, attractors, and hidden units, not a magical rebirth after a total death of neural nets.
Backpropagation—efficient reverse-mode differentiation through multilayer nets—made hidden-unit learning a community tool in 1986, after an earlier paper trail in automatic differentiation and ordered derivatives that was not invented out of nowhere that year.
In the 1990s, VC theory, soft-margin support-vector machines, and kernel methods offered a statistical-learning alternative to under-regularized multilayer nets—strong baselines with clearer capacity control, not a permanent replacement for neural networks.
Latest
A stopping layer that tries to decide when a sequential diagnosis agent should diagnose or defer — with finite-sample style tests, and an honest ‘exploratory, not confirmatory’ framing.
An 85-task infra/eval benchmark spanning kernels, long-horizon repo work, and end-to-end optimization — topped by Claude Opus 5 at 36.53%, with a lot of headroom left.
A scientific-discovery methods paper that treats ‘how many independent laws?’ as the zeroth step — then shows multi-law fits beating single-equation symbolic regression on engineering and galactic data.
Self-hosted Llama extracts structured T-stage from radiology reports with source-text links — 90% accuracy vs a four-expert reference on 130 reports.
AI is more than a chat box on the web. Bitware News covers what the technology actually does — methods, medical ML, evaluation — with primary sources and caveats.