Tag
Eval
2026-09-24
The 2017 transformer reorganized sequence modeling around self-attention without recurrence; GPT and BERT are two objectives on that backbone; Kaplan and Chinchilla scaling laws are empirical regularities about compute, data, and parameters—not a slogan that bigger is always better.
2026-09-22
What moved in 2012 was not a sudden invention of deep nets: ImageNet-scale labeled data, GPU training, ReLU, and dropout made a large ConvNet win a public, comparable benchmark—AlexNet as the ImageNet moment, not the first deep network.
2026-09-13
A stopping layer that tries to decide when a sequential diagnosis agent should diagnose or defer — with finite-sample style tests, and an honest ‘exploratory, not confirmatory’ framing.
2026-09-13
An 85-task infra/eval benchmark spanning kernels, long-horizon repo work, and end-to-end optimization — topped by Claude Opus 5 at 36.53%, with a lot of headroom left.
2026-09-13
A living eval suite of 43 biomedical HDLSS datasets across 28 feature×sample operating points — tabular foundation models lead at the reference cell, but logistic regression stays uncomfortably close.
2026-09-13
AI is more than a chat box on the web. Bitware News covers what the technology actually does — methods, medical ML, evaluation — with primary sources and caveats.