Bitware News

Tag

Eval

2026-09-24

Transformers and Scaling

The 2017 transformer reorganized sequence modeling around self-attention without recurrence; GPT and BERT are two objectives on that backbone; Kaplan and Chinchilla scaling laws are empirical regularities about compute, data, and parameters—not a slogan that bigger is always better.

2026-09-22

The Deep Learning Turn

What moved in 2012 was not a sudden invention of deep nets: ImageNet-scale labeled data, GPU training, ReLU, and dropout made a large ConvNet win a public, comparable benchmark—AlexNet as the ImageNet moment, not the first deep network.

2026-09-13

Cros: risk-constrained stopping for clinical diagnosis agents

A stopping layer that tries to decide when a sequential diagnosis agent should diagnose or defer — with finite-sample style tests, and an honest ‘exploratory, not confirmatory’ framing.

2026-09-13

Φ-Bench: can LLMs engineer the infrastructure that runs them?

An 85-task infra/eval benchmark spanning kernels, long-horizon repo work, and end-to-end optimization — topped by Claude Opus 5 at 36.53%, with a lot of headroom left.

2026-09-13

TabBench-Bio: evaluating ML where biomedicine actually lives — wide tables, tiny n

A living eval suite of 43 biomedical HDLSS datasets across 28 feature×sample operating points — tabular foundation models lead at the reference cell, but logistic regression stays uncomfortably close.

2026-09-13

Why this site exists

AI is more than a chat box on the web. Bitware News covers what the technology actually does — methods, medical ML, evaluation — with primary sources and caveats.