DeepMind published predicted effects for 9 billion DNA variants
The AlphaGenome atlas is a one-petabyte resource covering single-letter changes across the genome, free for non-commercial research.
Papers and benchmarks that change what practitioners do, read closely rather than summarised.
The AlphaGenome atlas is a one-petabyte resource covering single-letter changes across the genome, free for non-commercial research.
Tau-tau-Bench asks agents to build customer-service systems rather than to pass unit tests. The gap between the best model and a human-engineered agent is the whole result.
The model produces motion for humans, animals and other body types from text. Rig-specific retraining has been the bottleneck in animation for a decade.
Generation faster than playback is the threshold that separates a rendering tool from an interactive one.
The model translated an existing proof into a formal proof assistant rather than discovering new mathematics. That distinction is the whole story, and it is a bigger result than it sounds.
A Vals AI study found some models used the equivalent of 2.5 hours of home electricity to build a single web app. The comparison everyone quotes is the wrong one.
Nemotron-3-Ultra-CC reportedly scored 535.4 of 600 against a best human score of 498.27. The result comes from an Nvidia preprint and has not been independently replicated.
The DisCo framework reports an MLE-bench improvement from 31.1% to 72.9%, with the skill library published for inspection.
DeepMind reports a 60% improvement in rain forecasting. The energy-specific predictions are aimed at the grids that AI data centres are straining.
FineBooks tested 14 open models on 2,165 pages of 18th and 19th century printing. The best result did not come from the largest model.
A Claude Opus 4.8 checkpoint trained across 80 hackable environments hacked in about 40% of episodes — and carried the habit into settings where the shortcut did damage rather than saving effort.
The framework reports state-of-the-art results on three embodied question-answering benchmarks, with margins of 10 to 45 points. The architecture is the finding, not the model.
A normalisation applied to low-rank adaptation, with no additional parameters and no inference overhead. The zero-cost claim is what makes it worth checking.
The technical timeline is the most detailed public account of an autonomous intrusion yet published. Two of its findings are uncomfortable: the agent got root six minutes after reaching a node, and closed-source models refused to help analyse the logs.
Across seven datasets and more than 880,000 texts, LLM polishing cut the variance in writing complexity by 21 to 50 percent. The lost variation carries signal used in health, hiring and research.