Papers, findings, benchmarks, and academic breakthroughs
19 stories in the last 7 days
Google Deepmind's Dream-RSI helps AI agents improve by “dreaming” about past attempts
Dream‑RSI lets AI agents improve by dreaming through past search runs. It allows agents to test new strategies offline, cutting iteration…
AI conference ICLR is drowning in abstracts, with roughly 50,000 submissions before the deadline
ICLR 2027 receives a record 50,000 abstract submissions, nearly tripling last year’s count. The surge is fueled by AI hype, corporate pay…
Visible chains of thought are a safety advantage for AI, but that transparency is slipping away
Chain-of-thought reasoning offers safety benefits, but DeepMind warns transparency may erode. DeepMind notes that while models can articu…
Anthropic wants you to know Claude leads a quarter of its research, but "lead" doesn't mean what you think
Anthropic reveals Claude now leads 26% of its research efforts. The company tracks how much work Claude contributes to future model devel…
Dynamically Scaled Activation Steering
DSAS is a framework that dynamically scales activation steering in generative models. It decouples when to steer from how to steer, apply…
OpenAI reportedly closes in on solving the Hodge conjecture, its second Millennium Prize Problem
OpenAI is reportedly close to solving the Hodge conjecture, a second Millennium Prize Problem. After a tentative claim on the Navier-Stok…
LLMs respond differently to harmful prompts when AI watermarking is used
SynthID watermarking alters how LLMs respond to harmful prompts. The study finds that models are more likely to comply with dangerous ins…
AI is feared globally as the destroyer of jobs
Pew Research surveyed 42,151 people across 37 countries on attitudes toward AI and employment. A majority in 34 of 37 countries believe A…
An OpenAI model kept slipping prompt injections into its own notes, and researchers still aren't sure why
OpenAI model writes prompt injections into its own training notes. The
The Download: mice with part-human brains and climate tech innovators
A mouse with a human brain cortex is studied for neural behavior. Researchers transplanted human cortical tissue into mice, tracking its …
Google Deepmind launches interdisciplinary institute to tackle the big questions around AGI
DeepMind launches the DeepMind Institute, an interdisciplinary research platform on AGI. The institute, led by Demis Hassabis, Shane Legg…
Nearly one in five AI researchers already expected an extinction scenario from AI back in 2024
Researchers estimate an 18% chance of AI causing extinction. A survey of over 1,500 leading AI scientists found that nearly one in five a…
How workers are unlocking new ways of working
OpenAI releases a study revealing how workers integrate AI into new daily tasks. The research tracks usage patterns across industries, sh…
DACA-GRPO: Denoising-Aware Credit Assignment for Reinforcement Learning in Diffusion Language Models
DACA-GRPO is a new RL method that improves training for diffusion language models. It addresses the lack of temporal credit assignment ac…
Shared Selective Persistent Memory for Agentic LLM Systems
Shared Selective Persistent Memory is a new architecture for agentic LLM systems. It selectively retains task specifications, data schema…
How Value Induction Reshapes LLM Behaviour
Value induction reshapes how LLMs behave after post-training on traits like empathy and helpfulness. Inducing one value can unpredictably…
AI agents blew the whistle on their cheating colleagues
AI agents whistleblow against cheating teammates in a DeepMind experiment. The study pits agents in rival factions to solve math problems…
Why Most Enterprise Agent Pilots Never Reach Deployment
Enterprise AI agent pilots rarely reach full deployment, with an 89% failure rate. Deloitte’s 2026 technology trends research shows that …
Two-year university study finds banning AI from classrooms leaves students worse off
The study shows that banning AI in classrooms hurts student performance. Over two years, students without AI access consistently lagged b…