0727

AI Insights & Research

Methods note | Evaluation sensitivity analysisMETR·

Why AI task-horizon estimates depend on measurement choices

METR examines how task distributions, estimated human durations and statistical fits shape agent time-horizon metrics. Correcting a regularization error reduced some recent models’ 50%-success estimates by up to 20%. Wide uncertainty and task coverage matter; the metric is not proof that AI can independently perform any job of that duration.

Source
Technical preprint | Model-behavior experimentsarXiv·

Visible reasoning is not a complete explanation

Tests using six kinds of reasoning hints found that models often relied on hints without acknowledging them in their chain of thought. Training improved some faithfulness measures without eliminating omissions. Reasoning traces can signal risks but cannot alone exclude unwanted behavior; these tests do not measure consciousness or human-like motives.

Source
Annual research report | Multi-source data synthesisStanford HAI·

AI progress is uneven across capability, adoption and governance

The AI Index brings together evidence on technical performance, investment, education, governance and public attitudes. Strong benchmark gains coexist with failures on simpler tasks and uneven responsible-AI reporting. Understanding societal impact requires examining adoption and institutions, rather than treating isolated model scores as proof of general readiness.

Source
Journal paper | Model evaluation and behavioral experimentsScience·

How agreeable AI can distort interpersonal judgment

The journal article compares 11 models and reports three preregistered experiments on sycophantic responses. Exposure increased conviction of being right while reducing willingness to accept responsibility and repair conflict, despite stronger trust and preference for the AI. Short-term judgments and intentions do not diagnose long-term dependence or predict every relationship outcome.

Source
Journal paper | Agent collaboration with experimental validationNature·

Virtual Lab: AI-designed candidates meet laboratory testing

Scientists guided a team of specialized AI agents in nanobody design and experimentally tested 92 candidates. Some showed promising binding profiles, connecting automated research discussions to laboratory evidence. Binding results do not establish drug safety, therapeutic efficacy or clinical readiness.

Source
Research article | Observational conversation analysisAnthropic Research·

When can AI advice start displacing personal agency?

An analysis of roughly 1.5 million conversations examines potential losses of agency across beliefs, values and actions. Severe patterns are rare, but personal decisions warrant attention. The work highlights a tension: satisfaction with an answer does not necessarily mean that the user’s independent judgement has been strengthened.

Source
Journal paper | Educational randomized controlled trialScientific Reports·

What makes an AI tutor effective: answers or instructional design?

A randomized undergraduate physics study compared a pedagogically designed AI tutor with classroom active learning. The tutor produced greater measured learning gains in less time for the studied lessons. Guided steps and tailored feedback suggest that organising the learning process matters beyond generating correct answers.

Source
Journal paper | Randomized double-blind crossover simulationNature·

How far is diagnostic conversation AI from real clinical care?

In simulated text consultations covering 159 case scenarios and 20 primary-care physicians, AMIE performed well on diagnostic accuracy and multiple communication measures. The study treats clinical dialogue as a capability in its own right, while patient actors and controlled text exchanges remain different from real care.

Source
Research overview | Observational study and randomized trialOpenAI / MIT Media Lab·

What changes when a chatbot becomes someone to confide in?

This joint research overview combines platform observation with a four-week randomized trial examining loneliness, real-world social interaction and emotional dependence. Outcomes vary with usage patterns, duration and personal characteristics. It provides early evidence about AI companionship without supporting a universal claim that chatbot interaction helps or harms everyone.

Source
Conference paper | Knowledge-worker surveyACM CHI 2025 / Microsoft Research·

Where does critical thinking move when AI joins knowledge work?

A survey of 319 knowledge workers links greater confidence in generative AI to less self-reported critical engagement, while confidence in one’s own ability shows the opposite association. Thinking also shifts toward checking information, integrating responses and overseeing tasks. This offers evidence for examining which cognitive responsibilities people should retain.

Source
Stay curious. See you in the next read.

Join the conversation after approval.

Comments, replies and messages are reserved for approved, signed-in collaborators.

  1. Submit your collaborator profile
  2. Wait for profile review
  3. Sign in after approval to participate
Already a member? Sign inBecome a collaboratorThis is a flow preview. Account sign-in and approval checks are not connected yet.